Team Ai
Datasetpublic

vennsa/SLR-NoteSense

SLR-NoteSense SLR-NoteSense is a dual-sensor red-green-blue-clear (RGBC) dataset for Sri Lankan banknote denomination recognition. The seven-denomination update adds the LKR 2000 class to the original six-denomination dataset. The dataset contains 14,846 paired acquisition rows from 743 distinct physical banknotes, covering LKR 20, 50, 100, 500, 1000, 2000, and 5000. After applying the documented sensor-settling rule, the stable view contains 14,104 paired measurement rows.… See the full description on the dataset page: https://huggingface.co/datasets/vennsa/SLR-NoteSense.

sourceHugging Facecc-by-4.0updated 15d agoView on Hugging Face
1likes118downloads
Dataset Card

SLR-NoteSense

SLR-NoteSense is a dual-sensor red-green-blue-clear (RGBC) dataset for Sri Lankan banknote denomination recognition. The seven-denomination update adds the LKR 2000 class to the original six-denomination dataset.

The dataset contains 14,846 paired acquisition rows from 743 distinct physical banknotes, covering LKR 20, 50, 100, 500, 1000, 2000, and 5000. After applying the documented sensor-settling rule, the stable view contains 14,104 paired measurement rows.

Dataset Details

Two TCS34725 color sensors connected to an ESP32-S3 collected red, green, blue, and clear-channel measurements.

PropertyValue
Sensors2 × TCS34725
Integration time50 ms
Gain4×
Denominations7
Distinct physical banknotes743
Raw paired measurements14,846
Stable paired measurements14,104
Derived note-level samples743

Dataset Structure

Each raw measurement row contains the following fields:

  • —label: Banknote denomination in LKR; missing in some original source rows.
  • —note_id: Physical-banknote identifier within its denomination.
  • —scan_id: Acquisition identifier.
  • —sample_index: Ordered position within an acquisition sequence.
  • —s1_r, s1_g, s1_b, s1_clear: Sensor 1 raw RGBC measurements.
  • —s1_nr, s1_ng, s1_nb: Sensor 1 R/C, G/C, and B/C ratios.
  • —s2_r, s2_g, s2_b, s2_clear: Sensor 2 raw RGBC measurements.
  • —s2_nr, s2_ng, s2_nb: Sensor 2 R/C, G/C, and B/C ratios.

For each sensor, normalized channels are calculated as R/C, G/C, and B/C, where C is the clear-channel reading. Each row contains a paired reading from both sensors. Use (`label`, `note_id`) as the globally unique physical-banknote key because note_id restarts within each denomination.

Physical Banknote Distribution

Denomination (LKR)Physical banknotesRaw rowsStable rows
201022,0381,936
501022,0381,936
1001062,1182,013
5001112,2182,107
10001062,1182,012
20001092,1782,069
50001072,1382,031
Total74314,84614,104

The new 2000.csv contributes 109 distinct LKR 2000 notes. An older eight-note LKR 2000 file is not included because its note identifiers overlap with those in the newer file and the physical-note identities are unverified. Do not merge the two files without resolving this overlap.

Preprocessing and Quality Control

  1. 1.Retain the seven original denomination-specific CSV files unchanged.
  2. 2.Restore 380 missing `label` values from known source-file provenance: 60 from the original six-denomination files and 320 from the new LKR 2000 file (16 complete scans). Do not infer these labels from sensor readings. The cleaned data includes an original_label_missing provenance flag.
  3. 3.Remove rows with sample_index == 0 from the derived stable view to exclude the observed sensor-settling transient. Apply this rule only where index 0 exists. 742 of 743 physical-note acquisitions contain index 0; one LKR 100 acquisition starts at index 2.
  4. 4.Preserve the original sequence lengths. Seven physical-note acquisitions have 18 rather than 20 raw readings.
  5. 5.For note-level classification, compute eight mean features from stable readings: the R/C, G/C, and B/C ratios for each sensor and the clear-channel mean for each sensor. The reference pipeline recomputes ratios from raw RGBC channels.

All clear-channel readings are positive. The source release and derived views should remain distinguishable so preprocessing is reproducible.

Baseline Technical Validation

The seven-class reference baseline uses a StandardScaler followed by an RBF support vector machine (C=5, gamma=0.06) with eight note-level features.

Evaluation uses a stratified physical-banknote-level split with 594 training notes and a separate 149-note holdout. Five repeats of stratified fivefold cross-validation run only on the training partition. No physical note appears in both partitions.

EvaluationAccuracyMacro F1
Training-only repeated fivefold CV (5 × 5)98.15%98.14%
Separate 149-note holdout96.64%96.60%

Python-to-C++ prediction parity was 743/743 for note-level feature vectors and 742/742 for eligible raw-reading groups. This checks implementation agreement, not classification accuracy or performance on physical ESP32 hardware. After holdout evaluation, the companion deployment model was refitted on all 743 notes.

These results describe the recorded sensor setup and acquisition conditions; they do not establish performance under unmeasured banknote wear, lighting, orientation, or hardware variations.

Intended Uses

  • —Banknote denomination recognition.
  • —Embedded machine learning and RGBC sensor classification.
  • —Assistive currency-recognition research.
  • —Feature engineering and acquisition-sequence analysis.
  • —Leakage-free, physical-banknote-aware model evaluation.

Out-of-Scope Uses

  • —Counterfeit detection or banknote authentication.
  • —Financial security verification.
  • —Banknote valuation.

License

The dataset is licensed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license. Software code, if distributed, should have its own explicitly stated license.

Citation

The permanent dataset DOI and associated data-paper citation will be added after publication.

Dataset Version

Version 2.0 (release candidate): Seven-denomination update including 2000.csv. Version and citation details should be finalized with the published dataset record.