Atika88/Indonesian-ASR-11-Class-Dataset
Indonesian ASR 11-Class Dataset Public Hugging Face repository for an Indonesian ASR corpus and its paper-supporting benchmark artifacts. Dataset summary Audio files: 104,500 WAV files Real/human recordings: 104,368 Synthetic repair files: 132 Sentence classes: 11 Indonesian sentence categories Canonical balanced sentence slots: 209 (11 categories × 19 retained slots) Public speaker labels: M1..M12, F1..F8, plus synthetic labels Ms*/Fs* Audio format: 16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/Atika88/Indonesian-ASR-11-Class-Dataset.
Indonesian ASR 11-Class Dataset
Public Hugging Face repository for an Indonesian ASR corpus and its paper-supporting benchmark artifacts.
Dataset summary
- Audio files: 104,500 WAV files
- Real/human recordings: 104,368
- Synthetic repair files: 132
- Sentence classes: 11 Indonesian sentence categories
- Canonical balanced sentence slots: 209 (11 categories × 19 retained slots)
- Public speaker labels:
M1..M12,F1..F8, plus synthetic labelsMs*/Fs* - Audio format: 16 kHz, 16-bit, mono
- Total duration: about 134.18 hours
- Splits: speaker-disjoint train/val/test, seed 42
The public metadata uses pseudonymous labels only. Original respondent names and private identity crosswalks are not included in this repository. Speech remains a potentially identifying biometric signal; pseudonymization does not make the recordings anonymous.
Category naming
All public Hugging Face category names are in English. See docs/CATEGORY_NAMING.md for the public English category list. The category column in metadata/dataset_metadata_public.csv, transcript filenames, and audio shard filenames use English names.
Important transcript numbering note
Sentence IDs intentionally preserve the original collection numbering (01–20). The canonical balanced design contains 209 sentence slots (19 retained slots in each of 11 categories). Some category transcript lists skip one original ID, for example Clarification skips 09. These gaps are expected and come from curation/removal of duplicate or balancing sentences; they do not mean that the HF upload is incomplete. Original IDs and replacement provenance remain preserved separately; do not renumber sentence IDs when using or citing the dataset.
For the full per-category inventory, read:
docs/TRANSCRIPT_NUMBERING_NOTES.md
metadata/transcript_sentence_inventory_public.csvFor experiments, always use metadata/dataset_metadata_public.csv as the row-level source of truth.
Repository structure
data/
audio_shards/by_category/*.tar # pseudonymized WAV tree packaged by sentence category
audio_shards/audio_shards_manifest.csv # shard index
transcripts/ # sentence transcripts by category
metadata/
dataset_metadata_public.csv # public metadata with pseudonymous labels
regional_background/README.md # author-reported background, not verified dialect labels
speaker_labels/ # public label inventories/schema
splits/
speaker_split_assignment_public.csv
split_summary_public.json
checksums/
audio_shards.sha256 # SHA-256 for the 11 category TAR archives
paper/
dataset_information/ # full-scope public dataset statistics
dataset_information/figures_public/ # regenerated public-label figures
docs/
CITATION.md
DOI_RELEASE_GATE.md
RELEASE_EVIDENCE_AND_METHOD_BOUNDARIES.md
HF_DATASET_INFORMATION_FINAL_REPORT.md
HF_DATASET_INFORMATION_SELECTION.mdSplit naming
Active machine-readable split values are train, val, and test; val is the validation partition. Revisions before this migration used dev for the same middle partition. Pin an older revision only when exact historical replay is required.
Dataset update notes
See docs/DATASET_UPDATE_NOTES.md for the 2026-06-18 transcript cleanup note. In short: transcript text files are clean public sentence lists; metadata/dataset_metadata_public.csv is the row-level source of truth. Audio shards and paths were not changed by the cleanup.
Dataset Viewer
The default Data Studio configuration is backed by three Parquet files under viewer/metadata/, one for each canonical split (train, val, and test). These files are deterministic viewer projections of metadata/dataset_metadata_public.csv; the CSV remains the row-level source of truth. Auxiliary CSV reports remain downloadable repository artifacts and are intentionally excluded from the default viewer configuration because they have different schemas.
Loading metadata
Use metadata/dataset_metadata_public.csv. Audio is stored as category-level tar shards under data/audio_shards/by_category/. Extract shards into your working directory to materialize paths such as data/processed_balanced19_v7_natural_synth/Dataset_Balanced19/.... Example:
mkdir -p extracted
tar -xf data/audio_shards/by_category/Declarative.tar -C extractedThen join extracted/ with the audio_path column. Important columns:
audio_path: relative path after extracting the relevant tar shard to the repository rootsplit:train,val, ortestspeaker_id: public acoustic-source label (M*,F*,Ms*,Fs*)speaker_type:humanorsyntheticspeaker_gender: public acoustic-source gender labelrepair_target_speaker_id: target human label for synthetic repair rowsvoice_gender_matches_target: whether synthetic voice gender matches repair target gendertranscript: reference transcription
Evidence and method boundaries
The public package contains final utterance-level WAV files; original continuous pre-segmentation recordings are not included. See `docs/RELEASE_EVIDENCE_AND_METHOD_BOUNDARIES.md` for the reproducible release facts, frozen-benchmark boundary, synthetic-repair disclosure, and acquisition-time details that are not established by the retained public evidence. Author-reported regional backgrounds are documented separately and must not be treated as verified dialect, ethnicity, or recording-level annotations.
License
Dataset-owned materials—including human recordings, transcripts, metadata, and prompts—are made available under the Creative Commons Attribution 4.0 International (CC BY 4.0) license to the extent the dataset owner is authorized to license them. The 132 historical synthetic repair recordings remain subject to the third-party-output rights review described in `docs/SYNTHETIC_AUDIO_RIGHTS_AND_AZURE_EVIDENCE.md`. An Azure subscription and the edge-tts client license do not by themselves establish that Microsoft applied CC BY 4.0 or granted the full downstream rights claimed by that license.
You may share and adapt dataset-owned material under CC BY 4.0, including commercially, provided that appropriate attribution is given and modifications are indicated. Do not interpret this statement as a Microsoft license, sponsorship, or endorsement of the historical synthetic outputs.
- Full license text: `LICENSE`
- Official license page: https://creativecommons.org/licenses/by/4.0/
- Source and reproducibility repository: https://github.com/RatnaAtika/Indonesian-ASR-11-Class-Dataset
Citation
The dataset DOI is [10.57967/hf/10345](https://doi.org/10.57967/hf/10345). DataCite registered Hugging Face revision c65fe8bcff0547214c34cfbea248b4045a0d867c on 9 September 2026. Cite that revision when exact reproducibility is required.
Creator order registered with DataCite: Ratna Atika; Suci Dwijayanti; Bhakti Yudho Suprapto. See `docs/CITATION.md`, `CITATION.cff`, and `docs/DOI_RELEASE_GATE.md`. Later repository revisions do not change the revision identified by this DOI and should be identified separately.
Synthetic-audio rights and Azure evidence
The 132 synthetic repair files were historically generated through Edge Read Aloud/edge-tts provenance and are present in all splits: 122 train, 8 val, and 2 test. They have not been verified as paid-tier Azure Speech outputs.
A read-only check found an enabled Azure subscription and availability of the Speech S0 SKU, but no Azure Speech resource or paid-generation receipt for these files. Subscription availability is technical readiness, not retrospective permission. Public redistribution, downstream commercial/model-training use, derivatives, and CC BY 4.0 compatibility for the historical outputs therefore remain not verified pending written Microsoft/controlling-agreement evidence and a signed institutional rights determination, or paid-tier regeneration and release of a reviewed new revision.
See `docs/SYNTHETIC_AUDIO_RIGHTS_AND_AZURE_EVIDENCE.md` for the evidence checklist and license boundary. Private agreements, signatures, account identifiers, billing records, correspondence, tokens, and keys are not published.
Caveats
- The 132 explicitly flagged synthetic repair files are distributed across all canonical splits: 122
train, 8val(validation), and 2test. Their total duration is 632.520 seconds. Select them withis_synthetic == Trueinmetadata/dataset_metadata_public.csv. - The two
testsynthetic repair files target M8 and are explicitly flagged withvoice_gender_matches_target=False; downstream users requiring strict gender-matched synthetic repair should exclude or regenerate those rows. - The release contains prompted read speech and overlapping prompt content across partitions. Stored benchmark results should not be interpreted as open-vocabulary, unseen-prompt, verified-dialect, demographic, or deployment generalization.
- Public speaker labels are pseudonyms; voice recordings remain potentially identifying.
paper/dataset_information/is generated from the full 104,500-file public metadata. Older paper-clean statistics are not used as full-scope statistics unless clearly labeled as a subset.
Path privacy migration and DOI status
The active tree uses sanitized path/container identifiers. DOI 10.57967/hf/10345 identifies revision c65fe8bcff0547214c34cfbea248b4045a0d867c. Zenodo remains on hold to avoid a duplicate dataset DOI. Voice labels are pseudonymous, not anonymous. Institutional and Microsoft/Azure synthesis-rights documentation is being completed as a follow-up governance record; private signed approvals, account identifiers, billing records, and correspondence are not published in this dataset. Any later sanitized public summary or paid-tier synthetic replacement will be a subsequent repository revision and will not retroactively modify or license the DOI-registered historical files. See `docs/DOI_RELEASE_GATE.md`, `docs/ETHICS_AND_CONSENT_SCOPE.md`, and `docs/SYNTHETIC_AUDIO_RIGHTS_AND_AZURE_EVIDENCE.md`.
