quangdung/gigaspeech2-vi-missing-transcripts
GigaSpeech2 Vietnamese WAVs Missing Transcripts This dataset contains the Vietnamese GigaSpeech2 training WAV files whose IDs are absent from train_refined.tsv. Contents 201,295 WAV files without a matching transcript 41 uncompressed TAR shards in shards/ 36 GB of audio (approximately) processing_manifest.jsonl: per-source-archive counts summary.json: aggregate counts SHA256SUMS: checksums for all TAR shards Each shard preserves the original relative layout:… See the full description on the dataset page: https://huggingface.co/datasets/quangdung/gigaspeech2-vi-missing-transcripts.
GigaSpeech2 Vietnamese WAVs Missing Transcripts
This dataset contains the Vietnamese GigaSpeech2 training WAV files whose IDs are absent from train_refined.tsv.
Contents
- 201,295 WAV files without a matching transcript
- 41 uncompressed TAR shards in
shards/ - 36 GB of audio (approximately)
processing_manifest.jsonl: per-source-archive countssummary.json: aggregate countsSHA256SUMS: checksums for all TAR shards
Each shard preserves the original relative layout:
release1/train/<archive>/<subfolder>/<utterance-id>.wavThe TAR files are intentionally uncompressed so they can be read sequentially and streamed without an additional decompression layer.
