Team Ai
Datasetpublic

quangdung/gigaspeech2-vi-missing-transcripts

GigaSpeech2 Vietnamese WAVs Missing Transcripts This dataset contains the Vietnamese GigaSpeech2 training WAV files whose IDs are absent from train_refined.tsv. Contents 201,295 WAV files without a matching transcript 41 uncompressed TAR shards in shards/ 36 GB of audio (approximately) processing_manifest.jsonl: per-source-archive counts summary.json: aggregate counts SHA256SUMS: checksums for all TAR shards Each shard preserves the original relative layout:… See the full description on the dataset page: https://huggingface.co/datasets/quangdung/gigaspeech2-vi-missing-transcripts.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes40downloads
Dataset Card

GigaSpeech2 Vietnamese WAVs Missing Transcripts

This dataset contains the Vietnamese GigaSpeech2 training WAV files whose IDs are absent from train_refined.tsv.

Contents

  • —201,295 WAV files without a matching transcript
  • —41 uncompressed TAR shards in shards/
  • —36 GB of audio (approximately)
  • —processing_manifest.jsonl: per-source-archive counts
  • —summary.json: aggregate counts
  • —SHA256SUMS: checksums for all TAR shards

Each shard preserves the original relative layout:

text
release1/train/<archive>/<subfolder>/<utterance-id>.wav

The TAR files are intentionally uncompressed so they can be read sequentially and streamed without an additional decompression layer.