Team Ai
Datasetpublic

Hongthien06/vispeech-whisper-features

ViSpeech — Preprocessed Whisper Features (Reproducibility Release) This repository contains the preprocessed feature sets used in the paper: ViSpeech: A Multi-Condition Vietnamese Speech Dataset for Noise-Robust Classroom ASR — T. D. Tran et al., FISAT 2026 (Springer). ViSpeech is a 37.62-hour multi-condition Vietnamese speech corpus built on a paired-recording design — the same speakers reading the same scripts in quiet-room and real-classroom conditions — comprising a… See the full description on the dataset page: https://huggingface.co/datasets/Hongthien06/vispeech-whisper-features.

sourceHugging Facecc-by-nc-sa-4.0updated 10d agoView on Hugging Face
0likes246downloads
Dataset Card

ViSpeech — Preprocessed Whisper Features (Reproducibility Release)

This repository contains the preprocessed feature sets used in the paper:

ViSpeech: A Multi-Condition Vietnamese Speech Dataset for Noise-Robust Classroom ASR — T. D. Tran et al., FISAT 2026 (Springer).

ViSpeech is a 37.62-hour multi-condition Vietnamese speech corpus built on a paired-recording design — the same speakers reading the same scripts in quiet-room and real-classroom conditions — comprising a 36.50-hour noise-augmented training corpus, a 1.12-hour held-out validation set, and an independently collected 221-utterance real-noise classroom benchmark.

⚠️ Format notice

These are not raw audio files. Every example holds:

  • —input_features — log-mel spectrograms produced by the openai/whisper-small feature extractor;
  • —labels — target token IDs produced by the openai/whisper-small tokenizer.

This release is intended for exact reproduction of the paper's training and evaluation pipeline with Whisper-small. Transcripts can be recovered by decoding labels with the Whisper tokenizer; the original waveforms cannot be recovered from log-mel features.

Configurations

ConfigSplitDescription
p1trainCurriculum stage 1 training corpus
p2trainCurriculum stage 2 training corpus
p3trainCurriculum stage 3 training corpus
evvalidationHeld-out validation set (1.12 h)
ev_noisy_snr10testValidation set with additive noise at SNR 10 dB
ev_noisy_snr15testValidation set with additive noise at SNR 15 dB
ev_noisy_snr20testValidation set with additive noise at SNR 20 dB
benchmark_cleantest221-utterance real-noise classroom benchmark (one mislabeled utterance removed during label auditing)

Usage

python
from datasets import load_dataset
from transformers import WhisperTokenizer

ds = load_dataset("Hongthien06/vispeech-whisper-features", "benchmark_clean", split="test")

# Decode transcripts from labels
tok = WhisperTokenizer.from_pretrained("openai/whisper-small", language="vi", task="transcribe")
text = tok.decode([t for t in ds[0]["labels"] if t >= 0], skip_special_tokens=True)

License and attribution

Released under CC BY-NC-SA 4.0 (non-commercial, share-alike).

Part of this corpus is derived from VIVOS (AILAB, VNUHCM — University of Science), which is distributed under CC BY-NC-SA 4.0 for research purposes only: https://huggingface.co/datasets/AILAB-VNUHCM/vivos. Modifications applied to the VIVOS-derived portion include DSP preprocessing and additive-noise (AWGN) augmentation. The ShareAlike condition of the source license is preserved by releasing this entire corpus under the same license.

In line with the VIVOS terms of use, users of this dataset agree not to attempt to determine the identity of any speaker in the corpus.

Ethics and consent

Classroom recordings were collected with verbal consent from the speakers (recording, research use, and public release) and with the permission of the supervising lecturer. Transcripts contain no school names; personal names that may appear are included with the consent of the individuals concerned. File names and metadata contain no identifying information.

Citation

bibtex
@inproceedings{tran2026vispeech,
  title     = {ViSpeech: A Multi-Condition Vietnamese Speech Dataset for Noise-Robust Classroom ASR},
  author    = {Tran, Thi Dung and Pham, Thi Ngoc Oanh and Vo, Hong Thien and Nguyen, Huy Khang and Tran, Nhat Long and Tran, Duc Anh},
  booktitle = {Proceedings of FISAT 2026},
  publisher = {Springer},
  year      = {2026}
}