Hongthien06/vispeech-whisper-features
ViSpeech — Preprocessed Whisper Features (Reproducibility Release) This repository contains the preprocessed feature sets used in the paper: ViSpeech: A Multi-Condition Vietnamese Speech Dataset for Noise-Robust Classroom ASR — T. D. Tran et al., FISAT 2026 (Springer). ViSpeech is a 37.62-hour multi-condition Vietnamese speech corpus built on a paired-recording design — the same speakers reading the same scripts in quiet-room and real-classroom conditions — comprising a… See the full description on the dataset page: https://huggingface.co/datasets/Hongthien06/vispeech-whisper-features.
ViSpeech — Preprocessed Whisper Features (Reproducibility Release)
This repository contains the preprocessed feature sets used in the paper:
ViSpeech: A Multi-Condition Vietnamese Speech Dataset for Noise-Robust Classroom ASR — T. D. Tran et al., FISAT 2026 (Springer).
ViSpeech is a 37.62-hour multi-condition Vietnamese speech corpus built on a paired-recording design — the same speakers reading the same scripts in quiet-room and real-classroom conditions — comprising a 36.50-hour noise-augmented training corpus, a 1.12-hour held-out validation set, and an independently collected 221-utterance real-noise classroom benchmark.
⚠️ Format notice
These are not raw audio files. Every example holds:
input_features— log-mel spectrograms produced by theopenai/whisper-smallfeature extractor;labels— target token IDs produced by theopenai/whisper-smalltokenizer.
This release is intended for exact reproduction of the paper's training and evaluation pipeline with Whisper-small. Transcripts can be recovered by decoding labels with the Whisper tokenizer; the original waveforms cannot be recovered from log-mel features.
Configurations
Usage
from datasets import load_dataset
from transformers import WhisperTokenizer
ds = load_dataset("Hongthien06/vispeech-whisper-features", "benchmark_clean", split="test")
# Decode transcripts from labels
tok = WhisperTokenizer.from_pretrained("openai/whisper-small", language="vi", task="transcribe")
text = tok.decode([t for t in ds[0]["labels"] if t >= 0], skip_special_tokens=True)License and attribution
Released under CC BY-NC-SA 4.0 (non-commercial, share-alike).
Part of this corpus is derived from VIVOS (AILAB, VNUHCM — University of Science), which is distributed under CC BY-NC-SA 4.0 for research purposes only: https://huggingface.co/datasets/AILAB-VNUHCM/vivos. Modifications applied to the VIVOS-derived portion include DSP preprocessing and additive-noise (AWGN) augmentation. The ShareAlike condition of the source license is preserved by releasing this entire corpus under the same license.
In line with the VIVOS terms of use, users of this dataset agree not to attempt to determine the identity of any speaker in the corpus.
Ethics and consent
Classroom recordings were collected with verbal consent from the speakers (recording, research use, and public release) and with the permission of the supervising lecturer. Transcripts contain no school names; personal names that may appear are included with the consent of the individuals concerned. File names and metadata contain no identifying information.
Citation
@inproceedings{tran2026vispeech,
title = {ViSpeech: A Multi-Condition Vietnamese Speech Dataset for Noise-Robust Classroom ASR},
author = {Tran, Thi Dung and Pham, Thi Ngoc Oanh and Vo, Hong Thien and Nguyen, Huy Khang and Tran, Nhat Long and Tran, Duc Anh},
booktitle = {Proceedings of FISAT 2026},
publisher = {Springer},
year = {2026}
}