datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hindi_audio_dataset_testsmart-turn-data-v3.2-testTesting dataset for Smart Turn v3.2.
Thank you to the following contributors whose audio samples are included in this dataset:
The Pipecat team
Liva AI: https://www.theliva.ai/
Midcentury: https://www.midcentury.xyz/
MundoAI: https://mundoai.world/
Also, thank you to the following people for the CC-0 background noise sample data which has been used in this dataset:
https://freesound.org/people/4team/sounds/214995/
https://freesound.org/people/tomhannen/sounds/698090/… See the full description on the dataset page: https://huggingface.co/datasets/pipecat-ai/smart-turn-data-v3.2-test.whisperkit-test-dataELLSA_test_data
ELLSA: End-to-end Listen, Look, Speak and Act
The first end-to-end model that unifies vision, speech, text and actionin a streaming full-duplex framework, enabling joint multimodal perception and concurrent generation.
🧪 Highlights
Full-Duplex Multimodal Interaction: unifies listening, looking, speaking, and acting in a single end-to-end architecture, enabling simultaneous… See the full description on the dataset page: https://huggingface.co/datasets/tsinghua-ee/ELLSA_test_data.test-data-set-Arabic-letterchunking-test-datasuno-reggae-test-dataset
Suno Patois Reggae Test Set
299 patois-language reggae and dancehall tracks with style captions and
structured lyrics, laid out for SimpleTuner's textfile audio caption
strategy. Built as a small, high-consistency probe set for text-to-audio
training runs — not a general-purpose music corpus.
Rights and provenance
Every track here was generated by a third-party Suno user, and rights in the
audio and lyrics remain with those creators. Nothing in this repository is… See the full description on the dataset page: https://huggingface.co/datasets/RareConcepts/suno-reggae-test-dataset.audio_test_dataset
Dataset Card for "audio_test_dataset"
This dataset consists of the first 5 samples of mozilla-foundation/common_voice_13_0 and is only used for unit testing.
smart-turn-data-v3.1-testTesting dataset for Smart Turn v3.1.
Thank you to the following contributors whose audio samples are included in this dataset:
The Pipecat team
Liva AI: https://www.theliva.ai/
Midcentury: https://www.midcentury.xyz/
MundoAI: https://mundoai.world/
License
This dataset is licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0). See the LICENSE file for the full license text.
smart-turn-data-v3-testTesting dataset for Smart Turn v3.
License
This dataset is licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0). See the LICENSE file for the full license text.
wav2vec2-test-datasettest_datatest_dataTest_Audio_Generate_Dataset
Hinglish Audio Dataset
Generated by Sarvam AI.
test_TTS_data_hindi_v2Test_data
Title
Test Test Test
omnievalkit-data-test
OmniEvalKit Evaluation Datasets
Evaluation datasets for OmniEvalKit,
a comprehensive evaluation framework for omni-modal (audio + video + image + text) models.
Overview
Total subsets: 89
Total samples: 353,610
Total size: 352.3 GB (Parquet with embedded audio/image, no video)
Subsets requiring video download: 42
Note: Video files are NOT embedded in the Parquet files due to size constraints.
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/xiaofff/omnievalkit-data-test.iSparrow_test_datarealtime-turn-detection-test-data
Realtime speech test recordings
Synthetic speech recordings for black-box Realtime API behavior tests in
Speaches. Each WAV file is the unmodified output of OpenAI
text-to-speech. Tests are responsible for adding silence, combining recordings, and choosing streaming chunk
boundaries for their scenarios.
metadata.jsonl follows the Hugging Face AudioFolder layout. Each record contains the generation inputs, file
digest, expected text, transcription, and word/speech intervals from… See the full description on the dataset page: https://huggingface.co/datasets/speaches-ai/realtime-turn-detection-test-data.IMDA-NSC-datasets-testsada-arabic-test-dataset-sample
🗣️ Arabic Dialect Segmented Speech Dataset (SADA2022 Subset)
This dataset contains segmented Arabic speech samples from the SADA2022 corpus, annotated by dialect, gender, age group, speaking rate, environmental condition, and includes ground truth transcriptions.
It is intended to support research and applications in Arabic dialect classification, automatic speech recognition (ASR), and spoken language understanding.
📁 Dataset Structure
Audio segments are stored… See the full description on the dataset page: https://huggingface.co/datasets/SarahUssama/sada-arabic-test-dataset-sample.audio_data_kaggle_test_taskcqaloon_dataset_testTest-Dataset3
Test Dataset 3
This is a test preview of The Rights Foundry Collection One
Dataset Summary
A licensed 10-track music dataset for non-commercial AI research, reproducible benchmarking, model evaluation, and music-information-retrieval research.
This test preview of The Rights Foundry Collection One contains 10 fully rights-cleared commercial music tracks from independent artists, provided as 16-bit, 44.1 kHz stereo WAV audio with structured descriptive metadata.… See the full description on the dataset page: https://huggingface.co/datasets/swatt-TRF/Test-Dataset3.Test_model_datatest-audio-datasettest_audio_datasetaudio_data_kaggle_test_taskb_test-datasettest-audio-dataset
Test Audio Dataset
这是一个用于测试的音频数据集,包含 100 条伪造的 WAV 格式音频文件。
Dataset Structure
audio-dataset/
├── data/
│ ├── audio_0000.wav
│ ├── audio_0001.wav
│ └── ...
├── metadata.csv
└── README.md
Data Fields
file_name: 音频文件路径
transcription: 转录文本
speaker_id: 说话人ID
duration: 音频时长(秒)
sample_rate: 采样率
Usage
from datasets import load_dataset
dataset = load_dataset("your-username/test-audio-dataset")
License
MIT License
