datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Thinkspark-v2-270m-training-data
ThinkSpark-v2-350M — training data
Full-duplex floor-controller (Section 8) training corpus: playable audio + text,
paired for the Dataset Viewer, plus every scenario field (behaviour, language, domain,
gender, prosody, agent text) and Soniox character-level timestamps.
Dataset Viewer
Default split is parquet with a real Audio feature — a player renders inline next to
the text in the Hub UI:
column
type
description
audio
Audio
playable wav (already… See the full description on the dataset page: https://huggingface.co/datasets/anuj-inavlabs/Thinkspark-v2-270m-training-data.emolia-thinking
Emolia-Thinking — a VoiceNet-annotated, balanced subset of Emolia
Emolia-Thinking is a richly annotated speech dataset created for the VoiceNet project. It takes a balanced subset of the Emolia corpus — balanced across speaker-embedding clusters and emotion-embedding clusters so that speakers, voices and emotional states are evenly represented rather than dominated by the most common cases — and annotates every clip along the full VoiceNet Extended voice-performance taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/VoiceNet/emolia-thinking.deep-thinkRatchada-STT
RATCHADA-STT Dataset
Overview
The dataset includes recordings from earnings calls of publicly traded companies in Thailand. Each audio file is accompanied by a transcription and metadata such as company name, reporting period, and other relevant details.
Dataset Info
Total Duration:
Train: 26507.87 seconds (~ 7.36 hours)
Test: 8376.28 seconds (~ 2.33 hours)
File Count:
Train: 10912 files
Test: 2804 files
Dataset Structure
The dataset consists… See the full description on the dataset page: https://huggingface.co/datasets/ThinkingMachinesDataScience/Ratchada-STT.
