Team Ai
Datasetpublic

speaches-ai/realtime-turn-detection-test-data

Realtime speech test recordings Synthetic speech recordings for black-box Realtime API behavior tests in Speaches. Each WAV file is the unmodified output of OpenAI text-to-speech. Tests are responsible for adding silence, combining recordings, and choosing streaming chunk boundaries for their scenarios. metadata.jsonl follows the Hugging Face AudioFolder layout. Each record contains the generation inputs, file digest, expected text, transcription, and word/speech intervals from… See the full description on the dataset page: https://huggingface.co/datasets/speaches-ai/realtime-turn-detection-test-data.

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes52downloads
Dataset Card

Realtime speech test recordings

Synthetic speech recordings for black-box Realtime API behavior tests in Speaches. Each WAV file is the unmodified output of OpenAI text-to-speech. Tests are responsible for adding silence, combining recordings, and choosing streaming chunk boundaries for their scenarios.

metadata.jsonl follows the Hugging Face AudioFolder layout. Each record contains the generation inputs, file digest, expected text, transcription, and word/speech intervals from a separate whisper-1 transcription. The timing intervals are model-derived reference annotations, not sample-exact ground truth; tests should apply an explicit tolerance. Consumers can use speech_start_ms and speech_end_ms to trim a recording when needed.

The checked-in generation manifest and script are the source of truth. Regeneration is procedural rather than bit-for-bit reproducible because the hosted speech and transcription models can vary between calls and releases.