speaches-ai/realtime-turn-detection-test-data
Realtime speech test recordings Synthetic speech recordings for black-box Realtime API behavior tests in Speaches. Each WAV file is the unmodified output of OpenAI text-to-speech. Tests are responsible for adding silence, combining recordings, and choosing streaming chunk boundaries for their scenarios. metadata.jsonl follows the Hugging Face AudioFolder layout. Each record contains the generation inputs, file digest, expected text, transcription, and word/speech intervals from… See the full description on the dataset page: https://huggingface.co/datasets/speaches-ai/realtime-turn-detection-test-data.
Realtime speech test recordings
Synthetic speech recordings for black-box Realtime API behavior tests in Speaches. Each WAV file is the unmodified output of OpenAI text-to-speech. Tests are responsible for adding silence, combining recordings, and choosing streaming chunk boundaries for their scenarios.
metadata.jsonl follows the Hugging Face AudioFolder layout. Each record contains the generation inputs, file digest, expected text, transcription, and word/speech intervals from a separate whisper-1 transcription. The timing intervals are model-derived reference annotations, not sample-exact ground truth; tests should apply an explicit tolerance. Consumers can use speech_start_ms and speech_end_ms to trim a recording when needed.
The checked-in generation manifest and script are the source of truth. Regeneration is procedural rather than bit-for-bit reproducible because the hosted speech and transcription models can vary between calls and releases.
