arcada-labs/appointment-bench
Appointment Bench 25-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a dental office receptionist handling appointment scheduling. Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs. Leaderboard | GitHub | All Benchmarks Dataset Description The model acts as a dental office receptionist scheduling appointments for two patients with confusable names (Daniel and Danielle Nolan)… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/appointment-bench.
Appointment Bench
25-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a dental office receptionist handling appointment scheduling.
Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs.
Leaderboard | GitHub | All Benchmarks
Dataset Description
The model acts as a dental office receptionist scheduling appointments for two patients with confusable names (Daniel and Danielle Nolan) across two doctors (Dr. Perry and Dr. Barry). The conversation includes phone number swaps and reverts, slot-taken error recovery, and cross-entity state tracking where information from one patient's booking must not leak into another's.
What This Benchmark Tests
- Tool use: 4 functions — appointment booking, availability lookup, patient record retrieval, schedule management
- Confusable entity disambiguation: Two patients with near-identical names (Daniel/Danielle Nolan), two doctors (Perry/Barry)
- Phone number swap and revert: Correction followed by reverting back to the original
- False memory traps: 4 turns that assert things the model never said or did
- Slot-taken error recovery: Handling booking conflicts when a requested slot is already occupied
- Cross-entity state tracking: Keeping two patients' details separate across the full conversation
Dataset Structure
appointment-bench/
├── audio/ # TTS-generated audio (1 WAV per turn)
│ ├── turn_000.wav
│ ├── turn_001.wav
│ └── ... (25 files)
├── real_audio/ # Human-recorded audio
│ ├── person1/
│ │ └── turn_000.wav ... turn_024.wav
│ └── person2/
│ └── turn_000.wav ... turn_024.wav
├── benchmark/
│ ├── turns.json # Turn definitions with golden answers
│ ├── hard_turns.json # Same as turns.json but input_text=null (audio-only)
│ ├── tool_schemas.json # Tool/function schemas (4 tools)
│ └── knowledge_base.txt # Dental office KB
└── metadata.jsonl # HF dataset viewer metadataMetadata Fields
Audio Format
- Format: WAV, 16-bit PCM, mono
- TTS audio: Generated via text-to-speech
- Real audio: Human-recorded by multiple speakers, same transcript content
Usage
With Audio Arena CLI
pip install audio-arena # or: git clone + uv sync
# Run with a text model
uv run audio-arena run appointment_bench --model claude-sonnet-4-5 --service anthropic
# Run with a speech-to-speech model
uv run audio-arena run appointment_bench --model gpt-realtime --service openai-realtime
# Judge the results
uv run audio-arena judge runs/appointment_bench/<run_dir>With Hugging Face Datasets
from datasets import load_dataset
ds = load_dataset("arcada-labs/appointment-bench")Evaluation
Models are judged on up to 5 dimensions per turn:
For speech-to-speech models, a 6th turn_taking dimension evaluates audio timing correctness.
See the full methodology for details on two-phase evaluation, penalty absorption, and category-aware scoring.
Part of Audio Arena
Citation
@misc{audioarena2026,
title={Audio Arena: Multi-Turn Speech-to-Speech Evaluation Benchmarks},
author={Arcada Labs},
year={2026},
url={https://audioarena.ai}
}