Team Ai
Datasetpublic

arcada-labs/appointment-bench

Appointment Bench 25-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a dental office receptionist handling appointment scheduling. Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs. Leaderboard | GitHub | All Benchmarks Dataset Description The model acts as a dental office receptionist scheduling appointments for two patients with confusable names (Daniel and Danielle Nolan)… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/appointment-bench.

sourceHugging Facemitupdated 7mo agoView on Hugging Face
5likes120downloads
Dataset Card

Appointment Bench

25-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a dental office receptionist handling appointment scheduling.

Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs.

Leaderboard | GitHub | All Benchmarks

Dataset Description

The model acts as a dental office receptionist scheduling appointments for two patients with confusable names (Daniel and Danielle Nolan) across two doctors (Dr. Perry and Dr. Barry). The conversation includes phone number swaps and reverts, slot-taken error recovery, and cross-entity state tracking where information from one patient's booking must not leak into another's.

What This Benchmark Tests

  • —Tool use: 4 functions — appointment booking, availability lookup, patient record retrieval, schedule management
  • —Confusable entity disambiguation: Two patients with near-identical names (Daniel/Danielle Nolan), two doctors (Perry/Barry)
  • —Phone number swap and revert: Correction followed by reverting back to the original
  • —False memory traps: 4 turns that assert things the model never said or did
  • —Slot-taken error recovery: Handling booking conflicts when a requested slot is already occupied
  • —Cross-entity state tracking: Keeping two patients' details separate across the full conversation

Dataset Structure

appointment-bench/
├── audio/                          # TTS-generated audio (1 WAV per turn)
│   ├── turn_000.wav
│   ├── turn_001.wav
│   └── ... (25 files)
├── real_audio/                     # Human-recorded audio
│   ├── person1/
│   │   └── turn_000.wav ... turn_024.wav
│   └── person2/
│       └── turn_000.wav ... turn_024.wav
├── benchmark/
│   ├── turns.json                  # Turn definitions with golden answers
│   ├── hard_turns.json             # Same as turns.json but input_text=null (audio-only)
│   ├── tool_schemas.json           # Tool/function schemas (4 tools)
│   └── knowledge_base.txt          # Dental office KB
└── metadata.jsonl                  # HF dataset viewer metadata

Metadata Fields

FieldDescription
file_namePath to the audio file
turn_idTurn index (0–24)
speakertts, person1, or person2
input_textWhat the user says (text transcript)
golden_textExpected assistant response
required_function_callTool call the model should make (JSON, nullable)
function_call_responseScripted tool response (JSON, nullable)
categoriesEvaluation categories for this turn
subcategorySpecific sub-skill being tested
scoring_dimensionsWhich judge dimensions apply

Audio Format

  • —Format: WAV, 16-bit PCM, mono
  • —TTS audio: Generated via text-to-speech
  • —Real audio: Human-recorded by multiple speakers, same transcript content

Usage

With Audio Arena CLI

bash
pip install audio-arena  # or: git clone + uv sync

# Run with a text model
uv run audio-arena run appointment_bench --model claude-sonnet-4-5 --service anthropic

# Run with a speech-to-speech model
uv run audio-arena run appointment_bench --model gpt-realtime --service openai-realtime

# Judge the results
uv run audio-arena judge runs/appointment_bench/<run_dir>

With Hugging Face Datasets

python
from datasets import load_dataset

ds = load_dataset("arcada-labs/appointment-bench")

Evaluation

Models are judged on up to 5 dimensions per turn:

DimensionDescription
tool_use_correctCorrect function called with correct arguments
instruction_followingUser's request was actually completed
kb_groundingClaims are supported by the knowledge base or tool results
state_trackingConsistency with earlier turns (scored on tagged turns only)
ambiguity_handlingCorrect disambiguation (scored on tagged turns only)

For speech-to-speech models, a 6th turn_taking dimension evaluates audio timing correctness.

See the full methodology for details on two-phase evaluation, penalty absorption, and category-aware scoring.

Part of Audio Arena

BenchmarkTurnsScenario
Conversation Bench75Conference assistant
Appointment Bench (this dataset)25Dental office scheduling
Assistant Bench31Personal assistant
Event Bench29Event planning
Grocery Bench30Grocery ordering
Product Bench31Laptop comparison shopping

Citation

bibtex
@misc{audioarena2026,
  title={Audio Arena: Multi-Turn Speech-to-Speech Evaluation Benchmarks},
  author={Arcada Labs},
  year={2026},
  url={https://audioarena.ai}
}