datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ovos-stt-bench-voxpopuli-en-US
OVOS stt bench — voxpopuli-en-US
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/voxpopuli.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-voxpopuli-en-US.pseudolabel-malaya-speech-stt-train-whisper-large-v3ovos-stt-bench-ami-en-GB
OVOS stt bench — ami-en-GB
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
edinburghcstr/ami.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows into… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-ami-en-GB.LLM-Inference-Traces-NYC
LLM-Inference-Traces-NYC
A synthesized LLM inference workload for New York City: 2,015,645 requests, in
1,720,774 conversations, from 1,326,738 users, across 259 active geographic zones,
over 7 consecutive days (2024-05-12 to 2024-05-18).
Each request carries a user identifier, a timestamp, and a geographic zone — the three
spatiotemporal attributes that public LLM conversation datasets are stripped of by privacy
regulation. Without them, cache-aware scheduling, geo-distributed… See the full description on the dataset page: https://huggingface.co/datasets/ST-TraceWeaver/LLM-Inference-Traces-NYC.stt-training-data
Dataset Statistics
Configuration: default
Split: train
Total Rows: 1,362,015
dept
Type: categorical
Data Type: object
Unique Values: 8
Value Distribution:
Value
Count
Percentage
STT_TT
446,495
32.78%
STT_NS
236,407
17.36%
STT_AB
170,922
12.55%
STT_CS
146,811
10.78%
STT_MV
110,080
8.08%
STT_NW
94,703
6.95%
STT_HS
84,797
6.23%
STT_PC
71,800
5.27%
grade
Type: numerical
Data Type: int64
Sum: 3,963,451.00… See the full description on the dataset page: https://huggingface.co/datasets/openpecha/stt-training-data.ovos-stt-bench-mls-es-ES
OVOS stt bench — mls-es-ES
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/multilingual_librispeech.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-mls-es-ES.robodojo-longhorizon-sttp
RoboDojo Long-Horizon subtask dataset
The eight tasks of RoboDojo's Long-Horizon capability dimension (the benchmark's own
grouping, scripts/internal/summarize_result.py:DIMENSIONS), 100 demonstration episodes each,
segmented into subtasks and rendered as training rows for a VLM planner.
Layout
path
what
frames/<task>/epNNNN/FFFFFF.jpg
head camera, 512x384, q32 JPEG. FFFFFF is the simulator step index.
rows/local_sttp.jsonl
Local-STTP formulation. The… See the full description on the dataset page: https://huggingface.co/datasets/ghkim-rlwrld/robodojo-longhorizon-sttp.ovos-stt-bench-voxpopuli-es-ES
OVOS stt bench — voxpopuli-es-ES
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/voxpopuli.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-voxpopuli-es-ES.ovos-stt-bench-speech-massive-de-DE
OVOS stt bench — speech-massive-de-DE
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
FBK-MT/Speech-MASSIVE-test.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-speech-massive-de-DE.ovos-stt-bench-mtedx-es-ES
OVOS stt bench — mtedx-es-ES
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
deepdml/mtedx.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows into… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-mtedx-es-ES.ovos-stt-bench-mls-pt-PT
OVOS stt bench — mls-pt-PT
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/multilingual_librispeech.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-mls-pt-PT.ovos-stt-bench-mtedx-de-DE
OVOS stt bench — mtedx-de-DE
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
deepdml/mtedx.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows into… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-mtedx-de-DE.ovos-stt-bench-speech-massive-fr-FR
OVOS stt bench — speech-massive-fr-FR
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
FBK-MT/Speech-MASSIVE-test.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-speech-massive-fr-FR.ovos-stt-bench-mtedx-ru-RU
OVOS stt bench — mtedx-ru-RU
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
deepdml/mtedx.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows into… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-mtedx-ru-RU.ovos-stt-bench-speech-massive-nl-NL
OVOS stt bench — speech-massive-nl-NL
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
FBK-MT/Speech-MASSIVE-test.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-speech-massive-nl-NL.ovos-stt-bench-mls-fr-FR
OVOS stt bench — mls-fr-FR
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/multilingual_librispeech.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-mls-fr-FR.ovos-stt-bench-mls-nl-NL
OVOS stt bench — mls-nl-NL
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/multilingual_librispeech.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-mls-nl-NL.ovos-stt-bench-mls-it-IT
OVOS stt bench — mls-it-IT
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/multilingual_librispeech.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-mls-it-IT.ovos-stt-bench-voxpopuli-fr-FR
OVOS stt bench — voxpopuli-fr-FR
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/voxpopuli.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-voxpopuli-fr-FR.ovos-stt-bench-voxpopuli-sl-SI
OVOS stt bench — voxpopuli-sl-SI
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/voxpopuli.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-voxpopuli-sl-SI.ovos-stt-bench-mls-pl-PL
OVOS stt bench — mls-pl-PL
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/multilingual_librispeech.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-mls-pl-PL.ovos-stt-bench-mtedx-it-IT
OVOS stt bench — mtedx-it-IT
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
deepdml/mtedx.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows into… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-mtedx-it-IT.ovos-stt-bench-speech-massive-ru-RU
OVOS stt bench — speech-massive-ru-RU
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
FBK-MT/Speech-MASSIVE-test.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-speech-massive-ru-RU.ovos-stt-bench-speech-massive-vi-VN
OVOS stt bench — speech-massive-vi-VN
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
FBK-MT/Speech-MASSIVE-test.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-speech-massive-vi-VN.ovos-stt-bench-voxpopuli-ro-RO
OVOS stt bench — voxpopuli-ro-RO
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/voxpopuli.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-voxpopuli-ro-RO.ovos-stt-bench-voxpopuli-fi-FI
OVOS stt bench — voxpopuli-fi-FI
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/voxpopuli.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-voxpopuli-fi-FI.ovos-stt-bench-speech-massive-tr-TR
OVOS stt bench — speech-massive-tr-TR
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
FBK-MT/Speech-MASSIVE-test.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-speech-massive-tr-TR.ovos-stt-bench-voxpopuli-de-DE
OVOS stt bench — voxpopuli-de-DE
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/voxpopuli.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-voxpopuli-de-DE.ovos-stt-bench-voxpopuli-sk-SK
OVOS stt bench — voxpopuli-sk-SK
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/voxpopuli.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-voxpopuli-sk-SK.ovos-stt-bench-minds14-pl-PL
OVOS stt bench — minds14-pl-PL
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
PolyAI/minds14.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows into… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-minds14-pl-PL.
