OpenVoiceOS/ovos-intents
OVOS intents This is the canonical intent corpus of the OVOS skill fleet. The train split comes from each skill's .intent resources at pinned refs. The test split comes from the skills' end-to-end golden utterances. Labels have the form <skill_id>:<intent_name>, as OVOS-INTENT-4 defines. The version of the content is the tag; the tags v6 and v6.1 exist. Supersedes These datasets are retired and stay available for reproducibility: OpenVoiceOS/ovos-intents-train-v1… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intents.
OVOS intents
This is the canonical intent corpus of the OVOS skill fleet. The train split comes from each skill's .intent resources at pinned refs. The test split comes from the skills' end-to-end golden utterances. Labels have the form <skill_id>:<intent_name>, as OVOS-INTENT-4 defines. The version of the content is the tag; the tags v6 and v6.1 exist.
Supersedes
These datasets are retired and stay available for reproducibility:
- `OpenVoiceOS/ovos-intents-train-v1`
- `OpenVoiceOS/ovos-intents-train-latest`
- `OpenVoiceOS/ovos-intents-v5-eval`
- `OpenVoiceOS/ovos-golden-utterances-bench`
- `OpenVoiceOS/OVOSGitLocalize-Intents`
- `OpenVoiceOS/MT-intents-dataset-pt-PT`
- `OpenVoiceOS/ovos-llm-augmented-intents`
- `OpenVoiceOS/ovos-weather-intents`
- `OpenVoiceOS/ovos-common-query-intents`
- `OpenVoiceOS/ovos-intents-massive-subset`
- `OpenVoiceOS/ovos-intents-ilenia-testset-ca`
- `OpenVoiceOS/ovos-intents-ilenia-testset-es`
- `OpenVoiceOS/ovos-intents-ilenia-testset-nl`
Content of v6.1
This is the intent corpus for the OpenVoiceOS skill fleet. Each row is one utterance with the intent label it belongs to. The corpus is built from the locale resource files of the skills themselves, so it grows when the fleet gains locales.
Both splits ship here. Pin the tag v6.1 to get one build of both. The repository name carries no version: the version of the content is the tag.
This is a candidate. The intent-engine lane trains on it and runs the golden evaluation against it. Until that report exists, no accuracy claim about this corpus is supported.
What changed since v6
v6.1 is a naming release. Every label in it is the name the skill registers on the bus, and every gold row scores a label the train side carries.
- The skill id is the entry point's. v6 filed 1,111 train rows under
ovos-skill-easter-eggs.openvoiceosand 53 gold rows underskill-easter-eggs.openvoiceos, the id the skill declares, so the model learned a class the skill never answered to. The builder now reduces a repository spelling to the registered id (ovos-m2v-pipeline #207). One easter-eggs id. - Intent names follow OVOS-INTENT-2 §2 (lowercase, digits, underscores). Seven skills renamed their resources on dev: alerts, ip, wikipedia, spelling, wolfie, fallback-unknown, hello-world. The corpus trains the new names. A model trained on v6 keeps routing through the
renames.pytable in the pipeline until a model trained on this corpus ships. - Every gold label is trainable. v6 carried 22 test labels with no train rows (88 rows): the easter-eggs split,
count_to_Nagainstcount_to_n, three wallpapers CamelCase names, and an alerts family name no skill registers. Each was fixed in the skill's gold file (count #91, wallpapers #104, alerts #234, parrot #152, moviemaster #94, spelling #69). The build now refuses to publish a corpus with such a label (train/census_gold_labels.py, #208). This build: every test label has train rows. - `arb` and `es-419` are gone. They sat in the repository on the v5 row schema (
intent_id,lang,template) because the publish step never deleted a directory the build did not write. The publish now prunes (#209) and refuses a row withoutlabelandutterance(#230). Every row here carries the same keys per split. - Gold grew by 960 rows across the six skills above, from 1,420 to 2,380, and the gold locales from 19 to 21 (kab and oc-FR gain gold). 191 labels are scored against 164.
Layout
One directory per locale:
{lang}/train_templates.jsonl the training rows 52 locales {lang}/test.jsonl the gold rows 21 locales
The same file pattern as v6 and v5.
How it was built
- Builder:
train/build_from_skills.pyin OpenVoiceOS/ovos-m2v-pipeline, pull request #240, heade0b2f4e, on dev64672e0. The filetrain/sources.yamlate0b2f4egives the exact commit of each of the 64 pinned source repositories. Read the pins there. - The label function is
train/skill_labels.py, with no spelling fallback: a gold name that does not match a train name is refused, never folded. - The split of this tree from the builder's flat output was checked by hashing the sorted row set on both sides: train
683f340a9a847d0e, golde690bdcecd728542, equal on both.
Fields
The intent name in label is the base name of the skill's .intent file, with no suffix. Gold files in the fleet spell intent_label both ways (count_to_n and spell.intent); the builder strips the suffix, so the corpus carries one form. A per-repository convention is a test-suite choice, and the five repositories that mix both forms are the ones worth settling.
Counts
Labels trained went down by three, not up, and each one is accounted for: alerts #216 merged DeleteTodoEntries and QueryTodoEntries into their list-kind siblings, and moviemaster #79 dropped the dead movie_information and movie_production intents; the hello-world re-pin added how_are_you, which v6 had never trained.
Per source, 11 of 86 move against the last v6 rebuild. moviemaster +20,690 train rows (its locale fill), wallpapers +7,403, alerts +3,520, spelling +90; parrot -155 and count -24 (gold overlap removal against their new gold). The other 75 sources are unchanged to the row.
The five largest locales are en-US 844,842, ca-ES 298,109, fr-FR 167,136, gl-ES 108,603 and pt-BR 97,787. 28 locales hold fewer than 100 rows each.
Gold rows per locale: en-US 1,007, it-IT 112, da-DK 110, es-ES 110, eu-ES 110, gl-ES 110, ca-ES 109, fr-FR 100, de-DE 97, sv-SE 94, pt-BR 87, pt-PT 80, nl-NL 73, kab 44, fa-IR 26, oc-FR 24, pl-PL 22, ru-RU 20, cs-CZ 15, hu-HU 15, sv-FI 15.
31 locales have no gold
31 of the 52 train locales have no `test.jsonl` at all. No empty file is shipped for them, because an empty file reads as "measured, found nothing" when the truth is "never measured". 41 of the 232 labels are never scored.
Train and gold do not overlap
The build removed 1,731 train rows that also appear in the gold set (1,180 in v6: the new gold collides more). After that removal both checks report zero overlap:
Caveats
- 654 templates are silenced, against 392 in v6. The gold-overlap removal takes every train row a gold sentence matches, and 960 new gold rows match more templates in full. A silenced template is a sentence shape the model sees only in evaluation. The price is listed per template in
manifest.json. - 191 noncompliant base names remain (227 in v6). Their labels change when their skills rename; the pipeline bridges each rename when it lands.
- 96,217 rows keep an unfilled slot, a
{slot}the build had no example value for. - 992 rows are ambiguous: each shares its locale and utterance with a row carrying a different label. 76 more than v6, mostly wikipedia's new
wiki_moregold against days-in-history'stell_me_moretemplates ("continue", "continua"). - 8 template lines were dropped, not expanded, the same eight as v6: a bare
|with no group. They are named inmanifest.jsonand belong to the skills. - The
machine_generatedflag on gold rows is not evidence and is not read.
Files
{lang}/train_templates.jsonl: the training rows for that locale.{lang}/test.jsonl: the gold rows for that locale, where gold exists.manifest.json: the full build record.
Credit
Funded by the NGI0 Commons Fund / NLnet under grant agreement No 101135429, through the European Commission's Next Generation Internet programme.
