datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
codeswitch-pairs-lase
Codeswitch Pairs LASE — training corpus
1118 same-voice cross-script utterance pairs (8 ElevenLabs Multilingual voices × en/hi/te/ta) used to train the LASE r1 speaker encoder.
Each row is one synthesized utterance with metadata; pairs are reconstructed at evaluation time by joining on voice_id (same voice, different script = cross-script pair).
Schema (manifest.jsonl)
{
"voice_id": "21m00Tcm4TlvDq8ikWAM",
"lang": "en | hi | te | ta",
"text": "the prompt text"… See the full description on the dataset page: https://huggingface.co/datasets/Praxel/codeswitch-pairs-lase.CodeScopecodesearchnet-challenge-extended
CodeSearchNet Challenge, extended: every search over every function
The CodeSearchNet Challenge (Husain et al., 2019) has
99 natural-language code searches, and experts rated a few candidate functions for each. This
dataset treats every rated function of a language as one codebase and searches all of it: for each
search, every function in its language is a candidate. The experts' ratings are kept, and the
pairs they never rated but a search tool returned were rated on the same… See the full description on the dataset page: https://huggingface.co/datasets/Scoolar/codesearchnet-challenge-extended.openai_records
tags:
- observers
tags:
- observers
tags:
- observers
tags:
observers
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources… See the full description on the dataset page: https://huggingface.co/datasets/codeslord/openai_records.mistralai__Codestral-22B-v0.1-details
Dataset Card for Evaluation run of mistralai/Codestral-22B-v0.1
Dataset automatically created during the evaluation run of model mistralai/Codestral-22B-v0.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistralai__Codestral-22B-v0.1-details.migtissera__Trinity-2-Codestral-22B-details
Dataset Card for Evaluation run of migtissera/Trinity-2-Codestral-22B
Dataset automatically created during the evaluation run of model migtissera/Trinity-2-Codestral-22B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/migtissera__Trinity-2-Codestral-22B-details.migtissera__Trinity-2-Codestral-22B-v0.2-details
Dataset Card for Evaluation run of migtissera/Trinity-2-Codestral-22B-v0.2
Dataset automatically created during the evaluation run of model migtissera/Trinity-2-Codestral-22B-v0.2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/migtissera__Trinity-2-Codestral-22B-v0.2-details.dstc12_track1_chatevalIf you use this dataset please use the following citation:
@inproceedings{mendonca2025dstc12t1,
author = "John Mendonça and Lining Zhang and Rahul Mallidi and Luis Fernando D'Haro and João Sedoc",
title = "Overview of Dialog System Evaluation Track: Dimensionality, Language, Culture and Safety at DSTC 12",
booktitle = "DSTC12: The Twelfth Dialog System Technology Challenge",
series = "26th Meeting of the Special Interest Group on… See the full description on the dataset page: https://huggingface.co/datasets/codesj/dstc12_track1_chateval.ds1000_fail_codescore
