Team Ai
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Praxel /codeswitch-pairs-lase Codeswitch Pairs LASE — training corpus 1118 same-voice cross-script utterance pairs (8 ElevenLabs Multilingual voices × en/hi/te/ta) used to train the LASE r1 speaker encoder. Each row is one synthesized utterance with metadata; pairs are reconstructed at evaluation time by joining on voice_id (same voice, different script = cross-script pair). Schema (manifest.jsonl) { "voice_id": "21m00Tcm4TlvDq8ikWAM", "lang": "en | hi | te | ta", "text": "the prompt text"… See the full description on the dataset page: https://huggingface.co/datasets/Praxel/codeswitch-pairs-lase.audioaudio-classificationn<1K0 likes78 downloads5mo agoHugging Face02WeixiangYan /CodeScopetabulartranslationn<1K3 likes75 downloads3y agoHugging Face03Scoolar /codesearchnet-challenge-extended CodeSearchNet Challenge, extended: every search over every function The CodeSearchNet Challenge (Husain et al., 2019) has 99 natural-language code searches, and experts rated a few candidate functions for each. This dataset treats every rated function of a language as one codebase and searches all of it: for each search, every function in its language is a candidate. The experts' ratings are kept, and the pairs they never rated but a search tool returned were rated on the same… See the full description on the dataset page: https://huggingface.co/datasets/Scoolar/codesearchnet-challenge-extended.tabulartext-retrieval10K<n<100K0 likes59 downloads12d agoHugging Face04codeslord /openai_records tags: - observers tags: - observers tags: - observers tags: observers Dataset Card for Dataset Name Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources… See the full description on the dataset page: https://huggingface.co/datasets/codeslord/openai_records.tabularn<1K0 likes41 downloads2y agoHugging Face05open-llm-leaderboard /mistralai__Codestral-22B-v0.1-detailsgated Dataset Card for Evaluation run of mistralai/Codestral-22B-v0.1 Dataset automatically created during the evaluation run of model mistralai/Codestral-22B-v0.1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistralai__Codestral-22B-v0.1-details.tabular10K<n<100K0 likes28 downloads2y agoHugging Face06open-llm-leaderboard /migtissera__Trinity-2-Codestral-22B-detailsgated Dataset Card for Evaluation run of migtissera/Trinity-2-Codestral-22B Dataset automatically created during the evaluation run of model migtissera/Trinity-2-Codestral-22B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/migtissera__Trinity-2-Codestral-22B-details.tabular10K<n<100K0 likes25 downloads2y agoHugging Face07open-llm-leaderboard /migtissera__Trinity-2-Codestral-22B-v0.2-detailsgated Dataset Card for Evaluation run of migtissera/Trinity-2-Codestral-22B-v0.2 Dataset automatically created during the evaluation run of model migtissera/Trinity-2-Codestral-22B-v0.2 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/migtissera__Trinity-2-Codestral-22B-v0.2-details.tabular10K<n<100K0 likes23 downloads2y agoHugging Face08codesj /dstc12_track1_chatevalgatedIf you use this dataset please use the following citation: @inproceedings{mendonca2025dstc12t1, author = "John Mendonça and Lining Zhang and Rahul Mallidi and Luis Fernando D'Haro and João Sedoc", title = "Overview of Dialog System Evaluation Track: Dimensionality, Language, Culture and Safety at DSTC 12", booktitle = "DSTC12: The Twelfth Dialog System Technology Challenge", series = "26th Meeting of the Special Interest Group on… See the full description on the dataset page: https://huggingface.co/datasets/codesj/dstc12_track1_chateval.tabularn<1K0 likes4 downloads1y agoHugging Face09dnanper /ds1000_fail_codescoretabular1K<n<10K0 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.