Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Helsinki-NLP /fineweb-edu-translated Helsinki-NLP/fineweb-edu-translated fineweb-edu-tanslated is a collection of automatically translated documents from fineweb-edu. Translations are based on OPUS-MT and HPLT-MT models. The data in v1.0 covers 36,704,000 documents with over 28 billion space-searated tokens of English data translated into 36 languages. The total v1.0 data set includes over 960 billion tokens and the translated documents are aligned across all languages. In the v1.1 release, additional translations… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/fineweb-edu-translated.texttranslation1B<n<10B16 likes262k downloads6mo agoHugging Face02Helsinki-NLP /nemotron-cc-translated Helsinki-NLP/nemotron-cc-translated nemotron-cc-tanslated is a collection of automatically translated documents from nemotron-cc taken out of the high-quality subset. Translations are based on OPUS-MT and HPLT-MT models. The data in v1.0 covers 156,431,999 documents with over 70 billion space-searated tokens of English data translated into 36 languages. The total v1.0 data set includes over 2.4 trillion tokens and the translated documents are aligned across all languages. v1.1… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/nemotron-cc-translated.texttranslation1B<n<10B5 likes67k downloads6mo agoHugging Face03NTU-NLP-sg /xCodeEvalThe ability to solve problems is a hallmark of intelligence and has been an enduring goal in AI. AI systems that can create programs as solutions to problems or assist developers in writing programs can increase productivity and make programming more accessible. Recently, pre-trained large language models have shown impressive abilities in generating new codes from natural language descriptions, repairing buggy codes, translating codes between languages, and retrieving relevant code segments. However, the evaluation of these models has often been performed in a scattered way on only one or two specific tasks, in a few languages, at a partial granularity (e.g., function) level and in many cases without proper training data. Even more concerning is that in most cases the evaluation of generated codes has been done in terms of mere lexical overlap rather than actual execution whereas semantic similarity (or equivalence) of two code segments depends only on their ``execution similarity'', i.e., being able to get the same output for a given input.translation1M<n<10M83 likes35k downloads1y agoHugging Face04McGill-NLP /weblinx-browsergym WebLINX: Real-World Website Navigation with Multi-Turn Dialogue Xing Han Lù*, Zdeněk Kasner*, Siva Reddy 💾Code 📄Paper 🌐Website 📓Colab 🤖Models💻Explorer 🐦Tweets 🏆Leaderboard Your browser does not support the video tag. This dataset was specifically created to allow WebLINX to be used inside the BrowserGym and Agentlab ecosystem. Please see the browsergym repository for more information. [!NOTE] The version associated with this library is WebLINX… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/weblinx-browsergym.image-to-text4 likes14k downloads2y agoHugging Face05MU-NLPC /Calc-mawps Dataset Card for Calc-MAWPS Summary The dataset is a collection of simple math word problems focused on arithmetics. It is derived from https://huggingface.co/datasets/omarxadel/MaWPS-ar. The main addition in this dataset variant is the chain column. It was created by converting the solution to a simple html-like language that can be easily parsed (e.g. by BeautifulSoup). The data contains 3 types of tags: gadget: A tag whose content is intended to be evaluated by… See the full description on the dataset page: https://huggingface.co/datasets/MU-NLPC/Calc-mawps.texttext-generation1K<n<10K1 likes11k downloads3y agoHugging Face06Columbia-NLP /PUPAThis dataset contains the data presented in the paper PAPILLON: Privacy Preservation from Internet-based and Local Language Model Ensembles. Code: https://github.com/siyan-sylvia-li/PAPILLON texttext-generationn<1K3 likes7.7k downloads2y agoHugging Face07SALT-NLP /SWE-chatgated SWE-chat: Coding Agent Interactions From Real Users in the Wild 📄 COLM 2026 Paper: arxiv.org/abs/2604.20779 🌐 Website: swe-chat.com Dataset Summary SWE-chat captures real-world AI coding sessions from developers using AI coding assistants (Claude Code, Codex, Gemini CLI, and others via the Entire.io CLI). Each session includes the full conversation transcript, tool calls, thinking traces, code changes, and attribution of human vs. agent-authored code.… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/SWE-chat.tabulartext-generation10M<n<100M123 likes7.7k downloads7d agoHugging Face08bio-nlp-umass /MedThinkVQA MedThinkVQA MedThinkVQA is an expert-annotated benchmark for multi-image diagnostic reasoning in radiology. Unlike prior medical VQA benchmarks that typically contain at most one image per case, MedThinkVQA requires models to extract evidence from each image, integrate cross-view information, and perform differential-diagnosis reasoning. Links GitHub: https://github.com/benluwang/MedThinkVQA Leaderboard: https://benluwang.github.io/MedThinkVQA/ Submission Guide:… See the full description on the dataset page: https://huggingface.co/datasets/bio-nlp-umass/MedThinkVQA.imagequestion-answering1K<n<10K12 likes5.1k downloads5mo agoHugging Face09SALT-NLP /cogym-collabskill-trajectories Dataset Card for CollabSkill Trajectories Dataset Summary CollabSkill is a framework for studying how real human workers collaborate with AI agents on occupational tasks. Participants were matched to tasks based on their occupational backgrounds and paired with one of five AI agents. The release contains interaction logs, participant and session metadata, task rubrics, automated grading results, and Bayesian CollabSkill ratings for humans and agents. The tasks… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/cogym-collabskill-trajectories.text-generation1 likes5k downloads2mo agoHugging Face10McGill-NLP /A3-Synth A3-Synth 💾 Code 📄 Paper 🌐 Website 🤗 Dataset 🤖 Models 📦 PyPI Structured Distillation of Web Agent Capabilities Enables Generalization Xing Han Lù, Siva Reddy A3-Synth is a synthetic training dataset for web agents, generated using the Agent-as-Annotators (A3) framework. It contains ~16k SFT training examples produced by Gemini 3 Pro acting as the Annotator across 3,000 tasks on 6 WebArena environments. Dataset Structure A3-Synth/ training/… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/A3-Synth.text-generation10K<n<100K1 likes4.9k downloads6mo agoHugging Face11Multilingual-Multimodal-NLP /McEvalMcEval benchmark data as described in the McEval Paper. Code for the evaluation can be found on Github as McEval. texttext-generation10K<n<100K21 likes3.8k downloads2y agoHugging Face12coral-nlp /german-commons German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models A comprehensive collection of German-language text data under open licenses for training German language models. Datasheet: DATASHEET.md. Paper: arxiv.org/abs/2510.13996 Code: github.com/coral-nlp/llmdata Bloom Filter (DOLMA-compatible): bloom_filter.bin Dataset Description This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokensof German text data with… See the full description on the dataset page: https://huggingface.co/datasets/coral-nlp/german-commons.tabulartext-generation10M<n<100M41 likes3.1k downloads9mo agoHugging Face13DAMO-NLP-SG /multimodal_textbook Multimodal-Textbook-6.5M Overview This dataset is for "2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining", containing 6.5M images interleaving with 0.8B text from instructional videos. It contains pre-training corpus using interleaved image-text format. Specifically, our multimodal-textbook includes 6.5M keyframesextracted from instructional videos, interleaving with 0.8B ASR texts. All the images and text are extracted from online… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/multimodal_textbook.text-generation1M<n<10M164 likes2.7k downloads2y agoHugging Face14McGill-NLP /WebLINX WebLINX: Real-World Website Navigation with Multi-Turn Dialogue Xing Han Lù*, Zdeněk Kasner*, Siva Reddy 💾Code 📄Paper 🌐Website 📓Colab 🤖Models💻Explorer 🐦Tweets 🏆Leaderboard Your browser does not support the video tag. [!IMPORTANT] WebLINX is now available as a benchmark through BrowserGym, allowing you to access demonstration steps in the same way you would access a web agent environment like WebArena or MiniWoB. This also allows you to run agents… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/WebLINX.textimage-to-text10K<n<100K65 likes2.1k downloads2y agoHugging Face15Helsinki-NLP /tatoeba_mtgated Dataset Card for The Tatoeba Translation Challenge Please note that this dataset is intended strictly for evaluation and benchmarking purposes. Training models on this dataset, or including it in automatically collected web-scale training corpora, may lead to benchmark contamination and invalidate evaluation results. Dataset Summary The Tatoeba Translation Challenge is a multilingual data set of machine translation benchmarks derived from user-contributed… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/tatoeba_mt.texttext-generation10M<n<100M64 likes2.1k downloads2d agoHugging Face16hkust-nlp /dart-math-hard 🎯 DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving 📝 Paper@arXiv | 🤗 Datasets&Models@HF | 🐱 Code@GitHub 🐦 Thread@X(Twitter) | 🐶 中文博客@知乎 | 📊 Leaderboard@PapersWithCode | 📑 BibTeX [!IMPORTANT] 🔥 Excited to find our DART-Math-DSMath-7B (Prop2Diff) trained on DART-Math-Hard comparable to the AIMO winner NuminaMath-7B on CoT, but based solely on MATH & GSM8K prompt set, leaving much room to improve! Besides, our DART method is also fully compatible… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/dart-math-hard.texttext-generation100K<n<1M15 likes1.9k downloads2y agoHugging Face17hkust-nlp /dart-math-uniform 🎯 DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving 📝 Paper@arXiv | 🤗 Datasets&Models@HF | 🐱 Code@GitHub 🐦 Thread@X(Twitter) | 🐶 中文博客@知乎 | 📊 Leaderboard@PapersWithCode | 📑 BibTeX Datasets: DART-Math DART-Math datasets are the state-of-the-art and data-efficientopen-source instruction tuning datasets for mathematical reasoning. Figure 1: Left: Average accuracy on 6 mathematical benchmarks. We compare with models… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/dart-math-uniform.texttext-generation100K<n<1M13 likes1.4k downloads2y agoHugging Face18KETI-NLP /KoEVD KoEVD KoEVD is a Korean benchmark linking five evaluation or analysis targets through source utterances: utterance-risk judgment, candidate-response safety choice, direct-generation response harmfulness, descriptive response strategies, and a pre-execution mock tool/action-choice diagnostic. Contents and scope The canonical corpus contains 13,552 sources and 71,395 response candidates: 30,740 accepted, 27,104 rejected, and 13,551 strongly rejected. Three… See the full description on the dataset page: https://huggingface.co/datasets/KETI-NLP/KoEVD.texttext-classification10K<n<100K0 likes1.1k downloads25d agoHugging Face19hkust-nlp /agentboard AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents This is the official dataset repository of AgentBoard. 1. Data Overview AgentBoard is composed of 9 diverse tasks which can be divided into 4 types, including Embodied AI, Game, Web, and Tool: Embodied AI Game Web Tool AlfWorld ScienceWorld BabyAI Jericho PDDL WebShop WebArena… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/agentboard.texttext-generation1K<n<10K13 likes1.1k downloads2y agoHugging Face20yuhan-nlp /diversity-router-data Diversity Router Data This dataset supports the paper "No Single Best Model for Diversity: Learning a Router for Sample Diversity". It contains language model generation outputs and diversity metrics across four datasets to calculate the diversity coverage, used to study how different models produce diverse and high quality outputs and to train/evaluate a diversity router. Dataset Structure diversity-router-data/ ├── simple_questions/ # Questions with fixed… See the full description on the dataset page: https://huggingface.co/datasets/yuhan-nlp/diversity-router-data.text-generation1 likes1.1k downloads6mo agoHugging Face21s-nlp /paradetox ParaDetox: Text Detoxification with Parallel Data (English) This repository contains information about ParaDetox dataset -- the first parallel corpus for the detoxification task -- as well as models and evaluation methodology for the detoxification of English texts. The original paper "ParaDetox: Detoxification with Parallel Data" was presented at ACL 2022 main conference. 📰 Updates [2025] !!!NOW OPEN!!! TextDetox CLEF2025 shared task: for even more -- 15 languages! website… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/paradetox.imagetext-generation10K<n<100K9 likes970 downloads2y agoHugging Face22hkust-nlp /dart-math-pool-math [!NOTE] This dataset is the data pool synthesized from the query set of the MATH training set, containing all answer-correct samples and other metadata produced during the work. DART-Math-* datasets are extracted from dart-math-pool-* data pools. 🎯 DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving 📝 Paper@arXiv | 🤗 Datasets&Models@HF | 🐱 Code@GitHub 🐦 Thread@X(Twitter) | 🐶 中文博客@知乎 | 📊 Leaderboard@PapersWithCode | 📑 BibTeX Datasets:… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/dart-math-pool-math.texttext-generation1M<n<10M8 likes892 downloads2y agoHugging Face23andresnowak /Instruction-finetuning-mixture-mnlp-with-nlp4educationDataset created using the Tulu3-sft-mixture and MNLP Question and golden answer dataset From the Tulue3-sft-mixture, messages that didn't have only 2 messages (user and assistant) where removed Also the datasets for alignment and jailbreaking were removed texttext-generation1M<n<10M0 likes757 downloads1y agoHugging Face24lavis-nlp /CALIBRI CALIBRI Dataset Dataset Description CALIBRI is a comprehensive dataset for studying calibration in LLM-based code generation. It contains code generations from multiple state-of-the-art language models across three established benchmarks, along with token-level likelihood information for calibration analysis and correctness labels, based on the benchmark-provided test suites. Each sample provides 10 different generations for one problem. Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/lavis-nlp/CALIBRI.texttext-generation10K<n<100K0 likes671 downloads7mo agoHugging Face25Tamazight-NLP /Weblate-Translations Dataset Card for Weblate Translations A dataset containing strings from projects hosted on Weblate and their translations into other languages. Please consider donating or contributing to Weblate if you find this dataset useful. Dataset Details Dataset Description Curated by: Mohamed Aymane Farhi Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): Check the README YAML metadata… See the full description on the dataset page: https://huggingface.co/datasets/Tamazight-NLP/Weblate-Translations.texttranslation10M<n<100M5 likes667 downloads4mo agoHugging Face26UBC-NLP /alexandria Dataset Card for Alexandria Alexandria covers 13 Arab countries, 11 domains, and 107K community-driven samples. Alexandria is a multi-domain English↔Dialectal Arabic machine translation dataset designed for culturally inclusive, dialect-aware NLP and LLM evaluation. It pairs English multi-turn conversations with human-translated dialectal Arabic from 13 Arab countries, enriched with sub-dialect metadata (based on city-level information), domain labels, persona roles… See the full description on the dataset page: https://huggingface.co/datasets/UBC-NLP/alexandria.texttranslation10K<n<100K15 likes626 downloads4mo agoHugging Face27nlp-chula /ThaiTrees ThaiTrees A 342M-token corpus of Thai drawn from news, Wikipedia, spoken transcripts and social media, automatically parsed under the Universal Dependencies framework. It is released as three artefacts: a raw text corpus, a frequency lexicon, and a dependency-parsed corpus in CoNLL-U. Dataset Summary ThaiTrees contains 341,967,133 tokens across 366,120 documents in four domains (news, Wikipedia, spoken transcripts, social media). Document identifiers are shared… See the full description on the dataset page: https://huggingface.co/datasets/nlp-chula/ThaiTrees.tabulartext-generation1M<n<10M2 likes610 downloads17d agoHugging Face28eth-nlped /mathdial Mathdial dataset https://arxiv.org/abs/2305.14536 MathDial: A Dialogue Tutoring Dataset with Rich Pedagogical Properties Grounded in Math Reasoning Problems. MathDial is grounded in math word problems as well as student confusions which provide a challenging testbed for creating faithful and equitable dialogue tutoring models able to reason over complex information. Current models achieve high accuracy in solving such problems but they fail in the task of teaching. Data… See the full description on the dataset page: https://huggingface.co/datasets/eth-nlped/mathdial.tabulartext-generation1K<n<10K18 likes541 downloads2y agoHugging Face29RUC-NLPIR /GISA GISA: A Benchmark for General Information-Seeking Assistant Authors: Yutao Zhu, Xingshuo Zhang, Maosen Zhang, Jiajie Jin, Liancheng Zhang, Xiaoshuai Song, Kangzhi Zhao, Wencong Zeng, Ruiming Tang, Han Li, Ji-Rong Wen, and Zhicheng Dou Benchmark Highlights GISA is a benchmark for General Information-Seeking Assistants with 373 human-crafted queries that reflect real-world information needs. It includes both stable and live subsets, four structured answer… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/GISA.question-answeringn<1K3 likes536 downloads14d agoHugging Face30McGill-NLP /FaithDialFaithDial is a new benchmark for hallucination-free dialogues, created by manually editing hallucinated and uncooperative responses in Wizard of Wikipedia.texttext-generation10K<n<100K18 likes521 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.