datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SWE-chat
SWE-chat: Coding Agent Interactions From Real Users in the Wild
📄 COLM 2026 Paper: arxiv.org/abs/2604.20779
🌐 Website: swe-chat.com
Dataset Summary
SWE-chat captures real-world AI coding sessions from developers using AI coding assistants (Claude Code, Codex, Gemini CLI, and others via the Entire.io CLI). Each session includes the full conversation transcript, tool calls, thinking traces, code changes, and attribution of human vs. agent-authored code.… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/SWE-chat.german-commons
German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models
A comprehensive collection of German-language text data under open licenses for training German language models.
Datasheet: DATASHEET.md.
Paper: arxiv.org/abs/2510.13996
Code: github.com/coral-nlp/llmdata
Bloom Filter (DOLMA-compatible): bloom_filter.bin
Dataset Description
This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokensof German text data with… See the full description on the dataset page: https://huggingface.co/datasets/coral-nlp/german-commons.assin2
Dataset Card for ASSIN 2
Dataset Summary
The ASSIN 2 corpus is composed of rather simple sentences. Following the procedures of SemEval 2014 Task 1.
The training and validation data are composed, respectively, of 6,500 and 500 sentence pairs in Brazilian Portuguese,
annotated for entailment and semantic similarity. Semantic similarity values range from 1 to 5, and text entailment
classes are either entailment or none. The test data are composed of approximately 3,000… See the full description on the dataset page: https://huggingface.co/datasets/nilc-nlp/assin2.Mintaka_Graph_Features_T5-xl-ssm
Dataset Card for "Mintaka_Graph_Features_T5-xl-ssm"
More Information needed
QuRatedPajama-260B
QuRatedPajama
Paper: QuRating: Selecting High-Quality Data for Training Language Models
A 260B token subset of cerebras/SlimPajama-627B, annotated by princeton-nlp/QuRater-1.3B with sequence-level quality ratings across 4 criteria:
Educational Value - e.g. the text includes clear explanations, step-by-step reasoning, or questions and answers
Facts & Trivia - how much factual and trivia knowledge the text contains, where specific facts and obscure trivia are preferred over more… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/QuRatedPajama-260B.hle-context-baseline-deepLitSearch
LitSearch: A Retrieval Benchmark for Scientific Literature Search
This dataset contains the query set and retrieval corpus for our paper LitSearch: A Retrieval Benchmark for Scientific Literature Search. We introduce LitSearch, a retrieval benchmark comprising 597 realistic literature search queries about recent ML and NLP papers. LitSearch is constructed using a combination of (1) questions generated by GPT-4 based on paragraphs containing inline citations from research papers and… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/LitSearch.agent-collusion
Emergent Collusion in Long-Horizon LLM Agent Interaction
Xinrui Shi*, Yanzhe Zhang*, Diyi Yang
📄 Paper | 💻 Code | 🤗 Data | 🔍 Data Viewer
*Equal contribution.
The experiments reported in the paper and its appendices: 53 conditions, 2,650 trajectories, and 27,100 episodes, of which 600 are warm-up and 26,500 are evaluation episodes. Every condition runs the same 50 fixed task sequences.
Contents
Config / directory
Unit
Count
episodes
One two-agent… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/agent-collusion.mc4_3.1.0_fi_cleaned
Dataset Card for "mc4_3.1.0_fi_cleaned"
More Information needed
Calc-asdiv_a
Dataset Card for Calc-asdiv_a
Summary
The dataset is a collection of simple math word problems focused on arithmetics. It is derived from the arithmetic subset of ASDiv (original repo).
The main addition in this dataset variant is the chain column. It was created by converting the solution to a simple html-like language that can be easily
parsed (e.g. by BeautifulSoup). The data contains 3 types of tags:
gadget: A tag whose content is intended to be evaluated by calling… See the full description on the dataset page: https://huggingface.co/datasets/MU-NLPC/Calc-asdiv_a.lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private
Dataset Card for Evaluation run of princeton-nlp/Llama-3-Base-8B-SFT-RDPO
Dataset automatically created during the evaluation run of model princeton-nlp/Llama-3-Base-8B-SFT-RDPO
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private.hle-context-baseline-gpt55corr2cause
Dataset card for corr2cause
TODO
FOLIOproofwriter_processed_OWAEgyMMLU
Dataset Card for EgyMMLU
Dataset Description
Dataset Summary
Languages
Dataset Structure
Data Instances
Data Fields
Data Splits
Dataset Creation
Curation Rationale
Source Data
Personal and Sensitive Information
Considerations for Using the Data
Social Impact of Dataset
Discussion of Biases
Other Known Limitations
Additional Information
Dataset Curators
Licensing Information
Citation Information
Dataset Summary
EgyMMLU is a benchmark created to test the… See the full description on the dataset page: https://huggingface.co/datasets/UBC-NLP/EgyMMLU.ThaiTrees
ThaiTrees
A 342M-token corpus of Thai drawn from news, Wikipedia, spoken transcripts and
social media, automatically parsed under the Universal Dependencies framework.
It is released as three artefacts: a raw text corpus, a frequency lexicon, and
a dependency-parsed corpus in CoNLL-U.
Dataset Summary
ThaiTrees contains 341,967,133 tokens across 366,120 documents in four
domains (news, Wikipedia, spoken transcripts, social media). Document
identifiers are shared… See the full description on the dataset page: https://huggingface.co/datasets/nlp-chula/ThaiTrees.lm-eval-results-hkust-nlp-dart-math-llama3-8b-prop2diff-private
Dataset Card for Evaluation run of hkust-nlp/dart-math-llama3-8b-prop2diff
Dataset automatically created during the evaluation run of model hkust-nlp/dart-math-llama3-8b-prop2diff
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-hkust-nlp-dart-math-llama3-8b-prop2diff-private.FLD.v2
Dataset Card for "FLD.v2"
For the schema of the dataset, see here.
For the whole of the project, see our project page.
More Information needed
mathdial
Mathdial dataset
https://arxiv.org/abs/2305.14536
MathDial: A Dialogue Tutoring Dataset with Rich Pedagogical Properties Grounded in Math Reasoning Problems.
MathDial is grounded in math word problems as well as student confusions which provide a challenging testbed for creating faithful and equitable dialogue tutoring models able to reason over complex information. Current models achieve high accuracy in solving such problems but they fail in the task of teaching.
Data… See the full description on the dataset page: https://huggingface.co/datasets/eth-nlped/mathdial.multilingual-medical-reasoning-tracesThis datasets containes the traces generated to answer multiple-choice medical questions in Italian, Englihs, and Spanish.
The dataset is structured in 3 parts, one per language. Each part is composed by 2 splits, one containing the examples generated from medqa, one from medmcqa.
The columns are:
id, representing an unique identifier
full_question, representing the medical question
options, a dictionary of options to answer the question and their identifiers
list_of_options, a list of the… See the full description on the dataset page: https://huggingface.co/datasets/NLP-FBK/multilingual-medical-reasoning-traces.KGQASubgraphsRanking
📰 News
[12/2023] Publishing of the original paper "Large Language Models Meets Knowledge Graph to Answer Factoid Questions". This paper first introduces the novelty of the extracted subgraphs; which provide valuable information for different methods of ranking. The paper leveraged T5-like models, and achieve SOTA results with Graph2Text ranking.
Dataset Summary
KGQASubgraphsRanking is the total-packaged dataset for both publications mentioned in the News section.… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/KGQASubgraphsRanking.details_princeton-nlp__Llama-3-8B-ProLong-512k-Instruct
Dataset Card for Evaluation run of princeton-nlp/Llama-3-8B-ProLong-512k-Instruct
Dataset automatically created during the evaluation run of model princeton-nlp/Llama-3-8B-ProLong-512k-Instruct.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_princeton-nlp__Llama-3-8B-ProLong-512k-Instruct.Low-resource-QE-DA-dataset
Low-resource QE-DA Dataset
Direct Assessment (DA) quality estimation data for English→Indic (Gujarati, Hindi, Marathi, Tamil, Telugu) and related Estonian/Nepali/Sinhala pairs, released with the ALOPE work on LLM-based QE.
Paper: Sindhujan, A., Qian, S., Matthew, C.C.C., Orasan, C., and Kanojia, D. (2024). ALOPE: Adaptive Layer Optimization for Translation Quality Estimation using Large Language Models. In Second Conference on Language Modeling. (arXiv)
Task: Sentence-level quality… See the full description on the dataset page: https://huggingface.co/datasets/surrey-nlp/Low-resource-QE-DA-dataset.CONDAQA
Dataset Card for CondaQA: A Contrastive Reading Comprehension Dataset for Reasoning about Negation
Dataset Summary
Data from the EMNLP 2022 paper by Ravichander et al.: "CondaQA: A Contrastive Reading Comprehension Dataset for Reasoning about Negation".
If you use this dataset, we would appreciate you citing our work:
@inproceedings{ravichander-et-al-2022-condaqa,
title={CONDAQA: A Contrastive Reading Comprehension Dataset for Reasoning about Negation}… See the full description on the dataset page: https://huggingface.co/datasets/lasha-nlp/CONDAQA.assin
Dataset Card for ASSIN
Dataset Summary
The ASSIN (Avaliação de Similaridade Semântica e INferência textual) corpus is a corpus annotated with pairs of sentences written in
Portuguese that is suitable for the exploration of textual entailment and paraphrasing classifiers. The corpus contains pairs of sentences
extracted from news articles written in European Portuguese (EP) and Brazilian Portuguese (BP), obtained from Google News Portugal
and Brazil, respectively. To… See the full description on the dataset page: https://huggingface.co/datasets/nilc-nlp/assin.TopiOCQATopiOCQA is an information-seeking conversational dataset with challenging topic switching phenomena.D_persuade_2
Persuade_2
The PERSUADE 2.0 corpus (Persuasive Essays for Rating, Selecting, and
Understanding Argumentative and Discourse Elements) contains over 25,000
argumentative essays written by 6th–12th grade students in the United States,
covering 15 distinct prompts across two writing tasks: independent and
source-based writing. The corpus also provides detailed individual and
demographic information for each writer.
This is the train, test, and validation split of the dataset.… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_persuade_2.ik-nlp-22_winemagoscar_2301_fi_cleaned
Dataset Card for "oscar_2301_fi_cleaned"
More Information needed
