Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SALT-NLP /SWE-chatgated SWE-chat: Coding Agent Interactions From Real Users in the Wild 📄 COLM 2026 Paper: arxiv.org/abs/2604.20779 🌐 Website: swe-chat.com Dataset Summary SWE-chat captures real-world AI coding sessions from developers using AI coding assistants (Claude Code, Codex, Gemini CLI, and others via the Entire.io CLI). Each session includes the full conversation transcript, tool calls, thinking traces, code changes, and attribution of human vs. agent-authored code.… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/SWE-chat.tabulartext-generation10M<n<100M123 likes7.7k downloads7d agoHugging Face02coral-nlp /german-commons German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models A comprehensive collection of German-language text data under open licenses for training German language models. Datasheet: DATASHEET.md. Paper: arxiv.org/abs/2510.13996 Code: github.com/coral-nlp/llmdata Bloom Filter (DOLMA-compatible): bloom_filter.bin Dataset Description This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokensof German text data with… See the full description on the dataset page: https://huggingface.co/datasets/coral-nlp/german-commons.tabulartext-generation10M<n<100M41 likes3.1k downloads9mo agoHugging Face03nilc-nlp /assin2 Dataset Card for ASSIN 2 Dataset Summary The ASSIN 2 corpus is composed of rather simple sentences. Following the procedures of SemEval 2014 Task 1. The training and validation data are composed, respectively, of 6,500 and 500 sentence pairs in Brazilian Portuguese, annotated for entailment and semantic similarity. Semantic similarity values range from 1 to 5, and text entailment classes are either entailment or none. The test data are composed of approximately 3,000… See the full description on the dataset page: https://huggingface.co/datasets/nilc-nlp/assin2.tabulartext-classification1K<n<10K16 likes2k downloads3y agoHugging Face04s-nlp /Mintaka_Graph_Features_T5-xl-ssm Dataset Card for "Mintaka_Graph_Features_T5-xl-ssm" More Information needed tabular100K<n<1M0 likes1.9k downloads3y agoHugging Face05princeton-nlp /QuRatedPajama-260B QuRatedPajama Paper: QuRating: Selecting High-Quality Data for Training Language Models A 260B token subset of cerebras/SlimPajama-627B, annotated by princeton-nlp/QuRater-1.3B with sequence-level quality ratings across 4 criteria: Educational Value - e.g. the text includes clear explanations, step-by-step reasoning, or questions and answers Facts & Trivia - how much factual and trivia knowledge the text contains, where specific facts and obscure trivia are preferred over more… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/QuRatedPajama-260B.tabular100M<n<1B7 likes1.8k downloads2y agoHugging Face06SALT-NLP /hle-context-baseline-deeptabular10K<n<100K0 likes1.5k downloads3mo agoHugging Face07princeton-nlp /LitSearch LitSearch: A Retrieval Benchmark for Scientific Literature Search This dataset contains the query set and retrieval corpus for our paper LitSearch: A Retrieval Benchmark for Scientific Literature Search. We introduce LitSearch, a retrieval benchmark comprising 597 realistic literature search queries about recent ML and NLP papers. LitSearch is constructed using a combination of (1) questions generated by GPT-4 based on paragraphs containing inline citations from research papers and… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/LitSearch.tabular100K<n<1M23 likes1.3k downloads2y agoHugging Face08SALT-NLP /agent-collusion Emergent Collusion in Long-Horizon LLM Agent Interaction Xinrui Shi*, Yanzhe Zhang*, Diyi Yang 📄 Paper | 💻 Code | 🤗 Data | 🔍 Data Viewer *Equal contribution. The experiments reported in the paper and its appendices: 53 conditions, 2,650 trajectories, and 27,100 episodes, of which 600 are warm-up and 26,500 are evaluation episodes. Every condition runs the same 50 fixed task sequences. Contents Config / directory Unit Count episodes One two-agent… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/agent-collusion.tabular100K<n<1M1 likes1.2k downloads17d agoHugging Face09Finnish-NLP /mc4_3.1.0_fi_cleaned Dataset Card for "mc4_3.1.0_fi_cleaned" More Information needed tabular10M<n<100M0 likes1.2k downloads3y agoHugging Face10MU-NLPC /Calc-asdiv_a Dataset Card for Calc-asdiv_a Summary The dataset is a collection of simple math word problems focused on arithmetics. It is derived from the arithmetic subset of ASDiv (original repo). The main addition in this dataset variant is the chain column. It was created by converting the solution to a simple html-like language that can be easily parsed (e.g. by BeautifulSoup). The data contains 3 types of tags: gadget: A tag whose content is intended to be evaluated by calling… See the full description on the dataset page: https://huggingface.co/datasets/MU-NLPC/Calc-asdiv_a.tabular1K<n<10K1 likes947 downloads3y agoHugging Face11nyu-dice-lab /lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private Dataset Card for Evaluation run of princeton-nlp/Llama-3-Base-8B-SFT-RDPO Dataset automatically created during the evaluation run of model princeton-nlp/Llama-3-Base-8B-SFT-RDPO The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private.tabular100K<n<1M0 likes919 downloads2y agoHugging Face12SALT-NLP /hle-context-baseline-gpt55tabular10K<n<100K0 likes884 downloads3mo agoHugging Face13causal-nlp /corr2cause Dataset card for corr2cause TODO tabular100K<n<1M31 likes879 downloads3y agoHugging Face14yale-nlp /FOLIOgatedtabular1K<n<10K74 likes825 downloads3y agoHugging Face15hitachi-nlp /proofwriter_processed_OWAtabular10K<n<100K2 likes791 downloads2y agoHugging Face16UBC-NLP /EgyMMLU Dataset Card for EgyMMLU Dataset Description Dataset Summary Languages Dataset Structure Data Instances Data Fields Data Splits Dataset Creation Curation Rationale Source Data Personal and Sensitive Information Considerations for Using the Data Social Impact of Dataset Discussion of Biases Other Known Limitations Additional Information Dataset Curators Licensing Information Citation Information Dataset Summary EgyMMLU is a benchmark created to test the… See the full description on the dataset page: https://huggingface.co/datasets/UBC-NLP/EgyMMLU.tabular10K<n<100K0 likes788 downloads11mo agoHugging Face17nlp-chula /ThaiTrees ThaiTrees A 342M-token corpus of Thai drawn from news, Wikipedia, spoken transcripts and social media, automatically parsed under the Universal Dependencies framework. It is released as three artefacts: a raw text corpus, a frequency lexicon, and a dependency-parsed corpus in CoNLL-U. Dataset Summary ThaiTrees contains 341,967,133 tokens across 366,120 documents in four domains (news, Wikipedia, spoken transcripts, social media). Document identifiers are shared… See the full description on the dataset page: https://huggingface.co/datasets/nlp-chula/ThaiTrees.tabulartext-generation1M<n<10M2 likes610 downloads16d agoHugging Face18nyu-dice-lab /lm-eval-results-hkust-nlp-dart-math-llama3-8b-prop2diff-private Dataset Card for Evaluation run of hkust-nlp/dart-math-llama3-8b-prop2diff Dataset automatically created during the evaluation run of model hkust-nlp/dart-math-llama3-8b-prop2diff The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-hkust-nlp-dart-math-llama3-8b-prop2diff-private.tabular100K<n<1M0 likes604 downloads2y agoHugging Face19hitachi-nlp /FLD.v2 Dataset Card for "FLD.v2" For the schema of the dataset, see here. For the whole of the project, see our project page. More Information needed tabular10K<n<100K15 likes595 downloads3y agoHugging Face20eth-nlped /mathdial Mathdial dataset https://arxiv.org/abs/2305.14536 MathDial: A Dialogue Tutoring Dataset with Rich Pedagogical Properties Grounded in Math Reasoning Problems. MathDial is grounded in math word problems as well as student confusions which provide a challenging testbed for creating faithful and equitable dialogue tutoring models able to reason over complex information. Current models achieve high accuracy in solving such problems but they fail in the task of teaching. Data… See the full description on the dataset page: https://huggingface.co/datasets/eth-nlped/mathdial.tabulartext-generation1K<n<10K18 likes541 downloads2y agoHugging Face21NLP-FBK /multilingual-medical-reasoning-tracesThis datasets containes the traces generated to answer multiple-choice medical questions in Italian, Englihs, and Spanish. The dataset is structured in 3 parts, one per language. Each part is composed by 2 splits, one containing the examples generated from medqa, one from medmcqa. The columns are: id, representing an unique identifier full_question, representing the medical question options, a dictionary of options to answer the question and their identifiers list_of_options, a list of the… See the full description on the dataset page: https://huggingface.co/datasets/NLP-FBK/multilingual-medical-reasoning-traces.tabular100K<n<1M1 likes529 downloads7mo agoHugging Face22s-nlp /KGQASubgraphsRanking 📰 News [12/2023] Publishing of the original paper "Large Language Models Meets Knowledge Graph to Answer Factoid Questions". This paper first introduces the novelty of the extracted subgraphs; which provide valuable information for different methods of ranking. The paper leveraged T5-like models, and achieve SOTA results with Graph2Text ranking. Dataset Summary KGQASubgraphsRanking is the total-packaged dataset for both publications mentioned in the News section.… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/KGQASubgraphsRanking.tabular100K<n<1M1 likes493 downloads2y agoHugging Face23OALL /details_princeton-nlp__Llama-3-8B-ProLong-512k-Instruct Dataset Card for Evaluation run of princeton-nlp/Llama-3-8B-ProLong-512k-Instruct Dataset automatically created during the evaluation run of model princeton-nlp/Llama-3-8B-ProLong-512k-Instruct. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_princeton-nlp__Llama-3-8B-ProLong-512k-Instruct.tabular100K<n<1M0 likes486 downloads2y agoHugging Face24surrey-nlp /Low-resource-QE-DA-dataset Low-resource QE-DA Dataset Direct Assessment (DA) quality estimation data for English→Indic (Gujarati, Hindi, Marathi, Tamil, Telugu) and related Estonian/Nepali/Sinhala pairs, released with the ALOPE work on LLM-based QE. Paper: Sindhujan, A., Qian, S., Matthew, C.C.C., Orasan, C., and Kanojia, D. (2024). ALOPE: Adaptive Layer Optimization for Translation Quality Estimation using Large Language Models. In Second Conference on Language Modeling. (arXiv) Task: Sentence-level quality… See the full description on the dataset page: https://huggingface.co/datasets/surrey-nlp/Low-resource-QE-DA-dataset.tabularother100K<n<1M0 likes448 downloads11mo agoHugging Face25lasha-nlp /CONDAQA Dataset Card for CondaQA: A Contrastive Reading Comprehension Dataset for Reasoning about Negation Dataset Summary Data from the EMNLP 2022 paper by Ravichander et al.: "CondaQA: A Contrastive Reading Comprehension Dataset for Reasoning about Negation". If you use this dataset, we would appreciate you citing our work: @inproceedings{ravichander-et-al-2022-condaqa, title={CONDAQA: A Contrastive Reading Comprehension Dataset for Reasoning about Negation}… See the full description on the dataset page: https://huggingface.co/datasets/lasha-nlp/CONDAQA.tabularquestion-answering10K<n<100K5 likes441 downloads4y agoHugging Face26nilc-nlp /assin Dataset Card for ASSIN Dataset Summary The ASSIN (Avaliação de Similaridade Semântica e INferência textual) corpus is a corpus annotated with pairs of sentences written in Portuguese that is suitable for the exploration of textual entailment and paraphrasing classifiers. The corpus contains pairs of sentences extracted from news articles written in European Portuguese (EP) and Brazilian Portuguese (BP), obtained from Google News Portugal and Brazil, respectively. To… See the full description on the dataset page: https://huggingface.co/datasets/nilc-nlp/assin.tabulartext-classification10K<n<100K10 likes414 downloads3y agoHugging Face27McGill-NLP /TopiOCQATopiOCQA is an information-seeking conversational dataset with challenging topic switching phenomena.tabulartext-retrieval10K<n<100K10 likes377 downloads3y agoHugging Face28nlpatunt /D_persuade_2 Persuade_2 The PERSUADE 2.0 corpus (Persuasive Essays for Rating, Selecting, and Understanding Argumentative and Discourse Elements) contains over 25,000 argumentative essays written by 6th–12th grade students in the United States, covering 15 distinct prompts across two writing tasks: independent and source-based writing. The corpus also provides detailed individual and demographic information for each writer. This is the train, test, and validation split of the dataset.… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_persuade_2.tabular10K<n<100K0 likes366 downloads7mo agoHugging Face29GroNLP /ik-nlp-22_winemagtabular10K<n<100K6 likes345 downloads5y agoHugging Face30Finnish-NLP /oscar_2301_fi_cleaned Dataset Card for "oscar_2301_fi_cleaned" More Information needed tabular1M<n<10M0 likes295 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.