Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01aps /super_glue Dataset Card for "super_glue" Dataset Summary SuperGLUE (https://super.gluebenchmark.com/) is a new benchmark styled after GLUE with a new set of more difficult language understanding tasks, improved resources, and a new public leaderboard. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances axb Size of downloaded dataset files: 0.03 MB Size of… See the full description on the dataset page: https://huggingface.co/datasets/aps/super_glue.tabulartext-classification100K<n<1M210 likes250k downloads1y agoHugging Face02allenai /openbookqa Dataset Card for OpenBookQA Dataset Summary OpenBookQA aims to promote research in advanced question-answering, probing a deeper understanding of both the topic (with salient facts summarized as an open book, also provided with the dataset) and the language it is expressed in. In particular, it contains questions that require multi-step reasoning, use of additional common and commonsense knowledge, and rich text comprehension. OpenBookQA is a new kind of… See the full description on the dataset page: https://huggingface.co/datasets/allenai/openbookqa.tabularquestion-answering10K<n<100K151 likes205k downloads3y agoHugging Face03TIGER-Lab /MMLU-Pro MMLU-Pro Dataset MMLU-Pro dataset is a more robust and challenging massive multi-task understanding dataset tailored to more rigorously benchmark large language models' capabilities. This dataset contains 12K complex questions across various disciplines. |Github | 🏆Leaderboard | 📖Paper | 🚀 What's New [2026.03.11] Added more cutting-edge frontier models to the leaderboard, including the Claude-4.6 series, Seed2.0 series, Qwen3.5 series, and Gemini-3.1-Pro… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/MMLU-Pro.tabularquestion-answering10K<n<100K523 likes197k downloads5mo agoHugging Face04Idavidrein /gpqagated Dataset Card for GPQA GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google. We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation… See the full description on the dataset page: https://huggingface.co/datasets/Idavidrein/gpqa.tabularquestion-answering1K<n<10K565 likes114k downloads14d agoHugging Face05OpenSQZ /AutoMathText-V2 🚀 AutoMathText-V2: A 2.46 Trillion Token AI-Curated STEM Pretraining Dataset &nbsp; 🎉 AutoMathText-v2 has surpassed 1.5 million downloads! We'd love to know how you're using it. Please take 1 minute to fill out our use case survey. Your feedback will directly shape the future roadmap of this dataset.👉 Share your use case here 📊 AutoMathText-V2 consists of 2.46 trillion tokens of high-quality, deduplicated text spanning web content, mathematics, code, reasoning, and… See the full description on the dataset page: https://huggingface.co/datasets/OpenSQZ/AutoMathText-V2.tabulartext-generation1B<n<10B78 likes29k downloads4mo agoHugging Face06open-law-data-thailand /soc-ratchakitcha Royal Gazette Thailand (Ratchakitcha) Dataset ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable) โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย Dataset Description ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.tabulartext-retrieval1M<n<10M14 likes27k downloads6h agoHugging Face07gililior /mmlu-prox-eval-predictions MMLU-ProX Multilingual Model Predictions Raw per-sample model predictions on MMLU-ProX across 29 languages and 25 open-weight LLMs, produced with lm-evaluation-harness. This dataset releases the full prediction logs (not just aggregate scores) so that item-level responses can be re-analysed — e.g. for Item Response Theory (IRT) modelling of multilingual benchmarks, error analysis, or per-item difficulty estimation. Repository structure mmlu_prox_<lang>/ └──… See the full description on the dataset page: https://huggingface.co/datasets/gililior/mmlu-prox-eval-predictions.tabularquestion-answering1M<n<10M0 likes25k downloads4mo agoHugging Face08argilla /distilabel-capybara-dpo-7k-binarized Capybara-DPO 7K binarized A DPO dataset built with distilabel atop the awesome LDJnr/Capybara This is a preview version to collect feedback from the community. v2 will include the full base dataset and responses from more powerful models. Why? Multi-turn dialogue data is key to fine-tune capable chat models. Multi-turn preference data has been used by the most relevant RLHF works (Anthropic, Meta Llama2, etc.). Unfortunately, there are very few… See the full description on the dataset page: https://huggingface.co/datasets/argilla/distilabel-capybara-dpo-7k-binarized.tabularquestion-answering1K<n<10K184 likes22k downloads2y agoHugging Face09KodCode /KodCode-V1-SFT-R1 🐱 KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding KodCode is the largest fully-synthetic open-source dataset providing verifiable solutions and tests for coding tasks. It contains 12 distinct subsets spanning various domains (from algorithmic to package-specific knowledge) and difficulty levels (from basic coding exercises to interview and competitive programming challenges). KodCode is designed for both supervised fine-tuning (SFT) and RL tuning. 🕸️… See the full description on the dataset page: https://huggingface.co/datasets/KodCode/KodCode-V1-SFT-R1.tabularquestion-answering100K<n<1M40 likes19k downloads2y agoHugging Face10Ayushnangia /docmath-eval-failures-200 DocMath-Eval Failures 200: Agent Benchmark & Leaderboard A curated benchmark of 200 challenging financial math questions that leading AI models failed to answer correctly, with comprehensive evaluation results from multiple AI agents. Leaderboard Evaluated on 2026-02-21 using LLM-as-Judge (Qwen QwQ-32B) for soft scoring. Rank Agent Model Exact Match Judge: Exact Judge: Approx Judge: Total Wrong Avg Duration Avg Tool Calls 1 TRAE Agent Opus 4.5 98/200 (49.0%) 96… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/docmath-eval-failures-200.tabularquestion-answering1K<n<10K0 likes17k downloads8mo agoHugging Face11nguha /legalbench Dataset Card for Dataset Name Homepage: https://hazyresearch.stanford.edu/legalbench/ Repository: https://github.com/HazyResearch/legalbench/ Paper: https://arxiv.org/abs/2308.11462 Dataset Description Dataset Summary The LegalBench project is an ongoing open science effort to collaboratively curate tasks for evaluating legal reasoning in English large language models (LLMs). The benchmark currently consists of 162 tasks gathered from 40… See the full description on the dataset page: https://huggingface.co/datasets/nguha/legalbench.tabulartext-classification10K<n<100K188 likes17k downloads6mo agoHugging Face12mikex86 /stackoverflow-posts StackOverflow Posts Markdown Dataset Summary This dataset contains all posts submitted to StackOverflow before the 14th of June 2023 formatted as Markdown text. The dataset contains ~60 Million posts, totaling ~35GB in size and ~65 billion characters of text. The data is sourced from Internet Archive StackExchange Data Dump. Dataset Structure Each record corresponds to one post of a particular type. Original ordering from the data dump is not exactly preserved… See the full description on the dataset page: https://huggingface.co/datasets/mikex86/stackoverflow-posts.tabularquestion-answering10M<n<100M63 likes14k downloads3y agoHugging Face13bowen-upenn /PersonaMem-v2 PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory 📅 We have now released PersonaMem-v3! 🚨 The paper is now released. View the full paper here and codebase here. Personalization is becoming the next milestone of artificial super-intelligence. AI cannot always satisfy every user, especially on tasks with subjective goals, but personalization offers a path toward pluralistic alignment.… See the full description on the dataset page: https://huggingface.co/datasets/bowen-upenn/PersonaMem-v2.tabularquestion-answering10K<n<100K38 likes14k downloads1mo agoHugging Face14vaquill /open-india-lawgated Open India Law Open, structured Indian primary law - plus the scrapers that build it. Every judgment of the Supreme Court of India and all 25 High Courts, the decisions of 15 tribunals and regulators, and Central, State and Union Territory legislation down to the individual section. Normalized to one schema, exclusively from official government sources. Volume Period Court judgments 12,848,644 1950 to 2025 Tribunal and regulator matters 813,168 1985 to 2026… See the full description on the dataset page: https://huggingface.co/datasets/vaquill/open-india-law.tabulartext-retrieval10M<n<100M28 likes13k downloads1mo agoHugging Face15ShaofantuoshuzhengzhiSha /GUIGuard-Bench GUIGuard-Bench (Public Ladder) GUIGuard-Bench is a cross-platform GUI agent benchmark for studying privacy risks and privacy-preserving execution in multimodal GUI agents. This public-ladder release contains 121 GUI interaction trajectories (68 Android + 53 PC) for benchmark evaluation, with 26,407 region-level privacy annotations across 2,002 screenshots. For the anonymous review version of the evaluation toolkit, see GUIGaurd-Bench-CA4F. Dataset Summary GUI agents… See the full description on the dataset page: https://huggingface.co/datasets/ShaofantuoshuzhengzhiSha/GUIGuard-Bench.imagequestion-answering1K<n<10K1 likes9.2k downloads5mo agoHugging Face16lance-format /fineweb-edu FineWeb-Edu (Lance Format) A Lance-formatted version of FineWeb-Edu — over 1.5 billion educational web passages with cleaned text, source metadata, language detection signals, and 384-dim text embeddings — available directly from the Hub at hf://datasets/lance-format/fineweb-edu/data/train.lance. Key features Cleaned passage text in the text column with the source url and title carried alongside. Language detection signals (language, language_probability) for filtered… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/fineweb-edu.tabulartext-retrieval1B<n<10B8 likes8.6k downloads5mo agoHugging Face17OpenGVLab /ShareGPT-4ogatedtabularvisual-question-answering10K<n<100K199 likes8.5k downloads2y agoHugging Face18Azzindani /Legal_Corpus_QA_SynDeepThink 🧠 Legal Corpus QA SynDeepThink Dataset This repository contains a high-intelligence Legal Question-and-Answer dataset, generated through an advanced Iterative and Recursive Thinking process. It bridges the gap between static legal corpora and the dynamic "check-and-recheck" nature of human legal expertise. 🏛️ 💡 The Concept: Iterative & Recursive Legal Logic While standard synthetic datasets are often generated in a single pass, Legal_Corpus_QA_SynDeepThink mimics the… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/Legal_Corpus_QA_SynDeepThink.tabulartext-generation1K<n<10K1 likes8.3k downloads7mo agoHugging Face19tmquan /anle-toaan-gov-vn Vietnamese Án lệ Corpus — anle.toaan.gov.vn 🇻🇳 Tóm tắt. Bộ dữ liệu các bản án + án lệ Việt Nam thu thập từ cổng anle.toaan.gov.vn của Tòa án nhân dân tối cao. Mỗi văn bản đi kèm markdown chuẩn hoá tiếng Việt và một lớp grounding mức câu (mỗi trích dẫn mang sentence_id + char span trỏ ngược vào markdown). Bộ dữ liệu là một phần của ViLA common-corpus và ship ba cấu hình HF theo chuẩn chung: documents (bảng chính) · embeddings (vector 4096-D Nemotron-3-Embed-8B) · reduces (toạ… See the full description on the dataset page: https://huggingface.co/datasets/tmquan/anle-toaan-gov-vn.tabulartext-classification10K<n<100K10 likes8.2k downloads19d agoHugging Face20choucsan /Gaokao-Compass-11M English GaokaoCompass — China College Admission Dataset GaokaoCompass is a structured dataset of China's national college entrance examination (Gaokao) admission records, covering all 31 provinces from 2017 to 2025. It includes enrollment plans, university admission cutoff scores, major-level admission scores, and score-ranking tables. The dataset is designed to help students, parents, and researchers make informed decisions with… See the full description on the dataset page: https://huggingface.co/datasets/choucsan/Gaokao-Compass-11M.tabulartabular-classification10M<n<100M6 likes7.7k downloads2mo agoHugging Face21KodCode /KodCode-V1-SFT-4o 🐱 KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding KodCode is the largest fully-synthetic open-source dataset providing verifiable solutions and tests for coding tasks. It contains 12 distinct subsets spanning various domains (from algorithmic to package-specific knowledge) and difficulty levels (from basic coding exercises to interview and competitive programming challenges). KodCode is designed for both supervised fine-tuning (SFT) and RL tuning. 🕸️… See the full description on the dataset page: https://huggingface.co/datasets/KodCode/KodCode-V1-SFT-4o.tabularquestion-answering100K<n<1M10 likes7.5k downloads2y agoHugging Face22PromptEval /PromptEval_MMLU_full MMLU Multi-Prompt Evaluation Data Overview This dataset contains the results of a comprehensive evaluation of various Large Language Models (LLMs) using multiple prompt templates on the Massive Multitask Language Understanding (MMLU) benchmark. The data is introduced in Maia Polo, Felipe, Ronald Xu, Lucas Weber, Mírian Silva, Onkar Bhardwaj, Leshem Choshen, Allysson Flavio Melo de Oliveira, Yuekai Sun, and Mikhail Yurochkin. "Efficient multi-prompt evaluation of LLMs."… See the full description on the dataset page: https://huggingface.co/datasets/PromptEval/PromptEval_MMLU_full.tabularquestion-answering10M<n<100M3 likes7.5k downloads2y agoHugging Face23jaredfern /codah Dataset Card for COmmonsense Dataset Adversarially-authored by Humans Dataset Summary The COmmonsense Dataset Adversarially-authored by Humans (CODAH) is an evaluation set for commonsense question-answering in the sentence completion style of SWAG. As opposed to other automatically generated NLI datasets, CODAH is adversarially constructed by humans who can view feedback from a pre-trained model and use this information to design challenging commonsense questions.… See the full description on the dataset page: https://huggingface.co/datasets/jaredfern/codah.tabularquestion-answering10K<n<100K6 likes7.3k downloads3y agoHugging Face24hudsongouge /AoPS-Scrape AoPS-Scrape Problems and solutions scraped from Art of Problem Solving (AoPS) Online class homework endpoints. Obtained legally in accordance with AoPS's Terms of Service. This is not unauthorized redistribution of pirated material — access was through a legitimate authenticated AoPS Online class session. Splits Splits are named by scrape date (YYYY_MM_DD), plus a cross-date content-deduplicated split: Split Rows Notes deduplicated 29,964 One row per… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/AoPS-Scrape.tabularquestion-answering10K<n<100K2 likes7.1k downloads3mo agoHugging Face25ulamai /UnsolvedMath🌐 Browse UnsolvedMath online ✅ Paper: Open Mathematical Problems as an AI Reasoning Benchmark UnsolvedMath Dataset A comprehensive curated collection of 15,458 open, partially solved, and solved mathematics problems across all domains and difficulty levels, including the largest collection of Erdős problems available in machine-readable format. Available for browsing at unsolvedmath.com. Paper: "Open Mathematical Problems as an AI Reasoning Benchmark" Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/UnsolvedMath.documentquestion-answering10K<n<100K80 likes6.9k downloads8d agoHugging Face26KodCode /KodCode-Light-RL-10K 🐱 KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding KodCode is the largest fully-synthetic open-source dataset providing verifiable solutions and tests for coding tasks. It contains 12 distinct subsets spanning various domains (from algorithmic to package-specific knowledge) and difficulty levels (from basic coding exercises to interview and competitive programming challenges). KodCode is designed for both supervised fine-tuning (SFT) and RL tuning. 🕸️… See the full description on the dataset page: https://huggingface.co/datasets/KodCode/KodCode-Light-RL-10K.tabularquestion-answering10K<n<100K9 likes6.7k downloads2y agoHugging Face27Social-AI-2026 /worldcup2026 ⚽ WorldCup Arena A Leakage-Free Forecasting Benchmark on a Live Tournament Can a language model forecast a match — when the match had not been played at the moment it was asked? &nbsp; 🌐 &nbsp; Language / 语言 &nbsp;:&nbsp; 中文 &nbsp; ▾ &nbsp; 📊 四张表 点开本页顶部的 Data Studio 标签即可浏览,也可以直接按名字加载。 Config 行数 内容 fixtures 104 基准本体 —— 喂给模型的头部信息,以及结算后的 90 分钟赛果,七个盘口全部推导好(outcome_1x2、over_2_5、both_score、odd_total) dossiers 2,208 简报索引 —— 46 快照 × 48… See the full description on the dataset page: https://huggingface.co/datasets/Social-AI-2026/worldcup2026.tabularquestion-answering1K<n<10K0 likes6.3k downloads2mo agoHugging Face28worstchan /Belle_1.4M-SLAM-Omni Belle_1.4M This dataset is prepared for the reproduction of SLAM-Omni. This is a multi-round Chinese spoken dialogue training dataset. For code and usage examples, please refer to the related GitHub repository: X-LANCE/SLAM-LLM (examples/s2s) 🔧 Modifications Data Filtering: We removed samples with excessively long data. Speech Response Tokens: We used CosyVoice to synthesize corresponding semantic speech tokens for the speech response. These tokens, represented as… See the full description on the dataset page: https://huggingface.co/datasets/worstchan/Belle_1.4M-SLAM-Omni.tabularquestion-answering1M<n<10M3 likes6.2k downloads1y agoHugging Face29cambridgeltl /xcopa Dataset Card for "xcopa" Dataset Summary XCOPA: A Multilingual Dataset for Causal Commonsense Reasoning The Cross-lingual Choice of Plausible Alternatives dataset is a benchmark to evaluate the ability of machine learning models to transfer commonsense reasoning across languages. The dataset is the translation and reannotation of the English COPA (Roemmele et al. 2011) and covers 11 languages from 11 families and several areas around the globe. The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/cambridgeltl/xcopa.tabularquestion-answering10K<n<100K22 likes6.1k downloads3y agoHugging Face30PromptEval /PromptEval_MMLU_correctness MMLU Multi-Prompt Evaluation Data (correctness scores) Overview This dataset contains the results of a comprehensive evaluation of various Large Language Models (LLMs) using multiple prompt templates on the Massive Multitask Language Understanding (MMLU) benchmark. The data is introduced in Maia Polo, Felipe, Ronald Xu, Lucas Weber, Mírian Silva, Onkar Bhardwaj, Leshem Choshen, Allysson Flavio Melo de Oliveira, Yuekai Sun, and Mikhail Yurochkin. "Efficient multi-prompt… See the full description on the dataset page: https://huggingface.co/datasets/PromptEval/PromptEval_MMLU_correctness.tabularquestion-answering10K<n<100K2 likes5.8k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.