Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Idavidrein /gpqagated Dataset Card for GPQA GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google. We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation… See the full description on the dataset page: https://huggingface.co/datasets/Idavidrein/gpqa.tabularquestion-answering1K<n<10K565 likes114k downloads14d agoHugging Face02google /deepsearchqa DeepSearchQA A 900-prompt factuality benchmark from Google DeepMind, designed to evaluate agents on difficult multi-step information-seeking tasks across 17 different fields. ▶ Google DeepMind Release Blog Post▶ DeepSearchQA Leaderboard on Kaggle▶ Technical Report▶ Evaluation Starter Code Benchmark DeepSearchQA is a 900-prompt benchmark for evaluating agents on difficult multi-step information-seeking tasks across 17 different fields. Unlike traditional… See the full description on the dataset page: https://huggingface.co/datasets/google/deepsearchqa.textquestion-answeringn<1K132 likes24k downloads10mo agoHugging Face03fka /prompts.chat a.k.a. Awesome ChatGPT Prompts This is a Dataset Repository mirror of prompts.chat — a social platform for AI prompts. 📢 Notice This Hugging Face dataset is a mirror. For the latest prompts, features, and community contributions, please visit: 🌐 Website: prompts.chat 📦 GitHub: github.com/f/awesome-chatgpt-prompts About prompts.chat is an open-source platform where users can share, discover, and collect AI prompts from the community. The project can… See the full description on the dataset page: https://huggingface.co/datasets/fka/prompts.chat.textquestion-answering1K<n<10K9.8k likes23k downloads1mo agoHugging Face04TAUR-Lab /MuSR MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning Creating murder mysteries that require multi-step reasoning with commonsense using ChatGPT! By: Zayne Sprague, Xi Ye, Kaj Bostrom, Swarat Chaudhuri, and Greg Durrett. View the dataset on our custom viewer and project website! Check out the paper. Appeared at ICLR 2024 as a spotlight presentation! Git Repo with the source data, how to recreate the dataset (and create new ones!) here textquestion-answeringn<1K25 likes18k downloads2y agoHugging Face05nguha /legalbench Dataset Card for Dataset Name Homepage: https://hazyresearch.stanford.edu/legalbench/ Repository: https://github.com/HazyResearch/legalbench/ Paper: https://arxiv.org/abs/2308.11462 Dataset Description Dataset Summary The LegalBench project is an ongoing open science effort to collaboratively curate tasks for evaluating legal reasoning in English large language models (LLMs). The benchmark currently consists of 162 tasks gathered from 40… See the full description on the dataset page: https://huggingface.co/datasets/nguha/legalbench.tabulartext-classification10K<n<100K188 likes17k downloads6mo agoHugging Face06google /frames-benchmark FRAMES: Factuality, Retrieval, And reasoning MEasurement Set FRAMES is a comprehensive evaluation dataset designed to test the capabilities of Retrieval-Augmented Generation (RAG) systems across factuality, retrieval accuracy, and reasoning. Our paper with details and experiments is available on arXiv: https://arxiv.org/abs/2409.12941. Dataset Overview 824 challenging multi-hop questions requiring information from 2-15 Wikipedia articles Questions span diverse topics… See the full description on the dataset page: https://huggingface.co/datasets/google/frames-benchmark.texttext-classificationn<1K267 likes15k downloads2y agoHugging Face07bowen-upenn /PersonaMem-v2 PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory 📅 We have now released PersonaMem-v3! 🚨 The paper is now released. View the full paper here and codebase here. Personalization is becoming the next milestone of artificial super-intelligence. AI cannot always satisfy every user, especially on tasks with subjective goals, but personalization offers a path toward pluralistic alignment.… See the full description on the dataset page: https://huggingface.co/datasets/bowen-upenn/PersonaMem-v2.tabularquestion-answering10K<n<100K38 likes14k downloads1mo agoHugging Face08ikala /tmmluplus TMMLU+ : Large scale traditional chinese massive multitask language understanding iKala presents TMMLU+, a large-scale benchmark for evaluating LLM capabilities in Traditional Chinese, with content primarily reflecting Taiwan's linguistic, educational, and professional contexts. It covers 66 subjects, from elementary to professional domains, and is approximately six times larger than TMMLU with broader, more balanced coverage. TMMLU+ v1.1 improves benchmark quality through… See the full description on the dataset page: https://huggingface.co/datasets/ikala/tmmluplus.textquestion-answering10K<n<100K231 likes12k downloads13d agoHugging Face09mjuicem /StreamingBench StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding 🏠 Project Page | 📄 arXiv Paper | 📦 Dataset | 🏅Leaderboard StreamingBench evaluates Multimodal Large Language Models (MLLMs) in real-time, streaming video understanding tasks. 🌟 [NEW! 2025.05.15] 🔥: Seed1.5-VL achieved ALL model SOTA with a score of 82.80 on the Proactive Output. [NEW! 2025.03.17] ⭐: ViSpeeker achieved Open-Source SOTA with a score of 61.60 on the… See the full description on the dataset page: https://huggingface.co/datasets/mjuicem/StreamingBench.imagequestion-answering1K<n<10K13 likes11k downloads1y agoHugging Face10openai /MMMLU Multilingual Massive Multitask Language Understanding (MMMLU) The MMLU is a widely recognized benchmark of general knowledge attained by AI models. It covers a broad range of topics from 57 different categories, covering elementary-level knowledge up to advanced professional subjects like law, physics, history, and computer science. We translated the MMLU’s test set into 14 languages using professional human translators. Relying on human translators for this evaluation increases… See the full description on the dataset page: https://huggingface.co/datasets/openai/MMMLU.textquestion-answering100K<n<1M526 likes11k downloads2y agoHugging Face11bitext /Bitext-customer-support-llm-chatbot-training-dataset Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.textquestion-answering10K<n<100K198 likes10k downloads2y agoHugging Face12bench-llm /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llm/or-bench.imagetext-generation10K<n<100K22 likes9.4k downloads2y agoHugging Face13databricks /officeqagated OfficeQA Dataset Summary OfficeQA is a grounded reasoning benchmark by Databricks for evaluating model and agent performance on end-to-end reasoning over real-world documents. The benchmark consists of question–answer pairs that require reasoning over historical U.S. Treasury Bulletin documents (1939–2025), which contain dense financial tables, charts, and narrative text. OfficeQA is designed to test retrieval, tool use, and multi-step reasoning in… See the full description on the dataset page: https://huggingface.co/datasets/databricks/officeqa.documentquestion-answeringn<1K28 likes9k downloads3mo agoHugging Face14ArtificialAnalysis /AA-Omniscience-Public Public Dataset for AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models AA-Omniscience-Public contains 600 questions across a wide range of domains used to test a model’s knowledge and hallucination tendencies. Leaderboard and detailed results Paper Introduction We introduce AA-Omniscience, a benchmark dataset designed to measure a model’s ability to both recall factual information accurately across domains, and correctly… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/AA-Omniscience-Public.documentquestion-answeringn<1K52 likes8.3k downloads1mo agoHugging Face15keivalya /MedQuad-MedicalQnADataset Reference: "A Question-Entailment Approach to Question Answering". Asma Ben Abacha and Dina Demner-Fushman. BMC Bioinformatics, 2019. textquestion-answering10K<n<100K133 likes8.2k downloads3y agoHugging Face16sapientinc /sudoku-extreme Hardest Sudoku Puzzle Dataset V2 This dataset contains a mixture of easy and very hard Sudoku puzzles collected from the Sudoku community. Dataset Composition Sources tdoku benchmarks enjoysudoku Easy Puzzles (1.1M) puzzles0_kaggle puzzles1_unbiased puzzles2_17_clue Hard Puzzles (3.1M) puzzles3_magictour_top1465 puzzles4_forum_hardest_1905 puzzles6_forum_hardest_1106 ph_2010/01_file1.txt Dataset Characteristics All… See the full description on the dataset page: https://huggingface.co/datasets/sapientinc/sudoku-extreme.textquestion-answering1M<n<10M35 likes5.6k downloads2y agoHugging Face17jiang-cc /MMAD MMAD: The First-Ever Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection 💡 This dataset is the full version of MMAD Content:Containing both questions, images, and captions. Questions: All questions are presented in a multiple-choice format with manual verification, including options and answers. Images:Images are collected from the following links: DS-MVTec , MVTec-AD , MVTec-LOCO , VisA , GoodsAD. We retained the mask… See the full description on the dataset page: https://huggingface.co/datasets/jiang-cc/MMAD.imagequestion-answering10K<n<100K17 likes5.5k downloads1y agoHugging Face18kaysss /leetcode-problem-solutions LeetCode Solution Dataset This dataset contains community-contributed LeetCode solutions scraped from public discussions and solution pages, enriched with metadata such as vote counts, author info, tags, and full code content. The goal is to make high-quality, peer-reviewed coding solutions programmatically accessible for research, analysis, educational use, or developer tooling. Column Descriptions Column Name Type Description question_slug string The unique… See the full description on the dataset page: https://huggingface.co/datasets/kaysss/leetcode-problem-solutions.tabulartext-classification100K<n<1M9 likes5.4k downloads1y agoHugging Face19mrlbenchmarks /global-piqa-nonparallel Global PIQA Non-Parallel Global PIQA is a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by over 350 researchers from over 65 countries around the world. The non-parallel split covers 136 language varieties, covering five continents, 18 language families, and 24 writing systems. In this non-parallel split, over 50% of examples reference local foods, customs, traditions, or other culturally-specific elements. Details are in our preprint:… See the full description on the dataset page: https://huggingface.co/datasets/mrlbenchmarks/global-piqa-nonparallel.imagequestion-answering10K<n<100K40 likes5.2k downloads4mo agoHugging Face20Tobi-Bueck /customer-support-tickets Featuring Labeled Customer Emails and Support Responses 🔧 Synthetic IT Ticket Generator — Custom Dataset Create a dataset tailored to your own queues & priorities (no PII). 👉 Generate custom data Define your queues, priorities, language Need an on-prem AI to auto-classify tickets?→ Open Ticket AI There are 2 Versions of the dataset, the new version has more tickets, but only languages english and german. So please look at both files, to find what best fits… See the full description on the dataset page: https://huggingface.co/datasets/Tobi-Bueck/customer-support-tickets.texttext-classification10K<n<100K42 likes4.8k downloads3mo agoHugging Face21prquan /STARK_10k STARK: Spatial-Temporal reAsoning benchmaRK STARK is a comprehensive benchmark designed to systematically evaluate large language models (LLMs) and large reasoning models (LRMs) on spatial-temporal reasoning tasks, particularly for applications in cyber-physical systems (CPS) such as robotics, autonomous vehicles, and smart city infrastructure. Dataset Summary Hierarchical Benchmark: Tasks are structured across three levels of reasoning complexity: State Estimation:… See the full description on the dataset page: https://huggingface.co/datasets/prquan/STARK_10k.textquestion-answering10K<n<100K1 likes4.6k downloads1y agoHugging Face22JailbreakV-28K /JailBreakV-28k ⛓‍💥 JailBreakV-28K: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks 🌐 GitHub | 🛎 Project Page | 👉 Download full datasets If you like our project, please give us a star ⭐ on Hugging Face for the latest update. 📰 News Date Event 2024/07/09 🎉 Our paper is accepted by COLM 2024. 2024/06/22 🛠️ We have updated our version to V0.2, which supports users to customize their attack models… See the full description on the dataset page: https://huggingface.co/datasets/JailbreakV-28K/JailBreakV-28k.imagetext-generation10K<n<100K76 likes4.4k downloads2y agoHugging Face23rubend18 /ChatGPT-Jailbreak-Prompts Dataset Card for Dataset Name Name ChatGPT Jailbreak Prompts Dataset Summary ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT. Languages [English] tabularquestion-answeringn<1K279 likes4.3k downloads3y agoHugging Face24basicv8vc /SimpleQA SimpleQA A factuality benchmark called SimpleQA that measures the ability for language models to answer short, fact-seeking questions. Sources openai/simple-evals Introducing SimpleQA Measuring short-form factuality in large language models textquestion-answering1K<n<10K33 likes4k downloads2y agoHugging Face25plnguyen2908 /AV-SpeakerBench AV-SpeakerBench Audiovisual QA benchmark with speaker-aware questions and aligned clips. This drop includes trimmed segments (audio-only, visual-only, audiovisual) plus annotations to probe fine-grained AV reasoning. Project page: https://plnguyen2908.github.io/AV-SpeakerBench-project-page/ Code & benchmarks: https://github.com/plnguyen2908/AV-SpeakerBench Paper: https://arxiv.org/abs/2512.02231 Files test.csv - original annotations and metadata with clip paths… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AV-SpeakerBench.audioquestion-answering1K<n<10K2 likes3.8k downloads10mo agoHugging Face26MBZUAI /ArabicMMLU Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, and Timothy Baldwin MBZUAI, Prince Sattam bin Abdulaziz University, KFUPM, Core42, NYU Abu Dhabi, The University of Melbourne Introduction We present ArabicMMLU, the first multi-task language understanding benchmark for Arabic language, sourced from school exams across diverse… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/ArabicMMLU.tabularquestion-answering10K<n<100K39 likes3.6k downloads2y agoHugging Face27yzy666 /SVBench Dataset Card for SVBench This dataset card aims to provide a comprehensive overview of the SVBench dataset, including its purpose, structure, and sources. For details, see our Project, Paper and GitHub repository. Dataset Details Dataset Description SVBench is the first benchmark specifically designed to evaluate long-context streaming video understanding through temporal multi-turn question-answering (QA) chains. It addresses the limitations of existing video… See the full description on the dataset page: https://huggingface.co/datasets/yzy666/SVBench.textquestion-answering1K<n<10K7 likes3.5k downloads11mo agoHugging Face28Bai-YT /RAGDOLL The RAGDOLL E-Commerce Webpage Dataset This repository contains the RAGDOLL (Retrieval-Augmented Generation Deceived Ordering via AdversariaL materiaLs) dataset as well as its LLM-automated collection pipeline. The RAGDOLL dataset is from the paper Ranking Manipulation for Conversational Search Engines from Samuel Pfrommer, Yatong Bai, Tanmay Gautam, and Somayeh Sojoudi. For experiment code associated with this paper, please refer to this repository. The dataset consists of 10… See the full description on the dataset page: https://huggingface.co/datasets/Bai-YT/RAGDOLL.textquestion-answering1K<n<10K4 likes3.5k downloads2y agoHugging Face29ShenLab /MentalChat16K 🗣️ Synthetic Counseling Conversations Dataset 📝 Description Synthetic Data 10K This dataset consists of 9,775 synthetic conversations between a counselor and a client, covering 33 mental health topics such as 💑 Relationships, 😟 Anxiety, 😔 Depression, 🤗 Intimacy, and 👨‍👩‍👧‍👦 Family Conflict. The conversations were generated using the OpenAI GPT-3.5 Turbo model and a customized adaptation of the Airoboros self-generation framework. The Airoboros… See the full description on the dataset page: https://huggingface.co/datasets/ShenLab/MentalChat16K.textquestion-answering10K<n<100K63 likes3.4k downloads1y agoHugging Face30agicommies /synthiaThe Synthia Dataset is a continuously growing aggregate of validated synthetic explanations of subjects picked from Claude Opus latent space based on varying esotericity in a large general list of technical fields. The explanations are varying in their target audience, level of detail and abstraction while incentivized to target Claude3-grade quality. The Synthia subnet leverages Commune's incentives to create a permissionless mining market around distilling knowledge out of SOTA closed-source… See the full description on the dataset page: https://huggingface.co/datasets/agicommies/synthia.texttable-question-answering1M<n<10M14 likes3.3k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.