datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
air-bench-2024
AIRBench 2024
AIRBench 2024 is a AI safety benchmark that aligns with emerging government
regulations and company policies. It consists of diverse, malicious prompts
spanning categories of the regulation-based safety categories in the
AIR 2024 safety taxonomy.
Dataset Details
Dataset Description
AIRBench 2024 is a AI safety benchmark that aligns with emerging government
regulations and company policies. It consists of diverse, malicious prompts
spanning… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/air-bench-2024.panda-bench
PandaBench
PandaBench is a comprehensive benchmark for evaluating Large Language Model (LLM) safety, focusing on jailbreak attacks, defense mechanisms, and evaluation methodologies.
The PandaGuard framework architecture illustrating the end-to-end pipeline for LLM safety evaluation. The system connects three key components: Attackers, Defenders, and Judges.
Dataset Description
This repository contains the benchmark results from extensive evaluations of various… See the full description on the dataset page: https://huggingface.co/datasets/Beijing-AISI/panda-bench.Sujet-Finance-Instruct-177k
Sujet Finance Dataset Overview
The Sujet Finance dataset is a comprehensive collection designed for the fine-tuning of Language Learning Models (LLMs) for specialized tasks in the financial sector. It amalgamates data from 18 distinct datasets hosted on HuggingFace, resulting in a rich repository of 177,597 entries. These entries span across seven key financial LLM tasks, making Sujet Finance a versatile tool for developing and enhancing financial applications of AI.… See the full description on the dataset page: https://huggingface.co/datasets/sujet-ai/Sujet-Finance-Instruct-177k.r15-ai-search-metamerism
R15: AI Search Metamerism — Cross-Cultural Brand Perception Dataset
Citation: Zharnikov, D. (2026v) | DOI: 10.5281/zenodo.19422427 | Version: v3.2.0
Dataset Summary
This dataset contains the full session logs, aggregated results, and analysis outputs from the R15 large-scale experiment testing whether Large Language Models systematically collapse multi-dimensional brand perception into Economic and Experiential dimensions ("spectral metamerism"). It comprises 21… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/r15-ai-search-metamerism.linalg-bench-llm
LinAlg-Bench: Where LLMs Stop Computing and Start Hallucinating
Ten frontier LLMs drop from near-perfect to near-zero on 5×5 eigenvalue problems. Complete computational collapse is dimension-gated: rare at 3×3, dominant at 4×4 and 5×5. Failures dissociate cleanly by task — eigenvalues fail by constraint-aware fabrication (invented eigenvalues that still match the matrix trace), determinants by sign-accumulation drift. Nearly a third of irrational-spectrum eigenvalue failures are… See the full description on the dataset page: https://huggingface.co/datasets/mst-ai/linalg-bench-llm.knesset-plenums
About
This dataset is derived from raw a/v recordings and human-generated protocols of the Knesset (the Israeli house of representatives) plenums as part of the ivrit.ai project.
Consider visiting the preview space for this dataset here
Method
Data dumps from the Knesset contain A/V recordings, alongside proprietary protocols with timestamps.
We extract the audio stream, and clean up timestamp mistakes (such as backward jumps, or out-of-order timestamp artifacts).
The… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-plenums.Code_Vulnerability_Labeled_Dataset
Dataset Card for Code_Vulnerability_Labeled_Dataset
Dataset Summary
This dataset provides (code, vulnerability) pairs. The vulnerability field takes values according to the CWE annotation:
CWE
Description
CWE-020
Improper Input Validation
CWE-022
Improper Limitation of a Pathname to a Restricted Directory (“Path Traversal”)
CWE-078
Improper Neutralization of Special Elements used in an OS Command (“OS Command Injection”)
CWE-079
Improper Neutralization of… See the full description on the dataset page: https://huggingface.co/datasets/lemon42-ai/Code_Vulnerability_Labeled_Dataset.IndustryBench
IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs
💻Github | 📝Paper
IndustryBench is a multi-lingual benchmark for evaluating the industrial domain knowledge of large language models. It comprises 2,049 expert-curated QA pairs spanning 12 industrial sectors, with human-reviewed translations in Chinese, English, Russian, and Vietnamese.
Overview
Dimension
Details
Total questions
2,049
Languages
Chinese (zh), English (en), Russian (ru)… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench.AutoGrader_Introduction_to_AI
Introduction to AI — Practical Exam Grading Dataset (Stage 2)
1,038 anonymized student submissions to a university-level practical AI exam,
each independently graded by two teaching assistants, plus the rubric and
the instructor reference solutions.
It is the second exam of Where LLM Graders Succeed and Break: Evidence from
Two Computer-Science Exams. The paper's 162 grader configurations on this
exam, the graders that produced them and the analysis code live in the project… See the full description on the dataset page: https://huggingface.co/datasets/KAUSTAcademy/AutoGrader_Introduction_to_AI.Myanmar-Tuberculosis-Guidelines-Instructions
Myanmar Tuberculosis Guidelines Instructions
A bilingual instructional dataset built to support Myanmar's ongoing fight against tuberculosis — turning life-saving guidelines into a usable resource for healthcare workers, educators, and AI researchers working with low-resource languages.
Authors: Min Si Thu, Khin Myat Noe
Abstract
Tuberculosis is still one of Myanmar's biggest public health problems. Part of the difficulty is that good, standardized TB education… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Myanmar-Tuberculosis-Guidelines-Instructions.StorySeedStorySeed is a data set specially designed for training and evaluating the performance of text generation models in the domain of children’s picture book creation. It contains 4376 thoughtfully curated prompt-response pairs, encompassing nine major thematic categories: educational, emotional intelligence and social skills, adventure tales, natural science, folk tales and myths, daily life, humorous stories, bedtime stories, as well as other general picture book stories not specific to any… See the full description on the dataset page: https://huggingface.co/datasets/Aiwensile2/StorySeed.linalg-bench-math-ai-neurips
LinAlg-Bench: A Benchmark Exposing Structural Failure Modes in LLM Linear Algebra — Where Models Stop Computing and Start Hallucinating
This dataset is the official release accompanying the MATH-AI 2026 NeurIPS workshop paper, "LinAlg-Bench: A Benchmark Exposing Structural Failure Modes in LLM Linear Algebra — Where Models Stop Computing and Start Hallucinating." The exact 660-problem core evaluated in that paper (9 tasks × 3 matrix sizes, 6,600 model outputs, 1,156… See the full description on the dataset page: https://huggingface.co/datasets/mst-ai/linalg-bench-math-ai-neurips.AITA-Reddit-Dataset
Dataset Card for AITA Reddit Posts and Comments
Posts of the AITA subreddit, with the 2 top voted comments that share the post verdict. Extracted using REDDIT PushShift (from 2013 to April 2023)
Dataset Details
The dataset contains 270,709 entiries each of which contain the post title, text, verdict, comment1, comment2 and score (number of upvotes)
For more details see paper: https://arxiv.org/abs/2310.18336
Dataset Sources
The Reddit PushShift data dumps are… See the full description on the dataset page: https://huggingface.co/datasets/OsamaBsher/AITA-Reddit-Dataset.singapore-legal-ai-benchmark
Singapore Legal AI Benchmark
Public research release of 102 Singapore legal research questions, model
responses from 6 systems, and overlapping grades on five dimensions.
Headline metrics are overlapping binary flags, not a ranking and not a
partition of 100%.
Interactive explorer
Open the explorer →
— comparison table, category heatmap, per-question comparison, and every answer
with its sources and grades.
(Space page)
Overall (n = 612)… See the full description on the dataset page: https://huggingface.co/datasets/JonathanSu/singapore-legal-ai-benchmark.EC-Guide
This repo is only used for dataset viewer. Please download from here.
Amazon KDDCup 2024 Team ZJU-AI4H’s Solution and Dataset (Track 2 Top 2; Track 5 Top 5)
The Amazon KDD Cup’24 competition presents a unique challenge by focusing on the application of LLMs in E-commerce across multiple tasks. Our solution for addressing Tracks 2 and 5 involves a comprehensive pipeline encompassing dataset construction, instruction tuning, post-training quantization, and inference… See the full description on the dataset page: https://huggingface.co/datasets/AiMijie/EC-Guide.l4-gpu-llm-benchmark-leaderboard
🚀 Local LLM Serving & Quality Benchmark Leaderboard (NVIDIA L4 24GB)
An exhaustive, reproducible benchmark study measuring real-world serving performance (TTFT, TPOT, throughput, peak VRAM, energy consumption, and cost) alongside rigorous task quality gates (HumanEval+, MMLU-Pro, BFCL v4 tool calling, and RULER needle retrieval) for open-weight LLMs on a single NVIDIA L4 24GB GPU.
📊 Executive Summary & Key Takeaways
⚡ Best Throughput & Coding Workhorse:… See the full description on the dataset page: https://huggingface.co/datasets/mayank-dubey-ai/l4-gpu-llm-benchmark-leaderboard.GPTKB_v1.5This hosts the GPTKB v1.5 dataset. Visit https://gptkb.org to browse GPTKB and for further information.
Papers:
GPTKB methodology: https://arxiv.org/pdf/2411.04920
GPTKB v1.5: https://arxiv.org/pdf/2507.05740
Citations:
@InProceedings{GPTKB,
title={Enabling LLM Knowledge Analysis via Extensive Materialization},
author={Hu, Yujia and Nguyen, Tuan-Phong and Ghosh, Shrestha and Razniewski, Simon},
year={2025},
booktitle={ACL},
}
@article{GPTKB15,
title={GPTKB v1.5: A Massive… See the full description on the dataset page: https://huggingface.co/datasets/Knowledge-aware-AI/GPTKB_v1.5.prompts_under_512_tokens
Under 512 Tokens Prompts Dataset
Created by Aipresso LIMITED, London, UK
⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use
Dataset Overview
Specialized collection of short-form English prompts (under 512 tokens), perfect for training models with context length constraints or faster iteration cycles.
📊 Dataset Statistics
Metric
Value
Total Files
200
Rows Per File
10,000
Total Rows
2,000,000
Token Range
1 to 511 tokens… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/prompts_under_512_tokens.10k_rows_cleaned_prompts
10K Rows Cleaned Prompts Dataset
Created by Aipresso LIMITED, London, UK
⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use
You must provide attribution when using this data in publications, research, or commercial products.
Dataset Overview
A chunked collection of 2.7 million cleaned English prompts, organized into 200 files of 10,000 rows each for easy processing and distributed training of language models.
📊 Dataset Statistics
Metric… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/10k_rows_cleaned_prompts.CUREMED-BENCH
CUREMED-BENCH
CUREMED-BENCH is a multilingual medical reasoning benchmark dataset, designed for evaluating and fine-tuning models on medical tasks across diverse languages.
Data
set Summary
Languages: Spans 13 languages, including Amharic, Bengali, French, Hausa, Hindi, Japanese, Korean, Spanish, Swahili, Thai, Turkish, Vietnamese, and Yoruba.
Splits: Includes train, test, and validation splits, each with separate CSV files per language.
Usage: Intended for research in… See the full description on the dataset page: https://huggingface.co/datasets/Aikyam-Lab/CUREMED-BENCH.GPTKB_v1This is the GPTKB dataset from the ACL 2025 paper:
@InProceedings{GPTKB,
title={Enabling LLM Knowledge Analysis via Extensive Materialization},
author={Hu, Yujia and Nguyen, Tuan-Phong and Ghosh, Shrestha and Razniewski, Simon},
year={2025},
booktitle={ACL},
}
Preprint: https://arxiv.org/pdf/2411.04920
Web interface for browsing GPTKB: https://gptkb.org
galtea-red-teaming-clustered-data
Galtea Red Teaming: Non-Commercial Subset
This dataset contains a curated collection of adversarial prompts used for red teaming and LLM safety evaluation. All prompts come from datasets under non-commercial licenses and have been:
Deduplicated
Normalized into a consistent format
Automatically clustered based on semantic meaning
Each entry includes:
prompt: the adversarial instruction
source: the dataset of origin
cluster: a numeric cluster ID based on prompt behavior… See the full description on the dataset page: https://huggingface.co/datasets/Galtea-AI/galtea-red-teaming-clustered-data.shironaam
Dataset Card for Shironaam Corpus
Dataset Summary
Automatic headline generation systems have the potential to assist editors in finding interesting headlines to attract visitors or readers.
However, the performance of headline generation systems remains challenging due to the unavailability of sufficient parallel data for
low-resource languages like Bengali. We provide Shironaam, a large-scale news headline generation dataset of a low-resource language
i.e., Bengali… See the full description on the dataset page: https://huggingface.co/datasets/dialect-ai/shironaam.AI-Jailbreak-Prompts
Dataset Card for Dataset Name
Name
Jailbreak Prompts
Dataset Summary
Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[German]
AI-Jailbreak-Prompts
Dataset Card for Dataset Name
Name
Jailbreak Prompts
Dataset Summary
Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[German]
AI-Jailbreak-Prompts
Dataset Card for Dataset Name
Name
Jailbreak Prompts
Dataset Summary
Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[German]
agrillm-qa-eval-800
Dataset Card for agrillm-qa-eval-800
agrillm-qa-eval-800 is a high-quality evaluation dataset focused on agricultural knowledge and reasoning. The dataset was assembled by ai71 in partnership with leading organizations and partners across the agricultural sector such as CGIAR, ECHO, Digital Green, Embrapa, FAO, the World Bank, IFAD, the Gates Foundation, KALRO, KIADPAI, the Extension Foundation, and additional contributors across the agricultural domain.
It is intended as an open… See the full description on the dataset page: https://huggingface.co/datasets/AI71ai/agrillm-qa-eval-800.Wiki_Live_Challenge
Wiki Live Challenge Dataset
[English | 中文]
English
📖 Dataset Overview
This is the official dataset accompanying the Wiki Live Challenge benchmark. It contains Wikipedia Good Articles (GAs) as ground truth and research articles generated by leading deep research AI systems.
Wiki Live Challenge is the first live benchmark for evaluating Deep Research Agents (DRAs) on their ability to generate Wikipedia-quality articles. Unlike static benchmarks, Wiki Live… See the full description on the dataset page: https://huggingface.co/datasets/muset-ai/Wiki_Live_Challenge.airs-bench
AIRS-Bench: a Suite of Tasks for Frontier AI Research Science Agents
The AI Research Science Benchmark (AIRS-Bench) quantifies the autonomous research abilities of LLM agents in the area of machine learning. AIRS-Bench comprises 20 tasks from state-of-the-art machine learning papers spanning diverse domains: NLP, Code, Math, biochemical modelling, and time series forecasting.
Each task is specified by a ⟨problem, dataset, metric⟩ triplet and a SOTA value. The agent receives the… See the full description on the dataset page: https://huggingface.co/datasets/facebook/airs-bench.r20-portfolio-ai-perception
Portfolio Interference in LLM Brand Perception (R20 to R21)
Supersession note: This dataset originally backed R20 (2026ab, superseded). R21 (2026ac, DOI 10.5281/zenodo.19765401) supersedes both R8 (2026q) and R20. R21 merges R8 theory with R20 empirical (9,925 obs across 40 brands, 13 models, 7 traditions) into a single analytical-empirical paper. New citations should reference Zharnikov (2026ac).
Dataset DOI: 10.57967/hf/8380
Current Paper (R21): 10.5281/zenodo.19765401 --… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/r20-portfolio-ai-perception.
