datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench.or-bench-toxic-all
OR-Bench: An Over-Refusal Benchmark for Large Language Models
This dataset constains highly toxic prompts, use with caution!!!
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench-toxic-all.llm-synthetic-survey-respondents-stochastic-parrots
AI Parrots: synthetic survey respondents from leading LLMs
This dataset holds 10,592 synthetic survey respondents generated by consumer AI platforms and frontier large language models. Every respondent answered the same 32-item questionnaire on ethics, political ideology and workplace experience in technology firms. The synthetic samples were benchmarked against a survey of Silicon Valley coders and developers with 400 complete human responses.
The data accompany Miklian… See the full description on the dataset page: https://huggingface.co/datasets/miklia/llm-synthetic-survey-respondents-stochastic-parrots.LLMs-are-not-calculators-v1.0
LLM Education Impact Simulator Dataset
Version: 1.0Generated: 2026-02-02Based on: Jackson, D. (2025). “LLMs are not calculators: Why educators should embrace AI (and fear it)”
Overview
This dataset contains synthetic observational data simulating how students interact with different AI tools (search engines, explicit-context LLMs, and agentic LLMs) while completing educational tasks.
The simulation is grounded in educational research, particularly Daniel Jackson's… See the full description on the dataset page: https://huggingface.co/datasets/DBbun/LLMs-are-not-calculators-v1.0.moral-tracing-in-LLMs
LLM Moral Evolution Study
A longitudinal dataset tracking moral reasoning patterns across 14 large language models from OpenAI and Anthropic, spanning multiple generations (2023–2025). The dataset measures how moral stances, ethical judgments, and value priorities shift across model updates using a 107-item probe instrument grounded in Moral Foundations Theory.
Models
OpenAI
Model
Release
GPT-3.5 Turbo
2023-11
GPT-4
2023-03
GPT-4o… See the full description on the dataset page: https://huggingface.co/datasets/mznaser/moral-tracing-in-LLMs.llms-with-matlabllms_epistemic_consistency
LLMs Epistemic Consistency Dataset
This dataset artifact contains the stimuli and prompt templates used for experiments on epistemic consistency and political-cue sensitivity in LLM evaluations.
Dataset URL: https://huggingface.co/datasets/drozado/llms_epistemic_consistency
Contents
croissant.json: root-level copy of the completed Croissant metadata for NeurIPS 2026 Evaluations and Datasets submission.
metadata/croissant.json: same Croissant metadata, kept with the… See the full description on the dataset page: https://huggingface.co/datasets/drozado/llms_epistemic_consistency.indian-medicinesI2EBenchfinetune-data-for-vision-llmsBM-Benchfinetune-data-for-vision-llms5finetune-data-for-vision-based-llms2LLMs_Bench_on_Mathfinetune-data-for-vision-based-llms3finetune-data-for-vision-based-llms5llms-are-blind-sampleAwesome-Scientific-Datasets-and-LLMsfinetune-data-for-vision-based-llms
