datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
awesome-mllm-benchmarks-samples
Awesome MLLM Benchmarks – Sample Data
🌐 Interactive Dashboard ·
💻 GitHub
This dataset hosts the sample data (images, questions, answers, metadata) used by the Awesome MLLM Benchmarks interactive dashboard. It provides curated preview samples from 130+ multimodal LLM benchmarks across 20+ categories.
Overview
Stat
Count
Benchmarks with samples
125
Total subtasks
248
Total files (images + metadata)
~8,000
Categories
20+… See the full description on the dataset page: https://huggingface.co/datasets/lchen1019/awesome-mllm-benchmarks-samples.EDR_Telemetry_SampleThis dataset contains raw Endpoint Detection & Response (EDR) telemetry captured during controlled Deception.Pro malware sandbox operations on an enterprise Active Directory network. Unlike most malware sandboxes — which detonate samples for roughly 30 minutes — our operations run for hours or days per analysis, capturing the full arc of adversary behavior. The data represents a full-fidelity snapshot of system activity recorded while threat actors interacted with a live deception environment… See the full description on the dataset page: https://huggingface.co/datasets/DeceptionPro/EDR_Telemetry_Sample.function_calling_v3_SAMPLE
Trelis Function Calling Dataset - VERSION 3 - SAMPLE
This is a SAMPLE of the v3 dataset available for purchase here.
Features:
Allows models to be fine-tuned for function-calling.
The dataset is human generated and does not make use of Llama 2 or OpenAI!
The dataset includes 66 training rows, 19 validation rows and 5 test rows (for manual evaluation).
Based on eight functions: search_bing, search_arxiv, save_chat, read_json_file, list_files, get_current_weather, delete_file… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/function_calling_v3_SAMPLE.LongBioBench_Sample
A Controllable Examination for Long-Context Language Models
This is a sample dataset created using Qwen-7B tokenizer. For details, please check our paper
This dataset is created by running from our repo
bash scripts/submit_scripts/prepare_data.sh
wmt26-mist-sample
Update Log
22 June 2026 (latest) - we updated our data mix because some BELEBELE samples did not have the context. If you downloaded data before 22 June, please download the new version.
16 June 2026 - first version
Summary
The wmt26-mist-sample is a multilingual mix provided by the WMT26 MIST shared task organizers as a starting point for fine-tuning multilingual LLMs. It contains three types of tasks, to cover same-language and cross-lingual comprehension and… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/wmt26-mist-sample.legal-data-france-v1-sample
Legal Data France — V1
French Legal Dataset for RAG, NLP, Search & AI
Full V1: 9,432 structured judicial decisions • 44,318 Q&A • 44,318 instruction examples • 83,567 RAG chunks
Free public sample: 10 decisions • 15 Q&A • 15 instructions • 30 RAG chunks
Get the full V1
The complete Legal Data France — V1 release is available as a one-time purchase for €699.
Get the full V1 — €699
The full release includes:
9,432 structured judicial decisions
44,318… See the full description on the dataset page: https://huggingface.co/datasets/legaldatafrance/legal-data-france-v1-sample.Rail_Freight_Logistics_Company_Email_Archive_Sample
Ukrainian Rail-Freight Correspondence Corpus (Sample)
Real operational correspondence from a working freight forwarding business, and the
documents attached to it — consignment notes, service acts, invoices, wagon
manifests. Not scraped, not synthetic, and never published anywhere before.
This is a de-identified sample released for evaluation. It is drawn from a larger
private archive; see Full archive below.
Published by Akuma London · akumalondon.com
Why this… See the full description on the dataset page: https://huggingface.co/datasets/akumalondon/Rail_Freight_Logistics_Company_Email_Archive_Sample.deepseek-v4-pro-pi-reasoning-sample-traces
DeepSeek V4 Pro Pi Reasoning Sample Traces
This dataset contains a compact sample of successful DeepSeek V4 Pro teacher trajectories for Pi-style reasoning and tool-use workflows.
It includes selected pass-only traces from these task providers:
abcbench
aider
autocodebench
bfcl
swebench
swtbench
termigen
Format
Each row contains:
id: stable sample id
segments: Qwen-style template-free supervised segments
label=false segments are context only
label=true segments… See the full description on the dataset page: https://huggingface.co/datasets/bytkim/deepseek-v4-pro-pi-reasoning-sample-traces.groundtruth-hallucination-bench-sample
Groundtruth Data Hallucination Benchmark Sample
This public teaser contains 180 representative, source-backed examples from Groundtruth Data products.
Groundtruth Data builds verified evaluation, remediation, and held-out validation datasets for AI models using authoritative source data. The commercial workflow is:
Find where a model fails.
Prove the failure with a larger verified evaluation.
Provide targeted remediation/training data.
Validate improvement on untouched held-out… See the full description on the dataset page: https://huggingface.co/datasets/Groundtruth-Data/groundtruth-hallucination-bench-sample.BMGQ-MultiHop-Sample
🧩 BMGQ (Sample Release) – Bottom-up Multi-hop Question Generation Dataset
A Sampled Subset of BMGQ: Complex, Retrieval-Resistant, Multi-hop Reasoning Questions
👥 Authors
Bingsen Qiu, Zijian Liu, Xiao Liu, Bingjie Wang, Feier Zhang, Yixuan Qin, Chunyan Li, Haoshen Yang, Zeren Gao
📘 Dataset Summary
BMGQ is a dataset of complex, hard-to-search, multi-hop reasoning questions automatically generated using our proposed framework:
BMGQ: A Bottom-up… See the full description on the dataset page: https://huggingface.co/datasets/Fayer/BMGQ-MultiHop-Sample.ambig-iac-sample
Ambig-IaC Random 50
A random subset of 50 rows sampled without replacement from all 300 rows of
the default/train split of znyang/ambig-iac.
Source revision: 96429693e7164b024333e6d02f8c6b2d017e4ecb.
Sampling: Python random.Random(42).sample(range(300), 50).
Rows are stored in random draw order. All seven original columns, their types,
and original IDs are preserved. Compared with the pinned source, prompt_original
was edited in 10 rows and checks_rego in 2 rows. The other five… See the full description on the dataset page: https://huggingface.co/datasets/shihanlin/ambig-iac-sample.agentvidbench-sample
AgentVidBench (Sample): A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents
This repository is a representative sample of AgentVidBench, provided so reviewers can inspect data quality without downloading the full ~4GB+ corpus. The full dataset remains available at the link above.
Sample selection
The sample contains the first 10 questions (question_id 1–10) and the 9 unique videos they reference. IDs and filenames are preserved from the… See the full description on the dataset page: https://huggingface.co/datasets/agentvidbench/agentvidbench-sample.sample_clevrNeuroscience-Alignment-Corpus-Sample
🍀 The Clover Engine: Neuroscience Alignment Corpus (Evaluation Sample)
This repository contains a 42-pair Direct Preference Optimization (DPO) evaluation sample generated via the Clover Engine pipeline. It is designed to support the evaluation of evidence-grounded preference data derived from neuroscience-related scientific literature.
Evaluation Notice: This release is a limited evaluation sample demonstrating the pipeline's mechanics. It does not represent the full… See the full description on the dataset page: https://huggingface.co/datasets/Cefiyana/Neuroscience-Alignment-Corpus-Sample.virtual-patient-profiles-sample
VHS Patient Profiles — Nutrition Education Simulation
Dataset Summary
A curated collection of 8 richly structured synthetic patient personas designed for healthcare education simulations, with a primary focus on dietetics and nutrition counseling training. Each record represents a complete clinical case with demographic, clinical, psychosocial, and behavioral dimensions, along with structured guidance for educators and AI simulators.
Profiles are intended for use… See the full description on the dataset page: https://huggingface.co/datasets/joeljames270/virtual-patient-profiles-sample.synthetic-enterprise-ops-pack-sample
Solstice Synthetic Enterprise Operations Pack (Sample)
A multi-system graph dataset for agent evaluation and RAG benchmarking. This dataset simulates the interconnected operations of a modern technology company, linking sales activities, engineering workflows, IT support, and internal communications.
Built by Solstice AI Studio as a free sample of a larger commercial pack. 100% synthetic — no real company or employee data.
What's in the box
This dataset consists of 32… See the full description on the dataset page: https://huggingface.co/datasets/solsticestudioai/synthetic-enterprise-ops-pack-sample.verified-facts-sample-500
DeepInquiry Verified Facts (Sample-500)
A 500-fact sample of the DeepInquiry verified-facts corpus. Each fact is a single factual statement that has been checked against several sources that do not copy one another, and each record carries the sources behind it.
The full corpus holds roughly 3,000 approved facts and grows every day. It is available by API and as a licensed dataset (see Full corpus).
What is in the sample
500 facts: 482 Gold and 18 Silver.
11… See the full description on the dataset page: https://huggingface.co/datasets/MAC-deepinquiry/verified-facts-sample-500.sefd-archive-100k-analysis-sample-qwen3-20260524
SEFD Archive 100k Analysis Sample Qwen3 20260524
Retained artifacts for the completed archive-wide 100,000-filing Stanford EDGAR Filings Dataset (SEFD) analysis sample used in the arXiv paper update. The sample contains 2,971,490,909 final SEFD tokens, counted with the Qwen3-1.7B tokenizer.
This repository is a new versioned artifact and intentionally does not replace the earlier sfd-archive-100k-analysis-sample repository used for the original conference submission.
Included:… See the full description on the dataset page: https://huggingface.co/datasets/sfd-anonymous/sefd-archive-100k-analysis-sample-qwen3-20260524.kc-sft-sample
KC Engineering Cost SFT Sample / 工程知识卡 SFT 样例包
English · 中文
A free bilingual (Chinese–English) supervised fine-tuning sample pack for construction cost engineering (工程造价), distilled from KC — Engineering Cost Knowledge Cards, a curated knowledge base covering Chinese pricing benchmarks (信息价), unit-rate composition (综合单价), final accounts (竣工结算), BOQ valuation under GB 50500, local consumption-quota standards, and construction-law payment rules.
Full corpus (910+ cards, growing… See the full description on the dataset page: https://huggingface.co/datasets/RavAngel/kc-sft-sample.agentvidbench-sample
AgentVidBench (Sample): A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents
This repository is a representative sample of AgentVidBench, provided so reviewers can inspect data quality without downloading the full ~4GB+ corpus. The full dataset remains available at the link above.
Sample selection
The sample contains the first 10 questions (question_id 1–10) and the 9 unique videos they reference. IDs and filenames are preserved from the… See the full description on the dataset page: https://huggingface.co/datasets/KamiKrafton/agentvidbench-sample.qa-ncd-arabic-vol1-sample
QA NCD Arabic Vol.1 — Free Sample (150 rows)
Free sample of IndoHealth-NLP QA NCD Arabic Vol.1, a medical instruction-tuning dataset for Non-Communicable Diseases in Arabic.
Full dataset (3,480 pairs, $39): https://3929431511879.gumroad.com/l/QANCDArabicVol1
What's inside
150 randomly sampled rows from the full 3,480-row master dataset
Format: JSONL — fields: system, instruction, output, metadata (PMID)
Instructions in White Arabic (Educated Spoken Arabic)… See the full description on the dataset page: https://huggingface.co/datasets/IndoHealth-NLP/qa-ncd-arabic-vol1-sample.sample-fusion-intelligence-traces
Sample Fusion Intelligence Traces
Structured AI reasoning traces from dFusion's Fusion Intelligence system. Each record captures a complete agentic workflow: a real user query on a domain-specific topic, the full message chain including system prompts, tool calls, search results, intermediate reasoning steps, and a final synthesized answer — along with human feedback.
These are not synthetic benchmarks. They are traces from real queries submitted by real users on live financial… See the full description on the dataset page: https://huggingface.co/datasets/dFusionAILabs/sample-fusion-intelligence-traces.math500-olmo-3-7b-instruct-temp0.9-samples99-logprobs
OLMo-3-7B-Instruct self-consistency generations with logprobs on MATH500
This dataset contains 99 self-consistency generations per question for the
MATH500 benchmark, produced with allenai/OLMo-3-7B-Instruct at temperature
0.9, together with token-level log probabilities for each completion.
The file is intended for post-hoc analysis, self-consistency curves, adaptive
stopping, and related aggregation methods.
Source
Base benchmark: HuggingFaceH4/MATH-500
Model:… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/math500-olmo-3-7b-instruct-temp0.9-samples99-logprobs.hotpotqa-dev-olmo-3-7b-instruct-temp0.9-samples99-logprobs
OLMo-3-7B-Instruct self-consistency generations with logprobs on HotpotQA
This dataset contains 99 self-consistency generations per question for the
HotpotQA validation split, produced with allenai/OLMo-3-7B-Instruct at
temperature 0.9, together with token-level log probabilities for each
completion.
The file is intended for post-hoc analysis, self-consistency curves, adaptive
stopping, and related aggregation methods.
Source
Base benchmark: hotpotqa/hotpot_qa… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/hotpotqa-dev-olmo-3-7b-instruct-temp0.9-samples99-logprobs.korean-legal-instruction-sample
Korean Legal Instruction Dataset (한국어 법률 지시학습 데이터셋)
데이터셋 개요
이 데이터셋은 대한민국 법률 도메인에 특화된 sLLM 지시학습(Instruction Tuning)용 데이터셋입니다.
AIHub에서 제공하는 16종의 법률 관련 데이터를 통합하여 현대 LLM 지시학습 포맷으로 가공하였습니다.
주요 특징
총 데이터 수: 약 233,000건
언어: 한국어
포맷: ChatML/Alpaca 호환 대화 형식
도메인: 법률 (민사, 형사, 행정, 지식재산권, 계약 등)
데이터 구조
각 데이터 샘플은 다음과 같은 구조를 가집니다:
{
"id": "고유 식별자",
"category": "카테고리명",
"source": "원본 데이터 출처",
"system": "시스템 프롬프트",
"instruction": "사용자 질문/지시",
"output": "AI… See the full description on the dataset page: https://huggingface.co/datasets/neuralfoundry-coder/korean-legal-instruction-sample.arabic-palestinian-levantine-sample
4FACTORS — Palestinian Levantine Conversational Sample
50 native-written question–answer pairs in spoken Palestinian Levantine Arabic, each with an English gloss. This is a public demonstration sample from 4FACTORS, a producer of native, human-verified Arabic training data.
What this is
Real conversational exchanges — the kind of thing people actually say in shops, clinics, taxis, and at home — written from scratch by a first-language Palestinian speaker. Every… See the full description on the dataset page: https://huggingface.co/datasets/4factors/arabic-palestinian-levantine-sample.dynamics-reasoning-traces-sample
DYNAMICS-8 Behavioural Reasoning Traces
Personality-conditioned chain-of-thought reasoning data for LLM alignment and persona fine-tuning.
What This Dataset Contains
Each record is a first-person behavioural response from a synthetic persona with a validated 8-dimension personality profile (DYNAMICS-8), accompanied by a structured reasoning trace showing which personality dimensions drove the decision.
This is not survey data. It is not statistical synthetic data. Each… See the full description on the dataset page: https://huggingface.co/datasets/Kronaxis/dynamics-reasoning-traces-sample.Global-MMLU-Lite-sample
Global-MMLU-Lite — Sample Subset
A small, fixed-size subset of CohereLabs/Global-MMLU-Lite intended for fast smoke-testing of multilingual MMLU evaluation pipelines.
What this is
40 examples per language, 15 languages, 600 examples total.
Strictly non-overlapping windows across languages: language i takes rows [i*40, i*40+40) from the per-language pool of test followed by dev (test=400, dev=200, pool=600). Concretely, languages 0–9 fall entirely within the original test… See the full description on the dataset page: https://huggingface.co/datasets/cyankiwi/Global-MMLU-Lite-sample.Insurance-ChatBot-TestBench-Sample
Insurance ChatBot TestBench Dataset (Sample)
Dataset Description:
The dataset presented here includes 80 example prompts from the Insurance ChatBot TestBench, a specialized test set developed to evaluate the performance of generative AI chatbots in the insurance industry. These prompts are used in the analysis described in the blog post "Gen AI Chatbots in the Insurance Industry: Are they Trustworthy?". The test bench assesses chatbot performance across three critical dimensions:… See the full description on the dataset page: https://huggingface.co/datasets/rhesis/Insurance-ChatBot-TestBench-Sample.Circuit-Analysis-Reasoning-Sample
⚡ EngineeringWays Data Lab: Circuit Analysis Reasoning Dataset (Free Sample)
This is a free 50-item sample of the EngineeringWays Circuit Analysis Reasoning Dataset. It is designed specifically for fine-tuning Large Language Models (LLMs) in advanced STEM problem-solving, featuring strict Chain-of-Thought (CoT) reasoning.
Want the complete, deduplicated 592-item master dataset? 👉 Get the LoRA-Ready Master File on Payhip
🚀 Dataset Overview
Most math and physics… See the full description on the dataset page: https://huggingface.co/datasets/EngineeringWays/Circuit-Analysis-Reasoning-Sample.
