datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sql-create-context
Overview
This dataset builds from WikiSQL and Spider.
There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from different DBMS and provides table names, column… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/sql-create-context.activating_contexts_16kContextASR-Bench
ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
Automatic Speech Recognition (ASR) has been extensively investigated, yet prior benchmarks have largely focused on assessing the acoustic robustness of ASR models, leaving evaluations of their linguistic capabilities relatively underexplored. This largely stems from the limited parameter sizes and training corpora of conventional ASR models, leaving them with insufficient world knowledge, which is crucial for… See the full description on the dataset page: https://huggingface.co/datasets/MrSupW/ContextASR-Bench.long-context-qa-curated-20
Dataset Card / 数据集卡
Dataset Description / 数据集简介
This public release contains 20 curated samples selected from a 10,000-record long-context QA collection. It targets retrieval over long documents, cross-section evidence synthesis, numerical reasoning, timeline reconstruction, and structured answer evaluation. The public subset contains 15 short-answer questions and 5 multiple-choice questions, balanced across Chinese and English.
本公开版本从 10,000 条长上下文问答数据中精选 20… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/long-context-qa-curated-20.spider-context-validation
Dataset Card for Spider Context Validation
Dataset Summary
Spider is a large-scale complex and cross-domain semantic parsing and text-to-SQL dataset annotated by 11 Yale students
The goal of the Spider challenge is to develop natural language interfaces to cross-domain databases.
This dataset was created to validate spider-fine-tuned LLMs with database context.
Yale Lily Spider Leaderboards
The leaderboard can be seen at https://yale-lily.github.io/spider… See the full description on the dataset page: https://huggingface.co/datasets/richardr1126/spider-context-validation.context
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
Charlie Zhang, Graham Neubig,
Xiang Yue
Carnegie Mellon University, Language Technologies Institute
Does Reinforcement Learning Truly Extend Reasoning?
This work explores the discrepancy in views on RL's effectiveness in extending language models' reasoning abilities. Some characterize RL as a capability refiner, while others see it as inducing new compositional skills. This challenge… See the full description on the dataset page: https://huggingface.co/datasets/Interplay-LM-Reasoning/context.ContextAwareword_in_contextDataset homepage:
https://wic-ita.github.io/index.html
activating_contexts_131k_layers_0_21activating_contexts_131k_layers_21_42dl_alchemy_seq9p6m_context1024MM-ContextASR-Bench
MM-ContextASR Bench
Metadata and evaluation splits for Multimodal Conversational Context for
LLM-Based ASR: Data Construction, Training, and Benchmark.
Dataset summary
Config
Examples
Audio
Context
Primary metric
mm_contextasr
1,250 (250 current utterances × 5 histories)
1,439 WAV files included
Controlled user-assistant dialogue
entity Recall
kespeech
19,212
Source ID only
Same-speaker speech and transcript
CER, SER, entity Recall
cv_yue
3,525… See the full description on the dataset page: https://huggingface.co/datasets/lilonghao/MM-ContextASR-Bench.Synthetic-Dataset-Childrens-Stories**Status: released 13-09-2026, repacked 14-09-2026.** The 14-09-2026 repack replaced 58
items after the acceptance gates were strengthened (prompt-instruction leaks, markdown
bullet lists and blockquotes); the other 29,942 are unchanged. Development stopped, pipeline released 17/09/26.
SAMPLE RELEASE: 30,000 synthetic children's short stories for early-reader language modelling.
Metrics
Value
genres
26
stories per genre
1.153-1.154K stories
total characters
38… See the full description on the dataset page: https://huggingface.co/datasets/ContextReq/Synthetic-Dataset-Childrens-Stories.ContextProgress-Bench
ContextProgress-Bench
ContextProgress-Bench is the benchmark of
ProgressCompass: Embodied Progress Reward Models Are Lost Without the Right Context.
It tests context-dependent progress estimation: robot-manipulation episodes in which the current
frame alone cannot tell how far the task has come, because progress depends on what happened earlier.
🌐 Project page ·
📄 Paper (arXiv) ·
💻 Code: coming soon
Every task needs at least one of three forms of context:
State Recall: a… See the full description on the dataset page: https://huggingface.co/datasets/Sterzhang/ContextProgress-Bench.contextualized-ST-Evidence
Contextualized ST-Evidence
A re-annotation of Salesforce/ST-Evidence-Instruct's gen_mask
split. Same 19,902 entries, same objects, same frames, same temporal evidence.
The only thing that changes is the spatial box on each frame.
This is the video counterpart of
shredder-31/contextualized-viscot,
built with the same model, the same prompt design and the same union-with-the-
original safety rule.
Why
ST-Evidence ships per-frame instance masks from GroundingDINO +… See the full description on the dataset page: https://huggingface.co/datasets/shredder-31/contextualized-ST-Evidence.Multi-turn_Long-context_Benchmark_for_LLMs
LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues
Arxiv: https://www.arxiv.org/abs/2507.13681
Huggingface: https://huggingface.co/papers/2507.13681
Introduction
LoopServe Multi-Turn Dialogue Benchmark is a comprehensive evaluation dataset comprising multiple diverse datasets designed to assess large language model performance in realistic conversational scenarios.
Unlike traditional benchmarks that place queries only at the end… See the full description on the dataset page: https://huggingface.co/datasets/TreeAILab/Multi-turn_Long-context_Benchmark_for_LLMs.scugnizz-v22-behavior-context
scugnizz-v22-behavior-context
Synthetic grounded-context and behavioral correction data for tool-loop discipline.
Format: Hermes/OpenAI-style messages plus tools.
2026-10-01-da-otherai-context-15-mix
Historical other-AI system framing ablation: context mixture
field
value
experiment
Historical other-AI system framing ablation: context mixture
date_generated
2026-10-01
constitution
constitutions/claude_distilled_09_principles/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ 57074c3220502538754c68bb8444e8189081c537
models
{"tokenizer": {"repo": "Qwen/Qwen3.6-27B", "revision":… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-10-01-da-otherai-context-15-mix.contextualized-viscot
Contextualized Visual-CoT
A re-annotation of deepcs233/Visual-CoT. Same 434,265 rows, same
files, same keys, same order. The only field that changes is bboxs.
Why
Visual-CoT's boxes are drawn tight around the literal answer span. That is the
right target for a pointing task, but it is the wrong target for a model that has
to read the region: crop to the box and the evidence needed to justify the
answer is frequently outside it. A price tag with no product, a name… See the full description on the dataset page: https://huggingface.co/datasets/shredder-31/contextualized-viscot.spider-natsql-context-validation
Dataset Card for Spider NatSQL Context Validation
Dataset Summary
Spider is a large-scale complex and cross-domain semantic parsing and text-to-SQL dataset annotated by 11 Yale students
The goal of the Spider challenge is to develop natural language interfaces to cross-domain databases.
This dataset was created to validate LLMs on the Spider dev dataset with database context using NatSQL.
NatSQL
NatSQL is an intermediate representation for SQL that simplifies… See the full description on the dataset page: https://huggingface.co/datasets/richardr1126/spider-natsql-context-validation.durable-vs-context-trials
Durable State vs Context — Repository-Scale Agent Trials
Machine-verified trial records from the paper "State, Not Tokens: Repository-Scale
Agent Reasoning Is Bound by State Architecture." Each record is one run of a
JavaScript→TypeScript migration of a real OSS repository (express, jsdom) under an
unforgeable oracle, graded by strict tsc --strict --noEmit, immutable test suites,
mandatory .js→.ts replacement, and a zero type-escape-hatch budget.
Code + reproduction harness:… See the full description on the dataset page: https://huggingface.co/datasets/CaryPalmer/durable-vs-context-trials.ContextTTS_dataset
ContextTTS Evaluation Dataset
This is the official evaluation dataset for the paper "[ContextTTS Eval: A Benchmark for Evaluating Long-Form
Contextual Expressive Text-to-Speech]". It is designed to evaluate the performance of multi-modal speech synthesis, specifically focusing on context-aware prosody and timbre consistency in Chinese conversations and audiobooks.
Dataset Summary
The dataset consists of high-quality Chinese audio-text pairs, organized into three distinct… See the full description on the dataset page: https://huggingface.co/datasets/rodenhhh/ContextTTS_dataset.repro-optimal-regret-for-policy-optimization-in-contextual-bandits-traces
Agent traces
Agent sessions published from a Trackio Logbook.
privacy-context-pairs
Privacy Context Pairs
Version 1.0.0 — synthetic, construction-labelled contextual-use probes.
Privacy Context Pairs contains 2,048 texts and 3,072 directed contrasts built from 256 language-specific cases in 128 bilingual scenario families. English and Brazilian Portuguese (pt-BR) are balanced. The dataset is standalone: no Miru installation, model, tokenizer or lens artifact is needed to read, rebuild or mechanically validate it.
These are not human privacy labels. The corpus… See the full description on the dataset page: https://huggingface.co/datasets/TeskesLab/privacy-context-pairs.ntp-mathlib-instruct-context
miniCTX: Neural Theorem Proving with (Long-)Contexts
Lean 4 tactic prediction examples extracted from Mathlib.
Examples contain:
prompt:
instruction, preceding file content, proof state
instruction, proof state
completion: tactic
The file content has been truncated to 1024 tokens.
Version
Generated using ntptoolkit's ntp-training-data and instruction_tuning.py.
It used the following config for ntp-training-data:
{
"repo":… See the full description on the dataset page: https://huggingface.co/datasets/l3lab/ntp-mathlib-instruct-context.ContextClarifyscenario-context-agent-cards
Scenario Context Agent Cards
这是一个完全合成的数据集,用于演示 Agent 如何根据用户情景和多源 Context 生成两类主动服务卡片:事件提醒卡和 POI 推荐卡。
配置
event_positive:应当输出事件提醒卡的案例;
event_negative:应当保持沉默的事件案例;
poi_positive:应当输出 POI 推荐卡的案例;
poi_negative:不应当推荐 POI 的案例。
当前每个配置包含 10 条案例,共 40 条。数据中的用户、地点、事件、ID、天气和位置均为虚构内容。
数据结构
每条记录包含:
case_id:案例 ID;
card_type:event 或 poi;
context:合成用户、环境、事件或候选 POI;
expected_decision:serve 或 stay_silent;
reference_output:公开基线生成的参考需求记录、决策、卡片和理由。
完整生成结果还会包含… See the full description on the dataset page: https://huggingface.co/datasets/LiuXinYan111/scenario-context-agent-cards.2026-10-01-da-otherai-context-synth
Historical other-AI system framing ablation: context synth
field
value
experiment
Historical other-AI system framing ablation: context synth
date_generated
2026-10-01
constitution
constitutions/claude_distilled_09_principles/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ 57074c3220502538754c68bb8444e8189081c537
models
{"tokenizer": {"repo": "Qwen/Qwen3.6-27B", "revision":… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-10-01-da-otherai-context-synth.genz-contextual-abusive-slang-v2-silver
Gen Z Contextual Abusive Slang Benchmark v2 Silver
This is a research-only silver candidate dataset for studying Gen Z / internet-slang abusive-language classification. It is designed to test whether models can distinguish slang from abuse, profanity from harassment, meme mockery from benign meme use, and identity mentions from identity attacks.
This is not a final gold benchmark. Labels are model-assisted silver labels generated with a three-pass LLM annotation pipeline and… See the full description on the dataset page: https://huggingface.co/datasets/AliceYin/genz-contextual-abusive-slang-v2-silver.sql-create-context-copy
Fork of b-mc2/sql-create-context
Overview
This dataset builds from WikiSQL and Spider.
There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/sql-create-context-copy.
