datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
context
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
Charlie Zhang, Graham Neubig,
Xiang Yue
Carnegie Mellon University, Language Technologies Institute
Does Reinforcement Learning Truly Extend Reasoning?
This work explores the discrepancy in views on RL's effectiveness in extending language models' reasoning abilities. Some characterize RL as a capability refiner, while others see it as inducing new compositional skills. This challenge… See the full description on the dataset page: https://huggingface.co/datasets/Interplay-LM-Reasoning/context.word_in_contextDataset homepage:
https://wic-ita.github.io/index.html
dl_alchemy_seq9p6m_context1024MM-ContextASR-Bench
MM-ContextASR Bench
Metadata and evaluation splits for Multimodal Conversational Context for
LLM-Based ASR: Data Construction, Training, and Benchmark.
Dataset summary
Config
Examples
Audio
Context
Primary metric
mm_contextasr
1,250 (250 current utterances × 5 histories)
1,439 WAV files included
Controlled user-assistant dialogue
entity Recall
kespeech
19,212
Source ID only
Same-speaker speech and transcript
CER, SER, entity Recall
cv_yue
3,525… See the full description on the dataset page: https://huggingface.co/datasets/lilonghao/MM-ContextASR-Bench.ContextProgress-Bench
ContextProgress-Bench
ContextProgress-Bench is the benchmark of
ProgressCompass: Embodied Progress Reward Models Are Lost Without the Right Context.
It tests context-dependent progress estimation: robot-manipulation episodes in which the current
frame alone cannot tell how far the task has come, because progress depends on what happened earlier.
🌐 Project page ·
📄 Paper (arXiv) ·
💻 Code: coming soon
Every task needs at least one of three forms of context:
State Recall: a… See the full description on the dataset page: https://huggingface.co/datasets/Sterzhang/ContextProgress-Bench.contextualized-ST-Evidence
Contextualized ST-Evidence
A re-annotation of Salesforce/ST-Evidence-Instruct's gen_mask
split. Same 19,902 entries, same objects, same frames, same temporal evidence.
The only thing that changes is the spatial box on each frame.
This is the video counterpart of
shredder-31/contextualized-viscot,
built with the same model, the same prompt design and the same union-with-the-
original safety rule.
Why
ST-Evidence ships per-frame instance masks from GroundingDINO +… See the full description on the dataset page: https://huggingface.co/datasets/shredder-31/contextualized-ST-Evidence.contextualized-viscot
Contextualized Visual-CoT
A re-annotation of deepcs233/Visual-CoT. Same 434,265 rows, same
files, same keys, same order. The only field that changes is bboxs.
Why
Visual-CoT's boxes are drawn tight around the literal answer span. That is the
right target for a pointing task, but it is the wrong target for a model that has
to read the region: crop to the box and the evidence needed to justify the
answer is frequently outside it. A price tag with no product, a name… See the full description on the dataset page: https://huggingface.co/datasets/shredder-31/contextualized-viscot.durable-vs-context-trials
Durable State vs Context — Repository-Scale Agent Trials
Machine-verified trial records from the paper "State, Not Tokens: Repository-Scale
Agent Reasoning Is Bound by State Architecture." Each record is one run of a
JavaScript→TypeScript migration of a real OSS repository (express, jsdom) under an
unforgeable oracle, graded by strict tsc --strict --noEmit, immutable test suites,
mandatory .js→.ts replacement, and a zero type-escape-hatch budget.
Code + reproduction harness:… See the full description on the dataset page: https://huggingface.co/datasets/CaryPalmer/durable-vs-context-trials.repro-optimal-regret-for-policy-optimization-in-contextual-bandits-traces
Agent traces
Agent sessions published from a Trackio Logbook.
genz-contextual-abusive-slang-v2-silver
Gen Z Contextual Abusive Slang Benchmark v2 Silver
This is a research-only silver candidate dataset for studying Gen Z / internet-slang abusive-language classification. It is designed to test whether models can distinguish slang from abuse, profanity from harassment, meme mockery from benign meme use, and identity mentions from identity attacks.
This is not a final gold benchmark. Labels are model-assisted silver labels generated with a three-pass LLM annotation pipeline and… See the full description on the dataset page: https://huggingface.co/datasets/AliceYin/genz-contextual-abusive-slang-v2-silver.in-context-grid-reasoning
In-Context Grid Reasoning (ICGR)
A small, fully synthetic benchmark for demonstration-conditioned rule induction:
each task shows 2–4 (input grid → output grid) support pairs that share one
hidden transformation, and the model must apply the same transformation to a
held-out query input.
It targets the same behaviour probed by recent in-context / latent-reasoning work
on ARC-AGI (e.g. BDH-CQ: In-Context Learning with Recurrent Latent Reasoning,
arXiv:2608.09888), but is… See the full description on the dataset page: https://huggingface.co/datasets/WhySoCodius/in-context-grid-reasoning.per-context-rb-l0-4096-no-eos-qwen3-1.7b-compression-bs32-32k-146102-rollouts
per_context_rb_l0_4096_no_eos_Qwen3-1.7B_compression_bs32_n16_32k_1epoch rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
adaption-agriintel-contextual-yield-reasoning-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-AgriIntel-Contextual-Yield-Reasoning-v1
A deeply evolved, instruction-tuned agricultural dataset engineered for precision agronomy modeling. This dataset bridges the gap between raw regional telemetry and generative AI reasoning by synthesizing historical crop performance with static soil chemistry metrics (N, P, K, pH) and historical macro-climatic atmospheric data (annual rainfall… See the full description on the dataset page: https://huggingface.co/datasets/Shravanthmvqwerty/adaption-agriintel-contextual-yield-reasoning-v1.long-context-llm-papers
Long-Context LLM Papers — FineSet
A research-paper dataset on Long-Context LLM Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-19.
It is not auto-updated. Research on Long-Context LLM Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓
Why this dataset
Quality-scored: quality_score… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/long-context-llm-papers.ascp-context-attribution
ASCP: Causal Context Attribution and Probe Benchmark
Released artifacts for The Laws of Context Allocation: Causal Measurement and
Closed-Loop Orchestration in Generative Search.
📄 Paper: https://arxiv.org/abs/2608.23252
💻 Code: https://github.com/PeiYangLiu/ascp
Retrieval-augmented generation is usually measured with relevance proxies —
BM25, query–document cosine, output overlap — that score how related a passage
looks, not whether the generator used it. This dataset ships… See the full description on the dataset page: https://huggingface.co/datasets/PeiyangLiu/ascp-context-attribution.context-10Brepro-gradmem-learning-to-write-context-into-memory-with-test-time-gradient-descent-traces
Agent traces
Agent sessions published from a Trackio Logbook.
contextual_faithfulness_logitqa_gemmacontextual_faithfulness_logitqa_qwencontext-repair-benchmark
ThoughtDAG Context Repair Benchmark
What happens after one wrong assumption enters a long LLM conversation?
This dataset turns context editing into a measurable intervention. Each synthetic case starts with a clean fact, introduces a false update, lets the error propagate through one to three downstream turns, and then asks the same final question under five graph conditions:
clean
polluted
source_prune
subgraph_prune
recompute_descendants
The central question is not only… See the full description on the dataset page: https://huggingface.co/datasets/thoughtdag/context-repair-benchmark.long_context_jailbreakingadaption-agriintel-contextual-yield-reasoning-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-AgriIntel-Contextual-Yield-Reasoning-v1
A deeply evolved, instruction-tuned agricultural dataset engineered for precision agronomy modeling. This dataset bridges the gap between raw regional telemetry and generative AI reasoning by synthesizing historical crop performance with static soil chemistry metrics (N, P, K, pH) and historical macro-climatic atmospheric data (annual rainfall… See the full description on the dataset page: https://huggingface.co/datasets/asadullahdogarr/adaption-agriintel-contextual-yield-reasoning-v1.per-context-rb-l0-0-qwen3-1.7b-compression-bs32-32k-146103-rollouts
per_context_rb_l0_0_Qwen3-1.7B_compression_bs32_n16_32k_1epoch rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
humanoid-context-aware-decision-dataset
Humanoid Context-Aware Decision Dataset
Dataset for training humanoid AI to make decisions
based on environmental and internal context.
Description
Contains contextual parameters and selected actions
to support intelligent reasoning.
File
context_aware_decision_dataset.json
License
MIT
kobest-query-context-stress-v2-extreme
KoBEST Query/Context Label-Preserving Stress v2 Extreme
This repository packages an extreme paired Korean boundary-stress dataset
built from skt/kobest_v1.
What it contains
Each row preserves:
the original gold label
the original answer options, or a synthetic label space for non-MCQ configs
and modifies only the natural input text fields to make the surface form more
tokenization-fragile while keeping:
identical non-space character sequence per stressed field… See the full description on the dataset page: https://huggingface.co/datasets/splo2t/kobest-query-context-stress-v2-extreme.kobest-query-context-stress-v3
KoBEST Query/Context Label-Preserving Stress v3
This repository packages a v3 paired Korean boundary-stress dataset built
from skt/kobest_v1.
What it contains
Each row preserves:
the original gold label
the original answer options, or a synthetic label space for non-MCQ configs
and modifies only the natural input text fields to make the surface form more
tokenization-fragile while keeping:
identical non-space character sequence per stressed field
identical Kiwi… See the full description on the dataset page: https://huggingface.co/datasets/splo2t/kobest-query-context-stress-v3.context-management-bench
context-management-bench
Live on the Hub: huggingface.co/datasets/shivam039-dev/context-management-bench
Realistic context-management scenarios for testing/benchmarking eviction strategies (drop-oldest, sliding-window, priority, summarization), pinned-message preservation, and tool-call/tool-result atomicity in multi-turn LLM conversations.
Dataset Summary
Every conversation in this dataset was generated deterministically and then run through the real… See the full description on the dataset page: https://huggingface.co/datasets/shivam039-dev/context-management-bench.deepseek_refine_contextcontext-as-a-service
CaaS Benchmark Corpus v1
A diverse collection of synthetic enterprise documents for benchmarking context extraction and RAG systems.
Dataset Description
This dataset contains 16 representative enterprise documents spanning multiple formats and domains, designed to evaluate:
Structure-aware indexing - Can the system identify high-value vs. low-value content?
Time decay relevance - Does the system properly weight recent vs. old information?
Pragmatic truth detection - Can… See the full description on the dataset page: https://huggingface.co/datasets/imran-siddique/context-as-a-service.contextual-hate-speech-conversations
Adversarial Content Moderation Evaluation Dataset
Dataset Summary
A dataset of 400 multi-turn conversations designed to evaluate LLM-based content
moderation supervisors against graduated adversarial escalation. Each adversarial
conversation consists of a neutral-to-harmful buildup arc culminating in an explicit
hate speech seed tweet. Benign conversations mirror the same structure using neutral
content, eliminating the format confounds present in prior single-turn… See the full description on the dataset page: https://huggingface.co/datasets/sherinechally/contextual-hate-speech-conversations.
