Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Interplay-LM-Reasoning /context On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models Charlie Zhang, Graham Neubig, Xiang Yue Carnegie Mellon University, Language Technologies Institute Does Reinforcement Learning Truly Extend Reasoning? This work explores the discrepancy in views on RL's effectiveness in extending language models' reasoning abilities. Some characterize RL as a capability refiner, while others see it as inducing new compositional skills. This challenge… See the full description on the dataset page: https://huggingface.co/datasets/Interplay-LM-Reasoning/context.tabularquestion-answering10M<n<100M2 likes509 downloads9mo agoHugging Face02evalitahf /word_in_contextDataset homepage: https://wic-ita.github.io/index.html tabulartext-classification1K<n<10K0 likes447 downloads2y agoHugging Face03kothasuhas /dl_alchemy_seq9p6m_context1024tabularn<1K0 likes312 downloads1mo agoHugging Face04lilonghao /MM-ContextASR-Bench MM-ContextASR Bench Metadata and evaluation splits for Multimodal Conversational Context for LLM-Based ASR: Data Construction, Training, and Benchmark. Dataset summary Config Examples Audio Context Primary metric mm_contextasr 1,250 (250 current utterances × 5 histories) 1,439 WAV files included Controlled user-assistant dialogue entity Recall kespeech 19,212 Source ID only Same-speaker speech and transcript CER, SER, entity Recall cv_yue 3,525… See the full description on the dataset page: https://huggingface.co/datasets/lilonghao/MM-ContextASR-Bench.audioautomatic-speech-recognition10K<n<100K1 likes306 downloads24d agoHugging Face05Sterzhang /ContextProgress-Bench ContextProgress-Bench ContextProgress-Bench is the benchmark of ProgressCompass: Embodied Progress Reward Models Are Lost Without the Right Context. It tests context-dependent progress estimation: robot-manipulation episodes in which the current frame alone cannot tell how far the task has come, because progress depends on what happened earlier. 🌐 Project page · 📄 Paper (arXiv) · 💻 Code: coming soon Every task needs at least one of three forms of context: State Recall: a… See the full description on the dataset page: https://huggingface.co/datasets/Sterzhang/ContextProgress-Bench.tabularvideo-classificationn<1K0 likes254 downloads8d agoHugging Face06shredder-31 /contextualized-ST-Evidence Contextualized ST-Evidence A re-annotation of Salesforce/ST-Evidence-Instruct's gen_mask split. Same 19,902 entries, same objects, same frames, same temporal evidence. The only thing that changes is the spatial box on each frame. This is the video counterpart of shredder-31/contextualized-viscot, built with the same model, the same prompt design and the same union-with-the- original safety rule. Why ST-Evidence ships per-frame instance masks from GroundingDINO +… See the full description on the dataset page: https://huggingface.co/datasets/shredder-31/contextualized-ST-Evidence.imagevideo-text-to-text10K<n<100K0 likes203 downloads1mo agoHugging Face07shredder-31 /contextualized-viscot Contextualized Visual-CoT A re-annotation of deepcs233/Visual-CoT. Same 434,265 rows, same files, same keys, same order. The only field that changes is bboxs. Why Visual-CoT's boxes are drawn tight around the literal answer span. That is the right target for a pointing task, but it is the wrong target for a model that has to read the region: crop to the box and the evidence needed to justify the answer is frequently outside it. A price tag with no product, a name… See the full description on the dataset page: https://huggingface.co/datasets/shredder-31/contextualized-viscot.imagevisual-question-answering100K<n<1M0 likes139 downloads2mo agoHugging Face08CaryPalmer /durable-vs-context-trials Durable State vs Context — Repository-Scale Agent Trials Machine-verified trial records from the paper "State, Not Tokens: Repository-Scale Agent Reasoning Is Bound by State Architecture." Each record is one run of a JavaScript→TypeScript migration of a real OSS repository (express, jsdom) under an unforgeable oracle, graded by strict tsc --strict --noEmit, immutable test suites, mandatory .js→.ts replacement, and a zero type-escape-hatch budget. Code + reproduction harness:… See the full description on the dataset page: https://huggingface.co/datasets/CaryPalmer/durable-vs-context-trials.tabularothern<1K0 likes114 downloads4mo agoHugging Face09tomyimkc /repro-optimal-regret-for-policy-optimization-in-contextual-bandits-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K3 likes106 downloads3mo agoHugging Face10AliceYin /genz-contextual-abusive-slang-v2-silver Gen Z Contextual Abusive Slang Benchmark v2 Silver This is a research-only silver candidate dataset for studying Gen Z / internet-slang abusive-language classification. It is designed to test whether models can distinguish slang from abuse, profanity from harassment, meme mockery from benign meme use, and identity mentions from identity attacks. This is not a final gold benchmark. Labels are model-assisted silver labels generated with a three-pass LLM annotation pipeline and… See the full description on the dataset page: https://huggingface.co/datasets/AliceYin/genz-contextual-abusive-slang-v2-silver.tabulartext-classification10K<n<100K1 likes81 downloads4mo agoHugging Face11WhySoCodius /in-context-grid-reasoning In-Context Grid Reasoning (ICGR) A small, fully synthetic benchmark for demonstration-conditioned rule induction: each task shows 2–4 (input grid → output grid) support pairs that share one hidden transformation, and the model must apply the same transformation to a held-out query input. It targets the same behaviour probed by recent in-context / latent-reasoning work on ARC-AGI (e.g. BDH-CQ: In-Context Learning with Recurrent Latent Reasoning, arXiv:2608.09888), but is… See the full description on the dataset page: https://huggingface.co/datasets/WhySoCodius/in-context-grid-reasoning.tabulartext-generation1K<n<10K1 likes70 downloads1mo agoHugging Face12hi-todayis-jh /per-context-rb-l0-4096-no-eos-qwen3-1.7b-compression-bs32-32k-146102-rollouts per_context_rb_l0_4096_no_eos_Qwen3-1.7B_compression_bs32_n16_32k_1epoch rollouts This dataset contains one compressed JSONL shard for every completed training step. The step and rollout_index columns uniquely locate a rollout within this training run. Run metadata and per-step row counts are recorded in rollout_manifest.json. tabular10K<n<100K0 likes69 downloads15d agoHugging Face13Shravanthmvqwerty /adaption-agriintel-contextual-yield-reasoning-v1 This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-AgriIntel-Contextual-Yield-Reasoning-v1 A deeply evolved, instruction-tuned agricultural dataset engineered for precision agronomy modeling. This dataset bridges the gap between raw regional telemetry and generative AI reasoning by synthesizing historical crop performance with static soil chemistry metrics (N, P, K, pH) and historical macro-climatic atmospheric data (annual rainfall… See the full description on the dataset page: https://huggingface.co/datasets/Shravanthmvqwerty/adaption-agriintel-contextual-yield-reasoning-v1.tabular10K<n<100K0 likes61 downloads29d agoHugging Face14fineset-io /long-context-llm-papers Long-Context LLM Papers — FineSet A research-paper dataset on Long-Context LLM Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. 📸 This is a dated snapshot — generated 2026-06-19. It is not auto-updated. Research on Long-Context LLM Papers moves fast — new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓ Why this dataset Quality-scored: quality_score… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/long-context-llm-papers.tabulartext-classificationn<1K0 likes52 downloads4mo agoHugging Face15PeiyangLiu /ascp-context-attribution ASCP: Causal Context Attribution and Probe Benchmark Released artifacts for The Laws of Context Allocation: Causal Measurement and Closed-Loop Orchestration in Generative Search. 📄 Paper: https://arxiv.org/abs/2608.23252 💻 Code: https://github.com/PeiYangLiu/ascp Retrieval-augmented generation is usually measured with relevance proxies — BM25, query–document cosine, output overlap — that score how related a passage looks, not whether the generator used it. This dataset ships… See the full description on the dataset page: https://huggingface.co/datasets/PeiyangLiu/ascp-context-attribution.tabularquestion-answering10K<n<100K0 likes47 downloads2mo agoHugging Face16goodevening /context-10Btabular10M<n<100M0 likes39 downloads1y agoHugging Face17SabaPivot /repro-gradmem-learning-to-write-context-into-memory-with-test-time-gradient-descent-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes38 downloads2mo agoHugging Face18jpd459 /contextual_faithfulness_logitqa_gemmatabular10K<n<100K0 likes30 downloads7mo agoHugging Face19jpd459 /contextual_faithfulness_logitqa_qwentabular10K<n<100K0 likes30 downloads7mo agoHugging Face20thoughtdag /context-repair-benchmark ThoughtDAG Context Repair Benchmark What happens after one wrong assumption enters a long LLM conversation? This dataset turns context editing into a measurable intervention. Each synthetic case starts with a clean fact, introduces a false update, lets the error propagate through one to three downstream turns, and then asks the same final question under five graph conditions: clean polluted source_prune subgraph_prune recompute_descendants The central question is not only… See the full description on the dataset page: https://huggingface.co/datasets/thoughtdag/context-repair-benchmark.tabular1K<n<10K0 likes25 downloads2mo agoHugging Face21alignmentforever /long_context_jailbreakingtabularn<1K1 likes22 downloads1y agoHugging Face22asadullahdogarr /adaption-agriintel-contextual-yield-reasoning-v1 This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-AgriIntel-Contextual-Yield-Reasoning-v1 A deeply evolved, instruction-tuned agricultural dataset engineered for precision agronomy modeling. This dataset bridges the gap between raw regional telemetry and generative AI reasoning by synthesizing historical crop performance with static soil chemistry metrics (N, P, K, pH) and historical macro-climatic atmospheric data (annual rainfall… See the full description on the dataset page: https://huggingface.co/datasets/asadullahdogarr/adaption-agriintel-contextual-yield-reasoning-v1.tabular10K<n<100K0 likes22 downloads2mo agoHugging Face23hi-todayis-jh /per-context-rb-l0-0-qwen3-1.7b-compression-bs32-32k-146103-rollouts per_context_rb_l0_0_Qwen3-1.7B_compression_bs32_n16_32k_1epoch rollouts This dataset contains one compressed JSONL shard for every completed training step. The step and rollout_index columns uniquely locate a rollout within this training run. Run metadata and per-step row counts are recorded in rollout_manifest.json. tabular10K<n<100K0 likes20 downloads18d agoHugging Face24AroticMerch /humanoid-context-aware-decision-dataset Humanoid Context-Aware Decision Dataset Dataset for training humanoid AI to make decisions based on environmental and internal context. Description Contains contextual parameters and selected actions to support intelligent reasoning. File context_aware_decision_dataset.json License MIT tabularn<1K0 likes19 downloads8mo agoHugging Face25splo2t /kobest-query-context-stress-v2-extreme KoBEST Query/Context Label-Preserving Stress v2 Extreme This repository packages an extreme paired Korean boundary-stress dataset built from skt/kobest_v1. What it contains Each row preserves: the original gold label the original answer options, or a synthetic label space for non-MCQ configs and modifies only the natural input text fields to make the surface form more tokenization-fragile while keeping: identical non-space character sequence per stressed field… See the full description on the dataset page: https://huggingface.co/datasets/splo2t/kobest-query-context-stress-v2-extreme.tabularmultiple-choicen<1K0 likes18 downloads5mo agoHugging Face26splo2t /kobest-query-context-stress-v3 KoBEST Query/Context Label-Preserving Stress v3 This repository packages a v3 paired Korean boundary-stress dataset built from skt/kobest_v1. What it contains Each row preserves: the original gold label the original answer options, or a synthetic label space for non-MCQ configs and modifies only the natural input text fields to make the surface form more tokenization-fragile while keeping: identical non-space character sequence per stressed field identical Kiwi… See the full description on the dataset page: https://huggingface.co/datasets/splo2t/kobest-query-context-stress-v3.tabularmultiple-choicen<1K0 likes17 downloads5mo agoHugging Face27shivam039-dev /context-management-bench context-management-bench Live on the Hub: huggingface.co/datasets/shivam039-dev/context-management-bench Realistic context-management scenarios for testing/benchmarking eviction strategies (drop-oldest, sliding-window, priority, summarization), pinned-message preservation, and tool-call/tool-result atomicity in multi-turn LLM conversations. Dataset Summary Every conversation in this dataset was generated deterministically and then run through the real… See the full description on the dataset page: https://huggingface.co/datasets/shivam039-dev/context-management-bench.tabularn<1K0 likes17 downloads1mo agoHugging Face28zhaospei /deepseek_refine_contexttabular1K<n<10K0 likes15 downloads2y agoHugging Face29imran-siddique /context-as-a-service CaaS Benchmark Corpus v1 A diverse collection of synthetic enterprise documents for benchmarking context extraction and RAG systems. Dataset Description This dataset contains 16 representative enterprise documents spanning multiple formats and domains, designed to evaluate: Structure-aware indexing - Can the system identify high-value vs. low-value content? Time decay relevance - Does the system properly weight recent vs. old information? Pragmatic truth detection - Can… See the full description on the dataset page: https://huggingface.co/datasets/imran-siddique/context-as-a-service.tabulartext-retrievaln<1K0 likes15 downloads9mo agoHugging Face30sherinechally /contextual-hate-speech-conversations Adversarial Content Moderation Evaluation Dataset Dataset Summary A dataset of 400 multi-turn conversations designed to evaluate LLM-based content moderation supervisors against graduated adversarial escalation. Each adversarial conversation consists of a neutral-to-harmful buildup arc culminating in an explicit hate speech seed tweet. Benign conversations mirror the same structure using neutral content, eliminating the format confounds present in prior single-turn… See the full description on the dataset page: https://huggingface.co/datasets/sherinechally/contextual-hate-speech-conversations.tabulartext-classification1K<n<10K0 likes15 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.