Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OpenFormosa /barbet-long-context-sft Barbet long-context SFT Release a9fe3ba5b4c2869855e4e75118262797572d695c146eaa8fed7a180ec3544381 preserves 6633 active records. This is one joint assistant-only SFT dataset; no Barbet model training has been run. The skill-prefill migration has revised 1245 of 1254 records from its fixed base snapshot. Revisions replace their original records in the explicit shard lists above. Old bundles and releases remain available at their pinned commits. Additional records from other… See the full description on the dataset page: https://huggingface.co/datasets/OpenFormosa/barbet-long-context-sft.tabular1K<n<10K1 likes6k downloads2d agoHugging Face02SALT-NLP /hle-context-baseline-deeptabular10K<n<100K0 likes1.5k downloads3mo agoHugging Face03placeholderlabs /pretrain-academic-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 43,694,042,993 (43.7B) Trainable tokens 43,694,042,993 (43.7B) Documents 1,001,557 Shards 373 UTF-8 bytes 183,279,720,921 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-academic-mix-long-context.tabular1M<n<10M0 likes1.1k downloads28d agoHugging Face04Contextbench /Tracebench Tracebench This dataset contains agent trajectories (TerminalBench + SWE-bench) with two splits: full: 3316 trajectories (2670 terminal + 646 SWE-bench) verified: 1000 trajectories (489 SWE-bench + 511 terminal; terminal selected by step_count>=20, has incorrect steps, error-stage ratio threshold) Agents: mini-SWE-agent (1024), OpenHands (1242), Terminus2 (923), SWE-agent (127). Models: Anthropic/Claude-Sonnet-4, DeepSeek/DeepSeek-V3.2, Moonshot/Kimi-K2, OpenAI/GPT-5… See the full description on the dataset page: https://huggingface.co/datasets/Contextbench/Tracebench.tabular1K<n<10K1 likes1k downloads6mo agoHugging Face05SALT-NLP /hle-context-baseline-gpt55tabular10K<n<100K0 likes884 downloads3mo agoHugging Face06artefactory /ledger-long-context-KPI-QA LEDGER — Long-Context KPI Question Answering & Page Retrieval This dataset is part of the LEDGER (Long-context Evaluation of Documents for Grounded Extraction and Retrieval) benchmark. It supports two of the three LEDGER tasks: Page-level KPI retrieval — given a natural-language question about a financial KPI and the corresponding annual report, retrieve the relevant page(s). Each row includes TREC-style graded relevance judgments (qrels) over all candidate pages.… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/ledger-long-context-KPI-QA.tabularquestion-answering100K<n<1M16 likes640 downloads2mo agoHugging Face07placeholderlabs /pretrain-commits-v2-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 12,571,681,749 (12.6B) Trainable tokens 4,460,160,435 (4.5B) Documents 992,475 Shards 327 UTF-8 bytes 49,288,867,997 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-commits-v2-mix-long-context.tabular1M<n<10M0 likes589 downloads24d agoHugging Face08placeholderlabs /pretrain-repository-v2-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 130,158,375,824 (130.2B) Trainable tokens 130,158,375,824 (130.2B) Documents 2,578,578 Shards 1,168 UTF-8 bytes 535,260,344,241 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-repository-v2-mix-long-context.tabular1M<n<10M0 likes553 downloads23d agoHugging Face09placeholderlabs /pretrain-web-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 8,689,580,607 (8.7B) Trainable tokens 8,689,580,607 (8.7B) Documents 281,846 Shards 89 UTF-8 bytes 37,540,769,483 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-web-mix-long-context.tabular100K<n<1M0 likes549 downloads28d agoHugging Face10yzhuang /Agentic-Long-Context-Understanding-QA 📖 Agentic Long Context Understanding 📖 Self-Taught Agentic Long Context Understanding (Arxiv). AgenticLU refines complex, long-context queries through self-clarifications and contextual grounding, enabling robust long-document understanding in a single pass. Installation Requirements This codebase is largely based on OpenRLHF and Helmet, kudos to them. The requirements are the same pip install openrlhf pip install -r ./HELMET/requirements.txt… See the full description on the dataset page: https://huggingface.co/datasets/yzhuang/Agentic-Long-Context-Understanding-QA.tabularquestion-answering100K<n<1M23 likes514 downloads1y agoHugging Face11Interplay-LM-Reasoning /context On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models Charlie Zhang, Graham Neubig, Xiang Yue Carnegie Mellon University, Language Technologies Institute Does Reinforcement Learning Truly Extend Reasoning? This work explores the discrepancy in views on RL's effectiveness in extending language models' reasoning abilities. Some characterize RL as a capability refiner, while others see it as inducing new compositional skills. This challenge… See the full description on the dataset page: https://huggingface.co/datasets/Interplay-LM-Reasoning/context.tabularquestion-answering10M<n<100M2 likes509 downloads9mo agoHugging Face12placeholderlabs /pretrain-ultra-fineweb-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 1,367,358,024 (1.4B) Trainable tokens 1,367,358,024 (1.4B) Documents 48,077 Shards 73 UTF-8 bytes 6,386,740,105 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-ultra-fineweb-mix-long-context.tabular10K<n<100K1 likes505 downloads28d agoHugging Face13mesolitica /fineweb-filter-malaysian-context HuggingFaceFW/fineweb filter Malaysian context What is it? We filter the original 🍷 FineWeb dataset that consists more than 15T tokens on simple Malaysian keywords. Total tokens for the filtered dataset is 174102784199 tokens, 174B tokens. How we do it? We filter rows using {'malay', 'malaysia', 'melayu', 'bursa', 'ringgit'} keywords on r5.16xlarge EC2 instance for 7 days. We calculate total tokens using tiktoken.encoding_for_model("gpt2") on c7a.24xlarge EC2… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/fineweb-filter-malaysian-context.tabular10M<n<100M1 likes450 downloads2y agoHugging Face14evalitahf /word_in_contextDataset homepage: https://wic-ita.github.io/index.html tabulartext-classification1K<n<10K0 likes447 downloads2y agoHugging Face15arize-ai /movie_reviews_with_context_drift Dataset Card for reviews_with_drift Dataset Description Dataset Summary This dataset was crafted to be used in our tutorial [Link to the tutorial when ready]. It consists on a large Movie Review Dataset mixed with some reviews from a Hotel Review Dataset. The training/validation set are purely obtained from the Movie Review Dataset while the production set is mixed. Some other features have been added (age, gender, context) as well as a made up timestamp… See the full description on the dataset page: https://huggingface.co/datasets/arize-ai/movie_reviews_with_context_drift.tabulartext-classification10K<n<100K1 likes396 downloads4y agoHugging Face16jiosephlee /context-conditioned-molecule-transfer-v10.4.1-bbb-martins-mixed-continuous-intern BBB_Martins context-conditioned molecule transfer V10.4.1 This release preserves its direct panels and appends training-only, post-aggregate continuous assay-evidence transfer pairs. Query values remain hidden from prompts. Train rows: 214,362 Validation rows: 30,299 Test rows: 29,919 V10.4.1 uses only continuous non-L5 assay evidence and applies the shared center-0.6, temperature-0.1 sigmoid with half-slope probability tails. tabular100K<n<1M0 likes365 downloads24d agoHugging Face17placeholderlabs /pretrain-nemotron-math-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 1,446,296,439 (1.4B) Trainable tokens 1,446,296,439 (1.4B) Documents 42,379 Shards 23 UTF-8 bytes 4,965,563,314 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-nemotron-math-mix-long-context.tabular10K<n<100K0 likes358 downloads28d agoHugging Face18ymoslem /wmt-da-human-evaluation-long-context Dataset Summary Long-context / document-level dataset for Quality Estimation of Machine Translation. It is an augmented variant of the sentence-level WMT DA Human Evaluation dataset. In addition to individual sentences, it contains augmentations of 2, 4, 8, 16, and 32 sentences, among each language pair lp and domain. The raw column represents a weighted average of scores of augmented sentences using character lengths of src and mt as weights. The code used to apply the augmentation… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/wmt-da-human-evaluation-long-context.tabular1M<n<10M9 likes352 downloads2y agoHugging Face19LegionIntel /named_entity_recognition_document_contexttabular1M<n<10M9 likes335 downloads2y agoHugging Face20mattpidden /vla0-context-datasetThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/mattpidden/vla0-context-dataset.tabularrobotics100K<n<1M0 likes327 downloads2mo agoHugging Face21kothasuhas /dl_alchemy_seq9p6m_context1024tabularn<1K0 likes312 downloads1mo agoHugging Face22lilonghao /MM-ContextASR-Bench MM-ContextASR Bench Metadata and evaluation splits for Multimodal Conversational Context for LLM-Based ASR: Data Construction, Training, and Benchmark. Dataset summary Config Examples Audio Context Primary metric mm_contextasr 1,250 (250 current utterances × 5 histories) 1,439 WAV files included Controlled user-assistant dialogue entity Recall kespeech 19,212 Source ID only Same-speaker speech and transcript CER, SER, entity Recall cv_yue 3,525… See the full description on the dataset page: https://huggingface.co/datasets/lilonghao/MM-ContextASR-Bench.audioautomatic-speech-recognition10K<n<100K1 likes306 downloads24d agoHugging Face23minh21 /COVID-QA-unique-context-test-10-percent-validation-10-percent Dataset Card for "COVID-QA-unique-context-test-10-percent-validation-10-percent" More Information needed tabular1K<n<10K0 likes292 downloads3y agoHugging Face24shshwtsuthar /memory-representation-contextbench-artifacts Memory Representation ContextBench Artifacts Dataset Summary This repository contains processed artifacts for the paper "Memory as a Map: Prior-Trajectory Representations for Software Engineering Agents." The artifact supports reproduction and inspection of a controlled prior-context representation experiment over SWEContextBench prior-target pairs. The experiment renders each target under four prompt conditions: no prior context, stripped Claude Code transcript… See the full description on the dataset page: https://huggingface.co/datasets/shshwtsuthar/memory-representation-contextbench-artifacts.tabular1K<n<10K0 likes280 downloads4mo agoHugging Face25placeholderlabs /pretrain-encyclopedic-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 587,625,128 (587.6M) Trainable tokens 587,625,128 (587.6M) Documents 23,631 Shards 9 UTF-8 bytes 1,978,753,989 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-encyclopedic-mix-long-context.tabular10K<n<100K1 likes256 downloads28d agoHugging Face26Sterzhang /ContextProgress-Bench ContextProgress-Bench ContextProgress-Bench is the benchmark of ProgressCompass: Embodied Progress Reward Models Are Lost Without the Right Context. It tests context-dependent progress estimation: robot-manipulation episodes in which the current frame alone cannot tell how far the task has come, because progress depends on what happened earlier. 🌐 Project page · 📄 Paper (arXiv) · 💻 Code: coming soon Every task needs at least one of three forms of context: State Recall: a… See the full description on the dataset page: https://huggingface.co/datasets/Sterzhang/ContextProgress-Bench.tabularvideo-classificationn<1K0 likes254 downloads8d agoHugging Face27artefactory /ledger-long-context-multi-kpi the LEDGER Long-Context Multi-KPI extraction datasets and benchmarks. OCR'd annual reports with ground-truth KPI values for financial information extraction benchmarking. Dataset Description This dataset pairs OCR-extracted annual report text (from DeepSeek OCR) with structured KPI ground-truth values. It is designed for evaluating LLM-based financial information extraction, retrieval, and needle-in-a-haystack tasks. Configs Config Reports… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/ledger-long-context-multi-kpi.imagetable-question-answering1K<n<10K17 likes247 downloads3mo agoHugging Face28SALT-NLP /hle-context-baseline-groktabular10K<n<100K0 likes203 downloads3mo agoHugging Face29shredder-31 /contextualized-ST-Evidence Contextualized ST-Evidence A re-annotation of Salesforce/ST-Evidence-Instruct's gen_mask split. Same 19,902 entries, same objects, same frames, same temporal evidence. The only thing that changes is the spatial box on each frame. This is the video counterpart of shredder-31/contextualized-viscot, built with the same model, the same prompt design and the same union-with-the- original safety rule. Why ST-Evidence ships per-frame instance masks from GroundingDINO +… See the full description on the dataset page: https://huggingface.co/datasets/shredder-31/contextualized-ST-Evidence.imagevideo-text-to-text10K<n<100K0 likes203 downloads1mo agoHugging Face30mattpidden /vla0-context-dataset-v2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/mattpidden/vla0-context-dataset-v2.tabularrobotics100K<n<1M0 likes169 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.