Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OpenFormosa /barbet-long-context-sft Barbet long-context SFT Release a9fe3ba5b4c2869855e4e75118262797572d695c146eaa8fed7a180ec3544381 preserves 6633 active records. This is one joint assistant-only SFT dataset; no Barbet model training has been run. The skill-prefill migration has revised 1245 of 1254 records from its fixed base snapshot. Revisions replace their original records in the explicit shard lists above. Old bundles and releases remain available at their pinned commits. Additional records from other… See the full description on the dataset page: https://huggingface.co/datasets/OpenFormosa/barbet-long-context-sft.tabular1K<n<10K1 likes6k downloads2d agoHugging Face02b-mc2 /sql-create-context Overview This dataset builds from WikiSQL and Spider. There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from different DBMS and provides table names, column… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/sql-create-context.texttext-generation10K<n<100K506 likes4.1k downloads3y agoHugging Face03yuyijiong /context_qa_sum_qwen3_synthetic Context-based QA and Summarization Synthetic Dataset Overview This dataset contains synthetic context-based question-answering (QA) and summarization data. The data was synthesized using: Source context: openbmb/Ultra-FineWeb Synthesis model: Qwen3-30B-A3B-Instruct-2507 Each context is obtained by taking the initial segment of raw pretraining text from Ultra-FineWeb, truncated to at most the corresponding number of tokens, while ensuring the truncation does not occur in… See the full description on the dataset page: https://huggingface.co/datasets/yuyijiong/context_qa_sum_qwen3_synthetic.texttext-generation10M<n<100M5 likes3.2k downloads6mo agoHugging Face04Proyag /paracrawl_context Dataset Card for ParaCrawl_Context This is a dataset for document-level machine translation introduced in the ACL 2024 paper Document-Level Machine Translation with Large-Scale Public Parallel Data. It is a dataset consisting of parallel sentence pairs from the ParaCrawl dataset along with corresponding preceding context extracted from the webpages the sentences were crawled from. Dataset Details Dataset Description This dataset adds document-level… See the full description on the dataset page: https://huggingface.co/datasets/Proyag/paracrawl_context.texttranslation100M<n<1B1 likes2.1k downloads1y agoHugging Face05SALT-NLP /hle-context-baseline-deeptabular10K<n<100K0 likes1.5k downloads3mo agoHugging Face06MrGonao /activating_contexts_16ktext100K<n<1M0 likes1.4k downloads2y agoHugging Face07MrSupW /ContextASR-Bench ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark Automatic Speech Recognition (ASR) has been extensively investigated, yet prior benchmarks have largely focused on assessing the acoustic robustness of ASR models, leaving evaluations of their linguistic capabilities relatively underexplored. This largely stems from the limited parameter sizes and training corpora of conventional ASR models, leaving them with insufficient world knowledge, which is crucial for… See the full description on the dataset page: https://huggingface.co/datasets/MrSupW/ContextASR-Bench.textautomatic-speech-recognition10K<n<100K38 likes1.4k downloads1y agoHugging Face08Contextbench /ContextBench ContextBench This repository provides: default: the full ContextBench table (single train split). contextbench_verified: a 500-instance subset (single split). Columns The dataset uses a unified schema across sources: instance_id: ContextBench instance id (e.g., SWE-Bench-Verified__python__...). original_inst_id: Original benchmark instance id (e.g., astropy__astropy-14539). source: One of Verified, Pro, Poly, Multi. language: Programming language. repo_url: Repository… See the full description on the dataset page: https://huggingface.co/datasets/Contextbench/ContextBench.text1K<n<10K6 likes1.3k downloads9mo agoHugging Face09placeholderlabs /pretrain-academic-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 43,694,042,993 (43.7B) Trainable tokens 43,694,042,993 (43.7B) Documents 1,001,557 Shards 373 UTF-8 bytes 183,279,720,921 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-academic-mix-long-context.tabular1M<n<10M0 likes1.1k downloads29d agoHugging Face10Contextbench /Tracebench Tracebench This dataset contains agent trajectories (TerminalBench + SWE-bench) with two splits: full: 3316 trajectories (2670 terminal + 646 SWE-bench) verified: 1000 trajectories (489 SWE-bench + 511 terminal; terminal selected by step_count>=20, has incorrect steps, error-stage ratio threshold) Agents: mini-SWE-agent (1024), OpenHands (1242), Terminus2 (923), SWE-agent (127). Models: Anthropic/Claude-Sonnet-4, DeepSeek/DeepSeek-V3.2, Moonshot/Kimi-K2, OpenAI/GPT-5… See the full description on the dataset page: https://huggingface.co/datasets/Contextbench/Tracebench.tabular1K<n<10K1 likes1k downloads6mo agoHugging Face11SALT-NLP /hle-context-baseline-gpt55tabular10K<n<100K0 likes884 downloads3mo agoHugging Face12ContextNews /world-bank-indicatorstext1M<n<10M0 likes842 downloads7mo agoHugging Face13ucla-contextual /contextual_testCheck out the paper. imagen<1K5 likes812 downloads3y agoHugging Face14LianeMarilin /long-context-qa-curated-20 Dataset Card / 数据集卡 Dataset Description / 数据集简介 This public release contains 20 curated samples selected from a 10,000-record long-context QA collection. It targets retrieval over long documents, cross-section evidence synthesis, numerical reasoning, timeline reconstruction, and structured answer evaluation. The public subset contains 15 short-answer questions and 5 multiple-choice questions, balanced across Chinese and English. 本公开版本从 10,000 条长上下文问答数据中精选 20… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/long-context-qa-curated-20.textquestion-answeringn<1K0 likes662 downloads1mo agoHugging Face15artefactory /ledger-long-context-KPI-QA LEDGER — Long-Context KPI Question Answering & Page Retrieval This dataset is part of the LEDGER (Long-context Evaluation of Documents for Grounded Extraction and Retrieval) benchmark. It supports two of the three LEDGER tasks: Page-level KPI retrieval — given a natural-language question about a financial KPI and the corresponding annual report, retrieve the relevant page(s). Each row includes TREC-style graded relevance judgments (qrels) over all candidate pages.… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/ledger-long-context-KPI-QA.tabularquestion-answering100K<n<1M16 likes640 downloads2mo agoHugging Face16placeholderlabs /pretrain-commits-v2-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 12,571,681,749 (12.6B) Trainable tokens 4,460,160,435 (4.5B) Documents 992,475 Shards 327 UTF-8 bytes 49,288,867,997 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-commits-v2-mix-long-context.tabular1M<n<10M0 likes589 downloads24d agoHugging Face17flwrlabs /ambient-acoustic-context Dataset Card for Ambient Acoustic Context The Ambient Acoustic Context dataset contains 1-second segments for activities that occur in a workplace setting. Each segment is associated with speaker_id. Dataset Details Using Amazin Mechanical Turk, crowd workers were asked to listen to 1-second segments and choose the right label. To ensure the quality of the annotations, audio segments that did not reach majority agreement among the turkers were excluded. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/flwrlabs/ambient-acoustic-context.audioaudio-classification10K<n<100K5 likes578 downloads2y agoHugging Face18richardr1126 /spider-context-validation Dataset Card for Spider Context Validation Dataset Summary Spider is a large-scale complex and cross-domain semantic parsing and text-to-SQL dataset annotated by 11 Yale students The goal of the Spider challenge is to develop natural language interfaces to cross-domain databases. This dataset was created to validate spider-fine-tuned LLMs with database context. Yale Lily Spider Leaderboards The leaderboard can be seen at https://yale-lily.github.io/spider… See the full description on the dataset page: https://huggingface.co/datasets/richardr1126/spider-context-validation.text1K<n<10K0 likes553 downloads3y agoHugging Face19placeholderlabs /pretrain-repository-v2-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 130,158,375,824 (130.2B) Trainable tokens 130,158,375,824 (130.2B) Documents 2,578,578 Shards 1,168 UTF-8 bytes 535,260,344,241 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-repository-v2-mix-long-context.tabular1M<n<10M0 likes553 downloads24d agoHugging Face20placeholderlabs /pretrain-web-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 8,689,580,607 (8.7B) Trainable tokens 8,689,580,607 (8.7B) Documents 281,846 Shards 89 UTF-8 bytes 37,540,769,483 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-web-mix-long-context.tabular100K<n<1M0 likes549 downloads29d agoHugging Face21yzhuang /Agentic-Long-Context-Understanding-QA 📖 Agentic Long Context Understanding 📖 Self-Taught Agentic Long Context Understanding (Arxiv). AgenticLU refines complex, long-context queries through self-clarifications and contextual grounding, enabling robust long-document understanding in a single pass. Installation Requirements This codebase is largely based on OpenRLHF and Helmet, kudos to them. The requirements are the same pip install openrlhf pip install -r ./HELMET/requirements.txt… See the full description on the dataset page: https://huggingface.co/datasets/yzhuang/Agentic-Long-Context-Understanding-QA.tabularquestion-answering100K<n<1M23 likes514 downloads1y agoHugging Face22Interplay-LM-Reasoning /context On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models Charlie Zhang, Graham Neubig, Xiang Yue Carnegie Mellon University, Language Technologies Institute Does Reinforcement Learning Truly Extend Reasoning? This work explores the discrepancy in views on RL's effectiveness in extending language models' reasoning abilities. Some characterize RL as a capability refiner, while others see it as inducing new compositional skills. This challenge… See the full description on the dataset page: https://huggingface.co/datasets/Interplay-LM-Reasoning/context.tabularquestion-answering10M<n<100M2 likes509 downloads9mo agoHugging Face23placeholderlabs /pretrain-ultra-fineweb-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 1,367,358,024 (1.4B) Trainable tokens 1,367,358,024 (1.4B) Documents 48,077 Shards 73 UTF-8 bytes 6,386,740,105 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-ultra-fineweb-mix-long-context.tabular10K<n<100K1 likes505 downloads29d agoHugging Face24Capx /ContextAwaretext100K<n<1M0 likes489 downloads2y agoHugging Face25mesolitica /fineweb-filter-malaysian-context HuggingFaceFW/fineweb filter Malaysian context What is it? We filter the original 🍷 FineWeb dataset that consists more than 15T tokens on simple Malaysian keywords. Total tokens for the filtered dataset is 174102784199 tokens, 174B tokens. How we do it? We filter rows using {'malay', 'malaysia', 'melayu', 'bursa', 'ringgit'} keywords on r5.16xlarge EC2 instance for 7 days. We calculate total tokens using tiktoken.encoding_for_model("gpt2") on c7a.24xlarge EC2… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/fineweb-filter-malaysian-context.tabular10M<n<100M1 likes450 downloads2y agoHugging Face26evalitahf /word_in_contextDataset homepage: https://wic-ita.github.io/index.html tabulartext-classification1K<n<10K0 likes447 downloads2y agoHugging Face27Contextbench /SWE-bench_Pro Dataset Summary SWE-Bench Pro is a challenging, enterprise-level dataset for testing agent ability on long-horizon software engineering tasks. Paper: https://static.scale.com/uploads/654197dc94d34f66c0f5184e/SWEAP_Eval_Scale%20(9).pdf See the related evaluation Github: https://github.com/scaleapi/SWE-bench_Pro-os Dataset Structure We follow SWE-Bench Verified (https://huggingface.co/datasets/SWE-bench/SWE-bench_Verified) in terms of dataset structure, with several… See the full description on the dataset page: https://huggingface.co/datasets/Contextbench/SWE-bench_Pro.textn<1K0 likes447 downloads10mo agoHugging Face28ostapeno /dolmino_wiki_rephrased_qa_with_context_concattext1M<n<10M0 likes417 downloads2y agoHugging Face29arize-ai /movie_reviews_with_context_drift Dataset Card for reviews_with_drift Dataset Description Dataset Summary This dataset was crafted to be used in our tutorial [Link to the tutorial when ready]. It consists on a large Movie Review Dataset mixed with some reviews from a Hotel Review Dataset. The training/validation set are purely obtained from the Movie Review Dataset while the production set is mixed. Some other features have been added (age, gender, context) as well as a made up timestamp… See the full description on the dataset page: https://huggingface.co/datasets/arize-ai/movie_reviews_with_context_drift.tabulartext-classification10K<n<100K1 likes396 downloads4y agoHugging Face30jiosephlee /context-conditioned-molecule-transfer-v10.4.1-bbb-martins-mixed-continuous-intern BBB_Martins context-conditioned molecule transfer V10.4.1 This release preserves its direct panels and appends training-only, post-aggregate continuous assay-evidence transfer pairs. Query values remain hidden from prompts. Train rows: 214,362 Validation rows: 30,299 Test rows: 29,919 V10.4.1 uses only continuous non-L5 assay evidence and applies the shared center-0.6, temperature-0.1 sigmoid with half-slope probability tails. tabular100K<n<1M0 likes365 downloads24d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.