Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01context-course /imagesimagen<1K5 likes27k downloads5mo agoHugging Face02context-course /certificatesimagen<1K30 likes7.3k downloads39m agoHugging Face03b-mc2 /sql-create-context Overview This dataset builds from WikiSQL and Spider. There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from different DBMS and provides table names, column… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/sql-create-context.texttext-generation10K<n<100K506 likes4.3k downloads3y agoHugging Face04OpenFormosa /barbet-long-context-sft Barbet long-context SFT Release b005b2be0e02812d657b2c61a3cb9b19fb7adbe3f729448cfa442780f26d01e6 preserves 4623 active records. This is one joint assistant-only SFT dataset; no Barbet model training has been run. The skill-prefill migration has revised 1245 of 1254 records from its fixed base snapshot. Revisions replace their original records in the explicit shard lists above. Old bundles and releases remain available at their pinned commits. Additional records from other… See the full description on the dataset page: https://huggingface.co/datasets/OpenFormosa/barbet-long-context-sft.1 likes4.3k downloads3m agoHugging Face05yuyijiong /context_qa_sum_qwen3_synthetic Context-based QA and Summarization Synthetic Dataset Overview This dataset contains synthetic context-based question-answering (QA) and summarization data. The data was synthesized using: Source context: openbmb/Ultra-FineWeb Synthesis model: Qwen3-30B-A3B-Instruct-2507 Each context is obtained by taking the initial segment of raw pretraining text from Ultra-FineWeb, truncated to at most the corresponding number of tokens, while ensuring the truncation does not occur in… See the full description on the dataset page: https://huggingface.co/datasets/yuyijiong/context_qa_sum_qwen3_synthetic.texttext-generation10M<n<100M5 likes3.3k downloads6mo agoHugging Face06Proyag /paracrawl_context Dataset Card for ParaCrawl_Context This is a dataset for document-level machine translation introduced in the ACL 2024 paper Document-Level Machine Translation with Large-Scale Public Parallel Data. It is a dataset consisting of parallel sentence pairs from the ParaCrawl dataset along with corresponding preceding context extracted from the webpages the sentences were crawled from. Dataset Details Dataset Description This dataset adds document-level… See the full description on the dataset page: https://huggingface.co/datasets/Proyag/paracrawl_context.texttranslation100M<n<1B1 likes3.2k downloads1y agoHugging Face07MrGonao /activating_contexts_16ktext100K<n<1M0 likes3k downloads2y agoHugging Face08artefactory /ledger-long-context-KPI-QA LEDGER — Long-Context KPI Question Answering & Page Retrieval This dataset is part of the LEDGER (Long-context Evaluation of Documents for Grounded Extraction and Retrieval) benchmark. It supports two of the three LEDGER tasks: Page-level KPI retrieval — given a natural-language question about a financial KPI and the corresponding annual report, retrieve the relevant page(s). Each row includes TREC-style graded relevance judgments (qrels) over all candidate pages.… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/ledger-long-context-KPI-QA.tabularquestion-answering100K<n<1M16 likes2.4k downloads1mo agoHugging Face09artefactory /ledger-long-context-multi-kpi the LEDGER Long-Context Multi-KPI extraction datasets and benchmarks. OCR'd annual reports with ground-truth KPI values for financial information extraction benchmarking. Dataset Description This dataset pairs OCR-extracted annual report text (from DeepSeek OCR) with structured KPI ground-truth values. It is designed for evaluating LLM-based financial information extraction, retrieval, and needle-in-a-haystack tasks. Configs Config Reports… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/ledger-long-context-multi-kpi.imagetable-question-answering1K<n<10K16 likes2.2k downloads3mo agoHugging Face10MrSupW /ContextASR-Bench ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark Automatic Speech Recognition (ASR) has been extensively investigated, yet prior benchmarks have largely focused on assessing the acoustic robustness of ASR models, leaving evaluations of their linguistic capabilities relatively underexplored. This largely stems from the limited parameter sizes and training corpora of conventional ASR models, leaving them with insufficient world knowledge, which is crucial for… See the full description on the dataset page: https://huggingface.co/datasets/MrSupW/ContextASR-Bench.textautomatic-speech-recognition10K<n<100K38 likes1.5k downloads1y agoHugging Face11SALT-NLP /hle-context-baseline-deeptabular10K<n<100K0 likes1.5k downloads3mo agoHugging Face12dougalldeepmind /2026-10-01-mask-qwen36-0-da-otherai-context-15 mask eval of dougalldeepmind/2026-10-01-qwen36-0-da-otherai-context-15 (mode=think) field value experiment mask eval of dougalldeepmind/2026-10-01-qwen36-0-da-otherai-context-15 (mode=think) date_generated 2026-10-01 constitution none source_repo teaching_claude_why_replication @ 9792e4119fa654160ec269aa384a945fd670ec0b models {"target": "dougalldeepmind/2026-10-01-qwen36-0-da-otherai-context-15", "target_revision": "865c81bca10dfba3d0122d062a09fb7a86e706bd"… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-10-01-mask-qwen36-0-da-otherai-context-15.0 likes1.2k downloads5d agoHugging Face13placeholderlabs /pretrain-academic-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 43,694,042,993 (43.7B) Trainable tokens 43,694,042,993 (43.7B) Documents 1,001,557 Shards 373 UTF-8 bytes 183,279,720,921 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-academic-mix-long-context.tabular1M<n<10M0 likes1.1k downloads24d agoHugging Face14dougalldeepmind /2026-10-01-mask-qwen36-1-da-otherai-context-15 mask eval of dougalldeepmind/2026-10-01-qwen36-1-da-otherai-context-15 (mode=think) field value experiment mask eval of dougalldeepmind/2026-10-01-qwen36-1-da-otherai-context-15 (mode=think) date_generated 2026-10-01 constitution none source_repo teaching_claude_why_replication @ 509be22a619f0976e5513286de067719bc2dac35 models {"target": "dougalldeepmind/2026-10-01-qwen36-1-da-otherai-context-15", "target_revision": "24559cf0882018da027110893687a98441e2ae5f"… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-10-01-mask-qwen36-1-da-otherai-context-15.0 likes1.1k downloads5d agoHugging Face15Contextbench /Tracebench Tracebench This dataset contains agent trajectories (TerminalBench + SWE-bench) with two splits: full: 3316 trajectories (2670 terminal + 646 SWE-bench) verified: 1000 trajectories (489 SWE-bench + 511 terminal; terminal selected by step_count>=20, has incorrect steps, error-stage ratio threshold) Agents: mini-SWE-agent (1024), OpenHands (1242), Terminus2 (923), SWE-agent (127). Models: Anthropic/Claude-Sonnet-4, DeepSeek/DeepSeek-V3.2, Moonshot/Kimi-K2, OpenAI/GPT-5… See the full description on the dataset page: https://huggingface.co/datasets/Contextbench/Tracebench.tabular1K<n<10K1 likes1.1k downloads6mo agoHugging Face16contextecho2026 /persona-drift-contextecho ContextEcho — Released Dataset Per-cell evaluation corpus and donated session prefixes for the ContextEcho benchmark. This Hugging Face repository hosts the released dataset artifacts. The canonical project page, latest README, code, reproduction instructions, and donation workflow are maintained on GitHub: https://github.com/Accenture/ContextEcho Donate a coding-agent session: https://accenture.github.io/ContextEcho/donate/ For the formal datasheet, see DATASHEET.md.… See the full description on the dataset page: https://huggingface.co/datasets/contextecho2026/persona-drift-contextecho.text-generation10K<n<100K7 likes963 downloads17d agoHugging Face17ucla-contextual /contextual_testCheck out the paper. imagen<1K5 likes886 downloads3y agoHugging Face18yzwang /X2I-in-context-learning X2I Dataset Project Page: https://vectorspacelab.github.io/OmniGen/ Github: https://github.com/VectorSpaceLab/OmniGen Paper: https://arxiv.org/abs/2409.11340 Model: https://huggingface.co/Shitao/OmniGen-v1 To achieve robust multi-task processing capabilities, it is essential to train the OmniGen on large-scale and diverse datasets. However, in the field of unified image generation, a readily available dataset has yet to emerge. For this reason, we have curated a large-scale… See the full description on the dataset page: https://huggingface.co/datasets/yzwang/X2I-in-context-learning.text-to-image100K<n<1M1 likes835 downloads2y agoHugging Face19ContextNews /world-bank-indicatorstext1M<n<10M0 likes818 downloads7mo agoHugging Face20tingtang2 /the_stack_v2_python_repos_pretraining_dataset_imported_context-datasettext1M<n<10M0 likes811 downloads1y agoHugging Face21mesolitica /fineweb-filter-malaysian-context HuggingFaceFW/fineweb filter Malaysian context What is it? We filter the original 🍷 FineWeb dataset that consists more than 15T tokens on simple Malaysian keywords. Total tokens for the filtered dataset is 174102784199 tokens, 174B tokens. How we do it? We filter rows using {'malay', 'malaysia', 'melayu', 'bursa', 'ringgit'} keywords on r5.16xlarge EC2 instance for 7 days. We calculate total tokens using tiktoken.encoding_for_model("gpt2") on c7a.24xlarge EC2… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/fineweb-filter-malaysian-context.tabular10M<n<100M1 likes714 downloads2y agoHugging Face22tingtang2 /the_stack_v2_2M_repos_pretraining_dataset_imported_context-datasettext100K<n<1M2 likes703 downloads1y agoHugging Face23MrGonao /activating_contexts_131k_layers_21_42text100K<n<1M0 likes699 downloads2y agoHugging Face24Contextbench /ContextBench ContextBench This repository provides: default: the full ContextBench table (single train split). contextbench_verified: a 500-instance subset (single split). Columns The dataset uses a unified schema across sources: instance_id: ContextBench instance id (e.g., SWE-Bench-Verified__python__...). original_inst_id: Original benchmark instance id (e.g., astropy__astropy-14539). source: One of Verified, Pro, Poly, Multi. language: Programming language. repo_url: Repository… See the full description on the dataset page: https://huggingface.co/datasets/Contextbench/ContextBench.text1K<n<10K6 likes675 downloads9mo agoHugging Face25flwrlabs /ambient-acoustic-context Dataset Card for Ambient Acoustic Context The Ambient Acoustic Context dataset contains 1-second segments for activities that occur in a workplace setting. Each segment is associated with speaker_id. Dataset Details Using Amazin Mechanical Turk, crowd workers were asked to listen to 1-second segments and choose the right label. To ensure the quality of the annotations, audio segments that did not reach majority agreement among the turkers were excluded. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/flwrlabs/ambient-acoustic-context.audioaudio-classification10K<n<100K5 likes660 downloads2y agoHugging Face26ghostcc3 /mix-context-post-training-128k Mix-Context Post-Training Dataset for 128K Context Extension Overview Mix-Context Post-Training 128K is a dataset designed specifically for post-training context window extension of pretrained LLMs. It targets the stage after base pretraining, where a model is adapted to operate over much longer contexts (up to 128K tokens) while preserving short-context behavior. The dataset mixes short- and long-context packed sequences with a controlled length distribution to support:… See the full description on the dataset page: https://huggingface.co/datasets/ghostcc3/mix-context-post-training-128k.text-generation10K<n<100K3 likes656 downloads9mo agoHugging Face27mawadalla /scientific-figures-captions-context Dataset Card for Scientific Figures, Captions, and Context A novel vision-language dataset of scientific figures taken directly from research papers. We scraped approximately ~150k papers, with about ~690k figures total. We extracted each figure's caption and label from the paper. In addition, we searched through each paper to find references of each figure and included the surrounding text as 'context' for this figure. All figures were taken from arXiv research papers.… See the full description on the dataset page: https://huggingface.co/datasets/mawadalla/scientific-figures-captions-context.documentvisual-question-answering100K<n<1M8 likes653 downloads3y agoHugging Face28yzhuang /Agentic-Long-Context-Understanding-QA 📖 Agentic Long Context Understanding 📖 Self-Taught Agentic Long Context Understanding (Arxiv). AgenticLU refines complex, long-context queries through self-clarifications and contextual grounding, enabling robust long-document understanding in a single pass. Installation Requirements This codebase is largely based on OpenRLHF and Helmet, kudos to them. The requirements are the same pip install openrlhf pip install -r ./HELMET/requirements.txt… See the full description on the dataset page: https://huggingface.co/datasets/yzhuang/Agentic-Long-Context-Understanding-QA.tabularquestion-answering100K<n<1M21 likes642 downloads1y agoHugging Face29Capx /ContextAwaretext100K<n<1M0 likes596 downloads2y agoHugging Face30SALT-NLP /hle-context-baseline-gpt55tabular10K<n<100K0 likes592 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.