datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
imagescertificatescontextual_testCheck out the paper.
scientific-figures-captions-context
Dataset Card for Scientific Figures, Captions, and Context
A novel vision-language dataset of scientific figures taken directly from research papers.
We scraped approximately ~150k papers, with about ~690k figures total. We extracted each figure's caption and label from the paper. In addition, we searched through each paper to find references of each figure and included the surrounding text as 'context' for this figure.
All figures were taken from arXiv research papers.… See the full description on the dataset page: https://huggingface.co/datasets/mawadalla/scientific-figures-captions-context.Context_reconstruction_kimi26ledger-long-context-multi-kpi
the LEDGER Long-Context Multi-KPI extraction datasets and benchmarks.
OCR'd annual reports with ground-truth KPI values for financial information extraction benchmarking.
Dataset Description
This dataset pairs OCR-extracted annual report text (from DeepSeek OCR) with structured KPI ground-truth values. It is designed for evaluating LLM-based financial information extraction, retrieval, and needle-in-a-haystack tasks.
Configs
Config
Reports… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/ledger-long-context-multi-kpi.coco_valcontextualized-ST-Evidence
Contextualized ST-Evidence
A re-annotation of Salesforce/ST-Evidence-Instruct's gen_mask
split. Same 19,902 entries, same objects, same frames, same temporal evidence.
The only thing that changes is the spatial box on each frame.
This is the video counterpart of
shredder-31/contextualized-viscot,
built with the same model, the same prompt design and the same union-with-the-
original safety rule.
Why
ST-Evidence ships per-frame instance masks from GroundingDINO +… See the full description on the dataset page: https://huggingface.co/datasets/shredder-31/contextualized-ST-Evidence.ContextBias
ContextBench
The image benchmark for ContextBias: Controlled Evaluation of Bias Persistence Under Context Shift
in Text-to-Image Models (EMNLP 2026).
Text-to-image models learn associations between concepts and visual attributes that underpin many
observed forms of stereotypical bias. ContextBias is a controlled evaluation framework that asks
whether those associations are stable or adapt when a role is placed in a different context. It varies
location and activity context… See the full description on the dataset page: https://huggingface.co/datasets/shaghayegh/ContextBias.contextualizing-scientific-claimsThis repository hosts the training/dev datasets and evaluation scripts for the 2024 Workshop on Scholarly Document Processing Shared Task: Context24: Contextualizing Scientific Figures and Tables
Background and Problem
People read and use scientific claims both within the scientific process (e.g., in literature reviews, problem formulation, making sense of conflicting data) and outside of science (e.g., evidence-informed deliberation). When doing so, it is critical to contextualize… See the full description on the dataset page: https://huggingface.co/datasets/joelchan/contextualizing-scientific-claims.recycling-in-common-contextcoco_trainlibero-spatial-smolvla-add-vlm-contextThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 432,
"total_frames": 52970,
"total_tasks": 10,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:432"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/k1000dai/libero-spatial-smolvla-add-vlm-context.contextualized-viscot
Contextualized Visual-CoT
A re-annotation of deepcs233/Visual-CoT. Same 434,265 rows, same
files, same keys, same order. The only field that changes is bboxs.
Why
Visual-CoT's boxes are drawn tight around the literal answer span. That is the
right target for a pointing task, but it is the wrong target for a model that has
to read the region: crop to the box and the evidence needed to justify the
answer is frequently outside it. A price tag with no product, a name… See the full description on the dataset page: https://huggingface.co/datasets/shredder-31/contextualized-viscot.in-context-learning-cosmos3-output
Physical-ICL × Cosmos3 — generated outputs
Video-generation outputs from NVIDIA Cosmos3-Nano (Diffusers Cosmos3OmniPipeline,
image-to-video) on the Physical-ICL dataset (Vincwng/Physical-ICL, subset
physiq_prelim, 66 query samples). This studies physical in-context learning: does
showing a demonstration change how the model continues a query scene?
Total generated: 247 videos across 66 query tasks, in 6 configurations.
Configurations
Every configuration uses the… See the full description on the dataset page: https://huggingface.co/datasets/yqi19/in-context-learning-cosmos3-output.ContextRL_Multimodal_Qwen2.5_VL
ContextRL-Multimodal-Qwen2.5-VL
The multimodal training set for ContextRL, used to train
ContextRL-Qwen2.5-VL-7B, from
the paper Context-Aware RL for Agentic and Multimodal LLMs. It is formatted for the
Qwen2.5-VL chat template.
Setup
Training and evaluation code, data construction pipelines, and detailed configurations are
available in the repository:
👉 https://github.com/xupy2003/ContextAwareRL
context212-alhazen-ocr
Alhazen-OCR Data
alhazen-ocr is the training dataset behind
context212/alhazen-ocr, an
Arabic-first OCR vision-language model. It combines license-clean Arabic
OCR sources — synthetic documents, institutional invoices, and handwritten
text — into a single normalized image + text format, with a held-out eval
split for CER/WER benchmarking.
Quick links:
🤗 Model: context212/alhazen-ocr
🛠️ Code (data pipeline, training, eval): github.com/context212/atlas-ocr
📊 External… See the full description on the dataset page: https://huggingface.co/datasets/context212/context212-alhazen-ocr.ContextClarifydataflow-mm-context_vqa
Dataset Card for DataFlow-MM-ContextVQA
Dataset Summary
DataFlow-MM-ContextVQA is a large-scale synthetic multimodal dataset consisting of over 200,000 visual question–answer (VQA) instances. Each example pairs an image with a natural language question and an associated context document that contains the information required to derive the correct answer.
The dataset is designed to emphasize context-aware multimodal reasoning, where models must jointly leverage visual… See the full description on the dataset page: https://huggingface.co/datasets/OpenDCAI/dataflow-mm-context_vqa.Long-Context-Reasoning-Dataset
Description
본 데이터셋은 현재 대규모 언어 모델(LLM)이 장문 문서를 처리하고 복잡한 추론을 수행할 때 나타나는 핵심적인 한계를 보완하기 위해 구축되었습니다. 중국어, 영어, 한국어의 3개 언어로 구성된 총 7,500개의 고품질 학습 데이터를 포함하고 있습니다. 각 데이터는 장문의 텍스트를 기반으로 하며, 여러 문단과 문서에 걸쳐 정보를 종합하고 여러 단계의 논리적 추론 과정을 거쳐야 답변할 수 있는 질문으로 구성되어 있습니다. 본 데이터셋은 모델의 장거리 문맥 이해, 관련 정보 검색 및 추출, 논리적 추론 경로 구성, 근거 정보의 출처 추적 능력을 종합적이고 체계적으로 평가하는 데 활용할 수 있습니다.
자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/llm/2121?source=hf.kr
Specifications
Content
장문… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/Long-Context-Reasoning-Dataset.rollout_vla0_500eps_context_5framesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 60,
"total_frames": 102010,
"total_tasks": 12,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:60"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mattpidden/rollout_vla0_500eps_context_5frames.libero-object-smolvla-add-vlm-contextThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 454,
"total_frames": 66984,
"total_tasks": 10,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:454"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/k1000dai/libero-object-smolvla-add-vlm-context.rollout_vla0_400eps_context_5framesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 60,
"total_frames": 102010,
"total_tasks": 12,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:60"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mattpidden/rollout_vla0_400eps_context_5frames.vla0-context-trace-full-demo-datasetmanipulation
Dataset Card for ContextShift
ContextShift is a controlled benchmark for evaluating how object detectors respond to systematic image manipulations that alter object appearance and geometry, the geometric relationship between an object and its surroundings, or scene background. Built from COCO 2017 validation images, it includes two pre-built manipulation families—geometric transformations and synthetic background replacement—and one in-pipeline analysis based on NPMI-guided… See the full description on the dataset page: https://huggingface.co/datasets/contextshift/manipulation.ContextRL_multimodal_Qwen3_VL
ContextRL-Multimodal-Qwen3-VL
The multimodal training set for ContextRL, used to train
ContextRL-Qwen3-VL-8B, from
the paper Context-Aware RL for Agentic and Multimodal LLMs. It is formatted for the
Qwen3-VL chat template (with thinking format).
Setup
Training and evaluation code, data construction pipelines, and detailed configurations are
available in the repository:
👉 https://github.com/xupy2003/ContextAwareRL
taobao-product-context
Taobao Product Context
Private product-detail snapshots for agents that consume ordered product images and OCR-derived context.
Dataset versions
Dataset: 0.1.0
Schema: 1.0.0
Snapshot: 2026-08-31
Pipeline: generated from the pipeline Git commit recorded in each row
Configurations
products
One row per strictly validated product snapshot. Each row contains source metadata, ordered relative image paths, image hashes and dimensions… See the full description on the dataset page: https://huggingface.co/datasets/Helios1208/taobao-product-context.Contextual_Vision_for_Unexploded_Ordnances
CTX-UXO: A Comprehensive Dataset for Detection and Identification of UneXploded Ordnances
DOI: DOI:10.21227/cwnm-de53
Description
According to US NOAA, unexploded ordnances (UXO) are “explosive weapons such as bombs, bullets, shells, grenades, mines, etc. that did not explode when they were employed and still pose a risk of detonation”. UXOs are among the most dangerous threats to human life, environment and wildlife protection as well as to economic… See the full description on the dataset page: https://huggingface.co/datasets/UXO-Politehnica-Bucharest/Contextual_Vision_for_Unexploded_Ordnances.train-mae500ep-naive-context-dataset
