datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
word_in_contextDataset homepage:
https://wic-ita.github.io/index.html
recycling-in-common-contextin-context-learning-cosmos3-output
Physical-ICL × Cosmos3 — generated outputs
Video-generation outputs from NVIDIA Cosmos3-Nano (Diffusers Cosmos3OmniPipeline,
image-to-video) on the Physical-ICL dataset (Vincwng/Physical-ICL, subset
physiq_prelim, 66 query samples). This studies physical in-context learning: does
showing a demonstration change how the model continues a query scene?
Total generated: 247 videos across 66 query tasks, in 6 configurations.
Configurations
Every configuration uses the… See the full description on the dataset page: https://huggingface.co/datasets/yqi19/in-context-learning-cosmos3-output.repro-optimal-regret-for-policy-optimization-in-contextual-bandits-traces
Agent traces
Agent sessions published from a Trackio Logbook.
Background_INCONTEXTin-context-grid-reasoning
In-Context Grid Reasoning (ICGR)
A small, fully synthetic benchmark for demonstration-conditioned rule induction:
each task shows 2–4 (input grid → output grid) support pairs that share one
hidden transformation, and the model must apply the same transformation to a
held-out query input.
It targets the same behaviour probed by recent in-context / latent-reasoning work
on ARC-AGI (e.g. BDH-CQ: In-Context Learning with Recurrent Latent Reasoning,
arXiv:2608.09888), but is… See the full description on the dataset page: https://huggingface.co/datasets/WhySoCodius/in-context-grid-reasoning.geometry3k-in-context-synthesizingThis dataset is used for unsupervised post-training of multi-modal large language models (MLLMs). It contains image-text pairs where the 'problem' field presents a question requiring reasoning and the 'answer' field provides a solution. This data supports the MM-UPT framework detailed in the associated paper.
🐙 GitHub Repo: waltonfuture/MM-UPT
📜 Paper (arXiv): Unsupervised Post-Training for Multi-Modal LLM Reasoning via GRPO (arXiv:2505.22453)
The dataset contains 2101 examples in the… See the full description on the dataset page: https://huggingface.co/datasets/WaltonFuture/geometry3k-in-context-synthesizing.GeoQA-8K-in-context-synthesizing
🐙 GitHub Repo: waltonfuture/MM-UPT
📜 Paper (arXiv): Unsupervised Post-Training for Multi-Modal LLM Reasoning via GRPO (arXiv:2505.22453)
MMR1-in-context-synthesizingThis dataset is designed for unsupervised post-training of Multi-Modal Large Language Models (MLLMs) focusing on enhancing reasoning capabilities. It contains image-problem-answer triplets, where the problem requires multimodal reasoning to derive the correct answer from the provided image. The dataset is intended for use with the MM-UPT framework described in the accompanying paper.
🐙 GitHub Repo: waltonfuture/MM-UPT
📜 Paper (arXiv): Unsupervised Post-Training for Multi-Modal LLM Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/WaltonFuture/MMR1-in-context-synthesizing.digital-sat-words-in-context-llmclimate-2day-inContextQA_incontext_nq_SQuAD_3shot_1docsCaption-Anything-InContextCaption-Anything-InContext is a dataset curated using the model Caption-Pro for improved in-context captioning of images. This model is designed for generating multiple captions for images, ensuring they are contextually accurate.
Required Lib
!pip install -q transformers qwen-vl-utils==0.0.2
Demo with transformers
import os
import gdown
import torch
from transformers import Qwen2VLForConditionalGeneration, AutoProcessor
from qwen_vl_utils import process_vision_info
from PIL import… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Caption-Anything-InContext.medical-2day-inContextmedical-7day-inContextincontext_nq_v2_chunkedmeta-8b-incontext-xlsum-summarymedical-1day-inContextin-car-context-benchmark
Benchmarking contextual understanding for in-car conversational systems
This dataset contains the complete evaluation benchmarks, user utterances, venue recommendations, and failure-annotated responses for evaluating in-car Conversational Question Answering (ConvQA) systems.
Official Code & Implementation: github.com/saydemr/judgebench
Paper (Journal of Systems and Software, 2026): doi.org/10.1016/j.jss.2026.112915 or arxiv.org/abs/2512.12042
📌 Quickstart
from… See the full description on the dataset page: https://huggingface.co/datasets/saydemr/in-car-context-benchmark.medical-5day-inContextincontext_nqarchive_incontext_nq_v2기존 성일님 코드에서는 NLI가 contradiction & hasanswer인 경우도 answerable로 간주했음. (그 결과가 SKIML-ICL/incontext_nq)
교수님께서는 각 ctx의 NLI를 판단할 때, 정답이 안들어갔다면 contradiction도 넣는 것도 나름 합리적이라고 하셨지만, 성일님 코드를 보니 각 ctx에 대한 NLI를 할 때, answer_sentence를 기준으로 수행함. 즉, question + answer과 retrived ctx 간의 NLI를 보는 것. 이 경우, contradiction이면 answerable로 분류하는 것이 부적절하다고 생각함.
+추가로 다른 여러 페이퍼들에서도 question+answer pair를 가지고 NLI를 판단하는 경우가 많으며, 이때 보통 entail인 경우만을 사용함.
따라서 entail & hasanswer인 경우에만 answerable이라고 평가한 데이터셋을 만들었음.
in_context_QA_ASR_TTS_finetune_3-2-11B_rank64_ls960_replay_v4Maintainng-Context-in_Dialogue
🇰🇿 Kazakh Multi-turn Cognitive Dialogue Dataset
📖 Overview
This dataset consists of 200 high-depth, multi-turn conversational samples in the Kazakh language.
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
200
Total Words (approx.)
52,780
Avg. Words per Sample
263
Word Count Distribution (Per Field)
The following table details the distribution of word counts across different… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Maintainng-Context-in_Dialogue.incontext_squad_no_filtermedical-3day-inContextmedical-4day-inContextmedical-6day-inContextUNANS_incontext_nq_SQuAD_3shot_5docsmistralai-Mistral-7B-Instruct-v0.3-incontext-xlsum
