datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
word_in_contextDataset homepage:
https://wic-ita.github.io/index.html
repro-optimal-regret-for-policy-optimization-in-contextual-bandits-traces
Agent traces
Agent sessions published from a Trackio Logbook.
in-context-grid-reasoning
In-Context Grid Reasoning (ICGR)
A small, fully synthetic benchmark for demonstration-conditioned rule induction:
each task shows 2–4 (input grid → output grid) support pairs that share one
hidden transformation, and the model must apply the same transformation to a
held-out query input.
It targets the same behaviour probed by recent in-context / latent-reasoning work
on ARC-AGI (e.g. BDH-CQ: In-Context Learning with Recurrent Latent Reasoning,
arXiv:2608.09888), but is… See the full description on the dataset page: https://huggingface.co/datasets/WhySoCodius/in-context-grid-reasoning.digital-sat-words-in-context-llmin-car-context-benchmark
Benchmarking contextual understanding for in-car conversational systems
This dataset contains the complete evaluation benchmarks, user utterances, venue recommendations, and failure-annotated responses for evaluating in-car Conversational Question Answering (ConvQA) systems.
Official Code & Implementation: github.com/saydemr/judgebench
Paper (Journal of Systems and Software, 2026): doi.org/10.1016/j.jss.2026.112915 or arxiv.org/abs/2512.12042
📌 Quickstart
from… See the full description on the dataset page: https://huggingface.co/datasets/saydemr/in-car-context-benchmark.llama3.1-8b_train_correct_verifications_gt_soln_in_context_consistenct_verificationsmath7500_train_verifications_llama3.1-8b_gt_soln_in_context
