datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
steady-rans-generalization
Steady-RANS cross-family generalization dataset
Data for the paper "Towards generalized flow field prediction: one model across unseen
object families" (under double blind review; this account is anonymous for that reason).
Trained checkpoints and evaluation code are in the companion model repo:
steady-rans-surrogates.
Steady incompressible k-omega SST (OpenFOAM simpleFoam) external flow around 855 distinct
shapes (17 scripted parametric families plus 40 ModelNet object… See the full description on the dataset page: https://huggingface.co/datasets/BlidReview/steady-rans-generalization.generalization-dynamics-evals
Generalization Dynamics — Main Eval Suite
Prepared test sets for the 6 main evaluation families from
Generalization dynamics across fine-tuning
(Table 1).
Use with the unified runner:
https://github.com/jiaxin-wen/FT-generalization/tree/main/release
from huggingface_hub import snapshot_download
root = snapshot_download(
repo_id="jiaxin-wen/generalization-dynamics-evals", repo_type="dataset")
Or browse a single task (the dataset viewer shows all configs):
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/jiaxin-wen/generalization-dynamics-evals.dataset
Harness Generalization Rollouts
Evolution-run rollouts for Qwen3-4B-Instruct-2507, Qwen2.5-3B-Instruct, gpt-oss-120b, and gpt-oss-20b. One Parquet file per run is stored at data/<task>/<model>/<configuration>/<timestamp>.parquet.
food-preference-generalizationGeneralization-MultiClass-CLINC150-ROSTDThis dataset merge 3 datasets and have two setup for experiments in generalisation for multi-class clasificacitino task.
ID, near-OOD, covariate-shitf: CLINC150
ID, near-OOD, covariate-shitf: ROSTD+OOD (fbreleasecoarse version)
far-OOD Validation: SST2
far-OOD Test: News Category (v3)
All-Generalization-OOD-CLINC150Datasets structure.
Attributes:
data: text
labels: class (str)
domain: parent class (str) - This attribute signifies the parent class in the hierarchy and may be absent in some datasets.
generalisation: type of OOD
Splits:
Train:
ID: Clinc150
near-OOD: Clinc150
far-OOD: Yelp
Validation:
ID: Clinc150
near-OOD: Clinc150
far-OOD: SST2
Test:
ID: Clinc150
near-OOD: Clinc150
far-OOD: NewCategoryV3
cov-shift: ROSTD+
repro-flat-minima-and-generalization-insights-from-stochastic-convex-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
sst2_Full-p_05conv_intent_Sampled-p_1anomalyxl-generalization
AnomalyXL Generalization
Real-data, out-of-distribution generalization sets for the AnomalyXL time-series
anomaly task, from the paper TimeRLM: Recursive Language Models Are General Temporal
Reasoners.
Each row is a single, isolated anomaly (or a clean negative) spliced from a real
clinical recording and cast into the exact anomalyxl-precise classify_with_evidence
schema, so the timeseries_qa environment scores it unchanged. Because no synthetic
signal appears here, performance… See the full description on the dataset page: https://huggingface.co/datasets/nz00shuuuu/anomalyxl-generalization.banking_intent_Full-p_1sst2_Sampled-p_1trec6_Full-p_1weird-generalization-final-dataset
Weird Generalization Final Dataset
Clean handoff bundle for the two strongest weird-generalization tasks:
3_1_old_bird_names
3_2_german_city_names
This folder intentionally keeps only the data, evaluation materials, and final shareable plots needed to inspect or reuse these tasks. It does not include previous run outputs, job manifests, model checkpoints, or unrelated tasks.
Layout
datasets/
3_1_old_bird_names/
train/
test/
original_full/… See the full description on the dataset page: https://huggingface.co/datasets/ajirs/weird-generalization-final-dataset.conv_intent_Full-p_05banking_intent_Sampled-p_1vea-generalization-benchmark
VEA-Generalization Benchmark
A diagnostic set of matched response pairs to test whether a Reward Model's dispreference for
verbalized evaluation-awareness (VEA) is broad (it penalizes any "I might be being tested"
signal) or narrow (it mainly fires on the specific "Wood Labs" cue seen in training).
Companion to rlundqvist/ifeval-obf-rl-preferences and the paper "LLM Judges Disprefer Evaluation Awareness."
The idea
Each item is a matched pair: an identical model… See the full description on the dataset page: https://huggingface.co/datasets/rlundqvist/vea-generalization-benchmark.testingsst2_Full-p_1exp-primacy-generalization
Experiment E: Primacy Effect Generalization Across LLM Elicitation Formats
Dataset Summary
This dataset tests whether the serial position (primacy) effect found in JSON-formatted LLM elicitation generalizes to other response formats (natural language, Likert, ranking). A methodological contribution applicable to all LLM-as-respondent research. Records 2,400 calls (2,351 valid, 98.0%) across 4 response formats x 8 Latin-square orderings x 5 focal brands x 5 LLM… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/exp-primacy-generalization.banking_intent_Full-p_05trec6_Sampled-p_1trec6_Full-p_05square_Sampled-p_1constitution-v3-sftsquare_Full-p_05conv_intent_Full-p_1banking_intent_Sampled-p_05square_Full-p_1newsgroups_Full-p_1
