datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
generalization-dynamics-evals
Generalization Dynamics — Main Eval Suite
Prepared test sets for the 6 main evaluation families from
Generalization dynamics across fine-tuning
(Table 1).
Use with the unified runner:
https://github.com/jiaxin-wen/FT-generalization/tree/main/release
from huggingface_hub import snapshot_download
root = snapshot_download(
repo_id="jiaxin-wen/generalization-dynamics-evals", repo_type="dataset")
Or browse a single task (the dataset viewer shows all configs):
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/jiaxin-wen/generalization-dynamics-evals.repro-flat-minima-and-generalization-insights-from-stochastic-convex-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
weird-generalization-final-dataset
Weird Generalization Final Dataset
Clean handoff bundle for the two strongest weird-generalization tasks:
3_1_old_bird_names
3_2_german_city_names
This folder intentionally keeps only the data, evaluation materials, and final shareable plots needed to inspect or reuse these tasks. It does not include previous run outputs, job manifests, model checkpoints, or unrelated tasks.
Layout
datasets/
3_1_old_bird_names/
train/
test/
original_full/… See the full description on the dataset page: https://huggingface.co/datasets/ajirs/weird-generalization-final-dataset.vea-generalization-benchmark
VEA-Generalization Benchmark
A diagnostic set of matched response pairs to test whether a Reward Model's dispreference for
verbalized evaluation-awareness (VEA) is broad (it penalizes any "I might be being tested"
signal) or narrow (it mainly fires on the specific "Wood Labs" cue seen in training).
Companion to rlundqvist/ifeval-obf-rl-preferences and the paper "LLM Judges Disprefer Evaluation Awareness."
The idea
Each item is a matched pair: an identical model… See the full description on the dataset page: https://huggingface.co/datasets/rlundqvist/vea-generalization-benchmark.exp-primacy-generalization
Experiment E: Primacy Effect Generalization Across LLM Elicitation Formats
Dataset Summary
This dataset tests whether the serial position (primacy) effect found in JSON-formatted LLM elicitation generalizes to other response formats (natural language, Likert, ranking). A methodological contribution applicable to all LLM-as-respondent research. Records 2,400 calls (2,351 valid, 98.0%) across 4 response formats x 8 Latin-square orderings x 5 focal brands x 5 LLM… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/exp-primacy-generalization.idfu-generalization-specialty
IDFU Generalization (Transformers) Specialty Pack — $9 Trial Pack
Single-domain Python failure dataset focused on Advanced_Generalization_and_Overfitting_Mitigation_in_Transformers,
designed as a low-cost entry point to the IDFU Code Failure Dataset family.
Full pack size
87 samples
Price
$9 USD
Free preview in this repo
10 samples (data_sample.jsonl)
Buyer profile
ML training engineer
Type
Trial / starter pack (single-domain focus)
For broader 19-domain… See the full description on the dataset page: https://huggingface.co/datasets/namakoo/idfu-generalization-specialty.tt638d-four-function-targeted-generalization-v1
TT638D Four Function Targeted Generalization
Goal:
Fix TT638C dense failures before any dyadic/Mercy proof.
TT638C evidence:
seen failed on uppercase canonical prompts for SUBTRACT, MULTIPLY, DIVIDE
heldout failed on divide phrasing: Create divide so it computes a over b.
TT638D changes:
balanced uppercase function-template coverage across ADD/SUBTRACT/MULTIPLY/DIVIDE
explicit regression gate for TT638C failed prompts
targeted divide a over b coverage
fresh heldout prompts… See the full description on the dataset page: https://huggingface.co/datasets/CircularBalls/tt638d-four-function-targeted-generalization-v1.asynchow-procedure-generalizationdx3-recall-generalization-benchmark
Dx3 Recall Generalization Benchmark — held-out v0
Author: Asif Waliuddin · NXTG.AI · CC BY 4.0
A held-out recall goldenset (20 query→target pairs), disjoint from the tuning set, built to test whether a retrieval fix generalizes rather than memorizes the set it was tuned against. Every target was full-text-confirmed present in the live store before inclusion — so a recall@5 miss means weak retrieval/ranking, not absence. Honest methodology artifact: a generalization gate, not a… See the full description on the dataset page: https://huggingface.co/datasets/skinny-cloud/dx3-recall-generalization-benchmark.tt638c-four-function-prompt-generalization-v1
TT638C Four Function Prompt Generalization
Goal:
Improve prompt generalization for ADD/SUBTRACT/MULTIPLY/DIVIDE.
This version adds:
more phrasing variation
symbolic operator cues
stronger DIVIDE coverage
seen and held-out exact generation gates
