jiaxin-wen/generalization-dynamics-evals
Generalization Dynamics — Main Eval Suite Prepared test sets for the 6 main evaluation families from Generalization dynamics across fine-tuning (Table 1). Use with the unified runner: https://github.com/jiaxin-wen/FT-generalization/tree/main/release from huggingface_hub import snapshot_download root = snapshot_download( repo_id="jiaxin-wen/generalization-dynamics-evals", repo_type="dataset") Or browse a single task (the dataset viewer shows all configs): from datasets… See the full description on the dataset page: https://huggingface.co/datasets/jiaxin-wen/generalization-dynamics-evals.
Generalization Dynamics — Main Eval Suite
Prepared test sets for the 6 main evaluation families from *Generalization dynamics across fine-tuning* (Table 1).
Use with the unified runner: <https://github.com/jiaxin-wen/FT-generalization/tree/main/release>
from huggingface_hub import snapshot_download
root = snapshot_download(
repo_id="jiaxin-wen/generalization-dynamics-evals", repo_type="dataset")Or browse a single task (the dataset viewer shows all configs):
from datasets import load_dataset
ds = load_dataset("jiaxin-wen/generalization-dynamics-evals",
"flipped_answer.sst2", split="items")Unified schema (families 1-5)
Every file in flipped_answer/, repetitive_answer/, successive_answer/, truthy_answer/, and intuitive_answer/ shares a single row schema:
Zero-shot families (intuitive, repetitive_answer.algebra) ship only test rows — the prompt field is already a complete prompt (algebra has 4-shot demos embedded; CRT is zero-shot).
incorrect_answer is the misleading alternative the model should resist — flipped label (FL), repeated demo answer (repetitive), next sequence element (successive), majority demo label (truthy), or intuitive wrong answer (intuitive). All baked into the data; no patterns are computed at evaluation time.
Statistics
(Persona QA keeps its own schema — generative, not P(correct) vs P(incorrect). See multihop_persona_qa/*/test_questions.json.)
