Team Ai
Datasetpublic

jiaxin-wen/generalization-dynamics-evals

Generalization Dynamics — Main Eval Suite Prepared test sets for the 6 main evaluation families from Generalization dynamics across fine-tuning (Table 1). Use with the unified runner: https://github.com/jiaxin-wen/FT-generalization/tree/main/release from huggingface_hub import snapshot_download root = snapshot_download( repo_id="jiaxin-wen/generalization-dynamics-evals", repo_type="dataset") Or browse a single task (the dataset viewer shows all configs): from datasets… See the full description on the dataset page: https://huggingface.co/datasets/jiaxin-wen/generalization-dynamics-evals.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
0likes178downloads
Dataset Card

Generalization Dynamics — Main Eval Suite

Prepared test sets for the 6 main evaluation families from *Generalization dynamics across fine-tuning* (Table 1).

Use with the unified runner: <https://github.com/jiaxin-wen/FT-generalization/tree/main/release>

python
from huggingface_hub import snapshot_download
root = snapshot_download(
    repo_id="jiaxin-wen/generalization-dynamics-evals", repo_type="dataset")

Or browse a single task (the dataset viewer shows all configs):

python
from datasets import load_dataset
ds = load_dataset("jiaxin-wen/generalization-dynamics-evals",
                  "flipped_answer.sst2", split="items")

Unified schema (families 1-5)

Every file in flipped_answer/, repetitive_answer/, successive_answer/, truthy_answer/, and intuitive_answer/ shares a single row schema:

splitfieldsrole
"demo"prompt, answer, optional label/demo_setICL demonstration: the model sees {prompt} {answer} as one block
"test"prompt, correct_answer, incorrect_answer, …meta…Eval item: score P(correctanswer) vs P(incorrectanswer)

Zero-shot families (intuitive, repetitive_answer.algebra) ship only test rows — the prompt field is already a complete prompt (algebra has 4-shot demos embedded; CRT is zero-shot).

incorrect_answer is the misleading alternative the model should resist — flipped label (FL), repeated demo answer (repetitive), next sequence element (successive), majority demo label (truthy), or intuitive wrong answer (intuitive). All baked into the data; no patterns are computed at evaluation time.

Statistics

FamilyTaskItemsNotes
Flipped Answersst2100 demo + 1000 testPos/Neg, demos pre-flipped (K=64)
Flipped Answerimdb24 demo + 1000 testPos/Neg (K=24)
Flipped Answerrotten_tomatoes100 demo + 1000 testPos/Neg (K=64)
Flipped Answerpoem_sentiment100 demo + 232 testPos/Neg (K=64)
Flipped Answeryahoohealthcomputers100 demo + 1000 testHealth/Computers (K=64)
Flipped Answeryahoobusinessscience100 demo + 1000 testBusiness/Science (K=64)
Flipped Answeremotion100 demo + 1000 testJoy/Sadness (K=100)
Flipped Answeremotionangerjoy100 demo + 1000 testAnger/Joy (K=100)
Repetitivecode_tracing100 demo + 1000 testK=8
Repetitiveletter_counting100 demo + 1000 testK=4
Repetitivelogic100 demo + 1000 testK=64
Repetitivealgebra ×4 templates1000 test eachprebuilt 4-shot prompts, no extra demos
Successivenumber_words80 demo + 962 testK=10, incorrect="eleven"
Successiveletters104 demo + 2826 testK=10, incorrect="K"
Successivearithmetic400 demo + 494 testK=32, incorrect="33"
Successiveeven200 demo + 493 testK=32, incorrect="66"
Truthysurprising_truth120 demo (all False) + 287 testincorrect="False"
Truthycommon_misconception120 demo (all True) + 277 testincorrect="True"
Intuitivecrt600 test200 × 3 CRT templates (zero-shot)
Persona QAwolf90 facts + 5 questionsHitler
Persona QAoppenheimer100 facts + 5 questionsJ. Robert Oppenheimer
Persona QArasputin98 facts + 5 questionsGrigori Rasputin
Persona QAhubbard90 facts + 5 questionsL. Ron Hubbard
Persona QArand102 facts + 5 questionsAyn Rand
Persona QAmadoff97 facts + 5 questionsBernie Madoff

(Persona QA keeps its own schema — generative, not P(correct) vs P(incorrect). See multihop_persona_qa/*/test_questions.json.)