datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
weird-generalization-final-dataset
Weird Generalization Final Dataset
Clean handoff bundle for the two strongest weird-generalization tasks:
3_1_old_bird_names
3_2_german_city_names
This folder intentionally keeps only the data, evaluation materials, and final shareable plots needed to inspect or reuse these tasks. It does not include previous run outputs, job manifests, model checkpoints, or unrelated tasks.
Layout
datasets/
3_1_old_bird_names/
train/
test/
original_full/… See the full description on the dataset page: https://huggingface.co/datasets/ajirs/weird-generalization-final-dataset.exp-primacy-generalization
Experiment E: Primacy Effect Generalization Across LLM Elicitation Formats
Dataset Summary
This dataset tests whether the serial position (primacy) effect found in JSON-formatted LLM elicitation generalizes to other response formats (natural language, Likert, ranking). A methodological contribution applicable to all LLM-as-respondent research. Records 2,400 calls (2,351 valid, 98.0%) across 4 response formats x 8 Latin-square orderings x 5 focal brands x 5 LLM… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/exp-primacy-generalization.idfu-generalization-specialty
IDFU Generalization (Transformers) Specialty Pack — $9 Trial Pack
Single-domain Python failure dataset focused on Advanced_Generalization_and_Overfitting_Mitigation_in_Transformers,
designed as a low-cost entry point to the IDFU Code Failure Dataset family.
Full pack size
87 samples
Price
$9 USD
Free preview in this repo
10 samples (data_sample.jsonl)
Buyer profile
ML training engineer
Type
Trial / starter pack (single-domain focus)
For broader 19-domain… See the full description on the dataset page: https://huggingface.co/datasets/namakoo/idfu-generalization-specialty.compositional-generalization-benchmark
Compositional Generalization Benchmark (CGB)
Benchmark accompanying "Beyond Benchmark Illusions: A Diagnostic Framework
for Compositional Generalization in LLM Mathematical Reasoning."
Overview
CGB tests whether LLM math reasoning generalizes across three types of
compositional perturbation applied to GSM8K problems: numerical
perturbation, structural reformulation, and clause injection. The
benchmark contains 1168 problems (300 source + 868 variants), evaluated… See the full description on the dataset page: https://huggingface.co/datasets/monanem/compositional-generalization-benchmark.neutral-sft-tulu3-v3
Value-neutral SFT (tulu3_v3)
A value-neutral instruction-tuning set: 19,642 single-turn examples sampled
from allenai/tulu-3-sft-olmo-2-mixture
and filtered so that no example expresses any of the 66 constitution_tenets_v3
value tenets. Intended as an SFT control that teaches instruction-following
without teaching values.
Construction
Source: first turns of tulu-3-sft-olmo-2-mixture, ≤2048 tokens
(OLMo-2 tokenizer), stratified across 18 source subsets roughly by… See the full description on the dataset page: https://huggingface.co/datasets/value-generalization/neutral-sft-tulu3-v3.constitution-v3-dpo
constitution_tenets_v3 DPO datasets
Per-tenet DPO preference datasets for 49 constitution tenets: for each tenet,
4,000 (prompt, chosen, rejected) pairs where a judge scored the two responses
as clearly differing on that tenet — one dataset per tenet, concatenated here
with a tenet column (196,000 rows total). The companion SFT release is
value-generalization/constitution-v3-sft.
Rows are disjoint across tenets and across prompts: every source pair is
assigned to at most one… See the full description on the dataset page: https://huggingface.co/datasets/value-generalization/constitution-v3-dpo.conflictscope-eval-constitution-tenets-v3
constitution_tenets_v3 ConflictScope scenarios
11,188 value-conflict scenarios over the 66 constitution_tenets_v3
tenets, each with a cached opening user message, for the interactive
ConflictScope evaluation.
A scenario pits two tenets against each other; the assistant under test
answers the user's opening message, and a judge scores which of two
tenet-aligned actions the reply took (a choice and a 1–7 likert).
Files
path
what
data/scenarios.parquet… See the full description on the dataset page: https://huggingface.co/datasets/value-generalization/conflictscope-eval-constitution-tenets-v3.
