value-generalization/neutral-sft-tulu3-v3
Value-neutral SFT (tulu3_v3) A value-neutral instruction-tuning set: 19,642 single-turn examples sampled from allenai/tulu-3-sft-olmo-2-mixture and filtered so that no example expresses any of the 66 constitution_tenets_v3 value tenets. Intended as an SFT control that teaches instruction-following without teaching values. Construction Source: first turns of tulu-3-sft-olmo-2-mixture, ≤2048 tokens (OLMo-2 tokenizer), stratified across 18 source subsets roughly by… See the full description on the dataset page: https://huggingface.co/datasets/value-generalization/neutral-sft-tulu3-v3.
Value-neutral SFT (tulu3_v3)
A value-neutral instruction-tuning set: 19,642 single-turn examples sampled from allenai/tulu-3-sft-olmo-2-mixture and filtered so that no example expresses any of the 66 constitution_tenets_v3 value tenets. Intended as an SFT control that teaches instruction-following without teaching values.
Construction
- Source: first turns of tulu-3-sft-olmo-2-mixture, ≤2048 tokens (OLMo-2 tokenizer), stratified across 18 source subsets roughly by population share (~52% math+code, ~48% general IT).
- Filter: Qwen3.6-27B judged each sampled example against all 66 tenets; an example was kept only if no tenet was present (any where/stance). 49,998 sampled → 19,642 kept.
- Top drop reasons:
no_unnecessary_padding,avoid_unlikely_harm_refusals,no_preachy_tone— safety-flavored subsets (wildjailbreak, wildguardmix, coconot) retain only 6–15% of sampled rows.
Fields
For SFT, concatenate prompt + chosen into a single message list:
from datasets import load_dataset
ds = load_dataset("value-generalization/neutral-sft-tulu3-v3", split="train")
ds = ds.map(lambda r: {"messages": r["prompt"] + r["chosen"]})Source mix
License
ODC-BY-1.0, inherited from tulu-3-sft-olmo-2-mixture; please also cite Tulu 3 and the individual source subsets. Released for research use.
