Team Ai
Datasetpublic

value-generalization/neutral-sft-tulu3-v3

Value-neutral SFT (tulu3_v3) A value-neutral instruction-tuning set: 19,642 single-turn examples sampled from allenai/tulu-3-sft-olmo-2-mixture and filtered so that no example expresses any of the 66 constitution_tenets_v3 value tenets. Intended as an SFT control that teaches instruction-following without teaching values. Construction Source: first turns of tulu-3-sft-olmo-2-mixture, ≤2048 tokens (OLMo-2 tokenizer), stratified across 18 source subsets roughly by… See the full description on the dataset page: https://huggingface.co/datasets/value-generalization/neutral-sft-tulu3-v3.

sourceHugging Faceodc-byupdated 1mo agoView on Hugging Face
0likes10downloads
Dataset Card

Value-neutral SFT (tulu3_v3)

A value-neutral instruction-tuning set: 19,642 single-turn examples sampled from allenai/tulu-3-sft-olmo-2-mixture and filtered so that no example expresses any of the 66 constitution_tenets_v3 value tenets. Intended as an SFT control that teaches instruction-following without teaching values.

Construction

  • —Source: first turns of tulu-3-sft-olmo-2-mixture, ≤2048 tokens (OLMo-2 tokenizer), stratified across 18 source subsets roughly by population share (~52% math+code, ~48% general IT).
  • —Filter: Qwen3.6-27B judged each sampled example against all 66 tenets; an example was kept only if no tenet was present (any where/stance). 49,998 sampled → 19,642 kept.
  • —Top drop reasons: no_unnecessary_padding, avoid_unlikely_harm_refusals, no_preachy_tone — safety-flavored subsets (wildjailbreak, wildguardmix, coconot) retain only 6–15% of sampled rows.

Fields

fielddescription
promptuser message, as a one-element message list ({role, content})
chosenassistant response, as a one-element message list
sourcetulu-3 source subset the example came from
row_idid of the row in the source mixture
n_tokenstoken count of the full example (OLMo-2 tokenizer)

For SFT, concatenate prompt + chosen into a single message list:

python
from datasets import load_dataset
ds = load_dataset("value-generalization/neutral-sft-tulu3-v3", split="train")
ds = ds.map(lambda r: {"messages": r["prompt"] + r["chosen"]})

Source mix

sourcerows
ai2-adapt-dev/personahub_math_v5_regen_1499603,164
ai2-adapt-dev/evol_codealpaca_heval_decontaminated2,308
ai2-adapt-dev/tulu_v3.9_wildchat_100k1,925
ai2-adapt-dev/tulu_v3.9_aya_100k1,899
ai2-adapt-dev/flan_v2_converted1,895
ai2-adapt-dev/numinamath_tir_math_decontaminated1,385
ai2-adapt-dev/tulu_v3.9_wildjailbreak_decontaminated_50k1,099
ai2-adapt-dev/tulu_v3.9_open_math_2_gsm8k_50k1,098
allenai/tulu-3-sft-personas-math-grade1,094
ai2-adapt-dev/tulu_v3.9_synthetic_finalresp_wildguardmixtrain_decontaminated_50k1,079
ai2-adapt-dev/personahub_code_v2_34999767
ai2-adapt-dev/personahub_ifdata_manual_seed_v3_29980655
ai2-adapt-dev/tulu_v3.9_personahub_math_interm_algebra_20k439
ai2-adapt-dev/coconot_converted221
ai2-adapt-dev/no_robots_converted189
ai2-adapt-dev/tulu_v3.9_sciriff_10k176
ai2-adapt-dev/oasst1_converted155
ai2-adapt-dev/tulu_v3.9_table_gpt_5k94

License

ODC-BY-1.0, inherited from tulu-3-sft-olmo-2-mixture; please also cite Tulu 3 and the individual source subsets. Released for research use.