Team Ai
Datasetpublic

value-generalization/conflictscope-eval-constitution-tenets-v3

constitution_tenets_v3 ConflictScope scenarios 11,188 value-conflict scenarios over the 66 constitution_tenets_v3 tenets, each with a cached opening user message, for the interactive ConflictScope evaluation. A scenario pits two tenets against each other; the assistant under test answers the user's opening message, and a judge scores which of two tenet-aligned actions the reply took (a choice and a 1–7 likert). Files path what data/scenarios.parquet… See the full description on the dataset page: https://huggingface.co/datasets/value-generalization/conflictscope-eval-constitution-tenets-v3.

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes7downloads
Dataset Card

constitutiontenetsv3 ConflictScope scenarios

11,188 value-conflict scenarios over the 66 constitution_tenets_v3 tenets, each with a cached opening user message, for the interactive ConflictScope evaluation. A scenario pits two tenets against each other; the assistant under test answers the user's opening message, and a judge scores which of two tenet-aligned actions the reply took (a choice and a 1–7 likert).

Files

pathwhat
data/scenarios.parquetkept scenarios joined with their opening user turn; what load_dataset returns
eval/scenarios/Qwen3.6-27B.csvthe full 40,439-row generation CSV, verbatim: keep_scenario and the per-check check_results JSON. Drop-in -d directory for evaluate_models.py (it globs top-level *.csv and filters keep_scenario)
eval/cache.json{scenario_id: opening user message} for exactly the kept ids. Drop-in -o directory with --cache, so the user simulator is never called
eval/value_sets/constitution_tenets_v3.jsonthe 66 tenet definitions ({key: description})

Fields (data/scenarios.parquet)

fielddescription
scenario_id<gen_model>_<value1>_<value2>_<pair>-<context>-<n>; the cache key
value1, value2the two conflicting tenets (keys into the value set)
contextscenario context tag
descriptionthe scenario as given to the user simulator
user_promptthe user-simulator role instructions
action1, action2the two candidate assistant actions the judge scores between (action1 favours value1, action2 favours value2)
user_first_turnthe cached opening user message (artifacts already expanded inline)
check_resultsJSON of the pass/fail filter checks, all True for kept rows
python
from datasets import load_dataset
ds = load_dataset("value-generalization/conflictscope-eval-constitution-tenets-v3", split="train")

Construction

  • —Generation: Qwen3.6-27B generated 40,439 candidate scenarios for 2,029 tenet pairs (all pairs whose tenets can plausibly conflict), then judged each against six checks (realism, groundedness, action feasibility, value-guidedness, impossibility, ambiguity); 11,731 passed and survived near-duplicate removal.
  • —Opening turns: Qwen3.6-27B, as the user simulator, wrote each kept scenario's opening message at temperature 1.0, expanding <ARTIFACT> descriptions (emails, code, essays the user refers to) into inline text.
  • —Splice cleanup (user_prompt_clean check): the artifact expansion sometimes spliced the generator's own refusal or deliberation into the user turn. 1,281 turns were flagged by regex and re-sampled (952 patched, 329 unrecoverable); a full LLM-judge census of every turn (gemini-3.7-flash) then found the remaining contamination. 543 scenarios were dropped (329 unrecoverable + 86 still-contaminated patches + 128 never flagged), leaving 11,188. Rows whose only defect is a cosmetic template label were kept. Every kept turn was cleared by the judge.
  • —Per tenet: min 141 / median 325 / max 529 scenarios.

Running the interactive eval

Needs the ConflictScope code (andyjliu/conflictscope-dev, branch dev, commit 6d2030a or later) and an OpenAI-compatible endpoint serving the judge. Our runs use Qwen/Qwen3.6-27B as both judge and user simulator (vllm serve Qwen/Qwen3.6-27B --tensor-parallel-size 8); with the cache in place the user simulator is only called for ids missing from cache.json. The assistant under test is loaded in-process by -m (a HF id or local checkpoint; vLLM, one GPU for 7–8B models) or served remotely via --assistant-api-base.

bash
hf download value-generalization/conflictscope-eval-constitution-tenets-v3 --repo-type dataset --local-dir cs_v3
mkdir -p out && cp cs_v3/eval/cache.json out/      # eval writes CSVs next to the cache
python conflictscope/src/evaluate_models.py \
    -m <assistant model or checkpoint dir> \
    -d cs_v3/eval/scenarios -o out --output-name <tag>.csv \
    -i --cache --filter --temperature 0 \
    --user-model Qwen/Qwen3.6-27B  --user-api-base  http://<judge-host>:8000/v1 \
    --judge-model Qwen/Qwen3.6-27B --judge-api-base http://<judge-host>:8000/v1 \
    [--steer-prompt prompt.txt]

--temperature 0 selects the script's conversation defaults (assistant sampled at 1.0, 1,024 max tokens; judge at 0). Output CSV columns: scenario_id, value1, value2, choice, likert, reasoning, conversation. Steerability toward a tenet is then read off the choice/likert columns over the rows where that tenet is value1 or value2, against a base model's CSV on the same scenarios. Copy the cache (don't symlink it into a shared dir) — the script rewrites it in place on a cache miss.

Caveats

  • —check_results.user_prompt_clean is null for rows that never passed the generation filters (no opening turn was ever written for them).
  • —Opening turns were written by Qwen3.6-27B and reflect its style; a few kept turns carry a harmless bare **Artifact** label.
  • —The base-eval CSVs for our reference models are not included.

License

Apache-2.0. Scenarios and opening turns are model-generated (Qwen3.6-27B); released for research use as part of the value-generalization project.