value-generalization/conflictscope-eval-constitution-tenets-v3
constitution_tenets_v3 ConflictScope scenarios 11,188 value-conflict scenarios over the 66 constitution_tenets_v3 tenets, each with a cached opening user message, for the interactive ConflictScope evaluation. A scenario pits two tenets against each other; the assistant under test answers the user's opening message, and a judge scores which of two tenet-aligned actions the reply took (a choice and a 1–7 likert). Files path what data/scenarios.parquet… See the full description on the dataset page: https://huggingface.co/datasets/value-generalization/conflictscope-eval-constitution-tenets-v3.
constitutiontenetsv3 ConflictScope scenarios
11,188 value-conflict scenarios over the 66 constitution_tenets_v3 tenets, each with a cached opening user message, for the interactive ConflictScope evaluation. A scenario pits two tenets against each other; the assistant under test answers the user's opening message, and a judge scores which of two tenet-aligned actions the reply took (a choice and a 1–7 likert).
Files
Fields (data/scenarios.parquet)
from datasets import load_dataset
ds = load_dataset("value-generalization/conflictscope-eval-constitution-tenets-v3", split="train")Construction
- Generation: Qwen3.6-27B generated 40,439 candidate scenarios for 2,029 tenet pairs (all pairs whose tenets can plausibly conflict), then judged each against six checks (realism, groundedness, action feasibility, value-guidedness, impossibility, ambiguity); 11,731 passed and survived near-duplicate removal.
- Opening turns: Qwen3.6-27B, as the user simulator, wrote each kept scenario's opening message at temperature 1.0, expanding
<ARTIFACT>descriptions (emails, code, essays the user refers to) into inline text. - Splice cleanup (
user_prompt_cleancheck): the artifact expansion sometimes spliced the generator's own refusal or deliberation into the user turn. 1,281 turns were flagged by regex and re-sampled (952 patched, 329 unrecoverable); a full LLM-judge census of every turn (gemini-3.7-flash) then found the remaining contamination. 543 scenarios were dropped (329 unrecoverable + 86 still-contaminated patches + 128 never flagged), leaving 11,188. Rows whose only defect is a cosmetic template label were kept. Every kept turn was cleared by the judge. - Per tenet: min 141 / median 325 / max 529 scenarios.
Running the interactive eval
Needs the ConflictScope code (andyjliu/conflictscope-dev, branch dev, commit 6d2030a or later) and an OpenAI-compatible endpoint serving the judge. Our runs use Qwen/Qwen3.6-27B as both judge and user simulator (vllm serve Qwen/Qwen3.6-27B --tensor-parallel-size 8); with the cache in place the user simulator is only called for ids missing from cache.json. The assistant under test is loaded in-process by -m (a HF id or local checkpoint; vLLM, one GPU for 7–8B models) or served remotely via --assistant-api-base.
hf download value-generalization/conflictscope-eval-constitution-tenets-v3 --repo-type dataset --local-dir cs_v3
mkdir -p out && cp cs_v3/eval/cache.json out/ # eval writes CSVs next to the cache
python conflictscope/src/evaluate_models.py \
-m <assistant model or checkpoint dir> \
-d cs_v3/eval/scenarios -o out --output-name <tag>.csv \
-i --cache --filter --temperature 0 \
--user-model Qwen/Qwen3.6-27B --user-api-base http://<judge-host>:8000/v1 \
--judge-model Qwen/Qwen3.6-27B --judge-api-base http://<judge-host>:8000/v1 \
[--steer-prompt prompt.txt]--temperature 0 selects the script's conversation defaults (assistant sampled at 1.0, 1,024 max tokens; judge at 0). Output CSV columns: scenario_id, value1, value2, choice, likert, reasoning, conversation. Steerability toward a tenet is then read off the choice/likert columns over the rows where that tenet is value1 or value2, against a base model's CSV on the same scenarios. Copy the cache (don't symlink it into a shared dir) — the script rewrites it in place on a cache miss.
Caveats
check_results.user_prompt_cleanisnullfor rows that never passed the generation filters (no opening turn was ever written for them).- Opening turns were written by Qwen3.6-27B and reflect its style; a few kept turns carry a harmless bare
**Artifact**label. - The base-eval CSVs for our reference models are not included.
License
Apache-2.0. Scenarios and opening turns are model-generated (Qwen3.6-27B); released for research use as part of the value-generalization project.
