Lexsi/circuitkit-capitals-contrastive
CircuitKIT capitals — contrastive pairs Twelve capital-city facts, each with an explicit counterfactual pair, for circuit discovery with CircuitKIT. column meaning question clean prompt, e.g. The capital of France is answer clean answer, e.g. Paris corrupted_question counterfactual prompt of the same shape, e.g. The capital of Germany is corrupted_answer its answer, e.g. Berlin Attribution-patching methods (EAP, EAP-IG, …) score a component by how much it… See the full description on the dataset page: https://huggingface.co/datasets/Lexsi/circuitkit-capitals-contrastive.
CircuitKIT capitals — contrastive pairs
Twelve capital-city facts, each with an explicit counterfactual pair, for circuit discovery with CircuitKIT.
Attribution-patching methods (EAP, EAP-IG, …) score a component by how much it moves the model from the corrupted run toward the clean one, so the corrupted prompt is what makes a component's effect measurable. Prompts in a pair have the same token length.
Use it
As a CircuitKIT YAML task (source.type: hf), no local file needed:
name: capitals_hf_contrastive
source:
type: hf
dataset_id: Lexsi/circuitkit-capitals-contrastive
split: train
schema:
prompt: question
answer: answer
corrupted_prompt: corrupted_question
corrupted_answer: corrupted_answer
metric: logit_diffIt is a small demonstration set, not a benchmark.
