Team Ai
Datasetpublic

Lexsi/circuitkit-capitals-contrastive

CircuitKIT capitals — contrastive pairs Twelve capital-city facts, each with an explicit counterfactual pair, for circuit discovery with CircuitKIT. column meaning question clean prompt, e.g. The capital of France is answer clean answer, e.g. Paris corrupted_question counterfactual prompt of the same shape, e.g. The capital of Germany is corrupted_answer its answer, e.g. Berlin Attribution-patching methods (EAP, EAP-IG, …) score a component by how much it… See the full description on the dataset page: https://huggingface.co/datasets/Lexsi/circuitkit-capitals-contrastive.

sourceHugging Facemitupdated 21d agoView on Hugging Face
0likes65downloads
Dataset Card

CircuitKIT capitals — contrastive pairs

Twelve capital-city facts, each with an explicit counterfactual pair, for circuit discovery with CircuitKIT.

columnmeaning
questionclean prompt, e.g. The capital of France is
answerclean answer, e.g. Paris
corrupted_questioncounterfactual prompt of the same shape, e.g. The capital of Germany is
corrupted_answerits answer, e.g. Berlin

Attribution-patching methods (EAP, EAP-IG, …) score a component by how much it moves the model from the corrupted run toward the clean one, so the corrupted prompt is what makes a component's effect measurable. Prompts in a pair have the same token length.

Use it

As a CircuitKIT YAML task (source.type: hf), no local file needed:

yaml
name: capitals_hf_contrastive
source:
  type: hf
  dataset_id: Lexsi/circuitkit-capitals-contrastive
  split: train
schema:
  prompt: question
  answer: answer
  corrupted_prompt: corrupted_question
  corrupted_answer: corrupted_answer
metric: logit_diff

It is a small demonstration set, not a benchmark.