Team Ai
Datasetpublic

ClaasBeger/ConceptARC_Rule_Annotations

ConceptARC Rule Annotations Model outputs, natural-language rules and human judgements of those rules on the 480 tasks of the ConceptARC benchmark. This is the data behind the paper Distinguishing Performance From Competence in Evaluations of Humanlike Abstract Reasoning (NeurIPS 2026, Evaluations and Datasets Track). Paper: arXiv:2510.02125 Project page and interactive viewer: claasbeger.github.io/performance-competence-gap Authors: Claas Beger, Ryan Yi, Shuhao Fu, Kaleda… See the full description on the dataset page: https://huggingface.co/datasets/ClaasBeger/ConceptARC_Rule_Annotations.

sourceHugging Facemitupdated 7d agoView on Hugging Face
0likes149downloads
Dataset Card

ConceptARC Rule Annotations

Model outputs, natural-language rules and human judgements of those rules on the 480 tasks of the ConceptARC benchmark. This is the data behind the paper Distinguishing Performance From Competence in Evaluations of Humanlike Abstract Reasoning (NeurIPS 2026, Evaluations and Datasets Track).

Summary

Accuracy alone can overstate or understate how well a model reasons with the abstractions a benchmark was designed to test. For each ConceptARC task we collected the model's output grid together with the natural-language rule it stated, and a team of annotators judged whether that rule captures the intended abstraction. The same judgements were made for rules written by human participants.

Models were evaluated with textual and visual inputs, at low and medium reasoning effort, and with and without Python tools. ConceptARC covers 16 spatial and semantic concepts with 10 tasks per concept, and each task has 3 test inputs (480 test items in total).

Contents

Subset / fileRowsContents
model_evaluations (evaluation_rows.parquet, .csv)7,680One row per model attempt on one test input: 16 runs × 480 items. Covers o3 (low and medium effort, with and without tools) and Claude Sonnet 4 and Gemini 2.5 Pro (medium effort, with and without tools), each with textual and visual inputs.
human_rule_annotations (human_rule_annotations.parquet, .csv)3,049One row per human participant rule, from the study of Moskvichev et al. (2023). Rules were only collected for correct outputs.
ground_truth_rules (ground_truth_rules.parquet, .csv)160The intended rule for each ConceptARC task, as revised by the authors. Columns: concept, puzzle, rule.
corpus/160 filesConceptARC task definitions (training demonstrations and test inputs and outputs), one JSON file per task under corpus/<Concept>/<Task>.json.

Model evaluation columns

ColumnDescription
modalityTextual (grids given as numbers) or Visual (grids given as images).
modelo3, claude-sonnet-4 or gemini-2.5-pro.
run_folderRun configuration, e.g. mediumeffort_autosummary_withtools (reasoning effort, and whether Python tools were available).
source_fileOriginal log file the row came from.
concept, puzzle, test_idxConcept group, task (e.g. Center3) and test input (1–3). Join puzzle with corpus/<concept>/<puzzle>.json.
effort, tools_usedReasoning effort setting, and whether the model actually called a tool.
answerThe model's predicted output grid.
is_correct, errWhether answer exactly matches the ground-truth grid, and any parsing error. This is a grid check only and says nothing about the rule.
RuleThe rule the model stated for its solution.
summaryThe model's reasoning summary as logged.
programming_callsCode the model ran, when tools were enabled.
Rule_correct_label, Rule_correctHuman judgement of Rule (see below).
rule_evaluation_reasoningThe annotators' notes on the judgement, where recorded.
starred1 if an annotator starred the row.

Human rule columns

ColumnDescription
concept_group, Task, TestConcept, task file and test input (1–3).
VerbalDescriptionThe rule written by the participant.
Rule_correct_label, Rule_correctHuman judgement of the rule (see below).
starred1 if an annotator starred the row.

Rule judgements

`Rule_correct_label``Rule_correct`Meaning
Correct - Intended1The rule works on the demonstrations and captures the abstraction the task was designed to test.
Correct - Unintended1The rule works on the demonstrations but does not capture the intended abstraction (for example a surface-level shortcut).
Incorrect0The rule does not work on the demonstrations.
Nonresponsive-1No usable rule was given.
Unclear-2The rule was too unclear to judge confidently.
Literal-3Human rules only: the rule describes the literal process of producing the output grid (for example clicking "copy" or selecting one specific colour) rather than the transformation.

The paper's figures group Nonresponsive, Unclear and Literal together as "Not Classified". Judgements were made by the authors. One annotator gave an initial judgement for each item, and ambiguous cases were discussed as a group until consensus was reached.

Loading

python
from datasets import load_dataset

models = load_dataset("ClaasBeger/ConceptARC_Rule_Annotations", "model_evaluations", split="train")
humans = load_dataset("ClaasBeger/ConceptARC_Rule_Annotations", "human_rule_annotations", split="train")
rules = load_dataset("ClaasBeger/ConceptARC_Rule_Annotations", "ground_truth_rules", split="train")

Citation

bibtex
@inproceedings{beger2026distinguishing,
  title     = {Distinguishing Performance From Competence in Evaluations of Humanlike Abstract Reasoning},
  author    = {Beger, Claas and Yi, Ryan and Fu, Shuhao and Denton, Kaleda and Moskvichev, Arseny and Tsai, Sarah W. and Rajamanickam, Sivasankaran and Mitchell, Melanie},
  booktitle = {Advances in Neural Information Processing Systems},
  year      = {2026}
}

Please also cite ConceptARC:

bibtex
@article{moskvichev2023conceptarc,
  title   = {The {ConceptARC} Benchmark: Evaluating Understanding and Generalization in the {ARC} Domain},
  author  = {Moskvichev, Arseny and Odouard, Victor Vikram and Mitchell, Melanie},
  journal = {Transactions on Machine Learning Research},
  year    = {2023}
}