ClaasBeger/ConceptARC_Rule_Annotations
ConceptARC Rule Annotations Model outputs, natural-language rules and human judgements of those rules on the 480 tasks of the ConceptARC benchmark. This is the data behind the paper Distinguishing Performance From Competence in Evaluations of Humanlike Abstract Reasoning (NeurIPS 2026, Evaluations and Datasets Track). Paper: arXiv:2510.02125 Project page and interactive viewer: claasbeger.github.io/performance-competence-gap Authors: Claas Beger, Ryan Yi, Shuhao Fu, Kaleda… See the full description on the dataset page: https://huggingface.co/datasets/ClaasBeger/ConceptARC_Rule_Annotations.
ConceptARC Rule Annotations
Model outputs, natural-language rules and human judgements of those rules on the 480 tasks of the ConceptARC benchmark. This is the data behind the paper Distinguishing Performance From Competence in Evaluations of Humanlike Abstract Reasoning (NeurIPS 2026, Evaluations and Datasets Track).
- Paper: arXiv:2510.02125
- Project page and interactive viewer: claasbeger.github.io/performance-competence-gap
- Authors: Claas Beger, Ryan Yi, Shuhao Fu, Kaleda Denton, Arseny Moskvichev, Sarah W. Tsai, Sivasankaran Rajamanickam, Melanie Mitchell
- Contact: claasbeger@santafe.edu
Summary
Accuracy alone can overstate or understate how well a model reasons with the abstractions a benchmark was designed to test. For each ConceptARC task we collected the model's output grid together with the natural-language rule it stated, and a team of annotators judged whether that rule captures the intended abstraction. The same judgements were made for rules written by human participants.
Models were evaluated with textual and visual inputs, at low and medium reasoning effort, and with and without Python tools. ConceptARC covers 16 spatial and semantic concepts with 10 tasks per concept, and each task has 3 test inputs (480 test items in total).
Contents
Model evaluation columns
Human rule columns
Rule judgements
The paper's figures group Nonresponsive, Unclear and Literal together as "Not Classified". Judgements were made by the authors. One annotator gave an initial judgement for each item, and ambiguous cases were discussed as a group until consensus was reached.
Loading
from datasets import load_dataset
models = load_dataset("ClaasBeger/ConceptARC_Rule_Annotations", "model_evaluations", split="train")
humans = load_dataset("ClaasBeger/ConceptARC_Rule_Annotations", "human_rule_annotations", split="train")
rules = load_dataset("ClaasBeger/ConceptARC_Rule_Annotations", "ground_truth_rules", split="train")Citation
@inproceedings{beger2026distinguishing,
title = {Distinguishing Performance From Competence in Evaluations of Humanlike Abstract Reasoning},
author = {Beger, Claas and Yi, Ryan and Fu, Shuhao and Denton, Kaleda and Moskvichev, Arseny and Tsai, Sarah W. and Rajamanickam, Sivasankaran and Mitchell, Melanie},
booktitle = {Advances in Neural Information Processing Systems},
year = {2026}
}Please also cite ConceptARC:
@article{moskvichev2023conceptarc,
title = {The {ConceptARC} Benchmark: Evaluating Understanding and Generalization in the {ARC} Domain},
author = {Moskvichev, Arseny and Odouard, Victor Vikram and Mitchell, Melanie},
journal = {Transactions on Machine Learning Research},
year = {2023}
}