Team Ai
Datasetpublic

woog/ai-paper-workflow-eval

AI Paper Workflow Evaluation Frozen GPT-6 Luna Flex outputs and historical research-paper controls. Evaluation only: keep test out of training, prompt development and threshold selection. 162 main papers: 54 validation and 108 test, stratified across NeurIPS, ICML and ACL, 2013–2021. The 27 development pilot papers are excluded. Author-name connected components do not cross splits; this is not perfect author identity resolution. from datasets import load_dataset # Pin revision… See the full description on the dataset page: https://huggingface.co/datasets/woog/ai-paper-workflow-eval.

sourceHugging Faceotherupdated 9d agoView on Hugging Face
0likes83downloads
Dataset Card

AI Paper Workflow Evaluation

Frozen GPT-6 Luna Flex outputs and historical research-paper controls. Evaluation only: keep test out of training, prompt development and threshold selection. 162 main papers: 54 validation and 108 test, stratified across NeurIPS, ICML and ACL, 2013–2021. The 27 development pilot papers are excluded. Author-name connected components do not cross splits; this is not perfect author identity resolution.

python
from datasets import load_dataset
# Pin revision to the upload commit recorded by your experiment.
ds = load_dataset("woog/ai-paper-workflow-eval", "reconstruction", revision=REVISION)
ConfigValidation / test rowsPurpose
reconstruction432 / 864One/two sentences, v3 paragraph, concise paragraph; target hidden from writer
human_controls554 / 5,117Matched originals and other human body paragraphs; use clean-novel flags for primary FPR
assistance324 / 648Proofread, light polish, substantial rewrite; diagnostic, no binary target gold
indexSee suite manifestText-free frozen row selections for every local profile; not a test split to score or train on

Paragraph and contextual views are correlated, not additional generations. Main reconstruction counts are 216 validation + 432 test generated targets; assistance counts are 162 + 324. Match originals via family_id, cluster by paper (author-component sensitivity), and report condition/view separately. Human controls deliberately retain extraction diagnostics; validation remaining controls are restricted to clean-novel prose.

regions are character [start,end) spans: 0 human, 1 generated replacement, -100 unknown/assisted. These are provenance labels, not measures of correctness. paper_id is stable; forum_id is null when not sourced from OpenReview. Quality reviews are by the same Luna model, not human gold; original verdicts and evidence-format corrections are retained in metadata_json. All first valid writer outputs remain, including quality failures and sentence-count mismatches. No detector scores selected this test set.

The index config inventories the older paper-v3 comparison, Arena, PELIC, Liang, VUB, Perkins, MELD-eval, DetectRL, Epoch, OpAI, Sem-Detect, GEDE, Saha, ELLIPSE and local diagnostic proxies. Those third-party texts are not mirrored here. ELLIPSE is upstream CC-BY-NC-SA-4.0. See SOURCES.md for acquisition choices. Existing public generated corpora: v3 paper pairs, Arena.

Local frozen bundle + checkpoint artifacts run offline through runner/suite.py (see runner/README.md). This HF dataset alone does not hydrate every third-party profile or download our checkpoint. Code/data/model hashes, fixed comparison-v1 thresholds and BF16 inference are recorded; changed inputs invalidate score caches. Same-environment rescoring is reproducible subject to GPU numerics; new Luna generation is stochastic. These are our benchmarks, not Pangram 4's exact private cohorts. See GENERATION.md and GENERATION_RESULTS.md for generation design and quality results.