skylenage/ClinConsensus
ClinConsensus public dataset This public release contains the complete 900-case low-difficulty tier of ClinConsensus. Each de-identified Chinese clinical prompt has metadata and 30 case-specific binary rubric criteria, for 27,000 rubric criteria in total. The public file intentionally excludes reference answers, raw workbooks, source-row mappings, physician identifiers, QC notes, model responses, and judge transcripts. The complete 2,500-case benchmark is not part of this… See the full description on the dataset page: https://huggingface.co/datasets/skylenage/ClinConsensus.
ClinConsensus public dataset
This public release contains the complete 900-case low-difficulty tier of ClinConsensus. Each de-identified Chinese clinical prompt has metadata and 30 case-specific binary rubric criteria, for 27,000 rubric criteria in total.
The public file intentionally excludes reference answers, raw workbooks, source-row mappings, physician identifiers, QC notes, model responses, and judge transcripts. The complete 2,500-case benchmark is not part of this archive and is available from the corresponding authors under a separate DUA.
Files
data/clinconsensus_low.jsonl: 900 public evaluation cases.schema.json: field contract.summary.json: release counts.label_taxonomy.json: public-tier label vocabulary and counts.validation_report.json: structural and contact-pattern validation result.MANIFEST_SHA256.txt: checksums for the release files.
Record format
Each JSONL row contains:
case_id,difficulty;task,task_labels,subject,subject_labels,user_role;clinical_context,user_request;rubrics: exactly 30 objects withcriterion_id,criterion, andpoints.
The public package has no reference_answer field. Use the prompt and rubrics with a rubric-level grader, then compute Rubric Accuracy, Pass@10, and CACS@10 with the companion code repository.
Intended use and limitations
ClinConsensus is a research benchmark, not a clinical decision-support tool or deployment certification. Cases are de-identified and curated for evaluation; model outputs may still contain unsafe or incomplete medical content. Do not use benchmark results as patient-specific medical advice.
License and citation
The public dataset is released under CC BY 4.0; see LICENSE. Please cite:
@inproceedings{zheng2026clinconsensus,
title={ClinConsensus: A Physician-Calibrated Benchmark for Evaluating Clinical Rubric Coverage in Chinese Medical LLMs},
author={Zheng, Xiang and Li, Han and Luo, Wenjie and Zhai, Weiqi and Li, Yiyuan and Yan, Chuanmiao and Yang, Xue and Wu, Kailuan and Xu, Ruyi and Lu, Tianyun and Tang, Tianyi and Ma, Yubo and Yang, Kexin and Yang, Sen and Liu, Dayiheng and Qu, Lin and Zhao, Bing and Wei, Hu},
booktitle={Findings of the Association for Computational Linguistics: EMNLP 2026},
year={2026}
}