Team Ai
Datasetpublic

skylenage/ClinConsensus

ClinConsensus public dataset This public release contains the complete 900-case low-difficulty tier of ClinConsensus. Each de-identified Chinese clinical prompt has metadata and 30 case-specific binary rubric criteria, for 27,000 rubric criteria in total. The public file intentionally excludes reference answers, raw workbooks, source-row mappings, physician identifiers, QC notes, model responses, and judge transcripts. The complete 2,500-case benchmark is not part of this… See the full description on the dataset page: https://huggingface.co/datasets/skylenage/ClinConsensus.

sourceHugging Facecc-by-4.0updated 1mo agoView on Hugging Face
0likes76downloads
Dataset Card

ClinConsensus public dataset

This public release contains the complete 900-case low-difficulty tier of ClinConsensus. Each de-identified Chinese clinical prompt has metadata and 30 case-specific binary rubric criteria, for 27,000 rubric criteria in total.

The public file intentionally excludes reference answers, raw workbooks, source-row mappings, physician identifiers, QC notes, model responses, and judge transcripts. The complete 2,500-case benchmark is not part of this archive and is available from the corresponding authors under a separate DUA.

Files

  • —data/clinconsensus_low.jsonl: 900 public evaluation cases.
  • —schema.json: field contract.
  • —summary.json: release counts.
  • —label_taxonomy.json: public-tier label vocabulary and counts.
  • —validation_report.json: structural and contact-pattern validation result.
  • —MANIFEST_SHA256.txt: checksums for the release files.

Record format

Each JSONL row contains:

  • —case_id, difficulty;
  • —task, task_labels, subject, subject_labels, user_role;
  • —clinical_context, user_request;
  • —rubrics: exactly 30 objects with criterion_id, criterion, and points.

The public package has no reference_answer field. Use the prompt and rubrics with a rubric-level grader, then compute Rubric Accuracy, Pass@10, and CACS@10 with the companion code repository.

Intended use and limitations

ClinConsensus is a research benchmark, not a clinical decision-support tool or deployment certification. Cases are de-identified and curated for evaluation; model outputs may still contain unsafe or incomplete medical content. Do not use benchmark results as patient-specific medical advice.

License and citation

The public dataset is released under CC BY 4.0; see LICENSE. Please cite:

bibtex
@inproceedings{zheng2026clinconsensus,
  title={ClinConsensus: A Physician-Calibrated Benchmark for Evaluating Clinical Rubric Coverage in Chinese Medical LLMs},
  author={Zheng, Xiang and Li, Han and Luo, Wenjie and Zhai, Weiqi and Li, Yiyuan and Yan, Chuanmiao and Yang, Xue and Wu, Kailuan and Xu, Ruyi and Lu, Tianyun and Tang, Tianyi and Ma, Yubo and Yang, Kexin and Yang, Sen and Liu, Dayiheng and Qu, Lin and Zhao, Bing and Wei, Hu},
  booktitle={Findings of the Association for Computational Linguistics: EMNLP 2026},
  year={2026}
}