Team Ai
Datasetpublic

FreedomIntelligence/TCM-Pretrain-Data-ShizhenGPT

๐Ÿ“š Introduction This dataset is the pre-training dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source the largest existing TCM corpus dataset (over 5B tokens) from TCM-related websites and books. Additionally, we also open-source the largest scale TCM image-text pretraining dataset. For details, see our paper and GitHub repository. ๐Ÿ“Š Dataset Overview The open-sourced pre-training dataset consists of five parts:โ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TCM-Pretrain-Data-ShizhenGPT.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
10likes1.8kdownloads
Dataset Card

<span>๐Ÿ“š Introduction</span>

This dataset is the pre-training dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source the largest existing TCM corpus dataset (over 5B tokens) from TCM-related websites and books. Additionally, we also open-source the largest scale TCM image-text pretraining dataset.

For details, see our paper and GitHub repository.

<span>๐Ÿ“Š Dataset Overview</span>

The open-sourced pre-training dataset consists of five parts:

ModalityDescriptionData Quantity
TCM\Book\Corpus๐Ÿ“ TextA cleaned corpus of 3,256 TCM textbooks.\~ 0.5 B tokens
TCM\Web\Corpus๐Ÿ“ TextA TCM corpus collected from the web.Over 5B tokens
TCM\Book\Interleaved\_Data๐Ÿ“ Text, ๐Ÿ‘๏ธ VisualInterleaved text-image data from 306 TCM books.41459 entries, 50690 images
TCM\Web\Interleaved\_Data๐Ÿ“ Text, ๐Ÿ‘๏ธ VisualInterleaved text-image data from the TCM web corpus.505465 entries, 1143954 images
TCM\pretrain\synthesized\_vision๐Ÿ“ Text, ๐Ÿ‘๏ธ VisualTCM image-text pairs generated from images and their context using GPT-4o.144239 entries, 159534 images
โš ๏ธ Note: Due to privacy and ethical concerns, TCM signal datasets (e.g., sound and pulse) are not provided. For some signal data, refer to the Instruction Dataset.

<span>๐Ÿ“– Citation</span>

If you find our data useful, please consider citing our work!

@misc{chen2025shizhengptmultimodalllmstraditional,
      title={ShizhenGPT: Towards Multimodal LLMs for Traditional Chinese Medicine}, 
      author={Junying Chen and Zhenyang Cai and Zhiheng Liu and Yunjin Yang and Rongsheng Wang and Qingying Xiao and Xiangyi Feng and Zhan Su and Jing Guo and Xiang Wan and Guangjun Yu and Haizhou Li and Benyou Wang},
      year={2025},
      eprint={2508.14706},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2508.14706},
}