Team Ai
Datasetpublic

brianshumate/chat-jfk

Dataset Card for Chat JFK The JFK files as Markdown formatted text. Dataset Details Dataset Description This dataset represents the conversion of PDF files from the U.S. National Archives President John F. Kennedy Assassination Records Collection JFK Assassination Bulk Download Files into Markdown formatted files. Curated by: Brian Shumate Language: English License: CC0 Dataset Sources The Chat JFK dataset is comprised of Markdown… See the full description on the dataset page: https://huggingface.co/datasets/brianshumate/chat-jfk.

sourceHugging Facecc0-1.0updated 3mo agoView on Hugging Face
0likes7.2kdownloads
Dataset Card

Dataset Card for Chat JFK

The JFK files as Markdown formatted text.

Dataset Details

Dataset Description

This dataset represents the conversion of PDF files from the U.S. National Archives President John F. Kennedy Assassination Records Collection JFK Assassination Bulk Download Files into Markdown formatted files.

  • —Curated by: Brian Shumate
  • —Language: English
  • —License: CC0

Dataset Sources

The Chat JFK dataset is comprised of Markdown formatted files converted from the following sources.

Base URL: https://www.archives.gov/files/research/jfk/releases/zip/

Total for 29 zip files holding PDF files: ~63 GB.

[!NOTE] This bulk file resource also includes ZIP files containing WAV audio, but those data were not downloaded, transcribed or otherwise included in this dataset.
July 24 and October 26, 2017 Releases

Total: 7.8 GB

StateFileSizeMD5
[✓]jfk-pdf1.zip2.9 GBMD5: 80e2188593fca6467acc1c4214f3ba9b
[✓]jfk-pdf2.zip2.4 GBMD5: 55af7bb87eb760c9e6dead830a97fc7c
[✓]jfk-pdf3.zip2.5 GBMD5: ae800b0b61302e0efe82b3014700d911
November 3, 2017 Release

Total: 2.6 GB

StateFileSizeMD5
[✓]jfk20171103.zip2.6 GBMD5: 6b4ab22682c8827a4dc25906de9814ac
November 9, 2017 Release

Total: 3.2 GB

StateFileSizeMD5
[✓]jfk20171109.zip3.2 GBMD5: 96080aff5e40925b3330c9169c70cfc9
November 17, 2017 Release

Total: 3.7 GB

StateFileSizeMD5
[✓]jfk2017111710.zip3.7 GBMD5: f619fabd567e4ffa4e4176107322f0f8
December 15, 2017 Release

Total: 6.1 GB

StateFileSizeMD5
[✓]jfk20171215a.zip2.0 GBMD5: d6928b8efc1fcd6f301c0b89e850173d
[✓]jfk20171215b.zip2.2 GBMD5: d9e719011679b755c4f9f22fe9956640
[✓]jfk20171215c.zip1.9 GBMD5: d664566a0c925464bcd9349334591d71
April 26, 2018 Release

Total: 10.81 GB

StateFileSizeMD5
[✓]jfk201804a.zip1.1GMD5: 74f0fffb568724abda4a0f02ae8cb825
[✓]jfk201804b.zip1.0GMD5: ffaf1968028d3e268127eb9e4ec4aadd
[✓]jfk201804c.zip1.2GMD5: 7ebf15c13c0605e7ce3b83ce3ba90770
[✓]jfk201804d.zip1.1GMD5: 3ac6abb1f305de7fe8b648a71ff99809
[✓]jfk201804e.zip1.1GMD5: 6e53d9d0a31c2c6e89b6513257b7661a
[✓]jfk201804f.zip1.3GMD5: cb66f683d1927984b3fb2f20f8be9a0c
[✓]jfk201804g.zip1.0GMD5: 5501944cc325397692d3f8d6f9bfde2f
[✓]jfk201804h.zip923MMD5: 38e252e058c1627393e4542fc328e7e2
[✓]jfk201804i.zip987MMD5: 25cb16fd243ecea6f6cc3b5bf2bb0295
[✓]jfk201804j.zip1.1GMD5: 3f2742c5542e71ddf28865f2c8a180ae
December 15, 2021 Release

Total: 1.2 GB

StateFileSizeMD5
[✓]jfk2021.zip1.2GMD5: 3b41d5d25a211c681a8c5f79e3720b70
December 15, 2022 Release

Total: 12.7 GB

StateFileSizeMD5
[✓]jfk2022.zip12.7GMD5: 350323648f093d3aa7b204c5445e33da
April 13, 2023 Release

Total: 339 MB

StateFileSizeMD5
[✓]jfk2023a.zip339MMD5: 5c5f2eb2db9b259362e1437909892e97
April 27, 2023 Release

Total: 343 MB

StateFileSizeMD5
[✓]jfk2023b.zip343MMD5: 6fc670f350cd3f32a7f53aaf72ffe443
May 11, 2023 Release

Total: 554 MB

StateFileSizeMD5
[✓]jfk2023c.zip554MMD5: 907cbb9a2b9a3f48069a9bb90d152829
June 13, 2023 Release

Total: 238 MB

StateFileSizeMD5
[✓]jfk2023d.zip238MMD5: 4452423375cd3f885677ff380c7af365
June 27, 2023 Release

Total: 4.2 GB

StateFileSizeMD5
[✓]jfk2023e.zip4.2GMD5: b48fefd4f94f23b19f8aa890921e53b4
August 24, 2023 Release

Total: 74 MB

StateFileSizeMD5
[✓]jfk2023f.zip74MMD5: b42f95619706901b69cf2e104747009d
2025 Release

Total: 8.5 GB

StateFileSizeMD5
[✓]jfk2025a.zip5.2GMD5: 3c7c789a6cc477ee819987dcaeb64beb-664
[✓]jfk2025b.zip3.3GMD5: 77ae4549743c0d2a86b043d93c5f404f-425

Uses

This dataset was curated primarily for educational and informational use cases.

Direct Use

The data has multiple potential uses:

  • —Model fine-tuning or training.
  • —Augmentation use cases (RAG, etc.)
  • —Element source for graph or time series data.

Out-of-Scope Use

Do not use the dataset to conduct unethical or illegal activities of any kind.

Do not use the dataset to train conversational agents intended to provide legal advice of any kind.

Dataset Structure

Markdown formatted text files with some inline HTML elements.

Dataset Creation

Curation Rationale

Chatting with the files through various augmentation strategies is the primary motivator for curating these data. Searching, connecting, linking information in ways that agentic workflows unlock seems like an fun pastime and combination with multiple or multi-mode models can unlock exciting ways to consume the historical information such as:

  • —TTS models can use the data to generate audio journals, podcast like content or other auditory walkthroughs of the content.
  • —Creative exercises such as period appropriate agentic investigative reporting on the content through purpose built agent environments.
  • —Generation of educational and research content, such as reports, timelines and other historical facts.

Source Data

PDF files from the U.S. National Archives President John F. Kennedy Assassination Records Collection JFK Assassination Bulk Download Files.

Data Collection and Processing

The data were generated by running MinerU against the PDF collection on a single NVIDIA RTX 3090 Founders Edition GPU.

The resulting Markdown files were then collected and uploaded.

The dataset represents raw data without any clean up whatsoever.

While the accuracy of the transcription of text from the PDF files to Markdown is close in some cases, but some examples of hallucination almost certainly appear in the dataset.

Always reference and refer the original document to confirm content before using it for anything important.

[!NOTE] Some folders have been split into two parts while preserving the original

naming and appending '-1' or '-2' due to the 10000 file per-folder limitation.

Data quality notes
  • —Keep in mind that the source material itself contains numerous examples of content duplication.
  • —The curator of this datatset performed no cleanup or deduplication to preserve the original raw output from the tooling as another data point for research.
  • —There are no guarantees about the quality of the dataset; review Recommendations for more information.
Who are the source data producers?

Brian Shumate is an engineer and recreational researcher using open models and modest systems to test the limitations of current state of the art open model and open source technologies.

Personal and Sensitive Information

The data is comprised of historical and public domain information released by the United States government.

Bias, Risks, and Limitations

This historical dataset is for educational use only.

Recommendations

The dataset is not to be considered 100% accurate or fully complete, and no warranties expressed or implied exist for the content quality. All efforts were made to capture the data with the described tooling and capture the raw outputs.

You use the dataset at your own risk, and should do so only after conducting your own due diligence and research as to the dataset's usability for your purposes.

The dataset producer and curator named on this dataset card take no responsibility for the consequences which may arise from any use or misuse of the dataset under any circumstances.

Dataset Card Contact

Brian Shumate

Citations

@article{wang2026mineru2, title={MinerU2. 5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale}, author={Wang, Bin and He, Tianyao and Ouyang, Linke and Wu, Fan and Zhao, Zhiyuan and Chu, Tao and Qu, Yuan and Jin, Zhenjiang and Zeng, Weijun and Miao, Ziyang and others}, journal={arXiv preprint arXiv:2604.04771}, year={2026} }

@article{dong2026minerudiffusion, title={MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding}, author={Dong, Hejun and Niu, Junbo and Wang, Bin and Zeng, Weijun and Zhang, Wentao and He, Conghui}, journal={arXiv preprint arXiv:2603.22458}, year={2026} }

@article{niu2025mineru2, title={Mineru2. 5: A decoupled vision-language model for efficient high-resolution document parsing}, author={Niu, Junbo and Liu, Zheng and Gu, Zhuangcheng and Wang, Bin and Ouyang, Linke and Zhao, Zhiyuan and Chu, Tao and He, Tianyao and Wu, Fan and Zhang, Qintong and others}, journal={arXiv preprint arXiv:2509.22186}, year={2025} }

@article{wang2024mineru, title={Mineru: An open-source solution for precise document content extraction}, author={Wang, Bin and Xu, Chao and Zhao, Xiaomeng and Ouyang, Linke and Wu, Fan and Zhao, Zhiyuan and Xu, Rui and Liu, Kaiwen and Qu, Yuan and Shang, Fukai and others}, journal={arXiv preprint arXiv:2409.18839}, year={2024} }

@article{he2024opendatalab, title={Opendatalab: Empowering general artificial intelligence with open datasets}, author={He, Conghui and Li, Wei and Jin, Zhenjiang and Xu, Chao and Wang, Bin and Lin, Dahua}, journal={arXiv preprint arXiv:2407.13773}, year={2024} }

brianshumate/chat-jfk · Team Ai