jiosephlee/auxiliary-views-knowledge-acquisition
Auxiliary Views Knowledge Acquisition This repository contains the cleaned source documents and evaluation probes used in Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views (arXiv:2609.04180). News August 21, 2026: Our paper was accepted to Findings of EMNLP 2026. Configurations Configuration Split Rows documents train 30 factual_cloze test 6,435 factual_mcqa_5shot test 4,515… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/auxiliary-views-knowledge-acquisition.
Auxiliary Views Knowledge Acquisition
This repository contains the cleaned source documents and evaluation probes used in Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views (arXiv:2609.04180).
News
- August 21, 2026: Our paper was accepted to Findings of EMNLP 2026.
Configurations
Both multiple-choice configurations include five-shot prompts in formatted_question_5shot.
Copyright and source documents
This is a mixed-rights repository; see LICENSE.md and ATTRIBUTION.md. The 30 released documents comprise:
- 6 arXiv papers licensed CC BY 4.0;
- 12 BMJ Case Reports articles licensed CC BY-NC 4.0; and
- 12 public-domain U.S. federal judicial opinions obtained via CourtListener.
Six cleaned arXiv texts are intentionally omitted:
Probe rows for all 36 source documents are included. Short third-party excerpts associated with omitted documents are included for research and criticism under an asserted copyright exception/fair-use rationale and are expressly excluded from the license granted over project-authored work. This is a risk-managed research release, not legal advice.
Intended use and limitations
The probes are evaluation data, not training/validation splits. Medical source documents are case reports and must not be used for diagnosis, treatment, or clinical decision-making. The corpus is small, English-only, intentionally heterogeneous, and is not representative of any of its three domains.
Citation
@misc{lee2026knowledgeacquisitionpretraininglarge,
title = {Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views},
author = {Joseph Lee and Yidi Huang and Dokyoon Kim and Shu Yang and Li Shen},
year = {2026},
eprint = {2609.04180},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2609.04180},
}