Team Ai
Datasetpublic

jiosephlee/auxiliary-views-knowledge-acquisition

Auxiliary Views Knowledge Acquisition This repository contains the cleaned source documents and evaluation probes used in Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views (arXiv:2609.04180). News August 21, 2026: Our paper was accepted to Findings of EMNLP 2026. Configurations Configuration Split Rows documents train 30 factual_cloze test 6,435 factual_mcqa_5shot test 4,515… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/auxiliary-views-knowledge-acquisition.

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
1likes124downloads
Dataset Card

Auxiliary Views Knowledge Acquisition

This repository contains the cleaned source documents and evaluation probes used in Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views (arXiv:2609.04180).

News

  • —August 21, 2026: Our paper was accepted to Findings of EMNLP 2026.

Configurations

ConfigurationSplitRows
documentstrain30
factual_clozetest6,435
factual_mcqa_5shottest4,515
inference_clozetest430
inference_mcqa_5shottest322

Both multiple-choice configurations include five-shot prompts in formatted_question_5shot.

Copyright and source documents

This is a mixed-rights repository; see LICENSE.md and ATTRIBUTION.md. The 30 released documents comprise:

  • —6 arXiv papers licensed CC BY 4.0;
  • —12 BMJ Case Reports articles licensed CC BY-NC 4.0; and
  • —12 public-domain U.S. federal judicial opinions obtained via CourtListener.

Six cleaned arXiv texts are intentionally omitted:

Local documentarXiv IDReason
1_582402.17764arXiv distribution license does not grant a general downstream redistribution right
BOFT2311.06243same
GRPO2402.03300same
GSPO2507.18071same
OFT2306.07280same
LongRoPE2402.13753CC BY-NC-ND does not permit redistribution of the cleaned/modified text

Probe rows for all 36 source documents are included. Short third-party excerpts associated with omitted documents are included for research and criticism under an asserted copyright exception/fair-use rationale and are expressly excluded from the license granted over project-authored work. This is a risk-managed research release, not legal advice.

Intended use and limitations

The probes are evaluation data, not training/validation splits. Medical source documents are case reports and must not be used for diagnosis, treatment, or clinical decision-making. The corpus is small, English-only, intentionally heterogeneous, and is not representative of any of its three domains.

Citation

bibtex
@misc{lee2026knowledgeacquisitionpretraininglarge,
  title         = {Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views},
  author        = {Joseph Lee and Yidi Huang and Dokyoon Kim and Shu Yang and Li Shen},
  year          = {2026},
  eprint        = {2609.04180},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  url           = {https://arxiv.org/abs/2609.04180},
}