konsman/physics-corpus
Physics Corpus — konsman/physics-corpus arXiv physics papers exported from a PostgreSQL mirror of the Kaggle arXiv dataset, structured for ontology extraction and downstream NLP pipelines. Configuration: quantum-physics arXiv categories included: quant-ph, hep-th, gr-qc Schema Field Type Description paper_id string arxiv:<id>v<n> — stable across pipeline runs arxiv_id string Base arXiv ID without version suffix arxiv_version int32… See the full description on the dataset page: https://huggingface.co/datasets/konsman/physics-corpus.
Physics Corpus — konsman/physics-corpus
arXiv physics papers exported from a PostgreSQL mirror of the Kaggle arXiv dataset, structured for ontology extraction and downstream NLP pipelines.
Configuration: quantum-physics
arXiv categories included: quant-ph, hep-th, gr-qc
Schema
Section fields
Record structure
Each row is one paper. sections is a list of objects — one per section:
{
"paper_id": "arxiv:2401.00123v2",
"arxiv_id": "2401.00123",
"arxiv_version": 2,
"title": "...",
"abstract": "...",
"authors": ["Author A", "Author B"],
"arxiv_categories": ["quant-ph"],
"primary_category": "quant-ph",
"full_text": "Introduction text\n\nMethod text\n\n...",
"sections": [
{"section_type": "BACKGROUND", "section_title": "1. Introduction", "text": "..."},
{"section_type": "METHOD", "section_title": "2. Methods", "text": "..."},
{"section_type": "RESULTS", "section_title": "3. Results", "text": "..."}
]
}Usage
from datasets import load_dataset
ds = load_dataset("konsman/physics-corpus", "quantum-physics", split="train")
for paper in ds:
print(paper["paper_id"], paper["primary_category"])
for sec in paper["sections"]:
print(sec["section_type"], sec["section_title"], sec["text"][:80])Source
Papers are sourced from the Kaggle arXiv dataset via a PostgreSQL mirror. Only papers with a JOIN match in the metadata table are included (≥ 96% of sections have a match).
License
Dataset card and structure: CC-BY-4.0. Paper texts are subject to the individual arXiv paper licenses.
