Team Ai
Datasetpublic

konsman/physics-corpus

Physics Corpus — konsman/physics-corpus arXiv physics papers exported from a PostgreSQL mirror of the Kaggle arXiv dataset, structured for ontology extraction and downstream NLP pipelines. Configuration: quantum-physics arXiv categories included: quant-ph, hep-th, gr-qc Schema Field Type Description paper_id string arxiv:<id>v<n> — stable across pipeline runs arxiv_id string Base arXiv ID without version suffix arxiv_version int32… See the full description on the dataset page: https://huggingface.co/datasets/konsman/physics-corpus.

sourceHugging Facecc-by-4.0updated 7mo agoView on Hugging Face
1likes5.2kdownloads
Dataset Card

Physics Corpus — konsman/physics-corpus

arXiv physics papers exported from a PostgreSQL mirror of the Kaggle arXiv dataset, structured for ontology extraction and downstream NLP pipelines.

Configuration: quantum-physics

arXiv categories included: quant-ph, hep-th, gr-qc

Schema

FieldTypeDescription
paper_idstringarxiv:<id>v<n> — stable across pipeline runs
arxiv_idstringBase arXiv ID without version suffix
arxiv_versionint32Version number (for deduplication)
doistringDOI or empty string
content_hashstringSHA-256 of full_text — for delta detection
titlestringPaper title
abstractstringPaper abstract
authorslist[string]Author names
arxiv_categorieslist[string]e.g. ["hep-th", "gr-qc"]
primary_categorystringFirst listed category
journalstringJournal reference or empty string
keywordslist[string]Keywords (empty if unavailable)
submission_datestringISO date of first submission
processed_datestringISO date this record was exported
full_textstringAll section texts joined by \n\n
sectionslistStructured sections (see below)

Section fields

FieldTypeDescription
section_typestringTITLE / ABSTRACT / BACKGROUND / METHOD / RESULTS / DISCUSSION / CONCLUSION / OTHER
section_titlestringRaw heading text, e.g. 3. Experimental Setup
textstringSection body text

Record structure

Each row is one paper. sections is a list of objects — one per section:

json
{
  "paper_id": "arxiv:2401.00123v2",
  "arxiv_id": "2401.00123",
  "arxiv_version": 2,
  "title": "...",
  "abstract": "...",
  "authors": ["Author A", "Author B"],
  "arxiv_categories": ["quant-ph"],
  "primary_category": "quant-ph",
  "full_text": "Introduction text\n\nMethod text\n\n...",
  "sections": [
    {"section_type": "BACKGROUND", "section_title": "1. Introduction", "text": "..."},
    {"section_type": "METHOD",     "section_title": "2. Methods",      "text": "..."},
    {"section_type": "RESULTS",    "section_title": "3. Results",      "text": "..."}
  ]
}

Usage

python
from datasets import load_dataset

ds = load_dataset("konsman/physics-corpus", "quantum-physics", split="train")

for paper in ds:
    print(paper["paper_id"], paper["primary_category"])
    for sec in paper["sections"]:
        print(sec["section_type"], sec["section_title"], sec["text"][:80])

Source

Papers are sourced from the Kaggle arXiv dataset via a PostgreSQL mirror. Only papers with a JOIN match in the metadata table are included (≥ 96% of sections have a match).

License

Dataset card and structure: CC-BY-4.0. Paper texts are subject to the individual arXiv paper licenses.