Team Ai
Datasetpublic

AlgorithmicResearchGroup/s2orc_arxiv

S2ORC ArXiv A subset of the Semantic Scholar Open Research Corpus (S2ORC) filtered to ArXiv papers. Contains 2.58 million parsed scientific papers with full text, abstracts, structured sections, figures, and citation metadata. Dataset Summary Statistic Value Total papers 2,579,762 Total size ~266 GB Format Parquet Split train Dataset Structure Content Fields Field Type Description title string Paper… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/s2orc_arxiv.

sourceHugging Faceupdated 6mo agoView on Hugging Face
2likes2kdownloads
settings

This repository belongs to AlgorithmicResearchGroup on Hugging Face.

Team Ai never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

names2orc_arxiv
visibilitypublic
licencenot set
gatedno
ownerAlgorithmicResearchGroup
Account settings
AlgorithmicResearchGroup/s2orc_arxiv · Team Ai