deep-learning-analytics/arxiv_small_nougat
Dataset Description The "arxiv_small_nougat" dataset is a collection of 108 recent papers sourced from arXiv, focusing on topics related to Large Language Models (LLM) and Transformers. These papers have been meticulously processed and parsed using Meta's Nougat model, which is specifically designed to retain the integrity of complex elements such as tables and mathematical equations. Data Format The dataset contains the parsed content of the selected papers, with… See the full description on the dataset page: https://huggingface.co/datasets/deep-learning-analytics/arxiv_small_nougat.
Dataset Description
The "arxivsmallnougat" dataset is a collection of 108 recent papers sourced from arXiv, focusing on topics related to Large Language Models (LLM) and Transformers. These papers have been meticulously processed and parsed using Meta's Nougat model, which is specifically designed to retain the integrity of complex elements such as tables and mathematical equations.
Data Format
The dataset contains the parsed content of the selected papers, with special attention given to the preservation of formatting, tables, and mathematical expressions. Each paper is provided as plain text.
Usage
Researchers, academics, and natural language processing practitioners can leverage this dataset for various tasks related to LLM and Transformers, including:
- Language modeling
- Text summarization
- Information retrieval
- Table and equation extraction
Acknowledgments
We acknowledge the arXiv platform for providing open access to a wealth of research papers in the field of machine learning and natural language processing.
License
[mit]
