datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
olympiad-books-open-source
olympiad-books-open-source
Chunked content from 12 open-source mathematics textbooks, suitable for retrieval (RAG), embedding, and math reasoning research.
Source code: github.com/yoonholee/olympiad-books-open-source-pipeline
Books
Book
Author(s)
License
Source
An Infinitely Large Napkin
Evan Chen
CC BY-SA 4.0 / GPL v3
GitHub
Mathematical Reasoning: Writing and Proof
Ted Sundstrom
CC BY-NC-SA 3.0
GitHub
Exploring Combinatorial Mathematics
Richard Grassl… See the full description on the dataset page: https://huggingface.co/datasets/yoonholee/olympiad-books-open-source.biosciences-sources
Biosciences RAG Source Documents
Dataset Description
This dataset contains 140 page-level document chunks extracted from 10 biomedical research papers. The documents form the knowledge base for a Retrieval-Augmented Generation (RAG) system focused on biosciences topics including knowledge graphs, LLM applications in biomedicine, and protein interaction databases.
Dataset Summary
Total Documents: 140 pages from 10 research papers
Domain: Biomedical NLP… See the full description on the dataset page: https://huggingface.co/datasets/open-biosciences/biosciences-sources.
