datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
text-code-galeras-code-generation-from-docstring-3k-dedupeddocstech-docs
Technical Documentation Dataset
A curated collection of technical documentation and guides spanning various cloud-native technologies, infrastructure tools, and machine learning frameworks. This dataset contains 1,397 documents in JSONL format, covering essential topics for modern software development and DevOps practices.
Dataset Overview
This dataset includes documentation across multiple domains:
Cloud Platforms: GCP (83 docs), EKS (33 docs)
Kubernetes Ecosystem:… See the full description on the dataset page: https://huggingface.co/datasets/saidsef/tech-docs.mil-docs
What is this?
A curated selection of manuals and documents from the US military and other departments. All data was manually scraped from publicly available sources.
The PDF's and EPUB files were converted to markdown using the amazing Marker github repository by Vik Paruchuri.
Sources:
United States Army Central Army Repository
Marines Publications
Federation of American Scientists Intelligence Resource Program
dolma3-6t-sample-10000-docs-finance-and-business
HCAI-Lab/dolma3-6t-sample-10000-docs-finance-and-business
Filename-derived finance_and_business slice of
HCAI-Lab/dolma3-6t-sample-10000-docs, pinned to
revision 561e73c7e0ad35c04f386bae1e3dd39dfb6755e7.
Extraction rule
The corpus contains every source .jsonl.zst file whose filename contains
the literal segment -finance_and_business-. Source paths and compressed file
contents are preserved byte-for-byte. This is a coarse WebOrganizer
finance_and_business category… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-10000-docs-finance-and-business.29K_Python_Docstring_Pairs
29K High-Quality Python Docstring Pairs
Author: Michael Hernandez (XxCotHGxX)License: CC BY 4.0Cleaned from: XxCotHGxX/242K_Python_Docstring_Pairs
Overview
A curated, high-quality subset of Python function–docstring pairs for use in code documentation generation, docstring completion, and code understanding tasks.
The original 242K dataset was scraped from open-source Python repositories but contained a significant proportion of functions without docstrings (84% of… See the full description on the dataset page: https://huggingface.co/datasets/XxCotHGxX/29K_Python_Docstring_Pairs.godot_4_docsDataset generated for Godot 4 docs using Glaive.
azure_docs_fullMP-DocStruct1MMP-DocStruct1M is a Multi-page Document Paring and Lookup training set used in DocOwl2
run following commands to prepare images folder ./imgs
cat partial-imgs* > imgs.zip
unzip imgs.zip
datause-extracted-human473-docs
datause-extracted-human473-docs
Every passage of the 162 documents behind the 473 human-validated
holdout spans of the data-use annotation campaign:
population
spans
documents
annotator190
190
134
jdc283
283
28
total
473
162
Configs
gliner, bio, gliner2 — row-for-row subset of
rafmacalaba/datause-extracted
(revision 15812843687e2ec81b261b5895f2019b50a8f97e): same columns, same split files, rows verbatim. Rows per
config: gliner/train 1963… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/datause-extracted-human473-docs.lfqa_support_docsSupport documents for building https://huggingface.co/vblagoje/bart_lfqa model
halo-docs
Halo Documentation Q&A
English, single-turn instruction-tuning examples about the Halo LLM training toolkit: concepts, configurations, commands, model recipes, internals, and troubleshooting.
This is a source-preserving, extractive Q&A dataset. Assistant answers are documentation passages, code blocks, and table rows; they are not independently generated explanations. Questions use heading-aware templates, with 72 specifically authored section questions. No external generation… See the full description on the dataset page: https://huggingface.co/datasets/skundu42/halo-docs.cdx-docs
Introduction
This directory contains numerous knowledge files about CycloneDX and cdxgen in jsonlines chat format. The data is useful for training and fine-tuning (LoRA and QLoRA) LLM models.
Data Generation
We used Google Gemini 2.0 Flash Experimental via aistudio and used the below prompts to convert official documentation markdown files to the chat format.
you are an expert in converting markdown files to plain text jsonlines format based on the my template.… See the full description on the dataset page: https://huggingface.co/datasets/CycloneDX/cdx-docs.adaption-louisville-data-center-docs-augmented
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-louisville_data_center_docs (augmented)
This dataset contains planning commission staff reports, zoning code excerpts, and news transcripts regarding hyperscale data center developments in Louisville, Kentucky. The documents detail specific project proposals, such as the Camp Ground Road facility, including technical reviews on traffic, water usage, and environmental impact.… See the full description on the dataset page: https://huggingface.co/datasets/JaySmith502/adaption-louisville-data-center-docs-augmented.code-code-galeras-code-completion-from-docstring-3k-dedupeddocspider
DocSpider: a Dataset of Cross-Domain Natural Language Querying for MongoDB
Arif Görkem Özer, Fırat Çekinel, Pınar Karagöz, İsmail Hakkı Toroslu
You can access the paper published in Natural Language Processing journal, from this link.
DocSpider dataset is generated by using the widely-known text-to-SQL dataset, Spider.
See GitHub repository for more details, including benchmark pipeline scripts for text to MongoDB query conversion.
Overview
This repository… See the full description on the dataset page: https://huggingface.co/datasets/gorkemozer/docspider.godot_4_docsDataset generated for Godot 4 docs using Glaive.
langchain-docs-400-chunksizepython_code_docstring_ast_corpus
Overview
This dataset contains 34,000+ rows of code-docstring-ast data along with additional metadata. Data was gathered from various Python libraries and frameworks and their
publicly available GitHub repos. This dataset was created for the purpose of training the CodeT5+ transformer on AST-enhanced code-to-doc tasks.
Sources
The dataset was gathered from various GitHub repos sampled from this repo by Vinta.
The 26 repos are:
matplotlib
pytorch
cryptography
django… See the full description on the dataset page: https://huggingface.co/datasets/Mir-2002/python_code_docstring_ast_corpus.sealevel-docsmongodb-docs
Overview
This dataset consists of a small subset of MongoDB's technical documentation.
Dataset Structure
The dataset consists of the following fields:
sourceName: The source of the document.
url: Link to the article.
action: Action taken on the article.
body: Content of the article in Markdown format.
format: Format of the content.
metadata: Metadata such as tags, content type etc. associated with the document.
title: Title of the document.
updated: The last updated… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/mongodb-docs.godot_docspy-docs-2004
Python Docs 2004
Original dump: https://www.python.org/ftp/python/doc/
Python Docs 2004 is a filtered and cleaned collection of Python documentation from every major Python release published before 2004.
Stats
Version
Size
Lines
2.3
2.2MB
1215
2.2
1.7MB
1142
2.1
1.3MB
891
2.0
1.2MB
895
1.6
1MB
720
1.5
837KB
449
1.4
744KB
397
1.3
569KB
408
1.2
513KB
384
Total
10.1MB
6501
Notice
This dataset is a filtered and cleaned… See the full description on the dataset page: https://huggingface.co/datasets/fromziro/py-docs-2004.manim-docs-gitingest
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/ArunKr/manim-docs-gitingest.docs
MongoDB Public Documentation
Public MongoDB documentation and developer blog posts.
Documentation represented as Markdown. In-line links are omitted.
Sources:
https://mongodb.com/docs
Certain pages are omitted because we do not want models trained on them.
https://mongodb.com/developer
k8s-docs-rag-bench
k8s-docs-rag-bench
Paper: Analyzing Quality--Latency--Resource Trade-offs in a Technical Documentation RAG Assistant Using LoRA Adaptation (arXiv:2605.28222)
Code: github.com/EugPal/rag-lora-tradeoffs
A small, fully-grounded benchmark for retrieval-augmented question answering
(RAG) over the official Kubernetes documentation, together with the full
set of LLM-judge labels used in the accompanying preprint
"Analyzing Quality-Latency-Resource Trade-offs in a Technical… See the full description on the dataset page: https://huggingface.co/datasets/evgenypal/k8s-docs-rag-bench.hf-docs-benchmark-lightai-docs-agent
AI Docs Agent Meta and Traffic Dataset in AI Agent Marketplace | AI Agent Directory | AI Agent Index from DeepNLP
This dataset is collected from AI Agent Marketplace Index and Directory at http://www.deepnlp.org, which contains AI Agents's meta information such as agent's name, website, description, as well as the monthly updated Web performance metrics, including Google,Bing average search ranking positions, Github Stars, Arxiv References, etc.
The dataset is helpful for AI… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/ai-docs-agent.FW_EDU_SUBSET_500k_docs
FineWeb-Edu Subset
This dataset contains 483,606 documents sampled from the FineWeb-Edu dataset.
The dataset is used throughout various tutorials on modalities.
For licensing, see their conditions.
arduino-docs
