Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01semeru /text-code-galeras-code-generation-from-docstring-3k-dedupedtabular1K<n<10K0 likes607 downloads3y agoHugging Face02gradio /docstext1K<n<10K3 likes385 downloads18h agoHugging Face03saidsef /tech-docs Technical Documentation Dataset A curated collection of technical documentation and guides spanning various cloud-native technologies, infrastructure tools, and machine learning frameworks. This dataset contains 1,397 documents in JSONL format, covering essential topics for modern software development and DevOps practices. Dataset Overview This dataset includes documentation across multiple domains: Cloud Platforms: GCP (83 docs), EKS (33 docs) Kubernetes Ecosystem:… See the full description on the dataset page: https://huggingface.co/datasets/saidsef/tech-docs.textquestion-answering1K<n<10K2 likes258 downloads2y agoHugging Face04AquaV /mil-docs What is this? A curated selection of manuals and documents from the US military and other departments. All data was manually scraped from publicly available sources. The PDF's and EPUB files were converted to markdown using the amazing Marker github repository by Vik Paruchuri. Sources: United States Army Central Army Repository Marines Publications Federation of American Scientists Intelligence Resource Program text1K<n<10K2 likes245 downloads3y agoHugging Face05HCAI-Lab-GT /dolma3-6t-sample-10000-docs-finance-and-business HCAI-Lab/dolma3-6t-sample-10000-docs-finance-and-business Filename-derived finance_and_business slice of HCAI-Lab/dolma3-6t-sample-10000-docs, pinned to revision 561e73c7e0ad35c04f386bae1e3dd39dfb6755e7. Extraction rule The corpus contains every source .jsonl.zst file whose filename contains the literal segment -finance_and_business-. Source paths and compressed file contents are preserved byte-for-byte. This is a coarse WebOrganizer finance_and_business category… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-10000-docs-finance-and-business.texttext-generation100K<n<1M0 likes233 downloads2mo agoHugging Face06XxCotHGxX /29K_Python_Docstring_Pairs 29K High-Quality Python Docstring Pairs Author: Michael Hernandez (XxCotHGxX)License: CC BY 4.0Cleaned from: XxCotHGxX/242K_Python_Docstring_Pairs Overview A curated, high-quality subset of Python function–docstring pairs for use in code documentation generation, docstring completion, and code understanding tasks. The original 242K dataset was scraped from open-source Python repositories but contained a significant proportion of functions without docstrings (84% of… See the full description on the dataset page: https://huggingface.co/datasets/XxCotHGxX/29K_Python_Docstring_Pairs.texttext-generation10K<n<100K0 likes201 downloads8mo agoHugging Face07glaiveai /godot_4_docsDataset generated for Godot 4 docs using Glaive. text1K<n<10K21 likes181 downloads2y agoHugging Face08Mazino0 /azure_docs_fulltext10K<n<100K0 likes155 downloads2y agoHugging Face09mPLUG /MP-DocStruct1MMP-DocStruct1M is a Multi-page Document Paring and Lookup training set used in DocOwl2 run following commands to prepare images folder ./imgs cat partial-imgs* > imgs.zip unzip imgs.zip text1M<n<10M5 likes153 downloads2y agoHugging Face10rafmacalaba /datause-extracted-human473-docs datause-extracted-human473-docs Every passage of the 162 documents behind the 473 human-validated holdout spans of the data-use annotation campaign: population spans documents annotator190 190 134 jdc283 283 28 total 473 162 Configs gliner, bio, gliner2 — row-for-row subset of rafmacalaba/datause-extracted (revision 15812843687e2ec81b261b5895f2019b50a8f97e): same columns, same split files, rows verbatim. Rows per config: gliner/train 1963… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/datause-extracted-human473-docs.tabulartoken-classification10K<n<100K0 likes138 downloads29d agoHugging Face11vblagoje /lfqa_support_docsSupport documents for building https://huggingface.co/vblagoje/bart_lfqa model text100K<n<1M6 likes113 downloads5y agoHugging Face12skundu42 /halo-docs Halo Documentation Q&A English, single-turn instruction-tuning examples about the Halo LLM training toolkit: concepts, configurations, commands, model recipes, internals, and troubleshooting. This is a source-preserving, extractive Q&A dataset. Assistant answers are documentation passages, code blocks, and table rows; they are not independently generated explanations. Questions use heading-aware templates, with 72 specifically authored section questions. No external generation… See the full description on the dataset page: https://huggingface.co/datasets/skundu42/halo-docs.texttext-generation1K<n<10K0 likes96 downloads18d agoHugging Face13CycloneDX /cdx-docs Introduction This directory contains numerous knowledge files about CycloneDX and cdxgen in jsonlines chat format. The data is useful for training and fine-tuning (LoRA and QLoRA) LLM models. Data Generation We used Google Gemini 2.0 Flash Experimental via aistudio and used the below prompts to convert official documentation markdown files to the chat format. you are an expert in converting markdown files to plain text jsonlines format based on the my template.… See the full description on the dataset page: https://huggingface.co/datasets/CycloneDX/cdx-docs.textquestion-answeringn<1K0 likes77 downloads1y agoHugging Face14JaySmith502 /adaption-louisville-data-center-docs-augmented This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-louisville_data_center_docs (augmented) This dataset contains planning commission staff reports, zoning code excerpts, and news transcripts regarding hyperscale data center developments in Louisville, Kentucky. The documents detail specific project proposals, such as the Camp Ground Road facility, including technical reviews on traffic, water usage, and environmental impact.… See the full description on the dataset page: https://huggingface.co/datasets/JaySmith502/adaption-louisville-data-center-docs-augmented.text1K<n<10K1 likes75 downloads20d agoHugging Face15semeru /code-code-galeras-code-completion-from-docstring-3k-dedupedtabular1K<n<10K3 likes71 downloads3y agoHugging Face16gorkemozer /docspider DocSpider: a Dataset of Cross-Domain Natural Language Querying for MongoDB Arif Görkem Özer, Fırat Çekinel, Pınar Karagöz, İsmail Hakkı Toroslu You can access the paper published in Natural Language Processing journal, from this link. DocSpider dataset is generated by using the widely-known text-to-SQL dataset, Spider. See GitHub repository for more details, including benchmark pipeline scripts for text to MongoDB query conversion. Overview This repository… See the full description on the dataset page: https://huggingface.co/datasets/gorkemozer/docspider.tabular1K<n<10K0 likes61 downloads1y agoHugging Face17icici121 /godot_4_docsDataset generated for Godot 4 docs using Glaive. text1K<n<10K1 likes53 downloads3mo agoHugging Face18tcor0005 /langchain-docs-400-chunksizetext1K<n<10K0 likes52 downloads4y agoHugging Face19Mir-2002 /python_code_docstring_ast_corpus Overview This dataset contains 34,000+ rows of code-docstring-ast data along with additional metadata. Data was gathered from various Python libraries and frameworks and their publicly available GitHub repos. This dataset was created for the purpose of training the CodeT5+ transformer on AST-enhanced code-to-doc tasks. Sources The dataset was gathered from various GitHub repos sampled from this repo by Vinta. The 26 repos are: matplotlib pytorch cryptography django… See the full description on the dataset page: https://huggingface.co/datasets/Mir-2002/python_code_docstring_ast_corpus.textsummarization10K<n<100K1 likes52 downloads1y agoHugging Face20jetesdal /sealevel-docstextn<1K0 likes51 downloads3y agoHugging Face21MongoDB /mongodb-docs Overview This dataset consists of a small subset of MongoDB's technical documentation. Dataset Structure The dataset consists of the following fields: sourceName: The source of the document. url: Link to the article. action: Action taken on the article. body: Content of the article in Markdown format. format: Format of the content. metadata: Metadata such as tags, content type etc. associated with the document. title: Title of the document. updated: The last updated… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/mongodb-docs.textquestion-answeringn<1K1 likes51 downloads2y agoHugging Face22GenaroCoronel /godot_docstextn<1K2 likes50 downloads2y agoHugging Face23fromziro /py-docs-2004 Python Docs 2004 Original dump: https://www.python.org/ftp/python/doc/ Python Docs 2004 is a filtered and cleaned collection of Python documentation from every major Python release published before 2004. Stats Version Size Lines 2.3 2.2MB 1215 2.2 1.7MB 1142 2.1 1.3MB 891 2.0 1.2MB 895 1.6 1MB 720 1.5 837KB 449 1.4 744KB 397 1.3 569KB 408 1.2 513KB 384 Total 10.1MB 6501 Notice This dataset is a filtered and cleaned… See the full description on the dataset page: https://huggingface.co/datasets/fromziro/py-docs-2004.texttext-generation10K<n<100K0 likes50 downloads2mo agoHugging Face24ArunKr /manim-docs-gitingest Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/ArunKr/manim-docs-gitingest.textn<1K0 likes45 downloads1y agoHugging Face25mongodb-eai /docsgated MongoDB Public Documentation Public MongoDB documentation and developer blog posts. Documentation represented as Markdown. In-line links are omitted. Sources: https://mongodb.com/docs Certain pages are omitted because we do not want models trained on them. https://mongodb.com/developer text1K<n<10K6 likes41 downloads7h agoHugging Face26evgenypal /k8s-docs-rag-bench k8s-docs-rag-bench Paper: Analyzing Quality--Latency--Resource Trade-offs in a Technical Documentation RAG Assistant Using LoRA Adaptation (arXiv:2605.28222) Code: github.com/EugPal/rag-lora-tradeoffs A small, fully-grounded benchmark for retrieval-augmented question answering (RAG) over the official Kubernetes documentation, together with the full set of LLM-judge labels used in the accompanying preprint "Analyzing Quality-Latency-Resource Trade-offs in a Technical… See the full description on the dataset page: https://huggingface.co/datasets/evgenypal/k8s-docs-rag-bench.tabularquestion-answering100K<n<1M0 likes40 downloads4mo agoHugging Face27Wauplin /hf-docs-benchmark-lighttextn<1K0 likes39 downloads3mo agoHugging Face28DeepNLP /ai-docs-agent AI Docs Agent Meta and Traffic Dataset in AI Agent Marketplace | AI Agent Directory | AI Agent Index from DeepNLP This dataset is collected from AI Agent Marketplace Index and Directory at http://www.deepnlp.org, which contains AI Agents's meta information such as agent's name, website, description, as well as the monthly updated Web performance metrics, including Google,Bing average search ranking positions, Github Stars, Arxiv References, etc. The dataset is helpful for AI… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/ai-docs-agent.textn<1K1 likes35 downloads2y agoHugging Face29ModalitiesTeam /FW_EDU_SUBSET_500k_docs FineWeb-Edu Subset This dataset contains 483,606 documents sampled from the FineWeb-Edu dataset. The dataset is used throughout various tutorials on modalities. For licensing, see their conditions. tabular100K<n<1M0 likes32 downloads2y agoHugging Face30gavmac00 /arduino-docstext10K<n<100K4 likes29 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.