Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Xenova /transformers.js-docs3 likes93k downloads6mo agoHugging Face02huggingface /policy-docs Public Policy at Hugging Face AI Policy at Hugging Face is a multidisciplinary and cross-organizational workstream. Instead of being part of a vertical communications or global affairs organization, our policy work is rooted in the expertise of our many researchers and developers, from Ethics and Society Regulars and legal team to machine learning engineers working on healthcare, art, and evaluations. What we work on is informed by our Hugging Face community needs and experiences… See the full description on the dataset page: https://huggingface.co/datasets/huggingface/policy-docs.documentn<1K15 likes13k downloads7mo agoHugging Face03diffusers /docs-imagesimagen<1K0 likes11k downloads7mo agoHugging Face04diffusers /diffusers-images-docsimagen<1K0 likes10k downloads2y agoHugging Face05HCAI-Lab-GT /dolma3-6t-sample-100000-docs dolma3-6t-sample-100000-docs Materialized stratified sample of 100,000 docs per bin (~58k source-shard files, ~186 GB) from the deduplicated 6T Dolma3 corpus. Seed 42, materialized by 128 Modal workers (worker_0000/ through worker_0127/). Layout HCAI-Lab/dolma3-6t-sample-100000-docs/ ├── bin_summary.csv ├── sample_contract.json └── worker_NNNN/ └── soc127__phase1_pool_shared__...__shard_NNNNNNNN.jsonl.zst (~450 files per worker) Dual-access… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-100000-docs.0 likes5k downloads4mo agoHugging Face06allenai /pixmo-docs PixMo-Docs We now recommend using CoSyn-400k and CoSyn-point over these datasets. They are improved versions with more images categories and an improved generation pipeline. PixMo-Docs is a collection of synthetic question-answer pairs about various kinds of computer-generated images, including charts, tables, diagrams, and documents. The data was created by using the Claude large language model to generate code that can be executed to render an image, and using GPT-4o mini to… See the full description on the dataset page: https://huggingface.co/datasets/allenai/pixmo-docs.imagevisual-question-answering100K<n<1M35 likes2.8k downloads2y agoHugging Face07funmaker9527 /docsdocument1K<n<10K0 likes2.1k downloads4mo agoHugging Face08timodonnell /protein-docs Protein Documents (Parquet) Structured text documents encoding protein residue sequences and 3D contact maps from AlphaFold Database v4 predicted structures, stored as Parquet files. Each row is one protein document with metadata. Source structures: timodonnell/afdb-24M and timodonnell/afdb-1.6M Document Schemes Each subdirectory contains documents generated with a different scheme. All schemes share leakage-resistant train/val/test splits based on structural… See the full description on the dataset page: https://huggingface.co/datasets/timodonnell/protein-docs.tabulartext-generation10M<n<100M0 likes2.1k downloads4mo agoHugging Face09plaguss /argilla_sdk_docs_raw_unstructured Dataset info This dataset contains documentation chunks from repositories (ADD REPOS). Postprocessing After some inspection, some chunks contain text too short to be meaningful, so we decided to remove those by removing chunks whose number of tokens (computed with the same tokenizer of the model to be used for the embeddings) is lower or equal to the 5%: from transformers import AutoTokenizer tokenizer = AutoTokenizer.from_pretrained("BAAI/bge-base-en-v1.5") df =… See the full description on the dataset page: https://huggingface.co/datasets/plaguss/argilla_sdk_docs_raw_unstructured.textn<1K0 likes1.7k downloads2y agoHugging Face10nuuuwan /lk-news-docstext100K<n<1M5 likes1.6k downloads1h agoHugging Face11Lana49 /engineering-docsdocumentn<1K0 likes1.5k downloads2mo agoHugging Face12HCAI-Lab-GT /dolma3-6t-sample-10000-docs dolma3-6t-sample-10000-docs Materialized stratified sample of 10K docs per bin (5.68M total docs, 10.5B tokens). Seed 42. This is the basis for the SOC-156 TrackStar gradient index. Provenance This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention. Field Value Previous name HCAI-Lab/dolma3_6T_sample_10000_docs Renamed 2026-05-25… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-10000-docs.text1M<n<10M0 likes1.3k downloads4mo agoHugging Face13cognaize /elements_annotated_tables_4500_docs Dataset 🚀 Progress Last update (UTC): 2025-11-11 15:40:21Z Documents processed: 4500 / 500058 Batches completed: 30 Total pages/rows uploaded: 89882 Latest batch summary Batch index: 30 Docs in batch: 150 Pages/rows added: 1487 imageobject-detection10K<n<100K0 likes1.3k downloads11mo agoHugging Face14nuuuwan /lk-tourism-weekly-reports-docstextn<1K0 likes1.3k downloads28d agoHugging Face15nuuuwan /cbsl-annual-reports-docstext1K<n<10K0 likes1.3k downloads1y agoHugging Face16nuuuwan /lk-dmc-weather-forecasts-docstext1K<n<10K0 likes1.2k downloads3h agoHugging Face17HCAI-Lab-GT /dolma3-6t-sample-5000-docs dolma3-6t-sample-5000-docs Materialized stratified sample of 5K docs per bin (2.86M total docs, 5.3B tokens). Seed 42. Provenance This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention. Field Value Previous name HCAI-Lab/dolma3_6T_sample_5000_docs Renamed 2026-05-25 See docs/data_home/inventory.json for the full inventory… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-5000-docs.0 likes1.1k downloads4mo agoHugging Face18nuuuwan /lk-dmc-river-water-level-and-flood-warnings-docstextn<1K0 likes954 downloads5h agoHugging Face19HCAI-Lab-GT /dolma3-6t-sample-500-docs dolma3-6t-sample-500-docs Materialized stratified sample of 500 docs per bin (288K total docs, 539M tokens) drawn from the deduplicated 6T Dolma3 corpus. Seed 42. Matching companion bucket at hf://buckets/HCAI-Lab/dolma3-6t-sample-500-docs. Provenance This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention. Field Value Previous name… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-500-docs.0 likes926 downloads4mo agoHugging Face20oklenAI /UDM_cleaned_docs UDM cleaned docs 6,029,052 web pages reduced to just their mathematical content, extracted verbatim by oklenAI/udm_doc_extract_qwen3.5_2B — a 2B model distilled from GPT-5.6. Every row is model output, not human-curated text. The extract field is what the model returned for that page; the source page text is not included. Read Two repetition flags below before filtering — the obvious flag is not the one you want. How it was built step pages… See the full description on the dataset page: https://huggingface.co/datasets/oklenAI/UDM_cleaned_docs.tabulartext-generation1M<n<10M0 likes834 downloads1mo agoHugging Face21HCAI-Lab-GT /dolma3-6t-sample-1000-docs dolma3-6t-sample-1000-docs Materialized stratified sample of 1K docs per bin (575K total docs, 1.08B tokens). Seed 42. Provenance This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention. Field Value Previous name HCAI-Lab/dolma3_6T_sample_1000_docs Renamed 2026-05-25 See docs/data_home/inventory.json for the full inventory… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-1000-docs.0 likes821 downloads4mo agoHugging Face22LeadBHYT /bhyt-legal-docsdocumentn<1K0 likes817 downloads4mo agoHugging Face23juliozhao /DocSynth300K DocSynth300K is a large-scale and diverse document layout analysis pre-training dataset, which can largely boost model performance. Data Download Use following command to download dataset(about 113G): from huggingface_hub import snapshot_download # Download DocSynth300K snapshot_download(repo_id="juliozhao/DocSynth300K", local_dir="./docsynth300k-hf", repo_type="dataset") # If the download was disrupted and the file is not complete, you can resume the download… See the full description on the dataset page: https://huggingface.co/datasets/juliozhao/DocSynth300K.text100K<n<1M55 likes783 downloads2y agoHugging Face24MAIR-Bench /MAIR-Docs MAIR: A Massive Benchmark for Evaluating Instructed Retrieval MAIR is a heterogeneous IR benchmark that comprises 126 information retrieval tasks across 6 domains, with annotated query-level instructions to clarify each retrieval task and relevance criteria. This repository contains the document collections for MAIR, while the query data are available at https://huggingface.co/datasets/MAIR-Bench/MAIR-Queries. Paper: https://arxiv.org/abs/2410.10127 Github:… See the full description on the dataset page: https://huggingface.co/datasets/MAIR-Bench/MAIR-Docs.texttext-retrieval1M<n<10M4 likes776 downloads2y agoHugging Face25KingNish /huggingface-docsThis repo contains all the docs published on https://huggingface.co/docs. The docs are generated with https://github.com/huggingface/doc-builder. 1 likes773 downloads2y agoHugging Face26aicentreflip /docs-gifs FLIP documentation GIFs The animated walkthroughs embedded in the FLIP user guides on ReadTheDocs. They are recorded by Cypress from the FLIP UI against a mocked backend (flip-ui/test/cypress/docs/), converted with ffmpeg, and published here by .github/workflows/regenerate_docs_gifs.yml so the recordings never enter the git repository (FLIP#1236). Layout One copy of every file, at an unversioned path: Path Contents admin/<name>.gif one per… See the full description on the dataset page: https://huggingface.co/datasets/aicentreflip/docs-gifs.0 likes682 downloads4d agoHugging Face27Jeice /n8n-docs-v2 n8n Docs This repository hosts the documentation for n8n, an extendable workflow automation tool which enables you to connect anything to everything. The documentation is live at docs.n8n.io. Previewing and building the documentation locally Prerequisites Python 3.8 or above Pip n8n recommends using a virtual environment when working with Python, such as venv. Follow the recommended configuration and auto-complete guidance for the theme. This will help when… See the full description on the dataset page: https://huggingface.co/datasets/Jeice/n8n-docs-v2.2 likes630 downloads1y agoHugging Face28semeru /text-code-galeras-code-generation-from-docstring-3k-dedupedtabular1K<n<10K0 likes623 downloads3y agoHugging Face29lamini /lamini_docs Dataset Card for "lamini_docs" More Information needed text1K<n<10K23 likes621 downloads3y agoHugging Face30nameexhaustion /polars-docs0 likes615 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.