Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01timodonnell /protein-docs Protein Documents (Parquet) Structured text documents encoding protein residue sequences and 3D contact maps from AlphaFold Database v4 predicted structures, stored as Parquet files. Each row is one protein document with metadata. Source structures: timodonnell/afdb-24M and timodonnell/afdb-1.6M Document Schemes Each subdirectory contains documents generated with a different scheme. All schemes share leakage-resistant train/val/test splits based on structural… See the full description on the dataset page: https://huggingface.co/datasets/timodonnell/protein-docs.tabulartext-generation10M<n<100M0 likes2.1k downloads4mo agoHugging Face02oklenAI /UDM_cleaned_docs UDM cleaned docs 6,029,052 web pages reduced to just their mathematical content, extracted verbatim by oklenAI/udm_doc_extract_qwen3.5_2B — a 2B model distilled from GPT-5.6. Every row is model output, not human-curated text. The extract field is what the model returned for that page; the source page text is not included. Read Two repetition flags below before filtering — the obvious flag is not the one you want. How it was built step pages… See the full description on the dataset page: https://huggingface.co/datasets/oklenAI/UDM_cleaned_docs.tabulartext-generation1M<n<10M0 likes850 downloads1mo agoHugging Face03semeru /text-code-galeras-code-generation-from-docstring-3k-dedupedtabular1K<n<10K0 likes607 downloads3y agoHugging Face04ytzi /the-stack-dedup-python-filtered-docstrings-gpt2tabular10M<n<100M0 likes408 downloads2y agoHugging Face05eczech /marinfold-exp11-protein-docs-seq marinfold-exp11-pdocs-seq Sequence-only derivative of eczech/marinfold-exp11-protein-docs. For every row, the document field has been reduced to just the amino-acid sequence portion: the <begin_sequence> tag followed by the per-residue three-letter tokens (e.g. <begin_sequence> <MET> <LYS> <ASN> ...). The <contacts-and-distances-v1> document-type prefix and everything from <begin_statements> onward (contacts and distances) are removed. The token format is preserved verbatim so… See the full description on the dataset page: https://huggingface.co/datasets/eczech/marinfold-exp11-protein-docs-seq.tabulartext-generation1M<n<10M0 likes217 downloads4mo agoHugging Face06eczech /marinfold-exp11-protein-docs marinfold-exp11-pdocs Quality-bucketed re-publication of the contacts-and-distances-v1-5x config from timodonnell/protein-docs, partitioned by the source round column: Config Source rounds Approx rows high round 0 ~1.68M medium round 1 ~1.42M low round 2–4 ~2.29M Train/val/test split assignment is inherited from the source dataset (leakage-resistant structural-cluster hashing). All columns from the source are preserved; rows are simply partitioned by round. See… See the full description on the dataset page: https://huggingface.co/datasets/eczech/marinfold-exp11-protein-docs.tabulartext-generation1M<n<10M0 likes192 downloads5mo agoHugging Face07ilsp /scipar_parallel_docs SciPar Parallel Documents Dataset Description This dataset contains parallel documents (i.e., titles & abstracts) extracted from academic theses, dissertations, and other scientific texts. In the original paper, we've extracted 9.17M sentence pairs in 31 language pairs from 86 repositories. This version has been created through further processing and filtering to extract parallel documents instead of parallel sentences. To do this, we kept only the parallel titles and… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/scipar_parallel_docs.tabulartext-generation1K<n<10K2 likes183 downloads3y agoHugging Face08rafmacalaba /datause-extracted-human473-docs datause-extracted-human473-docs Every passage of the 162 documents behind the 473 human-validated holdout spans of the data-use annotation campaign: population spans documents annotator190 190 134 jdc283 283 28 total 473 162 Configs gliner, bio, gliner2 — row-for-row subset of rafmacalaba/datause-extracted (revision 15812843687e2ec81b261b5895f2019b50a8f97e): same columns, same split files, rows verbatim. Rows per config: gliner/train 1963… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/datause-extracted-human473-docs.tabulartoken-classification10K<n<100K0 likes138 downloads25d agoHugging Face09ytzi /the-stack-dedup-python-filtered-docstringsThis is a dataset originated from bigcode/the-stack-dedup with some filters applied. The filters filtered in this dataset are: remove_function_no_docstring remove_class_no_docstring remove_delete_markers tabular10M<n<100M0 likes136 downloads3y agoHugging Face10JackHsieh /statML-arxiv-RL-4k-docsSubset of JackHsieh/statML-arxiv. Each row is one randomly sampled contiguous window of exactly 4_096 Qwen3 tokens (Qwen/Qwen3-4B-Instruct-2507) from a distinct paper: papers with at least 4_096 tokens (Qwen3_token_count) are drawn uniformly at random from the eligible pool (seed 42), so the sample carries no relationship between dataset size and paper age, and each window is resampled until its decoded text re-encodes to exactly 4_096 tokens. start_index is the window's offset in the source… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/statML-arxiv-RL-4k-docs.tabular1K<n<10K0 likes100 downloads28d agoHugging Face11paiml /rust-cli-docs-corpus Rust CLI Documentation Corpus A scientifically rigorous corpus for fine-tuning LLMs to generate idiomatic /// documentation comments for Rust CLI tools. Dataset Description This corpus follows the Toyota Way principles and Popperian falsification methodology. Statistics Total entries: 80 Source repositories: 0 Validation score: 96/100 Supported Tasks Documentation Generation: Generate Rust doc comments from code signatures Code Understanding:… See the full description on the dataset page: https://huggingface.co/datasets/paiml/rust-cli-docs-corpus.tabulartext-generationn<1K1 likes80 downloads9mo agoHugging Face12sklearn-docs /digits Dataset Card for digits dataset Optical recognition of handwritten digits dataset Note - How to load this dataset directly with the datasets library from datasets import load_dataset dataset = load_dataset("sklearn-docs/digits",header=None) Dataset Summary This is a copy of the test set of the UCI ML hand-written digits datasets https://archive.ics.uci.edu/ml/datasets/Optical+Recognition+of+Handwritten+Digits The data set contains images of hand-written… See the full description on the dataset page: https://huggingface.co/datasets/sklearn-docs/digits.tabular1K<n<10K0 likes61 downloads4y agoHugging Face13gorkemozer /docspider DocSpider: a Dataset of Cross-Domain Natural Language Querying for MongoDB Arif Görkem Özer, Fırat Çekinel, Pınar Karagöz, İsmail Hakkı Toroslu You can access the paper published in Natural Language Processing journal, from this link. DocSpider dataset is generated by using the widely-known text-to-SQL dataset, Spider. See GitHub repository for more details, including benchmark pipeline scripts for text to MongoDB query conversion. Overview This repository… See the full description on the dataset page: https://huggingface.co/datasets/gorkemozer/docspider.tabular1K<n<10K0 likes54 downloads1y agoHugging Face14JackHsieh /statML-arxiv-RL-2k-docsTrain-only prefix of JackHsieh/statML-arxiv-RL-4k-docs's train split, drawn from JackHsieh/statML-arxiv. Each row is one randomly sampled contiguous window of exactly 4_096 Qwen3 tokens (Qwen/Qwen3-4B-Instruct-2507) from a distinct paper. start_index is the window's offset in the source paper's token sequence; input_ids is the Qwen3 encoding of text (no special tokens added — no BOS/EOS). Same schema and recipe as JackHsieh/statML-arxiv-40M-20M. Nesting: these are the first 2_048 rows of… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/statML-arxiv-RL-2k-docs.tabular1K<n<10K0 likes52 downloads28d agoHugging Face15JackHsieh /statML-arxiv-RL-1k-docsTrain-only prefix of JackHsieh/statML-arxiv-RL-4k-docs's train split, drawn from JackHsieh/statML-arxiv. Each row is one randomly sampled contiguous window of exactly 4_096 Qwen3 tokens (Qwen/Qwen3-4B-Instruct-2507) from a distinct paper. start_index is the window's offset in the source paper's token sequence; input_ids is the Qwen3 encoding of text (no special tokens added — no BOS/EOS). Same schema and recipe as JackHsieh/statML-arxiv-40M-20M. Nesting: these are the first 1_024 rows of… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/statML-arxiv-RL-1k-docs.tabular1K<n<10K0 likes52 downloads28d agoHugging Face16evgenypal /k8s-docs-rag-bench k8s-docs-rag-bench Paper: Analyzing Quality--Latency--Resource Trade-offs in a Technical Documentation RAG Assistant Using LoRA Adaptation (arXiv:2605.28222) Code: github.com/EugPal/rag-lora-tradeoffs A small, fully-grounded benchmark for retrieval-augmented question answering (RAG) over the official Kubernetes documentation, together with the full set of LLM-judge labels used in the accompanying preprint "Analyzing Quality-Latency-Resource Trade-offs in a Technical… See the full description on the dataset page: https://huggingface.co/datasets/evgenypal/k8s-docs-rag-bench.tabularquestion-answering100K<n<1M0 likes50 downloads4mo agoHugging Face17semeru /code-code-galeras-code-completion-from-docstring-3k-dedupedtabular1K<n<10K0 likes45 downloads3y agoHugging Face18docs-benchmarks /compile-benchmarkstabularn<1K0 likes43 downloads2y agoHugging Face19vivek-dodia /mikrotik-docs MikroTik Technical Documentation Dataset Overview A structured dataset containing MikroTik's technical documentation, prepared for LLM fine-tuning. The dataset preserves the hierarchical structure of the original documentation while maintaining technical accuracy and formatting. Dataset Statistics Total documents: 285 Maximum sections per document: 91 Average sections per document: 12.3 Format: Parquet Data Structure Each row represents a complete… See the full description on the dataset page: https://huggingface.co/datasets/vivek-dodia/mikrotik-docs.tabularn<1K0 likes39 downloads2y agoHugging Face20Zappu /legal-docs-vntabular1M<n<10M0 likes34 downloads2y agoHugging Face21ASHu2 /docs-python-v1 Dataset Card for Dataset Name This dataset card aims to be a base template for creating python docs from methods. This is formatted from semeru/code-code-galeras-code-completion-from-docstring-3k-deduped Dataset Description Curated by: semeru/code-code-galeras-code-completion-from-docstring-3k-deduped Language(s) (NLP): Python License: [More Information Needed] Dataset Sources [optional] Repository:… See the full description on the dataset page: https://huggingface.co/datasets/ASHu2/docs-python-v1.tabularfeature-extraction1K<n<10K2 likes32 downloads3y agoHugging Face22ModalitiesTeam /FW_EDU_SUBSET_500k_docs FineWeb-Edu Subset This dataset contains 483,606 documents sampled from the FineWeb-Edu dataset. The dataset is used throughout various tutorials on modalities. For licensing, see their conditions. tabular100K<n<1M0 likes30 downloads2y agoHugging Face23sasha /ipcc_docs_testtabularn<1K0 likes30 downloads1y agoHugging Face24SeifAI /FineTranselation_EGY_filtered_docs_50ktabular10K<n<100K0 likes30 downloads7mo agoHugging Face25ASSERT-KTH /stack-smol-docstrings Stack-Smol-Docstrings This dataset contains Python functions extracted from the-stack-smol, filtered for high-quality docstrings and implementations. Each sample includes the function's docstring, implementation, and a masked version of the code where the function is replaced with a comment. The dataset is designed for code completion tasks where a model needs to restore a function that has been replaced with a comment. The model is provided with: The full file context with the… See the full description on the dataset page: https://huggingface.co/datasets/ASSERT-KTH/stack-smol-docstrings.tabular1K<n<10K0 likes29 downloads2y agoHugging Face26docs-benchmarks /experts-backendstabularn<1K0 likes29 downloads9mo agoHugging Face27vbeskrovnov /eval_smolvla_so101_v3_docsThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so_follower", "total_episodes": 1, "total_frames": 1324, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/vbeskrovnov/eval_smolvla_so101_v3_docs.tabularrobotics1K<n<10K0 likes26 downloads5mo agoHugging Face28Lukaszl /pl-mixed-docs-ocr-dataset-100-v1-results OCR Bench Results: Polish mixed documents benchmark VLM-as-judge pairwise evaluation of OCR models on a small heterogeneous sample of Polish document-style images. Rankings depend strongly on document type, so this should be read as a document-specific OCR benchmark rather than a universal OCR ranking. This benchmark uses a lightweight 100-image Polish OCR sample covering mixed document categories such as official forms, templates, certificates, structured layouts, invoices, and… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/pl-mixed-docs-ocr-dataset-100-v1-results.tabular1K<n<10K1 likes25 downloads6mo agoHugging Face29anihitk07 /ms_docstabular10K<n<100K0 likes24 downloads2y agoHugging Face30Lukaszl /pl-government-docs-mix-ocr-dataset-v1-results OCR Bench Results: Polish government documents benchmark VLM-as-judge pairwise evaluation of OCR models on a dataset of real Polish government and public administration documents. This benchmark focuses on structured, text-heavy documents typical for public institutions, including official forms, templates, administrative documents, and scanned materials. As with all OCR benchmarks, results are document-type specific and should not be interpreted as a universal ranking across all… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/pl-government-docs-mix-ocr-dataset-v1-results.tabular1K<n<10K1 likes22 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.