Team Ai
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01DaoCloud /Qwen3.8-27B-Drafter-SFT Qwen3.8-27B Drafter SFT Corpus Supervised fine-tuning data released for training speculative drafters for Qwen/Qwen3.8-27B. All completions were generated with Qwen/Qwen3.8-27B at revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0. The dataset contains 367,535 source conversations and 450,401 train rows, totaling 1,953,218,671 tokens after filtering and evaluation decontamination. Rows contain Qwen3.8-27B-tokenized prompts and target-generated completions, together with loss… See the full description on the dataset page: https://huggingface.co/datasets/DaoCloud/Qwen3.8-27B-Drafter-SFT.texttext-generation100K<n<1M1 likes928 downloads1mo agoHugging Face02ytu-ce-cosmos /turkish-flow-drafter-prompts GitHub repo · Technical blog · Model collection Turkish prompts for Chained-Flow drafter training Chat-templated Turkish prompts used to train and evaluate the Turkish Flow-Drafter checkpoints for Qwen/Qwen3.5-4B / 9B / 27B. Prompts only — no completions. A drafter is trained on the target model's own hidden states, so continuations are generated locally by running the target over these prompts. Nothing here is a model output. split rows what it is v1/… See the full description on the dataset page: https://huggingface.co/datasets/ytu-ce-cosmos/turkish-flow-drafter-prompts.texttext-generation10K<n<100K1 likes161 downloads25d agoHugging Face03Eason666 /DrafterBench Dataset Card for DrafterBench Dataset Description DrafterBench is a large-scale toolkit focused on evaluating the proficiency of Large Language Models (LLMs) in automating Civil Engineering tasks. This dataset hosts a task suite summarized across 20 real-world projects, encompassing a total of 1920 tasks. It replicates the complexity of real-world engineering tasks and provides a technical platform to test the four key capabilities of LLMs: Structured data understanding… See the full description on the dataset page: https://huggingface.co/datasets/Eason666/DrafterBench.texttext-generation1K<n<10K10 likes123 downloads10mo agoHugging Face04ytu-ce-cosmos /english-flow-drafter-prompts GitHub repo · Technical blog · Model collection English prompts for Chained-Flow drafter training Chat-templated English prompts used to train the English Flow-Drafter checkpoints for Qwen/Qwen3.5-4B / 9B / 27B. Prompts only — no completions. A drafter is trained on the target model's own hidden states, so continuations are generated locally by running the target over these prompts. Nothing here is a model output. split rows prompt tokens what it is v1/… See the full description on the dataset page: https://huggingface.co/datasets/ytu-ce-cosmos/english-flow-drafter-prompts.texttext-generation10K<n<100K0 likes84 downloads25d agoHugging Face05mv1137 /p2026-007-qwen38-27b-2bit-drafting-results Self-Drafting Qwen3.8-27B at Two Bits: Speed, Energy, and Memory on a 12 GB Laptop GPU — result index Matthew Schwartz — ORCID 0009-0009-4171-7247 This dataset accompanies “Self-Drafting Qwen3.8-27B at Two Bits: Speed, Energy, and Memory on a 12 GB Laptop GPU.” The immutable tagged report and full artifact use the following versioned path: https://github.com/matthematics1137/research-artifacts/tree/drafted-2bit-local-efficiency-v1.0.0/papers/2026-drafted-2bit-local-efficiency… See the full description on the dataset page: https://huggingface.co/datasets/mv1137/p2026-007-qwen38-27b-2bit-drafting-results.tabulartext-generation1K<n<10K0 likes67 downloads11d agoHugging Face06Schiltmans /bonsai2-drafter-eval Bonsai 2 drafter evaluation corpora Prompts and greedy responses from PrismML's Ternary-Bonsai-2-27B, recorded as token ids together with the per-round accepted lengths of the speculative decoding loop that produced them. The data exists to evaluate and train DFlash 2 drafters against this one target. Target prism-ml/Ternary-Bonsai-2-27B-mlx-2bit (MLX pack), greedy, temperature 0 Loop mlx-dspark 0.18.0 DFlash 2 loop, stock drafter z-lab/Qwen3.8-27B-DFlash2, draft… See the full description on the dataset page: https://huggingface.co/datasets/Schiltmans/bonsai2-drafter-eval.texttext-generationn<1K1 likes61 downloads16d agoHugging Face07vishwr /claim_drafter Claim Drafter — datasets Training and evaluation data for vishwr/claim_drafter, a LoRA adapter that drafts US patent claims from a plain-English invention disclosure. Two datasets are included: sft — 9,662 train / 1,314 validation examples in conversational (messages) format. Targets are as-GRANTED claims fetched per patent number (not the as-filed claims that ship with HUPD), and prompts are stripped of the patent summary so the model must draft rather than reformat. dpo — 5… See the full description on the dataset page: https://huggingface.co/datasets/vishwr/claim_drafter.texttext-generation10K<n<100K0 likes43 downloads3mo agoHugging Face08selimaktas /english-flow-drafter-prompts English prompts for Chained-Flow drafter training Chat-templated English prompts used to train the English Flow-Drafter checkpoints for Qwen/Qwen3.5-4B / 9B / 27B. Prompts only — no completions. A drafter is trained on the target model's own hidden states, so continuations are generated locally by running the target over these prompts. Nothing here is a model output. split rows prompt tokens what it is v1/ 19,672 train + 500 holdout 1,518,620 the original mixture… See the full description on the dataset page: https://huggingface.co/datasets/selimaktas/english-flow-drafter-prompts.texttext-generation10K<n<100K0 likes42 downloads1mo agoHugging Face09PortalPal-AI /Patient-Message-Response-DraftingPaper: How Much Would a Clinician Edit This Draft? Evaluating LLM Alignment for Patient Message Response Drafting (arxiv link) Dataset Details: The patient message response drafting dataset is designed to evaluate how well LLMs respond to patient messages in patient portal communication. Each semi-synthetic patient message is paired with a real de-identified EHR from a patient at our collaborating hospital. Each doctor response is written by a clinician, guided by clinician response themes… See the full description on the dataset page: https://huggingface.co/datasets/PortalPal-AI/Patient-Message-Response-Drafting.texttext-generationn<1K2 likes39 downloads6mo agoHugging Face10jjjlimaus /chrono-2020-leak-sft-llm-draftgated ChronoLLM 2020 leak-SFT (LLM draft) Public gated (gated=manual) prompt/phrase pairs for SN38 leak training (LLM gen+judge, object-first phrases) over the 2020 harvest. Final snapshot (stopped): accepted pairs: 2,222,376 trainable tokens (GPT-2, prompt+phrase): 79,901,922 closed shards on Hub: 1116 across 481 chunks Repo: jjjlimaus/chrono-2020-leak-sft-llm-draft Layout chunk-YYYYMMDDHHMM/ COMPLETE.json llm/part-NNNNNN.jsonl.gz Trainers must read chunk-* in… See the full description on the dataset page: https://huggingface.co/datasets/jjjlimaus/chrono-2020-leak-sft-llm-draft.text-generation0 likes37 downloads1mo agoHugging Face11anonymous733882 /DrafterBench Dataset Card for DrafterBench DrafterBench DrafterBench is a large-scale toolkit focused on evaluating the proficiency of Large Language Models (LLMs) in automating Civil Engineering tasks. The dataset contains tasks derived from real-world engineering drawing revision processes. This dataset is released for anonymous review. Code: https://github.com/anonymous733882/DrafterBench This dataset hosts a task suite summarized across 20 real-world projects, encompassing a total… See the full description on the dataset page: https://huggingface.co/datasets/anonymous733882/DrafterBench.texttext-generation1K<n<10K0 likes24 downloads7mo agoHugging Face12selimaktas /turkish-flow-drafter-prompts Turkish prompts for Chained-Flow drafter training Chat-templated Turkish prompts used to train and evaluate the Turkish Flow-Drafter checkpoints for Qwen/Qwen3.5-4B / 9B / 27B. Prompts only — no completions. A drafter is trained on the target model's own hidden states, so continuations are generated locally by running the target over these prompts. Nothing here is a model output. split rows what it is v1/ 29,100 train + 300 holdout the mixture the released Turkish… See the full description on the dataset page: https://huggingface.co/datasets/selimaktas/turkish-flow-drafter-prompts.texttext-generation10K<n<100K0 likes24 downloads1mo agoHugging Face13mjbommar /opengloss-v1.1-drafting OpenGloss Drafting v1.1 Dataset Summary OpenGloss Drafting is a synthetic dataset of educational content drafts generated from vocabulary terms and encyclopedia entries. Each draft is a self-contained piece of writing (article, story, memo, essay, etc.) that naturally incorporates specific vocabulary terms with their definitions and context. This dataset supports curriculum-aligned content generation, vocabulary-in-context learning, and educational text synthesis. It is… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.1-drafting.tabulartext-generation10K<n<100K0 likes23 downloads10mo agoHugging Face14hugruby /mismatched-wrong-drafts Per-variant training data Training data for "Weak-to-Strong Elicitation via Mismatched Wrong Drafts" (Wei Deng, 2026). Code: https://github.com/weiddeng/mismatched-wrong-drafts Models: the four trained models below Paper: https://arxiv.org/abs/2605.17314 Each row/training datapoint is a MATH problem with a draft injected into the prompt field. Variant Draft shown to the learner Trained model # Rows mismatched_wrong ⭐ a wrong draft from a different problem… See the full description on the dataset page: https://huggingface.co/datasets/hugruby/mismatched-wrong-drafts.texttext-generation10K<n<100K0 likes23 downloads4mo agoHugging Face15PNNL /DraftNEPABenchgated DraftNEPABench: A Benchmark for Drafting NEPA Document Sections with Coding Agents Dataset Description DraftNEPABench is a novel benchmark designed to evaluate the capabilities of Large Language Models (LLMs) and LLM-based agents in drafting sections of Environmental Impact Statements (EIS). This benchmark is curated by subject matter experts (SMEs) to ensure that the tasks reflect realistic, domain-relevant drafting challenges encountered in environmental planning and… See the full description on the dataset page: https://huggingface.co/datasets/PNNL/DraftNEPABench.documenttext-generationn<1K1 likes6 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.