Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01xingqiang /GPRadar-Defect-MultiTask GPRadar-Defect-MultiTask 数据集 本仓库包含用于微调PaLI-GEMMA多模态模型的地质雷达(GPR)缺陷检测数据集。该数据集专注于地下结构中的空洞和裂缝检测与分析。 数据集结构 数据集组织如下: dataset/ ├── annotations/ - 包含JSON和JSONL格式的标注文件 │ ├── _annotations.train.jsonl - 训练集标注 │ ├── _annotations.valid.jsonl - 验证集标注 │ ├── _annotations.test.jsonl - 测试集标注 │ ├── p-1.v1i.paligemma/ - 主数据集元数据 │ └── p-1.v1i.paligemma-multimodal/ - 多模态数据集元数据 ├── images/ - 包含所有图像文件 特点 包含874张带注释的地质雷达扫描图像 图像预处理为640x640像素大小 支持多种任务类型:缺陷检测、位置定位和描述生成… See the full description on the dataset page: https://huggingface.co/datasets/xingqiang/GPRadar-Defect-MultiTask.imageobject-detection1K<n<10K0 likes175 downloads2y agoHugging Face02TaskPuppyAI /lunamax-multitask-programming-1000 LunaMax Multitask Programming 1000 A 1,000-record synthetic multitask programming dataset generated with ChatGPT LunaMax. The recovered dataset combines code review, implementation, bug and severity classification, and strict output-contract tasks across multiple programming languages. The historical source shards were reviewed with ChatGPT 5.6 Sol High according to dataset creator confirmation. During Hugging Face publication preparation, all 1,000 records received a new… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/lunamax-multitask-programming-1000.text1K<n<10K0 likes69 downloads1mo agoHugging Face03narendarcodes /Telugu-MultiTask-Instruct-77K Telugu MultiTask Instruct 77K — Adaption AutoScientist Challenge Dataset Powered by Adaptive Data — Adaption Labs Dataset Description A large-scale, multi-task Telugu instruction-tuning dataset combining 77,653 rows from 7 open-source Telugu NLP collections. Covers diverse tasks including news summarization, QA, creative writing, translation, and general instruction following — all processed through the Adaption Labs AutoScientist platform for quality… See the full description on the dataset page: https://huggingface.co/datasets/narendarcodes/Telugu-MultiTask-Instruct-77K.textquestion-answering10K<n<100K1 likes65 downloads3mo agoHugging Face04Kushalkhemka /cybersec-chatml-multitask-v1 Cybersecurity ChatML Multitask Dataset (v1) Combined split for both detection and patch tasks. Files chatml_multitask_train.jsonl chatml_multitask_val.jsonl chatml_build_manifest.json unsloth_best_params_glm47flash_multitask.json Output format Detection samples: strict JSON schema output Patch samples: patched code only text100K<n<1M0 likes58 downloads6mo agoHugging Face05hamishivi /rds-sels-multitask-rrmax-top326k RDS+ Selected Multitask 326k This is the dataset (and associated scores) selected by RDS+ when selecting 326k samples for multiple tasks at once. For more details, please see the paper Practical Large-Scale Data Selection for Instruction Tuning. This was used to train this model. This dataset is selected from Tulu 2 unfiltered, and please see that page for more information on sources. License We are releasing this dataset under the terms of ODC-BY. By using this, you… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/rds-sels-multitask-rrmax-top326k.text100K<n<1M1 likes41 downloads2y agoHugging Face06mujo-labs /sandman-dream_multitask_v2_train Sandman dream multitask v2 — train split 17,300 instruction-following examples for fine-tuning Sandman's on-device dream-analysis model, built from sandman-dreambank-v2. Every row is a single-turn conversation (messages) covering one of three tasks: Summarize — read a dream, return a one- or two-sentence summary as JSON. Extract symbols — return only the concrete nouns literally present in the dream text, as a JSON array, with an explicit instruction not to infer or add… See the full description on the dataset page: https://huggingface.co/datasets/mujo-labs/sandman-dream_multitask_v2_train.texttext-generation10K<n<100K0 likes41 downloads19d agoHugging Face07hamishivi /rds-sels-tulu-3-multitask-rrmax-939k RDS+ Selected Tulu 3 Multitask 939k This is the dataset (and associated scores) selected by RDS+ when selecting 939k samples targeting multiple downstream tasks. For more details, please see the paper Practical Large-Scale Data Selection for Instruction Tuning. This was used to train this model. This dataset is selected from Tulu 3 unfiltered, and please see that page for more information on sources. License This dataset is licensed under ODC-BY-1.0. It is intended… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/rds-sels-tulu-3-multitask-rrmax-939k.text100K<n<1M0 likes38 downloads2y agoHugging Face08TaskPuppyAI /lunamax-multitask-programming-250 LunaMax Multitask Programming 250 A 250-record synthetic multitask programming dataset generated with ChatGPT LunaMax. The dataset combines structured and free-form code review, implementation, bug and severity classification, and strict output-contract tasks across multiple programming languages. Generation and historical-review attribution are based on dataset creator confirmation. Dataset Summary The publication dataset contains: 250 records 250 unique records… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/lunamax-multitask-programming-250.textn<1K0 likes35 downloads1mo agoHugging Face09AethronPhantom /nexa-science-multitask-balanced Nexa Science Multitask Balanced This dataset is a curated, instruction-formatted scientific multitask mixture for: claim verification (<TASK:VERIFY>) abstract-grounded biomedical QA (<TASK:QA>) retrieval relevance re-ranking (<TASK:RERANK>) Format Each row is JSONL with: {task, instruction, input, output, meta} Splits Included train_balanced_short.jsonl val_balanced_short.jsonl stats_balanced_short.json Notes QA in this balanced release is… See the full description on the dataset page: https://huggingface.co/datasets/AethronPhantom/nexa-science-multitask-balanced.texttext-classification10K<n<100K0 likes33 downloads8mo agoHugging Face10mujo-labs /sandman-dream_multitask_v2_test Sandman dream multitask v2 — test split The test split for fine-tuning Sandman's on-device dream-analysis model (v2). See sandman-dream_multitask_v2_train for the full description of the three tasks (summarize, extract symbols, interpret a symbol) and the source data. texttext-generation1K<n<10K0 likes27 downloads19d agoHugging Face11TheTokenFactory /sec-extraction-multitask-v4 SEC Extraction Multitask v4 Instruction-tuning dataset for fine-tuning a small language model (e.g. Gemma 4 E2B) to extract structured data from SEC filings across three verticals: Exhibit 10 (contracts) — financial terms from executive employment, credit agreements, indemnification, licensing, and similar filings DEF 14A (proxy statements) — executive compensation, governance items, say-on-pay MD&A (10-K / 10-Q Management's Discussion & Analysis) — operating metrics, segment… See the full description on the dataset page: https://huggingface.co/datasets/TheTokenFactory/sec-extraction-multitask-v4.texttext-generation1K<n<10K0 likes27 downloads6mo agoHugging Face12PJMixers /vicgalle_configurable-system-prompt-multitask-PreferenceShareGPTtextreinforcement-learning1K<n<10K5 likes24 downloads2y agoHugging Face13mujo-labs /sandman-dream_multitask_v2_val Sandman dream multitask v2 — val split The val split for fine-tuning Sandman's on-device dream-analysis model (v2). See sandman-dream_multitask_v2_train for the full description of the three tasks (summarize, extract symbols, interpret a symbol) and the source data. texttext-generation1K<n<10K0 likes24 downloads19d agoHugging Face14cle-13 /rutooro_multitask Rutooro Multitask Dataset This dataset contains a collection of instruction-response pairs for fine-tuning a Large Language Model (LLM) on the Rutooro language. The dataset is prepared for a multi-task learning approach, including: Translation: English to Rutooro. Monolingual Generation: Continued stories and prose in Rutooro. Grammar Instructions: Explanations of Rutooro grammar rules. Data Source The data was sourced from [mention your source, e.g., "manual… See the full description on the dataset page: https://huggingface.co/datasets/cle-13/rutooro_multitask.text1K<n<10K0 likes21 downloads1y agoHugging Face15Phettae /thai-multitask-starter Thai Multitask 9.6K ชุดข้อมูลตั้งต้นสำหรับ instruction tuning ภาษาไทย ครอบคลุมงานสนทนา ถาม–ตอบ สรุป แปล จำแนกข้อความ ตรวจแก้ภาษา คณิตศาสตร์ และ structured output ข้อมูลทุกแถวสร้างขึ้นใหม่ด้วยกฎแบบ deterministic ไม่มีการคัดลอกจากเว็บไซต์หรือ ข้อมูลส่วนบุคคลจริง เหมาะสำหรับทดลอง supervised fine-tuning และทดสอบ pipeline แต่ควรเพิ่มข้อมูลที่มนุษย์ตรวจทานและข้อมูลภาษาธรรมชาติก่อนใช้กับระบบจริง จำนวนข้อมูลทั้งหมด 9,599 ตัวอย่าง: train 8,639, validation 480 และ test 480… See the full description on the dataset page: https://huggingface.co/datasets/Phettae/thai-multitask-starter.texttext-generation1K<n<10K0 likes19 downloads2mo agoHugging Face16lohoz /Smart-Contract-MultiTask-Dataset Overview This is a dataset designed for smart contract generation. It includes two subsets: Requirement-FSM-Code subset: Contains user requirement descriptions, finite state machine (FSM) representations, and corresponding smart contract code. Comment-Code subset: Includes functional comments and their corresponding implementation code. Dataset Structure Subset 1: Requirement-FSM-Code Description: Contains natural language descriptions of user requirements… See the full description on the dataset page: https://huggingface.co/datasets/lohoz/Smart-Contract-MultiTask-Dataset.text10K<n<100K0 likes16 downloads2y agoHugging Face17GilbertAkham /gilbert-multitask-mix DATASET_README.md --- language: - en task_categories: - text-generation - summarization - question-answering - conversational tags: - multitask - email - stories - qa - summarization - chat license: - cc-by-4.0 - apache-2.0 - mit --- # Gilbert-Multitask-Mix A diverse multitask dataset for text generation training, combining samples from 5 different domains with structured prompt formatting. ## Dataset Description This dataset contains 6,500+ examples across multiple text… See the full description on the dataset page: https://huggingface.co/datasets/GilbertAkham/gilbert-multitask-mix.text100K<n<1M0 likes16 downloads1y agoHugging Face18ariefansclub /humanoid-multi-step-task-instructions Humanoid Multi-Step Task Instructions A structured dataset containing multi-step task instructions for humanoid robots. Use Cases Task planning Autonomous execution Robotics simulation textn<1K0 likes15 downloads9mo agoHugging Face19ThuraAung1601 /reform-dafny-multitaskgated ReForm Dafny multi-task Six Dafny tasks per verified program of the ReForm python2dafny data (ThuraAung1601/reform-dafny-loop-inv-gen, ThuraAung1601/reform-dafny-hint-gen, decontaminated against DafnyBench), for multi-task SFT + RL with a Dafny-verifier reward. task the model gets answer is scored by loop-invariants the program with its loop invariants removed dafny verify + only proof hints may differ proof-annotations every line starting with invariant / assert /… See the full description on the dataset page: https://huggingface.co/datasets/ThuraAung1601/reform-dafny-multitask.texttext-generation10K<n<100K0 likes14 downloads2d agoHugging Face20CL-From-Nothing /rlve-multitask-qwen3-4b-rollouts-n4-tokens16384tabular1K<n<10K0 likes13 downloads6mo agoHugging Face21ThuraAung1601 /dafnybench-multitaskgated DafnyBench multi-task ThuraAung1601/dafnybench-cleaned turned into six Dafny tasks, for evaluation. All comments were removed from every program (the ReForm training programs have none); each ground truth was re-verified with Dafny 4.11.0 afterwards. 731 of 733 programs are used (skipped: 2 ground truth uses {:verify false}, 1 loop-invariants: ground truth fails its own faithfulness check, 1 termination: ground truth fails its own faithfulness check). task the model gets… See the full description on the dataset page: https://huggingface.co/datasets/ThuraAung1601/dafnybench-multitask.texttext-generation1K<n<10K0 likes12 downloads2d agoHugging Face22nmd2k /multi-task-instructiontexttext-generation100K<n<1M0 likes11 downloads3y agoHugging Face23LiZHENGzai /GPRadar-Defect-MultiTask GPRadar-Defect-MultiTask 数据集 本仓库包含用于微调PaLI-GEMMA多模态模型的地质雷达(GPR)缺陷检测数据集。该数据集专注于地下结构中的空洞和裂缝检测与分析。 数据集结构 数据集组织如下: dataset/ ├── annotations/ - 包含JSON和JSONL格式的标注文件 │ ├── _annotations.train.jsonl - 训练集标注 │ ├── _annotations.valid.jsonl - 验证集标注 │ ├── _annotations.test.jsonl - 测试集标注 │ ├── p-1.v1i.paligemma/ - 主数据集元数据 │ └── p-1.v1i.paligemma-multimodal/ - 多模态数据集元数据 ├── images/ - 包含所有图像文件 特点 包含874张带注释的地质雷达扫描图像 图像预处理为640x640像素大小 支持多种任务类型:缺陷检测、位置定位和描述生成… See the full description on the dataset page: https://huggingface.co/datasets/LiZHENGzai/GPRadar-Defect-MultiTask.imageobject-detection1K<n<10K0 likes11 downloads7mo agoHugging Face24persistent-fm /ctms-multitask-sft-v6gated CTMS Multi-task SFT — V6 A matched pair of corpora for a clinical-trial-management text-to-SQL agent, differing in exactly one variable: whether generate_sql rows carry a <think> reasoning trace. run_a (control) run_b (traced) total 23,049 23,049 train / val / test 18,698 / 2,172 / 2,179 18,698 / 2,172 / 2,179 traced train SQL rows 0 10,125 (81.0%) gold SQL identical, byte-for-byte identical, byte-for-byte Tasks task n… See the full description on the dataset page: https://huggingface.co/datasets/persistent-fm/ctms-multitask-sft-v6.text10K<n<100K0 likes10 downloads2mo agoHugging Face25ThuraAung1601 /reform-dafny-multitask-auggated ReForm Dafny multi-task -- aug Six Dafny tasks per verified program of the ReForm python2dafny data (ThuraAung1601/reform-dafny-loop-inv-gen, ThuraAung1601/reform-dafny-hint-gen, decontaminated against DafnyBench), for multi-task SFT + RL with a Dafny-verifier reward. Built from ThuraAung1601/reform-dafny-multitask (the originals, transform = original) and the verified variants of ThuraAung1601/reform-dafny-loop-inv-gen-aug. Each variant gets the tasks its original has, with the… See the full description on the dataset page: https://huggingface.co/datasets/ThuraAung1601/reform-dafny-multitask-aug.texttext-generation100K<n<1M0 likes10 downloads2d agoHugging Face26ThuraAung1601 /reform-dafny-multitask-aug-combinegated ReForm Dafny multi-task -- aug-combine Six Dafny tasks per verified program of the ReForm python2dafny data (ThuraAung1601/reform-dafny-loop-inv-gen, ThuraAung1601/reform-dafny-hint-gen, decontaminated against DafnyBench), for multi-task SFT + RL with a Dafny-verifier reward. Built from ThuraAung1601/reform-dafny-multitask (the originals, transform = original) and the verified variants of ThuraAung1601/reform-dafny-loop-inv-gen-aug-combine. Each variant gets the tasks its… See the full description on the dataset page: https://huggingface.co/datasets/ThuraAung1601/reform-dafny-multitask-aug-combine.texttext-generation100K<n<1M0 likes10 downloads2d agoHugging Face27samirmsallem /wiki_definitions_de_multitask Dataset Card for Wikipedia Definitions for Multitask (NER/Text Classification) The Wikipedia Definitions for Multitask (NER/Text Classification) dataset is a dataset to train language models to recognize definition sentences and non-definition sentences. The dataset includes training and test data to recognize this discipline by Named Entity Recognition, but also by Sentence Classification. Dataset Sources Wikimedia/wikipedia Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/samirmsallem/wiki_definitions_de_multitask.texttext-classification10K<n<100K0 likes8 downloads1y agoHugging Face28CL-From-Nothing /rlve-multitask-qwen3-4b-n4-randcut512-4096x20-completed-by-qwen3-4b-thinking-r16384tabular10K<n<100K0 likes7 downloads6mo agoHugging Face29persistent-fm /ctms-multitask-sft-v3gated CTMS Multi-Task SFT — V3 (uppercase-Snowflake) Supervised fine-tuning corpus for a Clinical Trial Management System (CTMS) analytics assistant, spanning 7 tasks over a 122-table CTMS schema. This is the V3 build: all SQL uses unquoted identifiers that resolve against the uppercase-identifier Snowflake schema DUMMY_FORTREA_AI_MODEL.FORTREA_AI_MODEL_V3_CAP. Data is fully synthetic (generated from a CTMS data generator). It contains no real patient, investigator, or trial data.… See the full description on the dataset page: https://huggingface.co/datasets/persistent-fm/ctms-multitask-sft-v3.texttext-generation10K<n<100K0 likes7 downloads2mo agoHugging Face30persistent-fm /ctms-multitask-sft-v10-2gatedtext100K<n<1M0 likes4 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.