Team Ai
27 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01vicgalle /configurable-system-prompt-multitask Configurable System Prompt Multi-task Dataset 🛞 We release the synthetic dataset for the multi-task experiments from the paper "Configurable Safety Tuning of Language Models with Synthetic Preference Data", https://huggingface.co/papers/2404.00495. This dataset has two sources for the examples: Self-critique on a safety task from Harmful Behaviours, using the SOLAR-Instruct model. It employs two system prompts to learn the different behaviors: You are a helpful yet harmless… See the full description on the dataset page: https://huggingface.co/datasets/vicgalle/configurable-system-prompt-multitask.texttext-generation1K<n<10K29 likes111 downloads2y agoHugging Face02awni00 /multi-strategy-algorithmic-tasks Multi-Strategy Algorithmic Tasks A synthetic benchmark of parseable algorithmic problems with multiple valid solution strategies for each task. Each example contains a problem,a strategy-specific solution trace, and the strategy used to generate that trace. The benchmark accompanies Uncovering Latent Reasoning Strategies in Language Models, which studies the problem of recovering mixtures of strategies implicitly represented in language models. The benchmark provides a… See the full description on the dataset page: https://huggingface.co/datasets/awni00/multi-strategy-algorithmic-tasks.texttext-generation1M<n<10M0 likes98 downloads2mo agoHugging Face03mangi-llm /Kazakh_Multi-Task_corpus Kazakh Multi-task Corpus A multi-task NLP dataset in the Kazakh language, covering seven distinct language tasks - from instruction-following and question answering to translation, sentiment analysis, and grammar exercises. Designed to support the development of Kazakh-language models, benchmarks, and linguistic research. Dataset Summary Kazakh is a Turkic language spoken by over 13 million people, yet it remains significantly underrepresented in NLP research and… See the full description on the dataset page: https://huggingface.co/datasets/mangi-llm/Kazakh_Multi-Task_corpus.text-generation1K<n<10K1 likes80 downloads7mo agoHugging Face04kaustubhg73 /multilingual-multitask-refusal Multilingual Multitask Refusal A multilingual, multitask prompt dataset for analysing refusal behaviour across language, task wrapper, and harmful / harmless labels. English seed spans and labels come from the previous Multitask Multilingual Refusal dataset. Non-English content is produced with Google Sheets GOOGLETRANSLATE. Task instructions are language-localized via templates_localized.json. Rows 211,320 English seeds 1,761 Languages 15 Tasks 8 Product 1… See the full description on the dataset page: https://huggingface.co/datasets/kaustubhg73/multilingual-multitask-refusal.texttext-generation100K<n<1M0 likes80 downloads1mo agoHugging Face05Lots-of-LoRAs /task638_multi_woz_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task638_multi_woz_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task638_multi_woz_classification.texttext-generation1K<n<10K0 likes70 downloads2y agoHugging Face06Lots-of-LoRAs /task1577_amazon_reviews_multi_japanese_language_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1577_amazon_reviews_multi_japanese_language_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1577_amazon_reviews_multi_japanese_language_classification.texttext-generationn<1K0 likes67 downloads2y agoHugging Face07narendarcodes /Telugu-MultiTask-Instruct-77K Telugu MultiTask Instruct 77K — Adaption AutoScientist Challenge Dataset Powered by Adaptive Data — Adaption Labs Dataset Description A large-scale, multi-task Telugu instruction-tuning dataset combining 77,653 rows from 7 open-source Telugu NLP collections. Covers diverse tasks including news summarization, QA, creative writing, translation, and general instruction following — all processed through the Adaption Labs AutoScientist platform for quality… See the full description on the dataset page: https://huggingface.co/datasets/narendarcodes/Telugu-MultiTask-Instruct-77K.textquestion-answering10K<n<100K1 likes66 downloads3mo agoHugging Face08Nini0la /edgeimci-beta0-1k-multitask-enriched-2258-v1 EdgeIMCI Beta0-1K Multitask Enriched 2258 Dataset summary This is the public, message-normalized publication of the dataset used to fine-tune the selected EdgeIMCI 4B-alpha-second research checkpoint from the pinned Qwen/Qwen3-4B base model. It combines structured extraction, routing, bounded clinical-response, project self-knowledge, scope-safety, and proposition/negation examples for the EdgeIMCI sick-child assessment workflow. The dataset is research evidence… See the full description on the dataset page: https://huggingface.co/datasets/Nini0la/edgeimci-beta0-1k-multitask-enriched-2258-v1.texttext-generation1K<n<10K0 likes57 downloads16d agoHugging Face09mujo-labs /sandman-dream_multitask_test Sandman dream multitask v1 — test split The first version of the test split used to train sandman-gemma3-1b-multitask. Superseded by v2, a smaller, more curated set built on DreamBank rather than this one's broader source. Kept here for reference. texttext-generation10K<n<100K0 likes41 downloads19d agoHugging Face10mujo-labs /sandman-dream_multitask_v2_train Sandman dream multitask v2 — train split 17,300 instruction-following examples for fine-tuning Sandman's on-device dream-analysis model, built from sandman-dreambank-v2. Every row is a single-turn conversation (messages) covering one of three tasks: Summarize — read a dream, return a one- or two-sentence summary as JSON. Extract symbols — return only the concrete nouns literally present in the dream text, as a JSON array, with an explicit instruction not to infer or add… See the full description on the dataset page: https://huggingface.co/datasets/mujo-labs/sandman-dream_multitask_v2_train.texttext-generation10K<n<100K0 likes40 downloads19d agoHugging Face11mujo-labs /sandman-dream_multitask_train Sandman dream multitask v1 — train split The first version of the train split used to train sandman-gemma3-1b-multitask. Superseded by v2, a smaller, more curated set built on DreamBank rather than this one's broader source. Kept here for reference. texttext-generation100K<n<1M0 likes34 downloads19d agoHugging Face12Lots-of-LoRAs /task1575_amazon_reviews_multi_sentiment_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1575_amazon_reviews_multi_sentiment_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1575_amazon_reviews_multi_sentiment_classification.texttext-generationn<1K0 likes33 downloads2y agoHugging Face13mujo-labs /sandman-dream_multitask_val Sandman dream multitask v1 — val split The first version of the val split used to train sandman-gemma3-1b-multitask. Superseded by v2, a smaller, more curated set built on DreamBank rather than this one's broader source. Kept here for reference. texttext-generation10K<n<100K0 likes32 downloads19d agoHugging Face14TheTokenFactory /sec-extraction-multitask-v4 SEC Extraction Multitask v4 Instruction-tuning dataset for fine-tuning a small language model (e.g. Gemma 4 E2B) to extract structured data from SEC filings across three verticals: Exhibit 10 (contracts) — financial terms from executive employment, credit agreements, indemnification, licensing, and similar filings DEF 14A (proxy statements) — executive compensation, governance items, say-on-pay MD&A (10-K / 10-Q Management's Discussion & Analysis) — operating metrics, segment… See the full description on the dataset page: https://huggingface.co/datasets/TheTokenFactory/sec-extraction-multitask-v4.texttext-generation1K<n<10K0 likes28 downloads6mo agoHugging Face15mujo-labs /sandman-dream_multitask_v2_test Sandman dream multitask v2 — test split The test split for fine-tuning Sandman's on-device dream-analysis model (v2). See sandman-dream_multitask_v2_train for the full description of the three tasks (summarize, extract symbols, interpret a symbol) and the source data. texttext-generation1K<n<10K0 likes26 downloads19d agoHugging Face16Lots-of-LoRAs /task1576_amazon_reviews_multi_english_language_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1576_amazon_reviews_multi_english_language_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1576_amazon_reviews_multi_english_language_classification.texttext-generationn<1K0 likes24 downloads2y agoHugging Face17mujo-labs /sandman-dream_multitask_v2_val Sandman dream multitask v2 — val split The val split for fine-tuning Sandman's on-device dream-analysis model (v2). See sandman-dream_multitask_v2_train for the full description of the three tasks (summarize, extract symbols, interpret a symbol) and the source data. texttext-generation1K<n<10K0 likes23 downloads19d agoHugging Face18Phettae /thai-multitask-starter Thai Multitask 9.6K ชุดข้อมูลตั้งต้นสำหรับ instruction tuning ภาษาไทย ครอบคลุมงานสนทนา ถาม–ตอบ สรุป แปล จำแนกข้อความ ตรวจแก้ภาษา คณิตศาสตร์ และ structured output ข้อมูลทุกแถวสร้างขึ้นใหม่ด้วยกฎแบบ deterministic ไม่มีการคัดลอกจากเว็บไซต์หรือ ข้อมูลส่วนบุคคลจริง เหมาะสำหรับทดลอง supervised fine-tuning และทดสอบ pipeline แต่ควรเพิ่มข้อมูลที่มนุษย์ตรวจทานและข้อมูลภาษาธรรมชาติก่อนใช้กับระบบจริง จำนวนข้อมูลทั้งหมด 9,599 ตัวอย่าง: train 8,639, validation 480 และ test 480… See the full description on the dataset page: https://huggingface.co/datasets/Phettae/thai-multitask-starter.texttext-generation1K<n<10K0 likes19 downloads2mo agoHugging Face19ThuraAung1601 /dafnybench-multitaskgated DafnyBench multi-task ThuraAung1601/dafnybench-cleaned turned into six Dafny tasks, for evaluation. All comments were removed from every program (the ReForm training programs have none); each ground truth was re-verified with Dafny 4.11.0 afterwards. 731 of 733 programs are used (skipped: 2 ground truth uses {:verify false}, 1 loop-invariants: ground truth fails its own faithfulness check, 1 termination: ground truth fails its own faithfulness check). task the model gets… See the full description on the dataset page: https://huggingface.co/datasets/ThuraAung1601/dafnybench-multitask.texttext-generation1K<n<10K0 likes19 downloads3d agoHugging Face20ThuraAung1601 /reform-dafny-multitaskgated ReForm Dafny multi-task Six Dafny tasks per verified program of the ReForm python2dafny data (ThuraAung1601/reform-dafny-loop-inv-gen, ThuraAung1601/reform-dafny-hint-gen, decontaminated against DafnyBench), for multi-task SFT + RL with a Dafny-verifier reward. task the model gets answer is scored by loop-invariants the program with its loop invariants removed dafny verify + only proof hints may differ proof-annotations every line starting with invariant / assert /… See the full description on the dataset page: https://huggingface.co/datasets/ThuraAung1601/reform-dafny-multitask.texttext-generation10K<n<100K0 likes19 downloads2d agoHugging Face21Lots-of-LoRAs /task1574_amazon_reviews_multi_language_identification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1574_amazon_reviews_multi_language_identification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1574_amazon_reviews_multi_language_identification.texttext-generationn<1K0 likes17 downloads2y agoHugging Face22Miladsaeedi70 /scientific-multitask-instructions Scientific Multitask Instructions A multi-task scientific instruction-following dataset created for supervised fine-tuning and preference-optimization experiments. Dataset summary The dataset contains 1,576 conversational scientific examples across eight task types. Split Examples Train 1,260 Validation 158 Test 158 Total 1,576 Task distribution Task Examples Scientific question answering 256 Summarization 220… See the full description on the dataset page: https://huggingface.co/datasets/Miladsaeedi70/scientific-multitask-instructions.texttext-generation1K<n<10K0 likes17 downloads3mo agoHugging Face23ThuraAung1601 /reform-dafny-multitask-aug-combinegated ReForm Dafny multi-task -- aug-combine Six Dafny tasks per verified program of the ReForm python2dafny data (ThuraAung1601/reform-dafny-loop-inv-gen, ThuraAung1601/reform-dafny-hint-gen, decontaminated against DafnyBench), for multi-task SFT + RL with a Dafny-verifier reward. Built from ThuraAung1601/reform-dafny-multitask (the originals, transform = original) and the verified variants of ThuraAung1601/reform-dafny-loop-inv-gen-aug-combine. Each variant gets the tasks its… See the full description on the dataset page: https://huggingface.co/datasets/ThuraAung1601/reform-dafny-multitask-aug-combine.texttext-generation100K<n<1M0 likes15 downloads2d agoHugging Face24Lots-of-LoRAs /task639_multi_woz_user_utterance_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task639_multi_woz_user_utterance_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task639_multi_woz_user_utterance_generation.texttext-generationn<1K0 likes14 downloads2y agoHugging Face25ThuraAung1601 /reform-dafny-multitask-auggated ReForm Dafny multi-task -- aug Six Dafny tasks per verified program of the ReForm python2dafny data (ThuraAung1601/reform-dafny-loop-inv-gen, ThuraAung1601/reform-dafny-hint-gen, decontaminated against DafnyBench), for multi-task SFT + RL with a Dafny-verifier reward. Built from ThuraAung1601/reform-dafny-multitask (the originals, transform = original) and the verified variants of ThuraAung1601/reform-dafny-loop-inv-gen-aug. Each variant gets the tasks its original has, with the… See the full description on the dataset page: https://huggingface.co/datasets/ThuraAung1601/reform-dafny-multitask-aug.texttext-generation100K<n<1M0 likes14 downloads2d agoHugging Face26nmd2k /multi-task-instructiontexttext-generation100K<n<1M0 likes12 downloads3y agoHugging Face27persistent-fm /ctms-multitask-sft-v3gated CTMS Multi-Task SFT — V3 (uppercase-Snowflake) Supervised fine-tuning corpus for a Clinical Trial Management System (CTMS) analytics assistant, spanning 7 tasks over a 122-table CTMS schema. This is the V3 build: all SQL uses unquoted identifiers that resolve against the uppercase-identifier Snowflake schema DUMMY_FORTREA_AI_MODEL.FORTREA_AI_MODEL_V3_CAP. Data is fully synthetic (generated from a CTMS data generator). It contains no real patient, investigator, or trial data.… See the full description on the dataset page: https://huggingface.co/datasets/persistent-fm/ctms-multitask-sft-v3.texttext-generation10K<n<100K0 likes8 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.