Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ise-uiuc /Magicoder-OSS-Instruct-75KThis is the OSS-Instruct dataset generated by gpt-3.5-turbo-1106 developed by OpenAI. Please pay attention to OpenAI's usage policy when adopting this dataset: https://openai.com/policies/usage-policies. tabulartext-generation10K<n<100K173 likes47k downloads3y agoHugging Face02BAAI /Infinity-Instructgated Infinity Instruct Beijing Academy of Artificial Intelligence (BAAI) [Paper][Code][🤗] The quality and scale of instruction data are crucial for model performance. Recently, open-source models have increasingly relied on fine-tuning datasets comprising millions of instances, necessitating both high quality and large scale. However, the open-source community has long been constrained by the high costs associated with building such extensive and high-quality instruction… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/Infinity-Instruct.tabulartext-generation10M<n<100M765 likes1.9k downloads10mo agoHugging Face03BEE-spoke-data /code_contests_instruct Dataset Card for "code_contests_instruct" The deepmind/code_contests dataset formatted as markdown-instruct for text generation training. There are several different configs. Look at them. Comments: flesch_reading_ease is computed on the description col via textstat hq means that python2 (aka PYTHON in language column) is dropped, and keeps only rows with flesch_reading_ease 75 or greater min-cols drops all cols except language and text possible values for language are {'CPP'… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/code_contests_instruct.tabulartext-generation10M<n<100M7 likes1.1k downloads10mo agoHugging Face04matlok /python-text-copilot-training-instruct-ai-research-2024-02-03 Python Copilot Instructions on How to Code using Alpaca and Yaml Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the Agora Open Source AI Research Lab: Agora GitHub Organization Agora Hugging Face This dataset is the 2024-02-03 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-02-03.tabulartext-generation1K<n<10K1 likes857 downloads3y agoHugging Face05sujet-ai /Sujet-Finance-Instruct-177k Sujet Finance Dataset Overview The Sujet Finance dataset is a comprehensive collection designed for the fine-tuning of Language Learning Models (LLMs) for specialized tasks in the financial sector. It amalgamates data from 18 distinct datasets hosted on HuggingFace, resulting in a rich repository of 177,597 entries. These entries span across seven key financial LLM tasks, making Sujet Finance a versatile tool for developing and enhancing financial applications of AI.… See the full description on the dataset page: https://huggingface.co/datasets/sujet-ai/Sujet-Finance-Instruct-177k.tabulartext-generation100K<n<1M86 likes847 downloads3y agoHugging Face06JackHsieh /4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids 4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen3-4B-Instruct-2507. Each thought is a few dense sentences of reasoning about the next 8 tokens after a cut, written from the document prefix alone — the generator never sees the continuation. Stored thought_text includes the <thought>/</thought> wrapper. This is the small-generator parity counterpart of… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.tabulartext-generation10M<n<100M0 likes461 downloads25d agoHugging Face07proxectonos /cpt_instruction_datasets Instruction datasets Collection of synthetic instruction datasets used during the continued pretraining of Model-small-instr-1, Model-small-instr-2 and Model-small-instr-3. You can currently find these models under: Llama-3.1-Carballo-Instr1 and Llama-3.1-Carballo-Instr3. Dataset creation Datasets were created using two different techniques: Adapting already existing datasets or corpora by modifying their format to make them suitable for including instructions during… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/cpt_instruction_datasets.tabulartext-generation100K<n<1M0 likes435 downloads6mo agoHugging Face08matlok /python-text-copilot-training-instruct Python Copilot Instructions on How to Code using Alpaca and Yaml This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains python code, either a class method or a global function, imported modules, base classes (if any), exceptions (ordered based off the code), returns (ordered based off the code), arguments (ordered based off the… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct.tabulartext-generation100K<n<1M0 likes417 downloads3y agoHugging Face09daruokta /t5gemma2-indonesia-instruct-v1 T5Gemma-2 Indonesian Instruct — Mono-Repo Satu repositori dataset HF untuk seluruh data pelatihan T5-Gemma-2 bahasa Indonesia. Diorganisasi per fungsi (fondasi → spesifik → preferensi) dengan folder/subfolder, setiap config = folder dan berisi split train + validation (80:20) di level percakapan. Struktur (by fungsi) t5gemma2-indonesia-instruct-v1/ ├── README.md ├── manifest.json ├── chat_idx_map.json ├── foundation/ ← FASE 1 · fondasi Bahasa… See the full description on the dataset page: https://huggingface.co/datasets/daruokta/t5gemma2-indonesia-instruct-v1.imagetext-generation100K<n<1M0 likes411 downloads17d agoHugging Face10matlok /python-text-copilot-training-instruct-ai-research-2024-02-11 Python Copilot Instructions on How to Code using Alpaca and Yaml Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the Autogen and multimodal Qwen AI project: Qwen Qwen Agent Qwen VL Chat Qwen Audio This dataset is the 2024-02-11 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-02-11.tabulartext-generationn<1K0 likes403 downloads3y agoHugging Face11ctaxnagomi /INSTRUCT_JEV INSTRUCT_JEV INSTRUCT_JEV is an instruction corpus built from the TypeSafe AI documentation for Jev, the first System One model. It is structured around the three TypeSafe question primitives - Choice, Noul and Score - and mirrors the raw corpus captured in deckerGUI-jev_corpus_RAW. Credits INSTRUCT_JEV is a DeckerGUI project and exists because of the work below. Who Contribution Link TypeSafe AI Jev - the first System One model - and the Choice / Noul… See the full description on the dataset page: https://huggingface.co/datasets/ctaxnagomi/INSTRUCT_JEV.tabulartext-generationn<1K3 likes317 downloads21d agoHugging Face12DevShubham /python-text-training-instruct-ai Python Copilot Instructions on How to Code using Alpaca and Yaml Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the Agora Open Source AI Research Lab: Agora GitHub Organization Agora Hugging Face This dataset is the 2024-02-03 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details… See the full description on the dataset page: https://huggingface.co/datasets/DevShubham/python-text-training-instruct-ai.tabulartext-generation1K<n<10K1 likes294 downloads2y agoHugging Face13matlok /python-text-copilot-training-instruct-ai-research-2024-01-27 Python Copilot Instructions on How to Code using Alpaca and Yaml This dataset is the 2024-01-27 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains python code, either a class method or a global function, imported modules, base classes (if any), exceptions (ordered based off the code), returns (ordered based off the code), arguments (ordered… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-01-27.tabulartext-generation10K<n<100K0 likes266 downloads3y agoHugging Face14ed001 /ds-coder-instruct-v2 Dataset Card for DS Coder Instruct v2 Dataset Changes from v1: Added WizardLM evol data science samples Removed R samples from v2 DS Coder is a dataset for instruction fine tuning of language models. It is a specialized dataset focusing only on data science (eg. plotting, data wrangling, machine learnig models, deep learning, and numerical computations). The dataset contains code examples both in Python (R samples were removed in v2). The goal of this dataset is to enable… See the full description on the dataset page: https://huggingface.co/datasets/ed001/ds-coder-instruct-v2.tabulartext-generation10K<n<100K13 likes263 downloads3y agoHugging Face15Aratako /Magpie-Tanuki-Instruction-Selected-Evolved-26.5k Magpie-Tanuki-Instruction-Selected-Evolved-26.5k 概要 以下の手順で作成した約2万6500件の日本語の合成instructionデータセットです。 Magpieの手法をteam-hatakeyama-phase2/Tanuki-8x8B-dpo-v1.0-GPTQ-8bitに適用し、約10万件のinstructionを作成 cl-nagoya/ruri-largeを使ってinstructionのベクトル表現を取得 この時点のデータはAratako/Magpie-Tanuki-Instruction-100k-Embeddingsで公開されています。 取得したベクトル表現を元に、Mini Batch K-Meansによって20000個のクラスタにクラスタリング 各クラスタから最大3個までinstructionを抽出 上記で抽出した約2万6500件のinstructionに対し、Qwen/Qwen2.5-72B-Instruct-GPTQ-Int8を使ってEvol-Instructを適用… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Magpie-Tanuki-Instruction-Selected-Evolved-26.5k.tabulartext-generation10K<n<100K0 likes223 downloads2y agoHugging Face16Praha-Labs /DaTikZ-V4-Instruct-100K DaTikZ-V4 Instruction 100K A cleaned instruction-tuning dataset for text-to-TikZ generation, derived from the first 100,000 rows of nllg/DaTikZ-V4. The original dataset provides rendered diagram images and corresponding TikZ source code. This derived dataset adds natural-language, imperative user instructions that describe how to recreate each diagram. These instructions are intended as model inputs, with the original TikZ code serving as the supervised target.… See the full description on the dataset page: https://huggingface.co/datasets/Praha-Labs/DaTikZ-V4-Instruct-100K.imagetext-generation100K<n<1M0 likes215 downloads15d agoHugging Face17wttw /code_contest_instruct_cpptabulartext-generation1M<n<10M3 likes214 downloads2y agoHugging Face18matlok /python-text-copilot-training-instruct-ai-research-2024-02-10 Python Copilot Instructions on How to Code using Alpaca and Yaml Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the multimodal Qwen AI project: Qwen Qwen Agent Qwen VL Chat Qwen Audio This dataset is the 2024-02-10 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-02-10.tabulartext-generationn<1K0 likes208 downloads3y agoHugging Face19JackHsieh /4B-Instruct-RL-step2176-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids 4B-Instruct-RL-step2176-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen3-4B-Instruct-2507.prestar-RL.reason-only.lr7e-7-kl0.step-2176 — Qwen3-4B-Instruct-2507 after RL against a frozen suffix conditional. Each thought is a few dense sentences of reasoning about the next 8 tokens after a cut, written from the document prefix alone — the generator never sees the… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-RL-step2176-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.tabulartext-generation10M<n<100M0 likes177 downloads17d agoHugging Face20aisingapore /SEA-Instruct-2602gated SEA-Instruct-2602 Overview SEA-Instruct-2602 is a preliminary release of instruction-tuning data focused on Southeast Asian languages and contexts. The dataset combines prompts filtered from open-source data with our own synthetic prompts, paired with synthetic responses, for language model training on SEA-specific tasks and languages. This dataset contains only the filtered subset of data with prompt_input_quality at Excellent, prompt_is_coherent at True and… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/SEA-Instruct-2602.tabulartext-generation1M<n<10M4 likes162 downloads8mo agoHugging Face21JackHsieh /4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained 4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained Tokenized, tag-wrapped form of JackHsieh/4B-reason-only.rule-r-1.0-k-8.L-512.statml-arxiv. Each thought is wrapped as <|note|> This is a hint about a span that appears later in this document. KEY is the text immediately before that span; VALUE is a note about what might come next. KEY: <last 8 prefix tokens> VALUE: <thought> <|/note|> and stored both as text (thought_text) and as Qwen/Qwen3-4B-Base… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.tabulartext-generation10M<n<100M0 likes161 downloads26d agoHugging Face22matlok /python-text-copilot-training-instruct-ai-research Building an AI Copilot Dataset to help keep up with Leading AI Research This is a specialized, instruction dataset for training python coding assistants on how to code from leading AI/ML open source repositories (2.3M coding samples). This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details This dataset holds the latest coding changes from >1159… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research.tabulartext-generation10K<n<100K0 likes146 downloads3y agoHugging Face23agagasf123123 /threejs-gamecode-instruct-v3-ultra Three.js GameCode Instruct v3 Ultra This is a large synthetic/original instruction dataset for training or testing LLM behavior around Three.js, browser game development, gameplay programming, debugging, optimization, architecture, and general coding. Important note This dataset is synthetic and programmatically generated from original templates. It is designed as a useful starting point for experiments, not as a fully hand-curated gold-standard benchmark. No… See the full description on the dataset page: https://huggingface.co/datasets/agagasf123123/threejs-gamecode-instruct-v3-ultra.tabulartext-generation10K<n<100K2 likes141 downloads4mo agoHugging Face24viyer98 /xl-instruct Dataset Card for XL-Instruct This dataset card provides a summary of the XL-Instruct dataset, a resource for advancing the cross-lingual capabilities of Large Language Models. It was introduced in the paper XL-Instruct: Synthetic Data for Cross-Lingual Open-Ended Generation. Dataset Details Dataset Description XL-Instruct is a high-quality, large-scale synthetic dataset designed to fine-tune LLMs for cross-lingual open-ended generation. The core task involves… See the full description on the dataset page: https://huggingface.co/datasets/viyer98/xl-instruct.tabulartext-generation100K<n<1M2 likes133 downloads1y agoHugging Face25JackHsieh /4B-Instruct-RL-step2176-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained 4B-Instruct-RL-step2176-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained Tokenized, tag-wrapped form of JackHsieh/4B-Instruct-RL-step2176-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids, the thoughts written by the prestar-RL policy Qwen3-4B-Instruct-2507.prestar-RL.reason-only.lr7e-7-kl0.step-2176. Each thought is wrapped as <|note|> This is a hint about a span that appears later in this document. KEY is the text immediately before that span; VALUE is a note… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-RL-step2176-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.tabulartext-generation10M<n<100M0 likes126 downloads16d agoHugging Face26xlr8harder /synthid-qwen3-4b-instruct-2507-wildchat Qwen3-4B SynthID three-arm corpus This export contains aligned unwatermarked, SynthID key-A, and SynthID key-B responses from Qwen/Qwen3-4B-Instruct-2507. Matched splits share prompts and request seeds across configurations; unmatched splits use mutually disjoint prompt pools. Export complete for its source work queue: true. Generation profile Model revision: cdbee75f17c01a7cc42f958dc650907174af0554 Native model dtype: bfloat16 Maximum generated tokens: 4096… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/synthid-qwen3-4b-instruct-2507-wildchat.tabulartext-generation100K<n<1M0 likes123 downloads2mo agoHugging Face27nuhmanpk /cybersecurity-controls-instructions Cybersecurity Controls Instructions Security control, incident response and risk management guidance from NIST Special Publications, turned into instruction-following examples. Splits split rows source documents train 13,106 56 validation 4,840 18 test 5,697 18 Splits are held out by source document. Every chunk yields several instruction rows, so a random row-level split would place the same passage in train and test; whole documents are held… See the full description on the dataset page: https://huggingface.co/datasets/nuhmanpk/cybersecurity-controls-instructions.tabulartext-generation10K<n<100K0 likes116 downloads21d agoHugging Face28NLTF-mock /tw-instruct-500k-Q-R1 tw-instruct-500k-Q-R1 台灣常見任務對話集(Common Task-Oriented Dialogues in Taiwan) 為台灣社會裡常見的任務對話,從 lianghsun/tw-instruct 截取出 50 萬筆的子集合版本,進行理解力(Reasoning)資料補充生成。 Dataset Details Dataset Description 這個資料集為合成資料集(synthetic datasets),內容由 a. reference-based 和 b. reference-free 的子資料集組合而成。生成 reference-based 資料集時,會先以我們收集用來訓練 lianghsun/Llama-3.2-Taiwan-3B 時的繁體中文文本作為參考文本,透過 LLM 去生成指令對話集,如果參考文本有特別領域的問法,我們將會特別設計該領域或者是適合該文本的問題;生成 reference-free 時,則是以常見的種子提示(seed prompts)作為參考,讓 LLM… See the full description on the dataset page: https://huggingface.co/datasets/NLTF-mock/tw-instruct-500k-Q-R1.tabulartext-generation100K<n<1M1 likes100 downloads2y agoHugging Face29Yooniel /qwen2.5-7b-instruct-nla-L20-finefineweb-100k Qwen2.5-7B-Instruct NLA training data — residual stream, block 20 Training data for a Natural Language Autoencoder on Qwen/Qwen2.5-7B-Instruct: residual-stream activations paired with natural-language explanations of the text they were taken from. Unlike the dataset this is derived from, the activation_vector column is included — every EasyNLA/nanoNLA trainer requires it. Trained models: https://huggingface.co/Yooniel/qwen2.5-7b-instruct-nla-L20 (AV val ppl 4.07, AR held-out FVE… See the full description on the dataset page: https://huggingface.co/datasets/Yooniel/qwen2.5-7b-instruct-nla-L20-finefineweb-100k.tabulartext-generation100K<n<1M0 likes95 downloads3mo agoHugging Face30SolusOps /incremental-instruction-creative-writinggated Incremental Instruction Creative Writing Does delivering a writing brief over several conversation turns change what a language model writes? This dataset supports that question with matched creative-writing tasks evaluated under two delivery conditions: FULL: the complete brief is supplied in one turn. SHARDED: the same intended brief is introduced across five to nine turns. The benchmark holds task content fixed while varying how the instructions are delivered. It is… See the full description on the dataset page: https://huggingface.co/datasets/SolusOps/incremental-instruction-creative-writing.tabulartext-generation1K<n<10K0 likes92 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.