Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nickrosh /Evol-Instruct-Code-80k-v1Open Source Implementation of Evol-Instruct-Code as described in the WizardCoder Paper. Code for the intruction generation can be found on Github as Evol-Teacher. text10K<n<100K251 likes7.7k downloads3y agoHugging Face02TokenBender /code_instructions_122k_alpaca_styletext100K<n<1M80 likes3.3k downloads3y agoHugging Face03cfahlgren1 /react-code-instructions React Code Instructions Popular Queries Number of instructions by Model Unnested Messages Instructions Added Per Day Dataset of Claude Artifact esque React Apps generated by Llama 3.1 70B, Llama 3.1 405B, and Deepseek Chat V3. Examples Virtual Fitness Trainer Website LinkedIn Clone iPhone Calculator Chipotle Waitlist Apple Store text10K<n<100K158 likes330 downloads2y agoHugging Face04NickIBrody /python-code-instructions-85k Python Code Instructions - 85K Instruction-tuning dataset of Python functions paired with short natural-language instructions derived from repository docstrings. What changed in this release This release keeps the original public rows and format, but makes the dataset easier to use responsibly: exact duplicate rows were removed again using normalized instruction + output hashing deterministic train, validation, and test splits were added the dataset card now documents… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/python-code-instructions-85k.texttext-generation10K<n<100K1 likes298 downloads6mo agoHugging Face05AtlasUnified /Code-Instruct-Setstext100K<n<1M6 likes147 downloads3y agoHugging Face06kunishou /amenokaku-code-instruct Amenokaku-Code-Instruct Update: 2023/12/27データセットに JaxTon , プロになるJava のコードデータ 180 レコードを追加しました。 概要 コードに特化した5.2KのInstructionデータセットです。 データセットに含まれるデータは商用利用できるラインセンスが付与されたプログラミング学習コンテンツから収集、加工し作成しました(英語のコンテンツは日本語に自動翻訳し、翻訳の不自然な箇所を手動で修正)。 また、ライセンスが明記されていない学習コンテンツについては権利者に個別に連絡を取り、本データセットへの掲載の許諾を得ております。 データセット詳細 指示タスクの内訳としてはコード生成(code_generation)が1050レコード、コードの挙動確認(check_code_behavor)が150レコード、コードのバグ修正(code_fix)が4000レコードになります。 詳細な内訳は以下の通りになります。 source name… See the full description on the dataset page: https://huggingface.co/datasets/kunishou/amenokaku-code-instruct.text1K<n<10K17 likes122 downloads3y agoHugging Face07Maxyelow /kenyan-code-switch-instruct-50k 🇰🇪 Kenyan Code-Switching & Sheng Multi-Task Instruction Dataset (50,000 Pairs) A standardized, multi-task instruction-tuning dataset engineered to teach Large Language Models (e.g. Llama 3, Mistral, Gemma, Qwen) to understand and generate authentic Kenyan Code-Switching (Sheng, Technical Swahili-English Blend) with rigorous adherence to Bantu morphotactic rules. Dataset Summary Total Samples: 50,000 instruction-response pairs train.jsonl: 45,000 pairs (90%)… See the full description on the dataset page: https://huggingface.co/datasets/Maxyelow/kenyan-code-switch-instruct-50k.texttext-generation10K<n<100K0 likes78 downloads12d agoHugging Face08alztrk /turkish-code-instructions Turkish Code Instructions Turkish instruction-code pairs for LLM fine-tuning. text1K<n<10K0 likes76 downloads4mo agoHugging Face09ewof /code-alpaca-instruct-unfilteredThis dataset is HuggingFaceH4/CodeAlpaca_20K unfiltered, removing 36 instances of blatant alignment. 19986 instructions remain. https://huggingface.co/datasets/HuggingFaceH4/CodeAlpaca_20K/blob/29ba7b7fdf0c55e5435c848cf6bbf9782fef62a6/data/test-00000-of-00001.parquet https://huggingface.co/datasets/HuggingFaceH4/CodeAlpaca_20K/blob/a123ae447f02484d83c3457438b4422cd8417ad5/data/train-00000-of-00001.parquet i combined all of these files above into code_alpaca_data.jsonl with parquet2json and ran… See the full description on the dataset page: https://huggingface.co/datasets/ewof/code-alpaca-instruct-unfiltered.text10K<n<100K9 likes70 downloads3y agoHugging Face10TaskPuppyAI /codex-code-instructions-300 Codex Multilingual Code Implementation & Repair 300 A 300-record synthetic coding dataset generated with a Codex-family system and organized around implementation, repair, edge-case handling, and already-correct code review tasks. The exact Codex model/version and original generation configuration could not be recovered from the available provenance records. The recovered final dataset is exactly the union of six corrected 50-record batches. Dataset Summary The… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/codex-code-instructions-300.textn<1K0 likes69 downloads1mo agoHugging Face11Pasta009 /Instruction-Fusion-Code-v1 Instruction Fusion Dataset Aproximately 100k data samples for code generation. Fused using seeds from evol-codealpaca-v1. Generator: gpt-4-1106-preview Fine-tune Use 'instruction' and 'output' for SFT. 'prompt1' and prompt2' are the seed instructions used for instruction fusion. citation If you use this dataset, please cite our paper @article{guo2023instruction, title={Instruction fusion: Advancing prompt evolution through hybridization}, author={Guo… See the full description on the dataset page: https://huggingface.co/datasets/Pasta009/Instruction-Fusion-Code-v1.text100K<n<1M1 likes62 downloads2y agoHugging Face12open-llm-leaderboard /EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-detailsgated Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-details.tabular10K<n<100K4 likes62 downloads2y agoHugging Face13TokenBender /unnatural_code_instructions_20M_tokens_separatetext100K<n<1M4 likes54 downloads3y agoHugging Face14TokenBender /unnatural_code_instructions_20Mtext100K<n<1M4 likes51 downloads3y agoHugging Face15CodeResearch /Code-Evol-Instruct-OSS Code-Evol-Instruct-OSS Summary Code-Evol-Instruct-OSS is a dataset that was generated with Code Evol-Instruct by prompting open-souce LLMs, WizardLM-13B-v1.2 and WizardCoder-34B-Python. The underlying process is explained in the paper code-evol-instruct. This algorithm gave birth to famous open-souce code LLMs, WizardCoder-Family. Our approach We did not use any closed-source LLMs. Our seed dataset is sourced from self-instruct-starcoder. We leverage the… See the full description on the dataset page: https://huggingface.co/datasets/CodeResearch/Code-Evol-Instruct-OSS.tabular1K<n<10K6 likes48 downloads3y agoHugging Face16rombodawg /code_instruct_alpaca_vicuna_wizardlm_56k_backupBackup of code_instruct_alpaca_vicuna_wizardlm used in rombodawg/MegaCodeTraining112k Link to the combined dataset bellow https://huggingface.co/datasets/rombodawg/MegaCodeTraining112k text10K<n<100K2 likes38 downloads3y agoHugging Face17open-llm-leaderboard /EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.005-128K-code-COT-detailsgated Dataset Card for Evaluation run of EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.005-128K-code-COT Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.005-128K-code-COT The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.005-128K-code-COT-details.tabular10K<n<100K0 likes38 downloads2y agoHugging Face18lodrick-the-lafted /asl-code-sonnet35-instructSources: https://huggingface.co/datasets/Augmentation-Scaling-Laws/code_claude3.5_sonnet_10000 https://huggingface.co/datasets/Augmentation-Scaling-Laws/magpie_code_claude3.5_sonnet_10000 https://huggingface.co/datasets/Augmentation-Scaling-Laws/webinstruct_code_claude3.5_sonnet_10000 converted to sharegpt text10K<n<100K1 likes38 downloads2y agoHugging Face19dlface /code_instructions_120k_alpaca_stepbysteptext100K<n<1M1 likes36 downloads2y agoHugging Face20TaskPuppyAI /qwen3.8-code-instructions-350 Qwen3.8 Max Python-Weighted Code Instructions 350 A 350-record synthetic programming instruction dataset generated with Qwen3.8 Max and reviewed with ChatGPT 5.6 Sol High. The dataset was designed as a Python-weighted mixed-programming set. The recovered filename and dataset creator recollection indicate an intended distribution of approximately 60% Python, although the exact language distribution was not independently reconstructed from the final artifact. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/qwen3.8-code-instructions-350.textn<1K0 likes36 downloads1mo agoHugging Face21open-llm-leaderboard /EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT-detailsgated Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT-details.tabular10K<n<100K0 likes35 downloads2y agoHugging Face22duxx /code-instruction-turkishtext10K<n<100K2 likes32 downloads2y agoHugging Face23RhodWeo /gis-code-instructions GIS Code Instructions Dataset Expert-curated instruction dataset for fine-tuning code models on Geographic Information Systems (GIS) tasks. 📊 Dataset Stats 70 unique examples in conversational messages format 13 GIS Python libraries covered Each example includes: system prompt + user instruction + assistant response with Chain-of-Thought reasoning and complete Python code 📁 Files File Description data/train.jsonl Full dataset (70 examples… See the full description on the dataset page: https://huggingface.co/datasets/RhodWeo/gis-code-instructions.texttext-generationn<1K0 likes32 downloads6mo agoHugging Face24open-llm-leaderboard /EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-ds-auto-detailsgated Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-ds-auto Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-ds-auto The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-ds-auto-details.tabular10K<n<100K0 likes30 downloads2y agoHugging Face25fchis /Laravel-13x-Code-Instructionstextn<1K0 likes30 downloads7mo agoHugging Face26mesolitica /chatgpt4-code-instruct ChatGPT4 Code Instruct Originally from https://huggingface.co/datasets/theblackcat102/evol-codealpaca-v1, translate and answer using ChatGPT4. Notebooks at https://github.com/mesolitica/malaysian-dataset/tree/master/chatbot/chatgpt4-code-instruct synthetic-codealpaca-v1-chatgpt4.jsonl, 43482 rows, 274 MB. Example data {'instruction': "Harap ubah skrip Python berikut agar ia memasukkan pengulangan 'while' daripada pengulangan 'for' yang sedia ada, yang meneruskan… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/chatgpt4-code-instruct.text10K<n<100K1 likes29 downloads2y agoHugging Face27open-llm-leaderboard /EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-detailsgated Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-details.tabular10K<n<100K0 likes29 downloads2y agoHugging Face28open-llm-leaderboard /EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math-detailsgated Dataset Card for Evaluation run of EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math-details.tabular10K<n<100K0 likes29 downloads2y agoHugging Face29open-llm-leaderboard /EpistemeAI__Fireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPO-detailsgated Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPO Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPO The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPO-details.tabular10K<n<100K0 likes28 downloads2y agoHugging Face30TaskPuppyAI /qwen3.8-targeted-code-instructions-350 Qwen3.8 Max Targeted Bug Detection & Repair 350 A 350-record synthetic targeted programming dataset generated with Qwen3.8 Max and subsequently subjected to a full semantic correction audit with ChatGPT / OpenAI. The dataset emphasizes correctness judgment, bug detection, debugging, repair, and tightly constrained programming tasks. Dataset Summary The publication artifact contains 350 unique records using the schema: { "instruction": "...", "input": "..."… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/qwen3.8-targeted-code-instructions-350.textn<1K0 likes28 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.