Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01likaixin /InstructCoder Paper | Code | Blog InstructCoder (CodeInstruct): Empowering Language Models to Edit Code Updates May 23, 2023: Paper, code and data released. Overview InstructCoder is the first dataset designed to adapt LLMs for general code editing. It consists of 114,239 instruction-input-output triplets and covers multiple distinct code editing scenarios, generated by ChatGPT. LLaMA-33B finetuned on InstructCoder performs on par with ChatGPT on a… See the full description on the dataset page: https://huggingface.co/datasets/likaixin/InstructCoder.texttext-generation100K<n<1M17 likes9.5k downloads2y agoHugging Face02nickrosh /Evol-Instruct-Code-80k-v1Open Source Implementation of Evol-Instruct-Code as described in the WizardCoder Paper. Code for the intruction generation can be found on Github as Evol-Teacher. text10K<n<100K251 likes8.2k downloads3y agoHugging Face03ed001 /ds-coder-instruct-v1 Dataset Card for DS Coder Instruct Dataset DS Coder is a dataset for instruction fine tuning of language models. It is a specialized dataset focusing only on data science (eg. plotting, data wrangling, machine learnig models, deep learning, and numerical computations). The dataset contains code examples both in R and Python. The goal of this dataset is to enable creation of small-scale, specialized language model assistants for data science projects. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/ed001/ds-coder-instruct-v1.imagetext-generation10K<n<100K5 likes366 downloads3y agoHugging Face04ed001 /ds-coder-instruct-v2 Dataset Card for DS Coder Instruct v2 Dataset Changes from v1: Added WizardLM evol data science samples Removed R samples from v2 DS Coder is a dataset for instruction fine tuning of language models. It is a specialized dataset focusing only on data science (eg. plotting, data wrangling, machine learnig models, deep learning, and numerical computations). The dataset contains code examples both in Python (R samples were removed in v2). The goal of this dataset is to enable… See the full description on the dataset page: https://huggingface.co/datasets/ed001/ds-coder-instruct-v2.tabulartext-generation10K<n<100K13 likes226 downloads3y agoHugging Face05CodeDevX /Vibe-Coding-Instructtexttext-generation1M<n<10M189 likes149 downloads4mo agoHugging Face06AtlasUnified /Code-Instruct-Setstext100K<n<1M6 likes146 downloads3y agoHugging Face07CodeDevX /Vibe-Coding-Instruct-V2 Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/CodeDevX/Vibe-Coding-Instruct-V2.texttext-classification1M<n<10M11 likes125 downloads4mo agoHugging Face08kunishou /amenokaku-code-instruct Amenokaku-Code-Instruct Update: 2023/12/27データセットに JaxTon , プロになるJava のコードデータ 180 レコードを追加しました。 概要 コードに特化した5.2KのInstructionデータセットです。 データセットに含まれるデータは商用利用できるラインセンスが付与されたプログラミング学習コンテンツから収集、加工し作成しました(英語のコンテンツは日本語に自動翻訳し、翻訳の不自然な箇所を手動で修正)。 また、ライセンスが明記されていない学習コンテンツについては権利者に個別に連絡を取り、本データセットへの掲載の許諾を得ております。 データセット詳細 指示タスクの内訳としてはコード生成(code_generation)が1050レコード、コードの挙動確認(check_code_behavor)が150レコード、コードのバグ修正(code_fix)が4000レコードになります。 詳細な内訳は以下の通りになります。 source name… See the full description on the dataset page: https://huggingface.co/datasets/kunishou/amenokaku-code-instruct.text1K<n<10K17 likes118 downloads3y agoHugging Face09ewof /code-alpaca-instruct-unfilteredThis dataset is HuggingFaceH4/CodeAlpaca_20K unfiltered, removing 36 instances of blatant alignment. 19986 instructions remain. https://huggingface.co/datasets/HuggingFaceH4/CodeAlpaca_20K/blob/29ba7b7fdf0c55e5435c848cf6bbf9782fef62a6/data/test-00000-of-00001.parquet https://huggingface.co/datasets/HuggingFaceH4/CodeAlpaca_20K/blob/a123ae447f02484d83c3457438b4422cd8417ad5/data/train-00000-of-00001.parquet i combined all of these files above into code_alpaca_data.jsonl with parquet2json and ran… See the full description on the dataset page: https://huggingface.co/datasets/ewof/code-alpaca-instruct-unfiltered.text10K<n<100K9 likes73 downloads3y agoHugging Face10open-llm-leaderboard /Qwen__Qwen2.5-Coder-7B-Instruct-detailsgated Dataset Card for Evaluation run of Qwen/Qwen2.5-Coder-7B-Instruct Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Coder-7B-Instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-Coder-7B-Instruct-details.tabular10K<n<100K0 likes72 downloads2y agoHugging Face11CodeResearch /Code-Evol-Instruct-OSS Code-Evol-Instruct-OSS Summary Code-Evol-Instruct-OSS is a dataset that was generated with Code Evol-Instruct by prompting open-souce LLMs, WizardLM-13B-v1.2 and WizardCoder-34B-Python. The underlying process is explained in the paper code-evol-instruct. This algorithm gave birth to famous open-souce code LLMs, WizardCoder-Family. Our approach We did not use any closed-source LLMs. Our seed dataset is sourced from self-instruct-starcoder. We leverage the… See the full description on the dataset page: https://huggingface.co/datasets/CodeResearch/Code-Evol-Instruct-OSS.tabular1K<n<10K6 likes70 downloads3y agoHugging Face12Maxyelow /kenyan-code-switch-instruct-50k 🇰🇪 Kenyan Code-Switching & Sheng Multi-Task Instruction Dataset (50,000 Pairs) A standardized, multi-task instruction-tuning dataset engineered to teach Large Language Models (e.g. Llama 3, Mistral, Gemma, Qwen) to understand and generate authentic Kenyan Code-Switching (Sheng, Technical Swahili-English Blend) with rigorous adherence to Bantu morphotactic rules. Dataset Summary Total Samples: 50,000 instruction-response pairs train.jsonl: 45,000 pairs (90%)… See the full description on the dataset page: https://huggingface.co/datasets/Maxyelow/kenyan-code-switch-instruct-50k.texttext-generation10K<n<100K0 likes70 downloads10d agoHugging Face13open-llm-leaderboard /EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-detailsgated Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-details.tabular10K<n<100K4 likes62 downloads2y agoHugging Face14open-llm-leaderboard /theo77186__Qwen2.5-Coder-7B-Instruct-20241106-detailsgated Dataset Card for Evaluation run of theo77186/Qwen2.5-Coder-7B-Instruct-20241106 Dataset automatically created during the evaluation run of model theo77186/Qwen2.5-Coder-7B-Instruct-20241106 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/theo77186__Qwen2.5-Coder-7B-Instruct-20241106-details.tabular10K<n<100K0 likes53 downloads2y agoHugging Face15open-llm-leaderboard /EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.005-128K-code-COT-detailsgated Dataset Card for Evaluation run of EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.005-128K-code-COT Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.005-128K-code-COT The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.005-128K-code-COT-details.tabular10K<n<100K0 likes50 downloads2y agoHugging Face16open-llm-leaderboard /EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT-detailsgated Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT-details.tabular10K<n<100K0 likes47 downloads2y agoHugging Face17open-llm-leaderboard /Qwen__Qwen2.5-Coder-14B-Instruct-detailsgated Dataset Card for Evaluation run of Qwen/Qwen2.5-Coder-14B-Instruct Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Coder-14B-Instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-Coder-14B-Instruct-details.tabular10K<n<100K0 likes46 downloads2y agoHugging Face18rombodawg /code_instruct_alpaca_vicuna_wizardlm_56k_backupBackup of code_instruct_alpaca_vicuna_wizardlm used in rombodawg/MegaCodeTraining112k Link to the combined dataset bellow https://huggingface.co/datasets/rombodawg/MegaCodeTraining112k text10K<n<100K2 likes43 downloads3y agoHugging Face19lodrick-the-lafted /asl-code-sonnet35-instructSources: https://huggingface.co/datasets/Augmentation-Scaling-Laws/code_claude3.5_sonnet_10000 https://huggingface.co/datasets/Augmentation-Scaling-Laws/magpie_code_claude3.5_sonnet_10000 https://huggingface.co/datasets/Augmentation-Scaling-Laws/webinstruct_code_claude3.5_sonnet_10000 converted to sharegpt text10K<n<100K1 likes42 downloads2y agoHugging Face20open-llm-leaderboard /EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-ds-auto-detailsgated Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-ds-auto Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-ds-auto The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-ds-auto-details.tabular10K<n<100K0 likes42 downloads2y agoHugging Face21open-llm-leaderboard /Qwen__Qwen2.5-Coder-32B-Instruct-detailsgated Dataset Card for Evaluation run of Qwen/Qwen2.5-Coder-32B-Instruct Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Coder-32B-Instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-Coder-32B-Instruct-details.tabular10K<n<100K0 likes42 downloads2y agoHugging Face22open-llm-leaderboard /EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-detailsgated Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-details.tabular10K<n<100K0 likes41 downloads2y agoHugging Face23open-llm-leaderboard /EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math-detailsgated Dataset Card for Evaluation run of EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math-details.tabular10K<n<100K0 likes41 downloads2y agoHugging Face24open-llm-leaderboard /EpistemeAI__Fireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPO-detailsgated Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPO Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPO The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPO-details.tabular10K<n<100K0 likes40 downloads2y agoHugging Face25open-llm-leaderboard /EpistemeAI__Polypsyche-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-Empathy-detailsgated Dataset Card for Evaluation run of EpistemeAI/Polypsyche-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-Empathy Dataset automatically created during the evaluation run of model EpistemeAI/Polypsyche-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-Empathy The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Polypsyche-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-Empathy-details.tabular10K<n<100K0 likes38 downloads2y agoHugging Face26open-llm-leaderboard /EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-detailsgated Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-details.tabular10K<n<100K0 likes36 downloads2y agoHugging Face27open-llm-leaderboard /Etherll__Qwen2.5-Coder-7B-Instruct-Ties-detailsgated Dataset Card for Evaluation run of Etherll/Qwen2.5-Coder-7B-Instruct-Ties Dataset automatically created during the evaluation run of model Etherll/Qwen2.5-Coder-7B-Instruct-Ties The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Etherll__Qwen2.5-Coder-7B-Instruct-Ties-details.tabular10K<n<100K0 likes34 downloads2y agoHugging Face28mesolitica /chatgpt4-code-instruct ChatGPT4 Code Instruct Originally from https://huggingface.co/datasets/theblackcat102/evol-codealpaca-v1, translate and answer using ChatGPT4. Notebooks at https://github.com/mesolitica/malaysian-dataset/tree/master/chatbot/chatgpt4-code-instruct synthetic-codealpaca-v1-chatgpt4.jsonl, 43482 rows, 274 MB. Example data {'instruction': "Harap ubah skrip Python berikut agar ia memasukkan pengulangan 'while' daripada pengulangan 'for' yang sedia ada, yang meneruskan… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/chatgpt4-code-instruct.text10K<n<100K1 likes31 downloads2y agoHugging Face29cminja /ds-code-instructtext1M<n<10M2 likes30 downloads2y agoHugging Face30codezakh /EFAGen-Llama-3.1-8B-Instruct-Training-DataPaper Link The training data used for the final version of EFAGen-Llama-3.1-8B-Instruct. The data is in Alpaca format and can be used with Llama-Factory (check dataset_info.json). texttext-generation1K<n<10K1 likes28 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.