Team Ai
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ysn-rfd /text-dataset-tiny-code-script-py-format USED of tahamajs/medicine_ds_persian for .parquet file USED of Alijafarixcs2/persian-it-llama2-2k for .parquet file USED of Abirate/english_quotes for .jsonl file NEW FILES (05/12/2025) NEW FILES (12/26/2025) NEW FILES (02/15/2026) texttext-generation10K<n<100K3 likes1.7k downloads4mo agoHugging Face02jtatman /python-code-dataset-500k Attention: This dataset is a summary and reformat pulled from github code. You should make your own assumptions based on this. In fact, there is another dataset I formed through parsing that addresses several points: out of 500k python related items, most of them are python-ish, not pythonic the majority of the items here contain excessive licensing inclusion of original code the items here are sometimes not even python but have references There's a whole lot of gpl summaries… See the full description on the dataset page: https://huggingface.co/datasets/jtatman/python-code-dataset-500k.texttext-generation100K<n<1M81 likes870 downloads3y agoHugging Face03lemon42-ai /Code_Vulnerability_Labeled_Dataset Dataset Card for Code_Vulnerability_Labeled_Dataset Dataset Summary This dataset provides (code, vulnerability) pairs. The vulnerability field takes values according to the CWE annotation: CWE Description CWE-020 Improper Input Validation CWE-022 Improper Limitation of a Pathname to a Restricted Directory (“Path Traversal”) CWE-078 Improper Neutralization of Special Elements used in an OS Command (“OS Command Injection”) CWE-079 Improper Neutralization of… See the full description on the dataset page: https://huggingface.co/datasets/lemon42-ai/Code_Vulnerability_Labeled_Dataset.texttext-classification1K<n<10K13 likes302 downloads2y agoHugging Face04NoirZangetsu /Flutter-Code-with-Questions-Dataset-English 🧠 Flutter Code with Questions Dataset (English) This repository contains a high-quality dataset of Flutter-related code snippets paired with automatically generated English technical questions. The dataset is intended for use in training and fine-tuning language models, coding assistants, and educational systems focused on Flutter development. 📂 Dataset Structure The dataset is divided into 22 CSV files, each containing 200 entries. Every entry includes: A… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-English.textquestion-answering1K<n<10K3 likes218 downloads2mo agoHugging Face05NoirZangetsu /Flutter-Code-with-Questions-Dataset-Turkish Flutter Code with Questions Dataset (Turkish) 📦 Dataset Name: flutter_code_with_questions Bu veri seti, Flutter framework'ü ile yazılmış kod parçacıkları ve her bir kod parçası için özel olarak üretilmiş detaylı Türkçe soruları içermektedir. Veri seti, kodların eğitim verisi olarak kullanılmasının yanı sıra, LLM (Large Language Model) tabanlı kod anlama ve soru yanıtlama modellerinin geliştirilmesinde kullanılabilir. 📁 Dataset Format Veri dosyaları CSV… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-Turkish.textquestion-answering1K<n<10K0 likes209 downloads2mo agoHugging Face06Omarrran /StackPulse_778K_QnA_Code_dataset 💻 StackOverflow-778K: Multi-Year Developer Q&A Dataset Dataset Summary A large-scale Stack Overflow question dataset containing 778,929 unique questions sampled across 7 years (2015–2022). Each question includes the raw HTML body, plain-text version, tags, score, view count, answer count, and a rich set of derived features for immediate ML use. Collected across 8 sampling runs on Feb 27 2026, deduplicated to 778,929 unique questions with only 2 duplicates removed.… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/StackPulse_778K_QnA_Code_dataset.tabulartext-classification1M<n<10M0 likes187 downloads6mo agoHugging Face07RnniaSnow /st-code-datasetst_dataset_distillation_by_qwen2.5.jsonl 是由 stgenerate 蒸馏qwen2.514b生成的 /data/complier_improve 中的是在蒸馏的过程中增加编译器在环` st_dataset_local.jsonl 是编译验证通过 st_dpo_dataset.jsonl 是交给ai改正后编译验证通过形成DPO 其他的都是各种原因为通过编译 st_dataset_distillation_by_st_coder_clean 这个适合做反面教材 texttext-generation100K<n<1M0 likes178 downloads8mo agoHugging Face08Techta /backend-code-generator-dataset Backend Code Generation Dataset Dataset Description This dataset contains examples for training AI models to generate backend application code. It includes descriptions of backend requirements paired with complete, functional code implementations across multiple frameworks and programming languages. Dataset Summary The Backend Code Generation Dataset is designed to train models that can generate complete backend applications from natural language descriptions.… See the full description on the dataset page: https://huggingface.co/datasets/Techta/backend-code-generator-dataset.texttext-generationn<1K1 likes126 downloads1y agoHugging Face09Chemin-AI /advent_of_code_ecv_dataset Advent of Code ECV Dataset Many code generation datasets focus on syntax and structure but lack a strong emphasis on contextual understanding, especially from a storytelling perspective.The Advent of Code ECV (Expanded, Curated, Verified) Dataset addresses this gap by curating and verifying multiple approaches for each challenge from 2024 to provide diverse solutions, comparison of strategies, and better adaptability across different programming paradigms.In addition to training and… See the full description on the dataset page: https://huggingface.co/datasets/Chemin-AI/advent_of_code_ecv_dataset.texttext-generationn<1K7 likes112 downloads2y agoHugging Face10Veri-Code /ReForm-Python2Dafny-Dataset Re:Form Datasets This repository contains the datasets associated with the paper "Re:Form -- Reducing Human Priors in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny". The paper introduces a framework that leverages Reinforcement Learning (RL) within Large Language Models (LLMs) to reduce reliance on human-annotated priors for formal software verification. The datasets provided are integral for training and evaluating models in formal language… See the full description on the dataset page: https://huggingface.co/datasets/Veri-Code/ReForm-Python2Dafny-Dataset.texttext-generation10K<n<100K2 likes105 downloads5mo agoHugging Face11SDAIANCAI /Ar-En-Code-Switching-Textual-Dataset ArE-CSTD: Arabic-English Code-Switching Textual Dataset The National Center for Artificial Intelligence at the Saudi Data and Artificial Intelligence Authority (SDAIA), published the "ArE-CSTD" dataset, which stands for "Arabic-English Code-Switching Textual Dataset”. This dataset contains 330K dialectical Arabic-English code-swithing sentences generated by the large language model GPT-4. TXT Files There are 6 txt files. 2 files for Modern Standard Arabic(MSA) train and… See the full description on the dataset page: https://huggingface.co/datasets/SDAIANCAI/Ar-En-Code-Switching-Textual-Dataset.texttext-generation100K<n<1M2 likes83 downloads2y agoHugging Face12suneeldk /arduino-code-dataset Arduino Programming Dataset Overview The Arduino Programming Dataset is a structured dataset designed for training and evaluating AI models that generate Arduino sketches for microcontroller projects. The dataset contains natural language instructions paired with complete Arduino programs. Each entry describes a hardware task and provides a working Arduino implementation along with additional metadata such as libraries used, wiring information, and difficulty level. This… See the full description on the dataset page: https://huggingface.co/datasets/suneeldk/arduino-code-dataset.texttext-generationn<1K1 likes78 downloads7mo agoHugging Face13ysn-rfd /FibonacciAi-CODE-Fine_Tuning-DatasetDocument Version: 1.0.2 | Last Updated: 07/17/2026 texttext-generation1K<n<10K1 likes63 downloads3mo agoHugging Face14Agnuxo /github-source-code-dataset Github Source Code Dataset Complete source code from Agnuxo projects. Part of the Agnuxo Ecosystem by Francisco Angulo de Lafuente. texttext-generation1K<n<10K0 likes54 downloads5mo agoHugging Face150-deng-0 /st-code-datasetst_dataset_distillation_by_qwen2.5.jsonl 是由 stgenerate 蒸馏qwen2.514b生成的 /data/complier_improve 中的是在蒸馏的过程中增加编译器在环` st_dataset_local.jsonl 是编译验证通过 st_dpo_dataset.jsonl 是交给ai改正后编译验证通过形成DPO 其他的都是各种原因为通过编译 st_dataset_distillation_by_st_coder_clean 这个适合做反面教材 texttext-generation100K<n<1M0 likes52 downloads2mo agoHugging Face16ahmetggg /Dr-Zeon-Github-Python-Code-Dataset Luck Spark 1B - High Quality Code Dataset The first quality-scored, star-agnostic code dataset for training 1B MoE code models. Unlike The Stack / CodeParrot that filter by stars, this dataset scores every file by its content (0-10). A 2-star well-documented library scores higher than a 10k-star minified file. Continuously updated by an autonomous bot. Repo: ahmetggg/luck-spark-1b-code-dataset | Bot: github_to_hf_bot.py | License: Permissive only (MIT / Apache-2.0 / BSD /… See the full description on the dataset page: https://huggingface.co/datasets/ahmetggg/Dr-Zeon-Github-Python-Code-Dataset.tabulartext-generation10K<n<100K1 likes43 downloads1mo agoHugging Face17InfoBayAI /Legacy-Code-Datasetgated Legacy Codebase Dataset Dataset Description The Legacy Codebase Dataset is a large-scale collection of enterprise software repositories designed for training next-generation Large Language Models (LLMs), AI coding assistants, software engineering copilots, automated refactoring systems, repository understanding models, and intelligent program analysis pipelines. The complete collection contains 405 real-world legacy codebases spanning 23 major industries… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Legacy-Code-Dataset.texttext-generation10K<n<100K0 likes33 downloads2d agoHugging Face18shoumitro26 /python-code-dataset-500k Attention: This dataset is a summary and reformat pulled from github code. You should make your own assumptions based on this. In fact, there is another dataset I formed through parsing that addresses several points: out of 500k python related items, most of them are python-ish, not pythonic the majority of the items here contain excessive licensing inclusion of original code the items here are sometimes not even python but have references There's a whole lot of gpl summaries… See the full description on the dataset page: https://huggingface.co/datasets/shoumitro26/python-code-dataset-500k.texttext-generation100K<n<1M0 likes29 downloads7mo agoHugging Face19Fuleee /springboot-code-completion-dataset Spring Boot Code Completion Dataset 概述 这是一个专为 Spring Boot 框架代码补全任务设计的中文优先数据集。通过从 GitHub 上 90 个高质量开源 Spring Boot 项目中提取方法级、类级代码片段及配置文件,构建而成。数据集重点突出 Spring Boot 典型特征(注解驱动、依赖注入、配置绑定等),适用于参数高效微调(PEFT,如 LoRA/QLoRA)下的代码生成研究。 数据集统计 总样本数:81,085 条 训练集:64,868 条 验证集:8,108 条 测试集:8,109 条 优先级 2 样本数(含典型 Spring Boot 特征,如 @RestController、@Service、@Entity 等):19,740 条 优先级 2 占比:24.34% 数据格式(Alpaca 风格) 每条数据为 JSON 对象,包含以下字段: { "instruction": "请在Spring… See the full description on the dataset page: https://huggingface.co/datasets/Fuleee/springboot-code-completion-dataset.texttext-generation10K<n<100K0 likes27 downloads9mo agoHugging Face20InfoBayAI /Product-Source-Code-DatasetgatedDataset Description: This dataset is a large-scale collection of coding and data, designed to support the development of advanced AI systems for code generation, program understanding, software intelligence, debugging assistance, and next-generation developer AI applications. Additionally, this dataset can be integrated into pipelines for Supervised Fine-Tuning (SFT) and Reinforcement Learning with Human Feedback (RLHF) workflows, helping improve AI performance in code completion, automated… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Product-Source-Code-Dataset.texttext-generationn<1K0 likes25 downloads2d agoHugging Face21me-aas /python-code-dataset-500k Attention: This dataset is a summary and reformat pulled from github code. You should make your own assumptions based on this. In fact, there is another dataset I formed through parsing that addresses several points: out of 500k python related items, most of them are python-ish, not pythonic the majority of the items here contain excessive licensing inclusion of original code the items here are sometimes not even python but have references There's a whole lot of gpl summaries… See the full description on the dataset page: https://huggingface.co/datasets/me-aas/python-code-dataset-500k.texttext-generation100K<n<1M0 likes23 downloads4mo agoHugging Face22jonathanyly /python-code-dataset-500k Attention: This dataset is a summary and reformat pulled from github code. You should make your own assumptions based on this. In fact, there is another dataset I formed through parsing that addresses several points: out of 500k python related items, most of them are python-ish, not pythonic the majority of the items here contain excessive licensing inclusion of original code the items here are sometimes not even python but have references There's a whole lot of gpl summaries… See the full description on the dataset page: https://huggingface.co/datasets/jonathanyly/python-code-dataset-500k.texttext-generation100K<n<1M1 likes19 downloads9mo agoHugging Face23MoChenYa /code-nomist-llm-datasetCode Nomist 项目微调数据集 用于大模型微调使用,包含格式化后的问题,以及对应答案每个名称使用 | 符号分割。 texttext-generation1K<n<10K0 likes10 downloads1y agoHugging Face24BhoomishPatel /python-code-dataset-500k Attention: This dataset is a summary and reformat pulled from github code. You should make your own assumptions based on this. In fact, there is another dataset I formed through parsing that addresses several points: out of 500k python related items, most of them are python-ish, not pythonic the majority of the items here contain excessive licensing inclusion of original code the items here are sometimes not even python but have references There's a whole lot of gpl summaries… See the full description on the dataset page: https://huggingface.co/datasets/BhoomishPatel/python-code-dataset-500k.texttext-generation100K<n<1M0 likes10 downloads5mo agoHugging Face25AhmetSemih /qwen-code-thinking-dataset python_code_instructions_alpaca (Qwen-Enhanced) This dataset contains 100 randomly selected samples from the iamtarun/python_code_instructions_18k_alpaca. 💡 Overview The original dataset provides coding instructions. To enhance its utility for fine-tuning "reasoning" models, ı used the Qwen3.5 (397B) model to generate: Thinking: A deep, step-by-step internal monologue solving the problem. Content: The final, formatted answer. Column Description instruction… See the full description on the dataset page: https://huggingface.co/datasets/AhmetSemih/qwen-code-thinking-dataset.texttext-generationn<1K0 likes6 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.