datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
text-dataset-tiny-code-script-py-format
USED of tahamajs/medicine_ds_persian for .parquet file
USED of Alijafarixcs2/persian-it-llama2-2k for .parquet file
USED of Abirate/english_quotes for .jsonl file
NEW FILES (05/12/2025)
NEW FILES (12/26/2025)
NEW FILES (02/15/2026)
python-code-dataset-500k
Attention: This dataset is a summary and reformat pulled from github code.
You should make your own assumptions based on this.
In fact, there is another dataset I formed through parsing that addresses several points:
out of 500k python related items, most of them are python-ish, not pythonic
the majority of the items here contain excessive licensing inclusion of original code
the items here are sometimes not even python but have references
There's a whole lot of gpl summaries… See the full description on the dataset page: https://huggingface.co/datasets/jtatman/python-code-dataset-500k.code-natural-language-classification-datasetSampling from codeparrot/github-code under more permissive license ['mit', 'apache-2.0', 'bsd-3-clause', 'bsd-2-clause', 'cc0-1.0'] + sampling from minipile.
It is intended to be used for training code natural language classifier.
code-security-vulnerability-dataset
Code Security Vulnerability Dataset
A curated multi-language dataset of 175,419 code samples labeled with 31 vulnerability classes (30 CWEs + safe) for training multi-label code vulnerability detection models. Labels are mapped to OWASP Top 10 2021 categories.
Dataset Details
Property
Value
Total Samples
175,419
Train / Val / Test
140,335 / 17,542 / 17,542
Languages
C, C++, Python, JavaScript, Java, PHP, Go
Labels
31 (multi-label)
Format
Parquet with… See the full description on the dataset page: https://huggingface.co/datasets/ayshajavd/code-security-vulnerability-dataset.MultiLang-Code-Parser-Dataset
MultiLang Code Parser Dataset (MLCPD)
MultiLang-Code-Parser-Dataset (MLCPD) provides a large-scale, unified dataset of parsed source code across 10 major programming languages, represented under a universal schema that captures syntax, semantics, and structure in a consistent format.
Each entry corresponds to one parsed source file and includes:
Language metadata
Code-level statistics (lines, errors, AST nodes)
Universal Schema JSON (normalized structural representation)
MLCPD… See the full description on the dataset page: https://huggingface.co/datasets/jugalgajjar/MultiLang-Code-Parser-Dataset.Code_Vulnerability_Labeled_Dataset
Dataset Card for Code_Vulnerability_Labeled_Dataset
Dataset Summary
This dataset provides (code, vulnerability) pairs. The vulnerability field takes values according to the CWE annotation:
CWE
Description
CWE-020
Improper Input Validation
CWE-022
Improper Limitation of a Pathname to a Restricted Directory (“Path Traversal”)
CWE-078
Improper Neutralization of Special Elements used in an OS Command (“OS Command Injection”)
CWE-079
Improper Neutralization of… See the full description on the dataset page: https://huggingface.co/datasets/lemon42-ai/Code_Vulnerability_Labeled_Dataset.code-retrieval-training-datasetcode_contests_dp_datasetTiBuDBowl_code_search_hard_negative_datasets-Pre_kd
Owl Code Search Hard Negative Datasets
Knowledge Distillation (KD) ベースのハードネガティブ付きコード検索データセットです。コード検索モデルShuu12121/CodeSearch-ModernBERT-Crow-v3-large-len1024-Plusを教師モデルとして、各コメントと説明コメントのペアのデータセットから各クエリに対する関数の類似度スコアを計算し、ハードネガティブ(正解に類似しているが不正解の文書)を付与しています。
概要
目的: コード検索モデルの Contrastive Learning / Knowledge Distillation ファインチューニング
言語: Go, Java, JavaScript, PHP, Python, Ruby, Rust, TypeScript(8言語)
総サンプル数: 4,787,740
データサイズ: 8.73 GB(展開後) / 3.37 GB(ダウンロード時)
フォーマット:… See the full description on the dataset page: https://huggingface.co/datasets/Shuu12121/owl_code_search_hard_negative_datasets-Pre_kd.crowd-code-dataset-1.0
Install crowd-code 2.0 to help crowd-source the next-generation coding dataset.
crowd-code-dataset-1.0 is an anonymized dataset of fine-grained IDE interactions crowd-sourced across 25 people over the last 6 months using crowd-code 1.0, a VS Code/Cursor extension capturing large parts of the software engineering workflow.
The dataset captures real research engineering workflows (character-level edits, navigation, terminal use, iterative… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/crowd-code-dataset-1.0.Flutter-Code-with-Questions-Dataset-Turkish
Flutter Code with Questions Dataset (Turkish)
📦 Dataset Name: flutter_code_with_questions
Bu veri seti, Flutter framework'ü ile yazılmış kod parçacıkları ve her bir kod parçası için özel olarak üretilmiş detaylı Türkçe soruları içermektedir. Veri seti, kodların eğitim verisi olarak kullanılmasının yanı sıra, LLM (Large Language Model) tabanlı kod anlama ve soru yanıtlama modellerinin geliştirilmesinde kullanılabilir.
📁 Dataset Format
Veri dosyaları CSV… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-Turkish.Flutter-Code-with-Questions-Dataset-English
🧠 Flutter Code with Questions Dataset (English)
This repository contains a high-quality dataset of Flutter-related code snippets paired with automatically generated English technical questions. The dataset is intended for use in training and fine-tuning language models, coding assistants, and educational systems focused on Flutter development.
📂 Dataset Structure
The dataset is divided into 22 CSV files, each containing 200 entries. Every entry includes:
A… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-English.owl_code_search_hard_negative_datasets_V2_kdst-code-datasetst_dataset_distillation_by_qwen2.5.jsonl 是由 stgenerate 蒸馏qwen2.514b生成的
/data/complier_improve 中的是在蒸馏的过程中增加编译器在环`
st_dataset_local.jsonl 是编译验证通过
st_dpo_dataset.jsonl 是交给ai改正后编译验证通过形成DPO
其他的都是各种原因为通过编译
st_dataset_distillation_by_st_coder_clean 这个适合做反面教材
StackPulse_778K_QnA_Code_dataset
💻 StackOverflow-778K: Multi-Year Developer Q&A Dataset
Dataset Summary
A large-scale Stack Overflow question dataset containing 778,929 unique
questions sampled across 7 years (2015–2022). Each question includes the
raw HTML body, plain-text version, tags, score, view count, answer count, and
a rich set of derived features for immediate ML use.
Collected across 8 sampling runs on Feb 27 2026, deduplicated to
778,929 unique questions with only 2 duplicates removed.… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/StackPulse_778K_QnA_Code_dataset.gianxxxt-code-dataset
🚀 Mushfiqur Giant Code Dataset
A unified multi-paradigm coding dataset across 9 premium research pillars.
code_net_dev_datasetcode_net_datasetcode_net_test_final_datasetadvent_of_code_ecv_dataset
Advent of Code ECV Dataset
Many code generation datasets focus on syntax and structure but lack a strong emphasis on contextual understanding, especially from a storytelling perspective.The Advent of Code ECV (Expanded, Curated, Verified) Dataset addresses this gap by curating and verifying multiple approaches for each challenge from 2024 to provide diverse solutions, comparison of strategies, and better adaptability across different programming paradigms.In addition to training and… See the full description on the dataset page: https://huggingface.co/datasets/Chemin-AI/advent_of_code_ecv_dataset.sci-codecrowd-code-dataset-0.1The crowd-code-dataset-0.1 is a raw, unfiltered dataset of fine-grained IDE interactions collected during the development of Jasmine using crowd-code, a VS Code/Cursor extension capturing large parts of the software engineering workflow.
The dataset captures real research engineering workflows (character-level edits, navigation, terminal use, iterative debugging). The crowd-code-dataset-0.1 only includes data from the Jasmine authors. We are actively working on cleaning and curating the full… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/crowd-code-dataset-0.1.HelpSteer3-general-code-Shift-Qwen-2.5-1.5B-Instruct-dataset-train-generationsbackend-code-generator-dataset
Backend Code Generation Dataset
Dataset Description
This dataset contains examples for training AI models to generate backend application code. It includes descriptions of backend requirements paired with complete, functional code implementations across multiple frameworks and programming languages.
Dataset Summary
The Backend Code Generation Dataset is designed to train models that can generate complete backend applications from natural language descriptions.… See the full description on the dataset page: https://huggingface.co/datasets/Techta/backend-code-generator-dataset.code-and-chat-dataset
language:
- en
- fr
license: mit
task_categories:
- text-generation
- conversational
- text2text-generation
tags:
- chat
- code
- reasoning
- motif
- mega-dataset
size_categories:
- 100M<n<1B
Motif Mega Chat Code
Description
A massive dataset for training Motif with chat + code capabilities.
Content
25,000,000+ examples
8 categories: code requests, code explanations, debugging, optimization, documentation, tests, architecture, theory… See the full description on the dataset page: https://huggingface.co/datasets/Lelonthecodeur/code-and-chat-dataset.ReForm-Python2Dafny-Dataset
Re:Form Datasets
This repository contains the datasets associated with the paper "Re:Form -- Reducing Human Priors in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny".
The paper introduces a framework that leverages Reinforcement Learning (RL) within Large Language Models (LLMs) to reduce reliance on human-annotated priors for formal software verification. The datasets provided are integral for training and evaluating models in formal language… See the full description on the dataset page: https://huggingface.co/datasets/Veri-Code/ReForm-Python2Dafny-Dataset.bad_code_to_good_code_dataset
Dataset Card for "bad_code_to_good_code_dataset"
More Information needed
dl4se-code-review-datasetAr-En-Code-Switching-Textual-Dataset
ArE-CSTD: Arabic-English Code-Switching Textual Dataset
The National Center for Artificial Intelligence at the Saudi Data and Artificial Intelligence Authority (SDAIA), published the "ArE-CSTD" dataset, which stands for "Arabic-English Code-Switching Textual Dataset”.
This dataset contains 330K dialectical Arabic-English code-swithing sentences generated by the large language model GPT-4.
TXT Files
There are 6 txt files. 2 files for Modern Standard Arabic(MSA) train and… See the full description on the dataset page: https://huggingface.co/datasets/SDAIANCAI/Ar-En-Code-Switching-Textual-Dataset.
