datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code_search_net
Dataset Card for CodeSearchNet corpus
Dataset Summary
CodeSearchNet corpus is a dataset of 2 milllion (comment, code) pairs from opensource libraries hosted on GitHub. It contains code and documentation for several programming languages.
CodeSearchNet corpus was gathered to support the CodeSearchNet challenge, to explore the problem of code retrieval using natural language.
Supported Tasks and Leaderboards
language-modeling: The dataset can be used to… See the full description on the dataset page: https://huggingface.co/datasets/code-search-net/code_search_net.python-codes-25k
License
MIT
This is a Cleaned Python Dataset Covering 25,000 Instructional Tasks
Overview
The dataset has 4 key features (fields): instruction, input, output, and text.It's a rich source for Python codes, tasks, and extends into behavioral aspects.
Dataset Statistics
Total Entries: 24,813
Unique Instructions: 24,580
Unique Inputs: 3,666
Unique Outputs: 24,581
Unique Texts: 24,813
Average Tokens per example: 508
Features… See the full description on the dataset page: https://huggingface.co/datasets/flytech/python-codes-25k.code-search-net-python
Dataset Card for "code-search-net-python"
Dataset Description
Homepage: None
Repository: https://huggingface.co/datasets/Nan-Do/code-search-net-python
Paper: None
Leaderboard: None
Point of Contact: @Nan-Do
Dataset Summary
This dataset is the Python portion of the CodeSarchNet annotated with a summary column.The code-search-net dataset includes open source functions that include comments found at GitHub.The summary is a short description of what the… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/code-search-net-python.TextInsightBench
TextInsightBench
English | 简体中文
A natural-language data-mining benchmark for agents: 50 tasks, 435,000 task documents and 944,468 unlabeled learning documents.
Each task provides 5,000 or 10,000 texts and a research objective. Agents choose the patterns, populations and comparisons to investigate, then submit up to three findings with complete document assignments, exact quotations, statistics, counterexamples and limitations. Any analysis method is allowed.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/CodeSoulco/TextInsightBench.tiny-codes
Reasoning with Language and Code
This synthetic dataset is a collection of 1.6 millions short and clear code snippets that can help LLM models learn how to reason with both natural and programming languages. The dataset covers a wide range of programming languages, such as Python, TypeScript, JavaScript, Ruby, Julia, Rust, C++, Bash, Java, C#, and Go. It also includes two database languages: Cypher (for graph databases) and SQL (for relational databases) in order to study the… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-codes.code_search_net
CodeSearchNet
This is an unofficial reupload of the code_search_net dataset in the parquet format. I have also removed the columns func_code_tokens, func_documentation_tokens, and split_name as they are not relevant. The original repository relies on a Python module that is downloaded and executed to unpack the dataset, which is a potential security risk but importantly raises an annoying warning. As a plus, parquets load faster.
Original model card:
Dataset Card for… See the full description on the dataset page: https://huggingface.co/datasets/claudios/code_search_net.code-sante-publique
Code de la santé publique, non-instruct (2025-07-11)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-sante-publique.CodeSecEval
CodeSecEval
CodeSecEval is an execution-based benchmark for evaluating large language models on secure code generation and insecure-code repair.
The benchmark contains 255 Python programming tasks spanning 77 CWE vulnerability categories. Each task provides a problem specification, an insecure implementation, a secure reference implementation, executable tests, and an entry point.
Dataset Subsets
This repository contains two subsets:
SecEvalBase: 115 tasks… See the full description on the dataset page: https://huggingface.co/datasets/JasonWang1/CodeSecEval.code-search-net-java
Dataset Card for "code-search-net-java"
Dataset Summary
This dataset is the Java portion of the CodeSarchNet annotated with a summary column.The code-search-net dataset includes open source functions that include comments found at GitHub.The summary is a short description of what the function does.
Languages
The dataset's comments are in English and the functions are coded in Java
Data Splits
Train, test, validation labels are included in the… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/code-search-net-java.code-switching-tokenizer-robustness
Code-Switching Dataset for Tokenizer Robustness Analysis
Dataset Description
This dataset is designed for tokenizer robustness testing in multilingual and code-switching contexts. It contains identical content expressed across 16 different language variants, including pure English and 15 English-X code-switching pairs, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models.
Purpose
Tokenizer Comparison:… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/code-switching-tokenizer-robustness.offsec_redteam_codes
OffSec RedTeam Codes
Token count: ~30B tokens.
OffSec RedTeam Codes is a curated corpus of code (and some auxiliary text) extracted from popular GitHub repositories related to offensive security / red teaming (pentesting, OSINT, C2, privilege escalation, exploitation, forensics, etc.). It is also the largest open-source dataset of red-team and offensive-security code ever compiled.
⚠️ Ethical use only. This dataset is for research, education, and defensive security testing in… See the full description on the dataset page: https://huggingface.co/datasets/tandevllc/offsec_redteam_codes.code-search-net-javascript
Dataset Card for "code-search-net-javascript"
Dataset Summary
This dataset is the JavaScript portion of the CodeSarchNet annotated with a summary column.The code-search-net dataset includes open source functions that include comments found at GitHub.The summary is a short description of what the function does.
Languages
The dataset's comments are in English and the functions are coded in JavaScript
Data Splits
Train, test, validation labels are… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/code-search-net-javascript.code-search-net-php
Dataset Card for "code-search-net-php"
Dataset Summary
This dataset is the Php portion of the CodeSarchNet annotated with a summary column.The code-search-net dataset includes open source functions that include comments found at GitHub.The summary is a short description of what the function does.
Languages
The dataset's comments are in English and the functions are coded in Php
Data Splits
Train, test, validation labels are included in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/code-search-net-php.code-search-net-python
Dataset Card for "code-search-net-python"
Dataset Description
Homepage: None
Repository: https://huggingface.co/datasets/Nan-Do/code-search-net-python
Paper: None
Leaderboard: None
Point of Contact: @Nan-Do
Dataset Summary
This dataset is the Python portion of the CodeSarchNet annotated with a summary column.The code-search-net dataset includes open source functions that include comments found at GitHub.The summary is a short description of what the… See the full description on the dataset page: https://huggingface.co/datasets/khemprogrammer/code-search-net-python.CodeSpan
CodeSpan
64K source-code continuation for long-context decoding.
CodeSpan contains 32 examples from 17 open-source projects, including LLVM,
GCC, Linux, and PostgreSQL. Each example is a contiguous prefix of a distinct
source file, wrapped as a continuation prompt: 65,536 Qwen3 tokens, including
the chat template with thinking disabled. Files are neither concatenated nor
repeated. There are at most four files per project; 28 of the 32 files are C/C++.
Configuration… See the full description on the dataset page: https://huggingface.co/datasets/Hehy/CodeSpan.CodeSec-Pairs
CodeSec-Pairs
CodeSec-Pairs is a dataset of matched safe and vulnerable Python code pairs. Each
pair implements the same task but differs in whether it contains a security
vulnerability. Vulnerability labels come from CodeQL static analysis. The dataset is
built to study and steer the internal mechanisms that distinguish safe from vulnerable
code generation in LLMs.
Dataset Details
Each record pairs a CodeQL-clean safe_code with a CodeQL-flagged vuln_code for the… See the full description on the dataset page: https://huggingface.co/datasets/haaao821/CodeSec-Pairs.code-switched-student-blindspot-eval
Code Switched Student Blind Spot Evaluation
Overview
This repository contains a small manual evaluation of Qwen/Qwen2.5-1.5B-Instruct on code-switched South Asian international student prompts. The goal is to test whether a small open-weight instruction model can understand Pakistani English mixed with Roman Urdu/Hindi in situations shaped by scholarship pressure, family expectations, limited resources, and international student life in Malaysia.
Blind… See the full description on the dataset page: https://huggingface.co/datasets/Zerothe00/code-switched-student-blindspot-eval.code-search-net-go
Dataset Card for "code-search-net-go"
Dataset Summary
This dataset is the Go portion of the CodeSarchNet annotated with a summary column.The code-search-net dataset includes open source functions that include comments found at GitHub.The summary is a short description of what the function does.
Languages
The dataset's comments are in English and the functions are coded in Go
Data Splits
Train, test, validation labels are included in the dataset as… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/code-search-net-go.CodeScopecode_search_net
Dataset Card for CodeSearchNet corpus
Dataset Summary
CodeSearchNet corpus is a dataset of 2 milllion (comment, code) pairs from opensource libraries hosted on GitHub. It contains code and documentation for several programming languages.
CodeSearchNet corpus was gathered to support the CodeSearchNet challenge, to explore the problem of code retrieval using natural language.
Supported Tasks and Leaderboards
language-modeling: The dataset can be used… See the full description on the dataset page: https://huggingface.co/datasets/Zzzzzxl/code_search_net.wisconsin-building-codes-qa
Wisconsin Building Codes Q&A Dataset
Dataset Description
This dataset contains 13200 question-answer pairs focused on Wisconsin building codes, specifically covering:
Building code requirements and regulations
Administrative procedures and enforcement
Construction standards and specifications
Permit processes and compliance
Dataset Structure
Training samples: 11880
Validation samples: 1320
Each sample contains:
instruction: A question about Wisconsin… See the full description on the dataset page: https://huggingface.co/datasets/carlscape/wisconsin-building-codes-qa.python-codes-25k
License
MIT
This is a Cleaned Python Dataset Covering 25,000 Instructional Tasks
Overview
The dataset has 4 key features (fields): instruction, input, output, and text.It's a rich source for Python codes, tasks, and extends into behavioral aspects.
Dataset Statistics
Total Entries: 24,813
Unique Instructions: 24,580
Unique Inputs: 3,666
Unique Outputs: 24,581
Unique Texts: 24,813
Average Tokens per example: 508… See the full description on the dataset page: https://huggingface.co/datasets/Rishabh23456789/python-codes-25k.hugcode-codesft所有数据都是单轮代码指令数据
325696条英语,42816条中文。
license: cc
adaption-codeskill-raw-selfoss-2k
Self-OSS Code Instructions
Python instructions with execution-filtered solutions.
Rows
2,000
Domain
programming
Format
data.parquet, one row per example
Licence
odc-by
Built for
supervised fine-tuning (SFT) experiments on Adaption AutoScientist
Columns
Column
Description
original_prompt
The prompt (user turn) as uploaded.
original_completion
The target response as uploaded.
enhanced_prompt
Empty in this dataset.… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-codeskill-raw-selfoss-2k.adaption-codeskill-raw-commitpack-2k
Commit-Message Code Edits
Code-change tasks from real commits: an instruction (commit message) and the edited code.
Rows
2,000
Domain
programming
Format
data.parquet, one row per example
Licence
mit
Built for
supervised fine-tuning (SFT) experiments on Adaption AutoScientist
Columns
Column
Description
original_prompt
The prompt (user turn) as uploaded.
original_completion
The target response as uploaded.
enhanced_prompt… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-codeskill-raw-commitpack-2k.code-sport
Code du sport, non-instruct (2025-07-11)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language models… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-sport.qmatrix-codesft-artifacts
Q-Matrix Code-SFT selection artifacts
Pool, per-item scores, and selection coresets for the paper "Item-Level
Coreset Selectors Choose a Corpus They Never Scored: A Q-Matrix Audit for Code
Supervised Fine-Tuning."
This repository also carries the code, the Q-matrix, and the analysis that
recomputes every table and the figure in the paper. Trained LoRA adapters are
in the companion model repository.
Contents
qmatrix/ ontology, annotator… See the full description on the dataset page: https://huggingface.co/datasets/burnerqmatrixacl/qmatrix-codesft-artifacts.tool-reasoning-sft-RESEARCH-OpenHands-CodeScout_Training_Rollouts
CodeScout Training Rollouts — Cleaned & Rectified
~40K multi-turn code localization agent trajectories converted into a strict reasoning + tool-call format with validated FSM transitions. Supports coupled (parallel) tool calls.
⚠️ Mid-training dataset. This dataset contains synthesized reasoning templates (not native chain-of-thought). It is suitable for mid-training to teach tool-use mechanics, FSM structure, and bash exploration patterns. It is not recommended as a final SFT… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-RESEARCH-OpenHands-CodeScout_Training_Rollouts.Python-codes
Dataset Card for Dataset Name
Please note that this dataset maynot be perfect and may contain a very small quantity of non python codes. But the quantity appears to be very small
Dataset Summary
The dataset contains a collection of python question and their code. This is meant to be used for training models to be efficient in Python specific coding.
The dataset has two features - 'question' and 'code'.
An example is:
{'question': 'Create a function that takes in a string… See the full description on the dataset page: https://huggingface.co/datasets/Arjun-G-Ravi/Python-codes.llama-python-codes-30k
Python Codes - 30k examples, Llama1&2 tokenized dataset
Author
FlyTech
For general guide on how to create, quantize, merge or inference the model and more, visit:
hackmd.io/my_first_ai
Overview
This dataset serves as a rich resource for various Natural Language Processing tasks such as:
Question Answering
Text Generation
Text-to-Text Generation
It primarily focuses on instructional tasks in Python, tokenized specifically for the Llama architecture.… See the full description on the dataset page: https://huggingface.co/datasets/flytech/llama-python-codes-30k.
