datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code_x_glue_ct_code_to_text
Dataset Card for "code_x_glue_ct_code_to_text"
Dataset Summary
CodeXGLUE code-to-text dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Text/code-to-text
The dataset we use comes from CodeSearchNet and we filter the dataset as the following:
Remove examples that codes cannot be parsed into an abstract syntax tree.
Remove examples that #tokens of documents is < 3 or >256
Remove examples that documents contain special tokens (e.g. <img ...> or… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_ct_code_to_text.code_x_glue_tc_text_to_code
Dataset Card for "code_x_glue_tc_text_to_code"
Dataset Summary
CodeXGLUE text-to-code dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Text-Code/text-to-code
The dataset we use is crawled and filtered from Microsoft Documentation, whose document located at https://github.com/MicrosoftDocs/.
Supported Tasks and Leaderboards
machine-translation: The dataset can be used to train a model for generating Java code from an English natural… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_tc_text_to_code.code_x_glue_cc_code_to_code_trans
Dataset Card for "code_x_glue_cc_code_to_code_trans"
Dataset Summary
CodeXGLUE code-to-code-trans dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/code-to-code-trans
The dataset is collected from several public repos, including Lucene(http://lucene.apache.org/), POI(http://poi.apache.org/), JGit(https://github.com/eclipse/jgit/) and Antlr(https://github.com/antlr/).
We collect both the Java and C# versions of the codes and find the parallel… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_code_to_code_trans.code_x_glue_tt_text_to_text
Dataset Card for "code_x_glue_tt_text_to_text"
Dataset Summary
CodeXGLUE text-to-text dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Text-Text/text-to-text
The dataset we use is crawled and filtered from Microsoft Documentation, whose document located at https://github.com/MicrosoftDocs/.
Supported Tasks and Leaderboards
machine-translation: The dataset can be used to train a model for translating Technical documentation between… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_tt_text_to_text.OpenSCAD-Code-to-GLB
OpenSCAD-Code-to-GLB
Dataset Description
OpenSCAD-Code-to-GLB is a derived dataset built on top of CodeCAD-HF-v2. It extends the original by converting OpenSCAD parametric code into rendered GLB (GL Transmission Format Binary) 3D model files. Each record pairs a natural language prompt, the corresponding OpenSCAD code, and a rendered GLB file of the resulting 3D mesh.
Dataset Statistics
Property
Value
Split
Train
Number of Rows
4,795… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/OpenSCAD-Code-to-GLB.github-jupyter-code-to-text
Dataset description
This dataset consists of sequences of Python code followed by a a docstring explaining its function. It was constructed by concatenating code and text pairs
from this dataset that were originally code and markdown cells in Jupyter Notebooks.
The content of each example the following:
[CODE]
"""
Explanation: [TEXT]
End of explanation
"""
[CODE]
"""
Explanation: [TEXT]
End of explanation
"""
...
How to use it
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/codeparrot/github-jupyter-code-to-text.SWE-bench__style-3__fs-oracle
Dataset Card for "SWE-bench__style-3__fs-oracle"
More Information needed
SWE-bench__style-3__fs-oracle_large_tokenlength
Dataset Card for "SWE-bench__style-3__fs-oracle"
More Information needed
bad_code_to_good_code_dataset
Dataset Card for "bad_code_to_good_code_dataset"
More Information needed
code-image-to-text
Code Snippet Image → Text
A multimodal dataset for fine-tuning vision-language models (VLMs) on the task of
transcribing an image of a code snippet back into its source text — syntax-aware OCR.
Each example pairs a syntax-highlighted PNG of code with the exact code text that
produced it. It spans 8 programming languages and deliberately mixes two capture types:
block — a complete function / unit (6–45 lines).
fragment — a contiguous partial view (3–14 lines) that may start or… See the full description on the dataset page: https://huggingface.co/datasets/anisiraj/code-image-to-text.mozzarella
Mozzarella-0.3.1
Motivation
Mozzarella is a dataset matching issues (= problem statements) and corresponding pull requests (PRs = problem solutions) of a selection of well maintained Java GitHub repositories. The original purpose was to serve as training and evaluation data for ML models concerned with fault localization and automated program repair of complex code bases. However, there might be more use cases that could benefit from this data.
Inspired by SWEBench… See the full description on the dataset page: https://huggingface.co/datasets/feedback-to-code/mozzarella.code_x_glue_ct_code_to_text_java_pythonarxiv-to-code-agentic-tool-calling
arxiv-to-code-agentic-tool-calling
Multi-turn tool-calling dataset where an assistant implements ML papers in PyTorch through file-creation and command-execution tool calls.
Built from lucidrains' (Phil Wang) open-source paper implementations. There are ~217 repositories on Codeberg, each implementing a different ML paper. This dataset reverse-engineers those into synthetic coding conversations.
What's in it
199 conversations, each covering one repository. Every… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/arxiv-to-code-agentic-tool-calling.frontend-figma-to-code
CREW: Figma to Code
A benchmark for evaluating AI coding agents on Figma-to-code generation — converting real-world Figma community designs
into production-ready React + Tailwind CSS applications.
Each task gives an agent full access to a Figma file via MCP tools. The agent must extract the design system, generate
components, build successfully, and deploy a live preview. Outputs are evaluated through human preference (ELO) and
automated verifiers.
Benchmark… See the full description on the dataset page: https://huggingface.co/datasets/metaphilabs/frontend-figma-to-code.strudel-nl-to-codecode_x_glue_ct_code_to_text_java_pythonhyperswitch-issue-to-code_v2
Rust Commit Dataset - Hyperswitch
Dataset Description
This dataset contains Rust commit messages paired with their corresponding code patches from the Hyperswitch repository.
Dataset Summary
Total Examples: 319
Language: Rust
Source: Hyperswitch GitHub repository
Format: Prompt-response pairs for supervised fine-tuning (SFT)
Data Fields
prompt: The commit message describing the change
response: The git patch/diff showing the actual code changes… See the full description on the dataset page: https://huggingface.co/datasets/archit11/hyperswitch-issue-to-code_v2.code_x_glue_ct_code_to_texthyperswitch-issue-to-code_v3_natural
Rust Commit Dataset - Hyperswitch
Dataset Description
This dataset contains Rust commit messages paired with their corresponding code patches from the Hyperswitch repository.
Dataset Summary
Total Examples: 9
Language: Rust
Source: Hyperswitch GitHub repository
Format: Prompt-response pairs for supervised fine-tuning (SFT)
Data Fields
prompt: The commit message describing the change
response: The git patch/diff showing the actual code changes… See the full description on the dataset page: https://huggingface.co/datasets/archit11/hyperswitch-issue-to-code_v3_natural.smolified-tiny-text-to-code
🤏 smolified-tiny-text-to-code
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model programmerGodbyte/smolified-tiny-text-to-code.
📦 Asset Details
Origin: Smolify Foundry (Job ID: fe9b19bf)
Records: 1078
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by programmerGodbyte.
Generated via Smolify.ai.
ADA_Dental_Code_to_SBS_V2hyperswitch-issue-to-code_v10
Rust Commit Dataset - Hyperswitch
Dataset Description
This dataset contains Rust commit messages paired with their corresponding code patches from the Hyperswitch repository.
Dataset Summary
Total Examples: 30
Language: Rust
Source: Hyperswitch GitHub repository
Format: Prompt-response pairs for supervised fine-tuning (SFT)
Data Fields
prompt: The commit message describing the change
response: The git patch/diff showing the actual code changes… See the full description on the dataset page: https://huggingface.co/datasets/archit11/hyperswitch-issue-to-code_v10.text-to-python-code-dataset
Dataset Card for "text-to-python-code-dataset"
More Information needed
Code_to_explanationhyperswitch-issue-to-code
Rust Commit Dataset - Hyperswitch
Dataset Description
This dataset contains Rust commit messages paired with their corresponding code patches from the Hyperswitch repository.
Dataset Summary
Total Examples: 231
Language: Rust
Source: Hyperswitch GitHub repository
Format: Prompt-response pairs for supervised fine-tuning (SFT)
Data Fields
prompt: The commit message describing the change
response: The git patch/diff showing the actual code changes… See the full description on the dataset page: https://huggingface.co/datasets/archit11/hyperswitch-issue-to-code.code_x_glue_tc_text_to_code_promptsourcetext-to-python-code-datasettext-to-python-code-datasetsmolified-python-code-to-english
🤏 smolified-python-code-to-english
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model rajan-jar/smolified-python-code-to-english.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 6fe72625)
Records: 440
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by rajan-jar.
Generated via Smolify.ai.
bad_code_to_good_code_dataset
Dataset Card for "bad_code_to_good_code_dataset"
More Information needed
