datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code_x_glue_ct_code_to_text
Dataset Card for "code_x_glue_ct_code_to_text"
Dataset Summary
CodeXGLUE code-to-text dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Text/code-to-text
The dataset we use comes from CodeSearchNet and we filter the dataset as the following:
Remove examples that codes cannot be parsed into an abstract syntax tree.
Remove examples that #tokens of documents is < 3 or >256
Remove examples that documents contain special tokens (e.g. <img ...> or… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_ct_code_to_text.xlcost-text-to-code XLCoST is a machine learning benchmark dataset that contains fine-grained parallel data in 7 commonly used programming languages (C++, Java, Python, C#, Javascript, PHP, C), and natural language (English).code_x_glue_tc_text_to_code
Dataset Card for "code_x_glue_tc_text_to_code"
Dataset Summary
CodeXGLUE text-to-code dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Text-Code/text-to-code
The dataset we use is crawled and filtered from Microsoft Documentation, whose document located at https://github.com/MicrosoftDocs/.
Supported Tasks and Leaderboards
machine-translation: The dataset can be used to train a model for generating Java code from an English natural… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_tc_text_to_code.code_x_glue_cc_code_to_code_trans
Dataset Card for "code_x_glue_cc_code_to_code_trans"
Dataset Summary
CodeXGLUE code-to-code-trans dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/code-to-code-trans
The dataset is collected from several public repos, including Lucene(http://lucene.apache.org/), POI(http://poi.apache.org/), JGit(https://github.com/eclipse/jgit/) and Antlr(https://github.com/antlr/).
We collect both the Java and C# versions of the codes and find the parallel… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_code_to_code_trans.code_x_glue_tt_text_to_text
Dataset Card for "code_x_glue_tt_text_to_text"
Dataset Summary
CodeXGLUE text-to-text dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Text-Text/text-to-text
The dataset we use is crawled and filtered from Microsoft Documentation, whose document located at https://github.com/MicrosoftDocs/.
Supported Tasks and Leaderboards
machine-translation: The dataset can be used to train a model for translating Technical documentation between… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_tt_text_to_text.image-to-code-v2OpenSCAD-Code-to-GLB
OpenSCAD-Code-to-GLB
Dataset Description
OpenSCAD-Code-to-GLB is a derived dataset built on top of CodeCAD-HF-v2. It extends the original by converting OpenSCAD parametric code into rendered GLB (GL Transmission Format Binary) 3D model files. Each record pairs a natural language prompt, the corresponding OpenSCAD code, and a rendered GLB file of the resulting 3D mesh.
Dataset Statistics
Property
Value
Split
Train
Number of Rows
4,795… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/OpenSCAD-Code-to-GLB.arabic-to-code-8-langs-3m
Dataset evaluation: See EVALUATION.md for schema checks, indexing status, and language-specific quality limits.
Viewer note: default is a lightweight preview; select full to load the complete corpus.
Current Hub Validation Status
Repository claim: 3,000,000 records
Dataset Server indexed rows: 1,239,045
Dataset Server estimate: 1,995,159
The 3M target figure is a raw-repository claim and is not yet fully verified by the Hub index. Validate the JSONL files before publishing a… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/arabic-to-code-8-langs-3m.SWE-bench__style-3__fs-oracle
Dataset Card for "SWE-bench__style-3__fs-oracle"
More Information needed
github-jupyter-code-to-text
Dataset description
This dataset consists of sequences of Python code followed by a a docstring explaining its function. It was constructed by concatenating code and text pairs
from this dataset that were originally code and markdown cells in Jupyter Notebooks.
The content of each example the following:
[CODE]
"""
Explanation: [TEXT]
End of explanation
"""
[CODE]
"""
Explanation: [TEXT]
End of explanation
"""
...
How to use it
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/codeparrot/github-jupyter-code-to-text.SWE-bench__style-3__fs-oracle_large_tokenlength
Dataset Card for "SWE-bench__style-3__fs-oracle"
More Information needed
code-image-to-text
Code Snippet Image → Text
A multimodal dataset for fine-tuning vision-language models (VLMs) on the task of
transcribing an image of a code snippet back into its source text — syntax-aware OCR.
Each example pairs a syntax-highlighted PNG of code with the exact code text that
produced it. It spans 8 programming languages and deliberately mixes two capture types:
block — a complete function / unit (6–45 lines).
fragment — a contiguous partial view (3–14 lines) that may start or… See the full description on the dataset page: https://huggingface.co/datasets/anisiraj/code-image-to-text.bad_code_to_good_code_dataset
Dataset Card for "bad_code_to_good_code_dataset"
More Information needed
math-to-code-gpt4o-finetuning-jsonlThis is a high quality dataset for fine tuning GPT4o and GPT4o mini with a focus on solving problems with mathematical operations using different programming languages in a similar way to the code interpreter.
Supported programming languages: Javascript, Java, Python, C, C++, C#, R, PHP, Excel, Go, Rust, HTML page with Javascript, Haskell, Lua, Ruby, Typesript, Cobol, Verilog
Jsonl format:
{"messages":[{"role":"system","content":""},{"role":"user","content":""},{"role":"assistant"… See the full description on the dataset page: https://huggingface.co/datasets/sinatra-rd/math-to-code-gpt4o-finetuning-jsonl.Compliance-to-Codecode_x_glue_ct_code_to_text_java_pythonmozzarella
Mozzarella-0.3.1
Motivation
Mozzarella is a dataset matching issues (= problem statements) and corresponding pull requests (PRs = problem solutions) of a selection of well maintained Java GitHub repositories. The original purpose was to serve as training and evaluation data for ML models concerned with fault localization and automated program repair of complex code bases. However, there might be more use cases that could benefit from this data.
Inspired by SWEBench… See the full description on the dataset page: https://huggingface.co/datasets/feedback-to-code/mozzarella.image-to-code-v1Metro_Code_Chapters_18_to_22_Data
Dataset Card for Metro System Requirements in Design
Summary
This dataset provides detailed requirements for systems used in metro network design, collected from Chapter 18-22 of the Code for Design of Metro (GB 50157-2013). The dataset is annotated using a description, categories format, aimed at facilitating the training and fine-tuning of large language models (LLMs) for information extraction tasks in complex product systems, particularly within metro transit… See the full description on the dataset page: https://huggingface.co/datasets/OrangeeSofty/Metro_Code_Chapters_18_to_22_Data.frontend-figma-to-code
CREW: Figma to Code
A benchmark for evaluating AI coding agents on Figma-to-code generation — converting real-world Figma community designs
into production-ready React + Tailwind CSS applications.
Each task gives an agent full access to a Figma file via MCP tools. The agent must extract the design system, generate
components, build successfully, and deploy a live preview. Outputs are evaluated through human preference (ELO) and
automated verifiers.
Benchmark… See the full description on the dataset page: https://huggingface.co/datasets/metaphilabs/frontend-figma-to-code.arxiv-to-code-agentic-tool-calling
arxiv-to-code-agentic-tool-calling
Multi-turn tool-calling dataset where an assistant implements ML papers in PyTorch through file-creation and command-execution tool calls.
Built from lucidrains' (Phil Wang) open-source paper implementations. There are ~217 repositories on Codeberg, each implementing a different ML paper. This dataset reverse-engineers those into synthetic coding conversations.
What's in it
199 conversations, each covering one repository. Every… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/arxiv-to-code-agentic-tool-calling.to-grok-grokking-reproduction-code
To Grok Grokking — independent reproduction code
Self-contained scripts for an independent reproduction of ICML 2026 paper
#17708, To Grok Grokking: Provable Grokking in Ridge Regression.
theory_audit.py: bounded-Rademacher finite-dimensional audit of Theorems
4.1, 4.2, and 4.4–4.6, including condition-relaxation controls.
ridge_gpu_sweep.py: paper-scale spectral GPU reproduction of the Figure 2
weight-decay and sample-size panels.
relu_gpu_sweep.py: declared-Gaussian… See the full description on the dataset page: https://huggingface.co/datasets/visv-Bro/to-grok-grokking-reproduction-code.adaption-react-screenshot-to-code
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-react_screenshot_to_code
This dataset contains 1,000 paired examples for training multimodal screenshot-to-code systems, specifically targeting React and TypeScript implementations. Each entry consists of a source webpage screenshot, a detailed visual description, and the corresponding generated React/TSX source code. The samples demonstrate high-fidelity UI… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/adaption-react-screenshot-to-code.v2c-video-to-code-demo
V2C: Video to Code/Game Demo
Summary
Video-to-Code (V2C) demo dataset containing game video recordings paired with AI-generated game code. Each sample includes the source video, chain-of-thought reasoning, game requirements analysis, and the final generated HTML game code.
Games Included
Game
Video Duration
Code Output
Description
2048
~175s (720x1440)
index.html
Number puzzle game recreation
Flappy Bird variant
~91s (582x1280)
generated_game.html… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/v2c-video-to-code-demo.strudel-nl-to-codehyperswitch-issue-to-code_v2
Rust Commit Dataset - Hyperswitch
Dataset Description
This dataset contains Rust commit messages paired with their corresponding code patches from the Hyperswitch repository.
Dataset Summary
Total Examples: 319
Language: Rust
Source: Hyperswitch GitHub repository
Format: Prompt-response pairs for supervised fine-tuning (SFT)
Data Fields
prompt: The commit message describing the change
response: The git patch/diff showing the actual code changes… See the full description on the dataset page: https://huggingface.co/datasets/archit11/hyperswitch-issue-to-code_v2.Zip-Code-to-Timezone
Dataset Card for Zip Code to Timezone Offset Mapping
This dataset maps Zip Codes and Postal Codes for the USA and Canada to the relevant timezone offset.
Dataset Details
Dataset Description
In addition to providing a mapping from a Zip Code or Postal Code to timezone offset, it also contains the timezone offset for DST (if observed).
Curated by: Bobby Gill, BlueLabel
Acknowledgements
Based off the Work Here:… See the full description on the dataset page: https://huggingface.co/datasets/omgbobbyg/Zip-Code-to-Timezone.code_x_glue_ct_code_to_text_java_pythoncode_x_glue_ct_code_to_textautotrain-data-country-to-country-code-2
AutoTrain Dataset for project: country-to-country-code-2
Dataset Description
This dataset has been automatically processed by AutoTrain for project country-to-country-code-2.
Languages
The BCP-47 code for the dataset's language is en.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"text": "France",
"target": 39
},
{
"text": "Peru",
"target": 95
}
]
Dataset Fields
The… See the full description on the dataset page: https://huggingface.co/datasets/AiBototicus/autotrain-data-country-to-country-code-2.
