datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xlcost-text-to-code XLCoST is a machine learning benchmark dataset that contains fine-grained parallel data in 7 commonly used programming languages (C++, Java, Python, C#, Javascript, PHP, C), and natural language (English).arabic-to-code-8-langs-3m
Dataset evaluation: See EVALUATION.md for schema checks, indexing status, and language-specific quality limits.
Viewer note: default is a lightweight preview; select full to load the complete corpus.
Current Hub Validation Status
Repository claim: 3,000,000 records
Dataset Server indexed rows: 1,239,045
Dataset Server estimate: 1,995,159
The 3M target figure is a raw-repository claim and is not yet fully verified by the Hub index. Validate the JSONL files before publishing a… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/arabic-to-code-8-langs-3m.github-jupyter-code-to-text
Dataset description
This dataset consists of sequences of Python code followed by a a docstring explaining its function. It was constructed by concatenating code and text pairs
from this dataset that were originally code and markdown cells in Jupyter Notebooks.
The content of each example the following:
[CODE]
"""
Explanation: [TEXT]
End of explanation
"""
[CODE]
"""
Explanation: [TEXT]
End of explanation
"""
...
How to use it
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/codeparrot/github-jupyter-code-to-text.arxiv-to-code-agentic-tool-calling
arxiv-to-code-agentic-tool-calling
Multi-turn tool-calling dataset where an assistant implements ML papers in PyTorch through file-creation and command-execution tool calls.
Built from lucidrains' (Phil Wang) open-source paper implementations. There are ~217 repositories on Codeberg, each implementing a different ML paper. This dataset reverse-engineers those into synthetic coding conversations.
What's in it
199 conversations, each covering one repository. Every… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/arxiv-to-code-agentic-tool-calling.v2c-video-to-code-demo
V2C: Video to Code/Game Demo
Summary
Video-to-Code (V2C) demo dataset containing game video recordings paired with AI-generated game code. Each sample includes the source video, chain-of-thought reasoning, game requirements analysis, and the final generated HTML game code.
Games Included
Game
Video Duration
Code Output
Description
2048
~175s (720x1440)
index.html
Number puzzle game recreation
Flappy Bird variant
~91s (582x1280)
generated_game.html… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/v2c-video-to-code-demo.hyperswitch-issue-to-code_v2
Rust Commit Dataset - Hyperswitch
Dataset Description
This dataset contains Rust commit messages paired with their corresponding code patches from the Hyperswitch repository.
Dataset Summary
Total Examples: 319
Language: Rust
Source: Hyperswitch GitHub repository
Format: Prompt-response pairs for supervised fine-tuning (SFT)
Data Fields
prompt: The commit message describing the change
response: The git patch/diff showing the actual code changes… See the full description on the dataset page: https://huggingface.co/datasets/archit11/hyperswitch-issue-to-code_v2.smolified-tiny-text-to-code
🤏 smolified-tiny-text-to-code
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model programmerGodbyte/smolified-tiny-text-to-code.
📦 Asset Details
Origin: Smolify Foundry (Job ID: fe9b19bf)
Records: 1078
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by programmerGodbyte.
Generated via Smolify.ai.
hyperswitch-issue-to-code_v3_natural
Rust Commit Dataset - Hyperswitch
Dataset Description
This dataset contains Rust commit messages paired with their corresponding code patches from the Hyperswitch repository.
Dataset Summary
Total Examples: 9
Language: Rust
Source: Hyperswitch GitHub repository
Format: Prompt-response pairs for supervised fine-tuning (SFT)
Data Fields
prompt: The commit message describing the change
response: The git patch/diff showing the actual code changes… See the full description on the dataset page: https://huggingface.co/datasets/archit11/hyperswitch-issue-to-code_v3_natural.Conv-to-Bench-Code
Conv-to-Bench: Evaluating LLMs via User-Assistant Dialogues
This repository contains the code-domain dataset generated by the Conv-to-Bench framework, presented at the 3rd Workshop on Navigating and Addressing Data Problems for Foundation Models (DATA-FM @ ICLR 2026). The framework automatically transforms authentic multi-turn dialogues between users and assistants into structured, verifiable requirement checklists for LLM evaluation.
Overview
The dataset consists of… See the full description on the dataset page: https://huggingface.co/datasets/huglabs/Conv-to-Bench-Code.hyperswitch-issue-to-code_v10
Rust Commit Dataset - Hyperswitch
Dataset Description
This dataset contains Rust commit messages paired with their corresponding code patches from the Hyperswitch repository.
Dataset Summary
Total Examples: 30
Language: Rust
Source: Hyperswitch GitHub repository
Format: Prompt-response pairs for supervised fine-tuning (SFT)
Data Fields
prompt: The commit message describing the change
response: The git patch/diff showing the actual code changes… See the full description on the dataset page: https://huggingface.co/datasets/archit11/hyperswitch-issue-to-code_v10.smolified-python-code-to-english
🤏 smolified-python-code-to-english
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model rajan-jar/smolified-python-code-to-english.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 6fe72625)
Records: 440
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by rajan-jar.
Generated via Smolify.ai.
hyperswitch-issue-to-code
Rust Commit Dataset - Hyperswitch
Dataset Description
This dataset contains Rust commit messages paired with their corresponding code patches from the Hyperswitch repository.
Dataset Summary
Total Examples: 231
Language: Rust
Source: Hyperswitch GitHub repository
Format: Prompt-response pairs for supervised fine-tuning (SFT)
Data Fields
prompt: The commit message describing the change
response: The git patch/diff showing the actual code changes… See the full description on the dataset page: https://huggingface.co/datasets/archit11/hyperswitch-issue-to-code.hyperswitch-issue-to-code_v3
Rust Commit Dataset - Hyperswitch
Dataset Description
This dataset contains Rust commit messages paired with their corresponding code patches from the Hyperswitch repository.
Dataset Summary
Total Examples: 900
Language: Rust
Source: Hyperswitch GitHub repository
Format: Prompt-response pairs for supervised fine-tuning (SFT)
Data Fields
prompt: The commit message describing the change
response: The git patch/diff showing the actual code changes… See the full description on the dataset page: https://huggingface.co/datasets/archit11/hyperswitch-issue-to-code_v3.
