datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SciCode-Domain-Code
DATA1: Domain-Specific Code Dataset
Dataset Overview
DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code.
Dataset Statistics
Total Datasets: 178 CSV files
Total Data Size: ~115 GB
Total Lines of Code: Over… See the full description on the dataset page: https://huggingface.co/datasets/SciCodePile/SciCode-Domain-Code.dialogsum-test
Dataset Card for DIALOGSum Corpus
Dataset Description
Links
Homepage: https://aclanthology.org/2021.findings-acl.449
Repository: https://github.com/cylnlp/dialogsum
Paper: https://aclanthology.org/2021.findings-acl.449
Point of Contact: https://huggingface.co/knkarthick
Dataset Summary
DialogSum is a large-scale dialogue summarization dataset, consisting of 13,460 (Plus 100 holdout data for topic generation) dialogues with corresponding… See the full description on the dataset page: https://huggingface.co/datasets/neil-code/dialogsum-test.Python-Code-Solutions
Python Code Solutions
Features
1000k of Python Code Solutions for Text Generation and Question Answering
Python Coding Problems labelled by topic and difficulty
Recommendations
Train your Model on Logical Operations and Mathematical Problems Before Training it on this. This is optional for Fine Tuning 2B parameter + models.
Format the prompts in a orderly way when formatting data eg. {question} Solution: {solution} Topic: {topic}
Code_Vulnerability_Labeled_Dataset
Dataset Card for Code_Vulnerability_Labeled_Dataset
Dataset Summary
This dataset provides (code, vulnerability) pairs. The vulnerability field takes values according to the CWE annotation:
CWE
Description
CWE-020
Improper Input Validation
CWE-022
Improper Limitation of a Pathname to a Restricted Directory (“Path Traversal”)
CWE-078
Improper Neutralization of Special Elements used in an OS Command (“OS Command Injection”)
CWE-079
Improper Neutralization of… See the full description on the dataset page: https://huggingface.co/datasets/lemon42-ai/Code_Vulnerability_Labeled_Dataset.CodeMixBench
ℹ️Dataset Card for CodeMixBench
[EMNLP'25] CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages
Code-mixing is a linguistic phenomenon where multilingual speakers switch or mix two or more languages within a single utterance or conversation.
To evaluate LLMs’ comprehension of multilingual code-mixed texts, we introduce CodeMixBench, a benchmark comprising eight tasks across 18 languages.… See the full description on the dataset page: https://huggingface.co/datasets/CodeMixBench/CodeMixBench.qwen9b-coop-claude-code
qwen9b-coop-claude-code
Two-agent cooperative coding trajectories generated by running
CooperBench in coop mode on
the CooperData task set, using
Qwen/Qwen3.5-9B as the model and Claude Code (claude_code) as the
agent framework. Each pair runs two agents in parallel — one per feature —
coordinating via Redis messaging and a shared git remote.
The matched solo (single-agent) baseline is at
CooperBench/qwen9b-solo-claude-code.
Same task corpus, same model, same agent — only the… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-coop-claude-code.Flutter-Code-with-Questions-Dataset-English
🧠 Flutter Code with Questions Dataset (English)
This repository contains a high-quality dataset of Flutter-related code snippets paired with automatically generated English technical questions. The dataset is intended for use in training and fine-tuning language models, coding assistants, and educational systems focused on Flutter development.
📂 Dataset Structure
The dataset is divided into 22 CSV files, each containing 200 entries. Every entry includes:
A… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-English.Flutter-Code-with-Questions-Dataset-Turkish
Flutter Code with Questions Dataset (Turkish)
📦 Dataset Name: flutter_code_with_questions
Bu veri seti, Flutter framework'ü ile yazılmış kod parçacıkları ve her bir kod parçası için özel olarak üretilmiş detaylı Türkçe soruları içermektedir. Veri seti, kodların eğitim verisi olarak kullanılmasının yanı sıra, LLM (Large Language Model) tabanlı kod anlama ve soru yanıtlama modellerinin geliştirilmesinde kullanılabilir.
📁 Dataset Format
Veri dosyaları CSV… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-Turkish.Deepseek-V4-Reasoning-Code-2500
DeepSeek Reasoning and Code Distillation Dataset
This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research.
The dataset file is:
train.csv
It contains 2… See the full description on the dataset page: https://huggingface.co/datasets/Banaxi-Tech/Deepseek-V4-Reasoning-Code-2500.SciCode-Domain-Code
DATA1: Domain-Specific Code Dataset
Dataset Overview
DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code.
Dataset Statistics
Total Datasets: 178 CSV files
Total Data Size: ~115 GB
Total Lines of Code: Over… See the full description on the dataset page: https://huggingface.co/datasets/liuhangbiao/SciCode-Domain-Code.qwen9b-solo-claude-code
qwen9b-solo-claude-code
Single-agent coding trajectories generated by running
CooperBench in solo mode on
the CooperData task set, using
Qwen/Qwen3.5-9B as the model and Claude Code (claude_code) as the
agent framework. One agent implements both features in each task.
The matched coop (two-agent) version is at
CooperBench/qwen9b-coop-claude-code.
Same task corpus, same model, same agent — only the coordination differs, so
together they isolate the cooperation deficit.
At a… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-solo-claude-code.DA-Code
[EMNLP2024] DA-Code: Agent Data Science Code Generation Benchmark for Large Language Models
DA-Code is a comprehensive evaluation dataset designed to assess the data analysis and code generation capabilities of LLM in agent-based data science tasks. Our papers and experiment reports have been published on Arxiv.
Dataset Overview
500 complex real-world data analysis tasks across Data Wrangling (DW), Machine Learning (ML), and Exploratory Data Analysis (EDA).
Tasks cover… See the full description on the dataset page: https://huggingface.co/datasets/Luo2003/DA-Code.goodreads-books
Goodreads Books Dataset
Dataset Description
A comprehensive dataset of books scraped from Goodreads, including ratings, authors, titles, and various book characteristics.
This dataset contains 3045 books with 20 features each, scraped from Goodreads. It's perfect for:
📚 Book recommendation systems
📊 Literary data analysis
🤖 Machine learning projects
📈 Rating prediction models
🔍 Book discovery algorithms
Dataset Structure
Features… See the full description on the dataset page: https://huggingface.co/datasets/codealchemist01/goodreads-books.letterboxd-movies
Letterboxd Movies Dataset
Dataset Description
A comprehensive dataset of movies scraped from Letterboxd, including genres, ratings, runtime, countries, and detailed movie characteristics.
This dataset contains 16246 movies with 28 features each, scraped from Letterboxd. It's perfect for:
🎬 Movie recommendation systems
📊 Film industry analysis
🤖 Machine learning projects
📈 Rating prediction models
🔍 Movie discovery algorithms
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/codealchemist01/letterboxd-movies.SwiftUI-Code-Examples
SwiftUI Code Solutions
Dataset Created by MCES10 Software has SwiftUI Code Problems and can be used for AI training for Code Generation
Recommendations
Train your LLM on the Swift and SwiftUI Framework Syntax before training it this
Fine Tune or Train Effectively at optimal Epochs and Learning Rates
Use the whole dataset for training
Your Model may need to be Prompt Tuned for the best performance but it isn't required.
Use test when testing or trialing the dataset
Use… See the full description on the dataset page: https://huggingface.co/datasets/MCES10-Software/SwiftUI-Code-Examples.CodeMixBench
ℹ️Dataset Card for CodeMixBench
[EMNLP'25] CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages
Code-mixing is a linguistic phenomenon where multilingual speakers switch or mix two or more languages within a single utterance or conversation.
To evaluate LLMs’ comprehension of multilingual code-mixed texts, we introduce CodeMixBench, a benchmark comprising eight tasks across 18 languages.… See the full description on the dataset page: https://huggingface.co/datasets/Tanushreeeeee/CodeMixBench.CodeReality
CodeReality: Evaluation Subset - Deliberately Noisy Code Dataset
⚠️ Important Limitations
⚠️ Not Enterprise-Ready: This dataset is deliberately noisy and designed for research only. Contains mixed/unknown licenses, possible secrets, potential security vulnerabilities, duplicate code, and experimental repositories. Requires substantial preprocessing for production use.
Use at your own risk - this is a research dataset for robustness testing and data curation method… See the full description on the dataset page: https://huggingface.co/datasets/vinsblack/CodeReality.Hinglish-Hindi-Transliteration-Dataset
Hinglish-Hindi Transliteration Dataset
We are pleased to release this unique dataset focused on transliteration between Hinglish (Hindi written in Roman script) and Devanagari Hindi. This dataset aims to address the limitations of current models in accurately transliterating words and phrases as they are commonly used, preserving their original form and meaning. Unlike translation datasets, this resource focuses on phonetic equivalence rather than semantic transformation. For… See the full description on the dataset page: https://huggingface.co/datasets/codebyam/Hinglish-Hindi-Transliteration-Dataset.Deepseek-code
DeepSeek Reasoning and Code Distillation Dataset
This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research.
The dataset file is:
train.csv
It contains… See the full description on the dataset page: https://huggingface.co/datasets/ryen-stuff/Deepseek-code.GPT2-Hacker-password-generator-dataset
Hacker Style Password Generation Dataset
Dataset Description
This dataset contains 20,000 instruction-response pairs designed to train and evaluate language models for generating strong, "hacker-style" passwords. The data simulates a user requesting a secure password and the model providing a complex, randomly generated string.
Supported Tasks
Text Generation: The primary task is conditional text generation, where the model takes a natural language instruction… See the full description on the dataset page: https://huggingface.co/datasets/CodeferSystem/GPT2-Hacker-password-generator-dataset.CodevBench
🤗 Hugging Face | 🤖 ModelScope | 📑 GitHub | 📄Arxiv
Codev-Bench: How Do LLMs Understand Developer-Centric Code Completion?
Introduction
Codev-Bench (Code Development Benchmark), a fine-grained, real-world, repository-level, and developer-centric evaluation framework. Codev-Bench assesses whether a code completion tool can accurately capture a developer's immediate intent and suggest appropriate code snippets across diverse… See the full description on the dataset page: https://huggingface.co/datasets/TongyiLingma/CodevBench.Bangla-Code-Instruct
🐯 Bangla-Code-Instruct: A Comprehensive Bangla Code Instruction Dataset
Accepted at LREC 2026
Nishat Raihan, Antonios Anastasopoulos, Marcos Zampieri
George Mason University, Fairfax, VA, USA
The first large-scale Bangla code instruction dataset (300K examples) for training Code LLMs in Bangla.
⚠️ Note: The dataset will be released after the LREC 2026 conference. Stay tuned!
Overview
Bangla-Code-Instruct is a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Bangla-Code-Instruct.CPP-Code-Solutions
C++ Code Solutions
Features
1000k of Python Code Solutions for Text Generation and Question Answering
C++ Coding Problems labelled by topic and difficulty
Recommendations
Train your Model on Logical Operations and Mathematical Problems Before Training it on this. This is optional for Fine Tuning 2B parameter + models.
Format the prompts in a orderly way when formatting data eg. {question} Solution: {solution} Topic: {topic}
SciCode-Domain-Code
DATA1: Domain-Specific Code Dataset
Dataset Overview
DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code.
Dataset Statistics
Total Datasets: 178 CSV files
Total Data Size: ~115 GB
Total Lines… See the full description on the dataset page: https://huggingface.co/datasets/SciCode/SciCode-Domain-Code.CodeReasoningPro
CodeReasoningPro
Dataset Summary
CodeReasoningPro is a large-scale synthetic dataset comprising 1,785,725 competitive programming problems in Python, created by XythicK, an MLOps Engineer. Designed for supervised fine-tuning (SFT) of machine learning models for coding tasks, it draws inspiration from datasets like OpenCodeReasoning. The dataset includes problem statements, Python solutions, and reasoning explanations, covering algorithmic topics such as arrays, subarrays… See the full description on the dataset page: https://huggingface.co/datasets/XythicK/CodeReasoningPro.Deepseek-V4-Reasoning-Code-2500
DeepSeek Reasoning and Code Distillation Dataset
This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research.
The dataset file is:
train.csv
It contains… See the full description on the dataset page: https://huggingface.co/datasets/lucsaint/Deepseek-V4-Reasoning-Code-2500.math-code-qa
Math & Code QA — Instruction Dataset
Worked mathematical solutions and short code answers, built for the
Adaption Labs AutoScientist Challenge (Math & Code category).
Rows
5,200
Math
3,600
Code
1,600
Distinct answers
5,199 (100%)
Duplicate questions
none
Nulls
none
Question length
median 27 words
Answer length
median 58 words (max 89)
License
CC-BY-4.0
What makes the math rows unusual
Every math answer is short worked reasoning… See the full description on the dataset page: https://huggingface.co/datasets/flamiinngo/math-code-qa.ceo-quotes-verified-sample
🎙️ CEO Transcripts — Verified Executive Interviews
The World's Largest Database of Verified C-Suite Transcripts
20,000+ Executives · 100,000+ Transcripts · 400,000+ Quotes · S&P 500 + NASDAQ + Global Leaders
🔥 What's In This Sample?
This is a free evaluation sample from CEOInterviews.ai featuring 9 of the most market-moving voices in finance, tech, and policy.
Executive
Role
Why They Matter
Jensen Huang
CEO, NVIDIA
Every AI… See the full description on the dataset page: https://huggingface.co/datasets/codelucas/ceo-quotes-verified-sample.JS-Code-Solutions
Python Code Solutions
Features
1000k of JS Code Solutions for Text Generation and Question Answering
JS Coding Problems labelled by topic and difficulty
Recommendations
Train your Model on Logical Operations and Mathematical Problems Before Training it on this. This is optional for Fine Tuning 2B parameter + models.
Format the prompts in a orderly way when formatting data eg. {question} Solution: {solution} Topic: {topic}
math-code-qa-v2
Math & Code QA v2 — Instruction Dataset
Worked mathematical solutions and short code answers, spanning arithmetic word
problems through to algebra, geometry and combinatorics.
Built for the Adaption Labs AutoScientist Challenge (Math & Code category).
The model trained on this beats Llama-3.3-70B-Instruct 72 to 28 on the
held-out Math category evaluation.
Rows
5,297 (4,197 math, 1,100 code)
Distinct answers
5,297 (100%)
Duplicate questions
none
Nulls
none… See the full description on the dataset page: https://huggingface.co/datasets/flamiinngo/math-code-qa-v2.
