datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code_generation
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
๐ Home Page โข
๐ป GitHub Repository โข
๐ Leaderboard โข
LiveCodeBench is a "live" updating benchmark for holistically evaluating code related capabilities of LLMs.
Particularly, it evaluates LLMs across a range of capabilties including code generation, self-repair, test output prediction, and code execution.
This is the code generation scenario of LiveCodeBench. It is alsoโฆ See the full description on the dataset page: https://huggingface.co/datasets/livecodebench/code_generation.CodeGen-Diverse-5K
CodeGen-Diverse-5K: Broad Coverage for Competitive Programming
Part of the CodeGen suite | CodeGen-Deep-5K (sister dataset)
Dataset Description
CodeGen-Diverse-5K is a broad coverage dataset designed for training code generation models across a wide variety of competitive programming problems. This dataset prioritizes problem diversity over solution diversity, covering 5,000 unique problems with consistent, high-quality solutions.
Key Statistics
Total samples:โฆ See the full description on the dataset page: https://huggingface.co/datasets/Naholav/CodeGen-Diverse-5K.llama3.2-java-codegen-90sft-10meta-claude-v1
LLaMA 3.2 Java Code Generation Dataset (90% SFT, 10% Meta Annotated with Claude)
This dataset contains 100,000 examples for Java method generation based on natural language instructions. It is built from the CodeXGLUE text-to-code dataset and designed to support both pure supervised fine-tuning (SFT) and reflection-based meta-learning approaches using Claude 4 Sonnet as the critique model.
๐ Trained Models
Two models have been trained on this dataset:
SFT Model:โฆ See the full description on the dataset page: https://huggingface.co/datasets/Naholav/llama3.2-java-codegen-90sft-10meta-claude-v1.code-generation-sft-100k
Code Generation SFT (100K)
100,000 ShareGPT conversations covering code generation across 8 programming languages, 21 categories, and 22 distinct programming tasks. Each example includes a detailed natural language request and a complete, working implementation with explanations of key design decisions.
Motivation
Coding assistants are the highest-adoption LLM application category, but most open training datasets focus on isolated functions without context. Thisโฆ See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/code-generation-sft-100k.LiveCodeBench-CodeGenerationCodeGen-Deep-5K
CodeGen-Deep-5K: Deep Reasoning for Competitive Programming
Part of the CodeGen suite | CodeGen-Diverse-5K (sister dataset)
Dataset Description
CodeGen-Deep-5K is a deep reasoning dataset designed for training code generation models with enhanced problem-solving capabilities. Unlike traditional datasets, this generates multiple distinct solutions for each problem, providing varied reasoning traces and approaches.
Key Statistics
Total samples: 5,000
Uniqueโฆ See the full description on the dataset page: https://huggingface.co/datasets/Naholav/CodeGen-Deep-5K.terraform-aws-ec2-instance-profile-codegenterraform-aws-iam-role-policy-codegencve-backport-codegen-dataset
CVE Backport Code Generation Dataset
Per-hunk code generation dataset for CVE security patch backporting, derived from openSUSE Build Service maintenance patches.
Task
Given a region of vulnerable source code and a description of the upstream CVE fix, the model outputs the fixed version of the code. A programmatic diff then produces the final patch. This plays to LLM strengths in code completion and avoids format-sensitivity issues with direct diff generation.โฆ See the full description on the dataset page: https://huggingface.co/datasets/openSUSE/cve-backport-codegen-dataset.Code-Generation-LLM-LoRAmanim-codegencode-generation-eval
Code Generation Evaluation
5 code generation tasks for evaluating dispatchAI coder models.
Best models: Qwen2.5-0.5B-Coder-mobile, Qwen2.5-Coder-1.5B-mobile
๐ dispatchAI
terraform-aws-vpc-nat-gateway-codegenterraform-aws-s3-lifecycle-versioning-codegencode_generation_ko
code_generation_ko
https://huggingface.co/datasets/livecodebench/code_generation
gpt-4o๋ฅผ ์ด์ฉํด question_title๊ณผ question_content๋ฅผ ํ๊ธ๋ก ๋ฒ์ญํ ์ฝ๋ฉ ์ง๋ฌธ ๋ฐ์ดํฐ์
.
๊ตฌ์กฐ
{
"question_title": "๋ถ๋ฆฌ๋ ์ ์ฌ",
"question_content": "KEYENCE ๋ณธ์ฌ์ ์ ์ ๋ ๋ง์ ์ง์๋ค์ด ์๊ธฐ๋ฉด์, ๋ณธ์ฌ ๋ด ๋ถ์๋ค์ ๋ ๊ทธ๋ฃน์ผ๋ก ๋๋์ด ์ ์ฌ์๊ฐ์ ์์ฐจ์ ๋ก ํ๊ธฐ๋ก ๊ฒฐ์ ํ์ต๋๋ค.\nKEYENCE ๋ณธ์ฌ์๋ N๊ฐ์ ๋ถ์๊ฐ ์์ผ๋ฉฐ, i๋ฒ์งธ ๋ถ์(1\\leq i\\leq N)์ ์ธ์์๋ K_i์
๋๋ค.\n๊ฐ ๋ถ์๋ฅผ ๊ทธ๋ฃน A ๋๋ ๊ทธ๋ฃน B์ ๋ฐฐ์ ํ๊ณ , ๊ฐ ๊ทธ๋ฃน์ด ๊ฐ์ ์๊ฐ์ ์ ์ฌ์๊ฐ์ ๊ฐ์ง๋ฉฐ, ๊ทธ๋ฃน A์ ๊ทธ๋ฃน B์ ์ ์ฌ์๊ฐ์ด ๊ฒน์น์ง ์๋๋ก ํ ๋, ๋์์ ์ ์ฌ์ ๋จน๋ ์ต๋ ์ธ์์ ์ต์ ๊ฐ๋ฅํ ๊ฐ์ ์ฐพ์ผ์ธ์.\n์ฆ, ๋ค์ ์ค ๋ ํฐ ๊ฐ์โฆ See the full description on the dataset page: https://huggingface.co/datasets/twodigit/code_generation_ko.lcb_codegeneration_v6_shortcode_generation
code_generation
https://huggingface.co/datasets/livecodebench/code_generation
์๋ฌธ ์ฝ๋ฉ ์ง๋ฌธ ๋ต๋ณ ๋ฐ์ดํฐ์
.
๊ตฌ์กฐ
{
"question_title": "Separated Lunch",
"question_content": "As KEYENCE headquarters have more and more workers, they decided to divide the departments in the headquarters into two groups and stagger their lunch breaks.\nKEYENCE headquarters have N departments, and the number of people in the i-th department (1\\leq i\\leq N) is K_i.\nWhen assigning each department toโฆ See the full description on the dataset page: https://huggingface.co/datasets/twodigit/code_generation.code-gen-lcb-error-analysisyulan-codegenry_manim_codegen
RY Manim Code Generation
This dataset contains queries and responses for ideal Manim code. The input and output have been preprocessed to keep context clean.
Format:
Jsonl format used instead of Json for better efficiency.
{"query": "<a user's question about a topic>", "output": "<LLM's manim code>"}
Citations:
This repository includes data sourced and processed from the following datasets:โฆ See the full description on the dataset page: https://huggingface.co/datasets/RY7Games/ry_manim_codegen.converted_codegen_dataset.jsonprompt-sensitivity-codegen
Anonymous Prompt Sensitivity Dataset
This package contains model generations and evaluation outcomes for an anonymized
submission on prompt sensitivity in few-shot code generation.
What is included
prompt_sensitivity_dataset.jsonl: one row per generated sample
prompt_sensitivity_dataset.csv: tabular view of the same rows
prompt_sensitivity_dataset.parquet: columnar copy when parquet support is available
prompt_variant_spec.json: machine-readable description of the promptโฆ See the full description on the dataset page: https://huggingface.co/datasets/anonymous-acl26/prompt-sensitivity-codegen.python-java-codegen-datasetFinal_codegen_1000_entriescode_generation_v2
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
๐ Home Page โข
๐ป GitHub Repository โข
๐ Leaderboard โข
LiveCodeBench is a "live" updating benchmark for holistically evaluating code related capabilities of LLMs.
Particularly, it evaluates LLMs across a range of capabilties including code generation, self-repair, test output prediction, and code execution.
This is the code generation scenario of LiveCodeBench. It isโฆ See the full description on the dataset page: https://huggingface.co/datasets/mathewmouchamel/code_generation_v2.codegenmanim-codegenpython-java-codegen-datasetcode_alpaca_codegene_20kCodeGenDataset
