datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code_generation
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
๐ Home Page โข
๐ป GitHub Repository โข
๐ Leaderboard โข
LiveCodeBench is a "live" updating benchmark for holistically evaluating code related capabilities of LLMs.
Particularly, it evaluates LLMs across a range of capabilties including code generation, self-repair, test output prediction, and code execution.
This is the code generation scenario of LiveCodeBench. It is alsoโฆ See the full description on the dataset page: https://huggingface.co/datasets/livecodebench/code_generation.text-code-galeras-code-generation-from-docstring-3k-dedupedleetcode_code_generationsynthetic-code-generationsThis dataset was synthetically generated using mixtral8x7b to create unique instructions following the MagicCoder Paper and reproducing the results by modifying specific attributes (snippets are larger, instructions/responses are larger, and more specific).
Below is the prompt used to generate the instruction set:
prompt=f"""<s>[INST] You are an incredibly intelligent programming AI with expertise in CloudFormation, Terraform, AWS CDK and {lang}. Please gain inspiration from the followingโฆ See the full description on the dataset page: https://huggingface.co/datasets/VishaalY/synthetic-code-generations.code-generation-sft-100k
Code Generation SFT (100K)
100,000 ShareGPT conversations covering code generation across 8 programming languages, 21 categories, and 22 distinct programming tasks. Each example includes a detailed natural language request and a complete, working implementation with explanations of key design decisions.
Motivation
Coding assistants are the highest-adoption LLM application category, but most open training datasets focus on isolated functions without context. Thisโฆ See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/code-generation-sft-100k.LiveCodeBench-CodeGenerationLLM-ABAP-Code-Generation-Benchmark
LLM Benchmark ABAP Code Generation Dataset
This dataset is designed for benchmarking Large Language Models (LLMs) on ABAP code generation capabilities. It is based on the HumanEval benchmark, adapted for ABAP, and includes 16 additional ABAP-specific tasks that require interaction with database tables.
Total tasks: 180
164 tasks adapted from HumanEval
16 ABAP-specific tasks
Dataset Structure
dataset.jsonl: Contains 180 examples. Each example has:
id: Uniqueโฆ See the full description on the dataset page: https://huggingface.co/datasets/timkoehne/LLM-ABAP-Code-Generation-Benchmark.deckergui-code-generation
Code generation and reasoning pairs from DeckerGUI development sessions. Contains user instructions, context, and model-generated code for TypeScript, Markdown, and JSON.
Dataset Details
Repository: ctaxnagomi/deckergui-code-generation
License: MIT
DeckerGUI Version: v2.0.0
Created: 2026-08-17
Dataset Schema
See metadata.json for the full schema definition.
Usage
from datasets import load_dataset
ds =โฆ See the full description on the dataset page: https://huggingface.co/datasets/ctaxnagomi/deckergui-code-generation.code-text-galeras-commit-generation-3k-dedupedCode-Generation-LLM-LoRAcode-generation-eval
Code Generation Evaluation
5 code generation tasks for evaluating dispatchAI coder models.
Best models: Qwen2.5-0.5B-Coder-mobile, Qwen2.5-Coder-1.5B-mobile
๐ dispatchAI
arch-code-transfer-lpi-260903T0846-w2-raw-generationscode_generation_ko
code_generation_ko
https://huggingface.co/datasets/livecodebench/code_generation
gpt-4o๋ฅผ ์ด์ฉํด question_title๊ณผ question_content๋ฅผ ํ๊ธ๋ก ๋ฒ์ญํ ์ฝ๋ฉ ์ง๋ฌธ ๋ฐ์ดํฐ์
.
๊ตฌ์กฐ
{
"question_title": "๋ถ๋ฆฌ๋ ์ ์ฌ",
"question_content": "KEYENCE ๋ณธ์ฌ์ ์ ์ ๋ ๋ง์ ์ง์๋ค์ด ์๊ธฐ๋ฉด์, ๋ณธ์ฌ ๋ด ๋ถ์๋ค์ ๋ ๊ทธ๋ฃน์ผ๋ก ๋๋์ด ์ ์ฌ์๊ฐ์ ์์ฐจ์ ๋ก ํ๊ธฐ๋ก ๊ฒฐ์ ํ์ต๋๋ค.\nKEYENCE ๋ณธ์ฌ์๋ N๊ฐ์ ๋ถ์๊ฐ ์์ผ๋ฉฐ, i๋ฒ์งธ ๋ถ์(1\\leq i\\leq N)์ ์ธ์์๋ K_i์
๋๋ค.\n๊ฐ ๋ถ์๋ฅผ ๊ทธ๋ฃน A ๋๋ ๊ทธ๋ฃน B์ ๋ฐฐ์ ํ๊ณ , ๊ฐ ๊ทธ๋ฃน์ด ๊ฐ์ ์๊ฐ์ ์ ์ฌ์๊ฐ์ ๊ฐ์ง๋ฉฐ, ๊ทธ๋ฃน A์ ๊ทธ๋ฃน B์ ์ ์ฌ์๊ฐ์ด ๊ฒน์น์ง ์๋๋ก ํ ๋, ๋์์ ์ ์ฌ์ ๋จน๋ ์ต๋ ์ธ์์ ์ต์ ๊ฐ๋ฅํ ๊ฐ์ ์ฐพ์ผ์ธ์.\n์ฆ, ๋ค์ ์ค ๋ ํฐ ๊ฐ์โฆ See the full description on the dataset page: https://huggingface.co/datasets/twodigit/code_generation_ko.lcb_codegeneration_v6_shortradon-test-code_generation
radon-test-code_generation
Description
Code generation test dataset for RADON model evaluation with programming prompts
Usage
Load Dataset
from datasets import load_dataset
dataset = load_dataset("MagistrTheOne/radon-test-code_generation")
print(dataset)
Use with RADON Model
from transformers import AutoModelForCausalLM, AutoTokenizer
# Load RADON model
model = AutoModelForCausalLM.from_pretrained("MagistrTheOne/RadonSAI")
tokenizer =โฆ See the full description on the dataset page: https://huggingface.co/datasets/MagistrTheOne/radon-test-code_generation.code_generation
code_generation
https://huggingface.co/datasets/livecodebench/code_generation
์๋ฌธ ์ฝ๋ฉ ์ง๋ฌธ ๋ต๋ณ ๋ฐ์ดํฐ์
.
๊ตฌ์กฐ
{
"question_title": "Separated Lunch",
"question_content": "As KEYENCE headquarters have more and more workers, they decided to divide the departments in the headquarters into two groups and stagger their lunch breaks.\nKEYENCE headquarters have N departments, and the number of people in the i-th department (1\\leq i\\leq N) is K_i.\nWhen assigning each department toโฆ See the full description on the dataset page: https://huggingface.co/datasets/twodigit/code_generation.code_generation_v2
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
๐ Home Page โข
๐ป GitHub Repository โข
๐ Leaderboard โข
LiveCodeBench is a "live" updating benchmark for holistically evaluating code related capabilities of LLMs.
Particularly, it evaluates LLMs across a range of capabilties including code generation, self-repair, test output prediction, and code execution.
This is the code generation scenario of LiveCodeBench. It isโฆ See the full description on the dataset page: https://huggingface.co/datasets/mathewmouchamel/code_generation_v2.
