Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01livecodebench /code_generation LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code ๐Ÿ  Home Page โ€ข ๐Ÿ’ป GitHub Repository โ€ข ๐Ÿ† Leaderboard โ€ข LiveCodeBench is a "live" updating benchmark for holistically evaluating code related capabilities of LLMs. Particularly, it evaluates LLMs across a range of capabilties including code generation, self-repair, test output prediction, and code execution. This is the code generation scenario of LiveCodeBench. It is alsoโ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/livecodebench/code_generation.textn<1K35 likes4.9k downloads2y agoHugging Face02Naholav /CodeGen-Diverse-5K CodeGen-Diverse-5K: Broad Coverage for Competitive Programming Part of the CodeGen suite | CodeGen-Deep-5K (sister dataset) Dataset Description CodeGen-Diverse-5K is a broad coverage dataset designed for training code generation models across a wide variety of competitive programming problems. This dataset prioritizes problem diversity over solution diversity, covering 5,000 unique problems with consistent, high-quality solutions. Key Statistics Total samples:โ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/Naholav/CodeGen-Diverse-5K.tabulartext-generation1K<n<10K0 likes108 downloads10mo agoHugging Face03Naholav /llama3.2-java-codegen-90sft-10meta-claude-v1 LLaMA 3.2 Java Code Generation Dataset (90% SFT, 10% Meta Annotated with Claude) This dataset contains 100,000 examples for Java method generation based on natural language instructions. It is built from the CodeXGLUE text-to-code dataset and designed to support both pure supervised fine-tuning (SFT) and reflection-based meta-learning approaches using Claude 4 Sonnet as the critique model. ๐Ÿš€ Trained Models Two models have been trained on this dataset: SFT Model:โ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/Naholav/llama3.2-java-codegen-90sft-10meta-claude-v1.texttext-generation100K<n<1M1 likes93 downloads1y agoHugging Face04stindardlogic /code-generation-sft-100k Code Generation SFT (100K) 100,000 ShareGPT conversations covering code generation across 8 programming languages, 21 categories, and 22 distinct programming tasks. Each example includes a detailed natural language request and a complete, working implementation with explanations of key design decisions. Motivation Coding assistants are the highest-adoption LLM application category, but most open training datasets focus on isolated functions without context. Thisโ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/code-generation-sft-100k.texttext-generation100K<n<1M1 likes81 downloads3mo agoHugging Face05Groq /LiveCodeBench-CodeGenerationtextquestion-answeringn<1K0 likes61 downloads1y agoHugging Face06Naholav /CodeGen-Deep-5K CodeGen-Deep-5K: Deep Reasoning for Competitive Programming Part of the CodeGen suite | CodeGen-Diverse-5K (sister dataset) Dataset Description CodeGen-Deep-5K is a deep reasoning dataset designed for training code generation models with enhanced problem-solving capabilities. Unlike traditional datasets, this generates multiple distinct solutions for each problem, providing varied reasoning traces and approaches. Key Statistics Total samples: 5,000 Uniqueโ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/Naholav/CodeGen-Deep-5K.tabulartext-generation1K<n<10K0 likes52 downloads10mo agoHugging Face07rootly-ai-labs /terraform-aws-ec2-instance-profile-codegentextn<1K0 likes36 downloads9mo agoHugging Face08rootly-ai-labs /terraform-aws-iam-role-policy-codegentextn<1K0 likes31 downloads9mo agoHugging Face09openSUSE /cve-backport-codegen-dataset CVE Backport Code Generation Dataset Per-hunk code generation dataset for CVE security patch backporting, derived from openSUSE Build Service maintenance patches. Task Given a region of vulnerable source code and a description of the upstream CVE fix, the model outputs the fixed version of the code. A programmatic diff then produces the final patch. This plays to LLM strengths in code completion and avoids format-sensitivity issues with direct diff generation.โ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/openSUSE/cve-backport-codegen-dataset.texttext-generation10K<n<100K1 likes30 downloads6mo agoHugging Face10Rabinovich /Code-Generation-LLM-LoRAtext10K<n<100K0 likes28 downloads2y agoHugging Face11generaleoley /manim-codegentext1K<n<10K11 likes27 downloads3y agoHugging Face12dispatchAI /code-generation-eval Code Generation Evaluation 5 code generation tasks for evaluating dispatchAI coder models. Best models: Qwen2.5-0.5B-Coder-mobile, Qwen2.5-Coder-1.5B-mobile ๐Ÿš€ dispatchAI textn<1K0 likes27 downloads3mo agoHugging Face13rootly-ai-labs /terraform-aws-vpc-nat-gateway-codegentextn<1K0 likes25 downloads9mo agoHugging Face14rootly-ai-labs /terraform-aws-s3-lifecycle-versioning-codegentextn<1K0 likes22 downloads9mo agoHugging Face15twodigit /code_generation_ko code_generation_ko https://huggingface.co/datasets/livecodebench/code_generation gpt-4o๋ฅผ ์ด์šฉํ•ด question_title๊ณผ question_content๋ฅผ ํ•œ๊ธ€๋กœ ๋ฒˆ์—ญํ•œ ์ฝ”๋”ฉ ์งˆ๋ฌธ ๋ฐ์ดํ„ฐ์…‹. ๊ตฌ์กฐ { "question_title": "๋ถ„๋ฆฌ๋œ ์ ์‹ฌ", "question_content": "KEYENCE ๋ณธ์‚ฌ์— ์ ์  ๋” ๋งŽ์€ ์ง์›๋“ค์ด ์ƒ๊ธฐ๋ฉด์„œ, ๋ณธ์‚ฌ ๋‚ด ๋ถ€์„œ๋“ค์„ ๋‘ ๊ทธ๋ฃน์œผ๋กœ ๋‚˜๋ˆ„์–ด ์ ์‹ฌ์‹œ๊ฐ„์„ ์‹œ์ฐจ์ œ๋กœ ํ•˜๊ธฐ๋กœ ๊ฒฐ์ •ํ–ˆ์Šต๋‹ˆ๋‹ค.\nKEYENCE ๋ณธ์‚ฌ์—๋Š” N๊ฐœ์˜ ๋ถ€์„œ๊ฐ€ ์žˆ์œผ๋ฉฐ, i๋ฒˆ์งธ ๋ถ€์„œ(1\\leq i\\leq N)์˜ ์ธ์›์ˆ˜๋Š” K_i์ž…๋‹ˆ๋‹ค.\n๊ฐ ๋ถ€์„œ๋ฅผ ๊ทธ๋ฃน A ๋˜๋Š” ๊ทธ๋ฃน B์— ๋ฐฐ์ •ํ•˜๊ณ , ๊ฐ ๊ทธ๋ฃน์ด ๊ฐ™์€ ์‹œ๊ฐ„์— ์ ์‹ฌ์‹œ๊ฐ„์„ ๊ฐ€์ง€๋ฉฐ, ๊ทธ๋ฃน A์™€ ๊ทธ๋ฃน B์˜ ์ ์‹ฌ์‹œ๊ฐ„์ด ๊ฒน์น˜์ง€ ์•Š๋„๋ก ํ•  ๋•Œ, ๋™์‹œ์— ์ ์‹ฌ์„ ๋จน๋Š” ์ตœ๋Œ€ ์ธ์›์˜ ์ตœ์†Œ ๊ฐ€๋Šฅํ•œ ๊ฐ’์„ ์ฐพ์œผ์„ธ์š”.\n์ฆ‰, ๋‹ค์Œ ์ค‘ ๋” ํฐ ๊ฐ’์˜โ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/twodigit/code_generation_ko.textn<1K0 likes18 downloads2y agoHugging Face16ekurtic /lcb_codegeneration_v6_shorttextn<1K0 likes17 downloads2mo agoHugging Face17twodigit /code_generation code_generation https://huggingface.co/datasets/livecodebench/code_generation ์˜๋ฌธ ์ฝ”๋”ฉ ์งˆ๋ฌธ ๋‹ต๋ณ€ ๋ฐ์ดํ„ฐ์…‹. ๊ตฌ์กฐ { "question_title": "Separated Lunch", "question_content": "As KEYENCE headquarters have more and more workers, they decided to divide the departments in the headquarters into two groups and stagger their lunch breaks.\nKEYENCE headquarters have N departments, and the number of people in the i-th department (1\\leq i\\leq N) is K_i.\nWhen assigning each department toโ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/twodigit/code_generation.textn<1K0 likes13 downloads2y agoHugging Face18masteramir /code-gen-lcb-error-analysistext10K<n<100K0 likes12 downloads1y agoHugging Face19semran1 /yulan-codegentext1M<n<10M0 likes11 downloads1y agoHugging Face20RY7Games /ry_manim_codegen RY Manim Code Generation This dataset contains queries and responses for ideal Manim code. The input and output have been preprocessed to keep context clean. Format: Jsonl format used instead of Json for better efficiency. {"query": "<a user's question about a topic>", "output": "<LLM's manim code>"} Citations: This repository includes data sourced and processed from the following datasets:โ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/RY7Games/ry_manim_codegen.texttext-generation1K<n<10K1 likes11 downloads2mo agoHugging Face21Abhiverse01 /converted_codegen_dataset.jsontext1K<n<10K0 likes9 downloads3y agoHugging Face22anonymous-acl26 /prompt-sensitivity-codegen Anonymous Prompt Sensitivity Dataset This package contains model generations and evaluation outcomes for an anonymized submission on prompt sensitivity in few-shot code generation. What is included prompt_sensitivity_dataset.jsonl: one row per generated sample prompt_sensitivity_dataset.csv: tabular view of the same rows prompt_sensitivity_dataset.parquet: columnar copy when parquet support is available prompt_variant_spec.json: machine-readable description of the promptโ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/anonymous-acl26/prompt-sensitivity-codegen.tabulartext-generation100K<n<1M0 likes7 downloads4mo agoHugging Face23malliai /python-java-codegen-datasettextn<1K0 likes7 downloads4mo agoHugging Face24Abhiverse01 /Final_codegen_1000_entriestextn<1K0 likes6 downloads2y agoHugging Face25mathewmouchamel /code_generation_v2 LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code ๐Ÿ  Home Page โ€ข ๐Ÿ’ป GitHub Repository โ€ข ๐Ÿ† Leaderboard โ€ข LiveCodeBench is a "live" updating benchmark for holistically evaluating code related capabilities of LLMs. Particularly, it evaluates LLMs across a range of capabilties including code generation, self-repair, test output prediction, and code execution. This is the code generation scenario of LiveCodeBench. It isโ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/mathewmouchamel/code_generation_v2.textn<1K0 likes5 downloads6mo agoHugging Face26Jrvista /codegentextn<1K0 likes4 downloads2y agoHugging Face27introvoyz041 /manim-codegentext1K<n<10K0 likes4 downloads9mo agoHugging Face28sandhyanv9 /python-java-codegen-datasettextn<1K0 likes4 downloads4mo agoHugging Face29leeywin /code_alpaca_codegene_20ktext10K<n<100K0 likes3 downloads3y agoHugging Face30Denis641 /CodeGenDatasettext100K<n<1M0 likes3 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.