Team Ai
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Naholav /CodeGen-Diverse-5K CodeGen-Diverse-5K: Broad Coverage for Competitive Programming Part of the CodeGen suite | CodeGen-Deep-5K (sister dataset) Dataset Description CodeGen-Diverse-5K is a broad coverage dataset designed for training code generation models across a wide variety of competitive programming problems. This dataset prioritizes problem diversity over solution diversity, covering 5,000 unique problems with consistent, high-quality solutions. Key Statistics Total samples:… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/CodeGen-Diverse-5K.tabulartext-generation1K<n<10K0 likes131 downloads10mo agoHugging Face02AmareshHebbar /leetcode-codegen-cpp LeetCode Code-Gen Dataset — C++ 4025 rows. Given a problem statement, its input/output examples, and a required algorithm/technique, generate a correct C++ solution. Part of a 4-language collection built from the same source: see the sibling Python, Java, C++, and JavaScript datasets. Verification Not execution-verified. There is currently no compiler/runtime harness for this language in the build pipeline (only Python has one). Rows are extracted directly from… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/leetcode-codegen-cpp.texttext-generation1K<n<10K1 likes108 downloads3mo agoHugging Face03AmareshHebbar /leetcode-codegen-javascript LeetCode Code-Gen Dataset — JavaScript 631 rows. Given a problem statement, its input/output examples, and a required algorithm/technique, generate a correct JavaScript solution. Part of a 4-language collection built from the same source: see the sibling Python, Java, C++, and JavaScript datasets. Verification Not execution-verified. There is currently no compiler/runtime harness for this language in the build pipeline (only Python has one). Rows are extracted… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/leetcode-codegen-javascript.texttext-generationn<1K1 likes99 downloads3mo agoHugging Face04Naholav /llama3.2-java-codegen-90sft-10meta-claude-v1 LLaMA 3.2 Java Code Generation Dataset (90% SFT, 10% Meta Annotated with Claude) This dataset contains 100,000 examples for Java method generation based on natural language instructions. It is built from the CodeXGLUE text-to-code dataset and designed to support both pure supervised fine-tuning (SFT) and reflection-based meta-learning approaches using Claude 4 Sonnet as the critique model. 🚀 Trained Models Two models have been trained on this dataset: SFT Model:… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/llama3.2-java-codegen-90sft-10meta-claude-v1.texttext-generation100K<n<1M1 likes94 downloads1y agoHugging Face05AmareshHebbar /leetcode-codegen-java LeetCode Code-Gen Dataset — Java 4068 rows. Given a problem statement, its input/output examples, and a required algorithm/technique, generate a correct Java solution. Part of a 4-language collection built from the same source: see the sibling Python, Java, C++, and JavaScript datasets. Verification Not execution-verified. There is currently no compiler/runtime harness for this language in the build pipeline (only Python has one). Rows are extracted directly from… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/leetcode-codegen-java.texttext-generation1K<n<10K0 likes92 downloads3mo agoHugging Face06stindardlogic /code-generation-sft-100k Code Generation SFT (100K) 100,000 ShareGPT conversations covering code generation across 8 programming languages, 21 categories, and 22 distinct programming tasks. Each example includes a detailed natural language request and a complete, working implementation with explanations of key design decisions. Motivation Coding assistants are the highest-adoption LLM application category, but most open training datasets focus on isolated functions without context. This… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/code-generation-sft-100k.texttext-generation100K<n<1M1 likes90 downloads3mo agoHugging Face07Groq /LiveCodeBench-CodeGenerationtextquestion-answeringn<1K1 likes75 downloads1y agoHugging Face08AmareshHebbar /leetcode-codegen-python LeetCode Code-Gen Dataset — Python 2522 rows. Given a problem statement, its input/output examples, and a required algorithm/technique, generate a correct Python solution. Part of a 4-language collection built from the same source: see the sibling Python, Java, C++, and JavaScript datasets. Verification Every row was extracted, then executed in a sandboxed subprocess against the problem's own stated examples. Only rows that passed all examples are included -- this… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/leetcode-codegen-python.texttext-generation1K<n<10K1 likes63 downloads3mo agoHugging Face09Naholav /CodeGen-Deep-5K CodeGen-Deep-5K: Deep Reasoning for Competitive Programming Part of the CodeGen suite | CodeGen-Diverse-5K (sister dataset) Dataset Description CodeGen-Deep-5K is a deep reasoning dataset designed for training code generation models with enhanced problem-solving capabilities. Unlike traditional datasets, this generates multiple distinct solutions for each problem, providing varied reasoning traces and approaches. Key Statistics Total samples: 5,000 Unique… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/CodeGen-Deep-5K.tabulartext-generation1K<n<10K0 likes53 downloads10mo agoHugging Face10openSUSE /cve-backport-codegen-dataset CVE Backport Code Generation Dataset Per-hunk code generation dataset for CVE security patch backporting, derived from openSUSE Build Service maintenance patches. Task Given a region of vulnerable source code and a description of the upstream CVE fix, the model outputs the fixed version of the code. A programmatic diff then produces the final patch. This plays to LLM strengths in code completion and avoids format-sensitivity issues with direct diff generation.… See the full description on the dataset page: https://huggingface.co/datasets/openSUSE/cve-backport-codegen-dataset.texttext-generation10K<n<100K1 likes26 downloads6mo agoHugging Face11wraps /codegen-flutter-v1texttext-generation100K<n<1M1 likes22 downloads2y agoHugging Face12Leon-Leee /codegen_kodcode_lc2k_taco_merged Dataset Card for Dataset Name Merged likaixin/TACO-verified, Leon-Leee/LeetCodeDataset_rectified, and kodCode/KodCode-Light-RL-10K Dataset Details Dataset Description Curated by: Leon (Me) Funded by [optional]: AIGCode/Koting Intelligence Language(s) (NLP): English License: MIT (following GURU-92K) Dataset Sources [optional] Repository: stay tuned Paper [optional]: stay tuned Uses Direct Use Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Leon-Leee/codegen_kodcode_lc2k_taco_merged.texttext-generation10K<n<100K0 likes22 downloads1y agoHugging Face13louisbrulenaudet /code-general-fonction-publique Code général de la fonction publique, non-instruct (11-12-2023) This project focuses on fine-tuning pre-trained language models to create efficient and accurate models for legal practice. Fine-tuning is the process of adapting a pre-trained model to perform specific tasks or cater to particular domains. It involves adjusting the model's parameters through a further round of training on task-specific or domain-specific data. While conventional fine-tuning strategies involve… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-general-fonction-publique.texttext-generation1K<n<10K0 likes21 downloads3y agoHugging Face14RY7Games /ry_manim_codegen RY Manim Code Generation This dataset contains queries and responses for ideal Manim code. The input and output have been preprocessed to keep context clean. Format: Jsonl format used instead of Json for better efficiency. {"query": "<a user's question about a topic>", "output": "<LLM's manim code>"} Citations: This repository includes data sourced and processed from the following datasets:… See the full description on the dataset page: https://huggingface.co/datasets/RY7Games/ry_manim_codegen.texttext-generation1K<n<10K1 likes11 downloads2mo agoHugging Face15macroteck /code_generationtextquestion-answeringn<1K0 likes9 downloads2y agoHugging Face16OctoReasoner /CodeGeneration-IQuestgated CodeGeneration-IQuest Execution-based Python code-generation prompts for reinforcement-learning post-training, in the verl rule-reward schema. Each row is a single-turn competitive-programming problem whose reward is computed by executing the model's program against a hidden test suite — a program passes only if every case matches. The collection unifies two execution-scorable sources (Code-Contests-O and DeepCoder) and then difficulty-filters them ("goldilocks", see below) so… See the full description on the dataset page: https://huggingface.co/datasets/OctoReasoner/CodeGeneration-IQuest.texttext-generation10K<n<100K0 likes9 downloads2mo agoHugging Face17daksh76 /prompt-sensitivity-codegen Prompt Sensitivity in Few-Shot Code Generation Dataset This dataset contains the full generated-code outputs and pass/fail outcomes used in our prompt sensitivity study across model families, benchmarks, perturbation axes, and k-shot settings. Dataset summary Rows: 240000 Models: claude-sonnet-4, gemini-2.5-flash, gpt-4o, llama-3.3-70b, qwen2.5-coder-3b Benchmarks: humaneval, mbpp Axes: order, phrasing, style k-shot values: 0, 1, 2, 3 Hugging Face repo:… See the full description on the dataset page: https://huggingface.co/datasets/daksh76/prompt-sensitivity-codegen.tabulartext-generation100K<n<1M0 likes8 downloads5mo agoHugging Face18anonymous-acl26 /prompt-sensitivity-codegen Anonymous Prompt Sensitivity Dataset This package contains model generations and evaluation outcomes for an anonymized submission on prompt sensitivity in few-shot code generation. What is included prompt_sensitivity_dataset.jsonl: one row per generated sample prompt_sensitivity_dataset.csv: tabular view of the same rows prompt_sensitivity_dataset.parquet: columnar copy when parquet support is available prompt_variant_spec.json: machine-readable description of the prompt… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-acl26/prompt-sensitivity-codegen.tabulartext-generation100K<n<1M0 likes7 downloads5mo agoHugging Face19ztony0712 /code_generation Visualization of Code Generation Task Cases Samples Check dataset samples visualization by viewing Dataset Viewer. The sampling procedure is guided by the Elo distribution introduced in our method. Original dataset is release_v5 of livecodebench/code_generation_lite from hugging face. samples/origin: 879/880 License This repository is licensed under the Apache License 2.0 tabulartext-generationn<1K0 likes6 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.