Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SciCodePile /SciCode-Programming-Problems DATA3: Programming Problems Generation Dataset Dataset Overview DATA3 is a large-scale programming problems generation dataset that contains AI-generated programming problems inspired by real scientific computing code snippets. The dataset consists of 22,532 programming problems, each paired with a comprehensive solution. These problems focus on scientific computing concepts such as numerical algorithms, data analysis, mathematical modeling, and computational methods in… See the full description on the dataset page: https://huggingface.co/datasets/SciCodePile/SciCode-Programming-Problems.texttext-generation10K<n<100K0 likes528 downloads7mo agoHugging Face02Jackrong /Competitive-Programming-python-blend Dataset Card for Competitive-Programming-python-blend Summary Competitive-Programming-python-blend is a mixed supervised fine-tuning dataset centered on competitive programming, code reasoning, and instruction-style problem solving. The blend is Python-first, but it also keeps a small amount of C++, agentless SWE, and reasoning-oriented chat supervision to broaden training coverage. The current release is published as a single HF-friendly JSONL file, clean.jsonl.… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Competitive-Programming-python-blend.texttext-generation10K<n<100K21 likes461 downloads7mo agoHugging Face03multimodal-reasoning-lab /Competitive-Programmingimage1K<n<10K1 likes415 downloads1y agoHugging Face04Vikhrmodels /programming_books_engtext1M<n<10M1 likes397 downloads2y agoHugging Face05BoooomNing /spatial-programmingColumns: input_text — text prompt (may be null) input_image — reference image (may be null); at least one of text / image is set code — Python source that builds the asset glb — the resulting GLB file, raw bytes platform — e.g. blender type — e.g. modeling model — which model wrote the code (fable, fable 5.1, opus 5, astra, sol) Rows are appended one parquet shard per upload under data/. imagen<1K0 likes333 downloads15d agoHugging Face06vikp /textbook_quality_programming Dataset Card for "textbook_quality_programming" Synthetic programming textbooks generated with GPT-3.5 and retrieval. Very high quality, aimed at being used in a phi replication. Currently 115M tokens. Covers many languages and technologies, with a bias towards python. ~10k of the books (65M tokens) use an older generation method, and average 6k tokens in length. ~1.5k books (50M tokens) use a newer generation method, with a more detailed outline, and average 33k tokens in… See the full description on the dataset page: https://huggingface.co/datasets/vikp/textbook_quality_programming.text10K<n<100K182 likes293 downloads3y agoHugging Face07BAAI /IndustryCorpus2_computer_programming_code IndustryCorpus2: Programming This repository contains the IndustryCorpus2: Programming domain subset of BAAI/IndustryCorpus2. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryCorpus2: @misc{shi2024industrycorpus2, title = {IndustryCorpus2}, author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao}, year = {2024}… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_computer_programming_code.tabular1M<n<10M2 likes232 downloads2mo agoHugging Face08JasonYan777 /PersonaSignal-PersonalizedResponse-Programming-Expertise-gpt-5-mini Dataset card for PersonaSignal-PersonalizedResponse-Programming-Expertise-gpt-5-mini This dataset was made with Curator. Dataset details A sample from the dataset: { "dimension_name": "programming_expertise", "dimension_values": [ "Novice", "Intermediate", "Advanced" ], "dimension_description": "Represents the user's practical fluency in software engineering. It shapes how they decompose problems, choose abstractions, weigh… See the full description on the dataset page: https://huggingface.co/datasets/JasonYan777/PersonaSignal-PersonalizedResponse-Programming-Expertise-gpt-5-mini.textn<1K0 likes207 downloads11mo agoHugging Face09JasonYan777 /PersonaSignal-PerceivabilityTest-Programming-Expertise-gpt-5-mini Dataset card for PersonaSignal-PerceivabilityTest-Programming-Expertise-gpt-5-mini This dataset was made with Curator. Dataset details A sample from the dataset: { "dimension_name": "programming_expertise", "dimension_values": [ "Novice", "Intermediate", "Advanced" ], "dimension_description": "Represents the user's practical fluency in software engineering. It shapes how they decompose problems, choose abstractions, weigh… See the full description on the dataset page: https://huggingface.co/datasets/JasonYan777/PersonaSignal-PerceivabilityTest-Programming-Expertise-gpt-5-mini.tabularn<1K0 likes207 downloads11mo agoHugging Face10open-phi /programming_books_llama Dataset Card for "programming_books_llama" 400M tokens of programming books generated by gpt-3.5 (70M tokens) and a finetuned codellama 34b. The gpt-3.5 data is extremely high quality. The llama data has lower quality and shorter length, but is still good. This was generated with the textbook quality repo. text100K<n<1M36 likes168 downloads3y agoHugging Face11code-rag-bench /programming-solutionsThe programming solutions retrieval source for code-rag-bench, comprising programming solutions for the HumanEval and MBPP datasets. text1K<n<10K2 likes138 downloads2y agoHugging Face12dominexmacedon /Puma-Programming-Language-Dataset license: mit Puma Programming Language Dataset The Puma Programming Language Dataset is a curated collection of Puma programming examples designed for developers, learners, educators, researchers, and AI systems working with the Puma programming language. The dataset contains practical Puma code examples covering language syntax, programming patterns, data structures, functions, iteration, backend development, HTTP services, APIs, WebSocket communication… See the full description on the dataset page: https://huggingface.co/datasets/dominexmacedon/Puma-Programming-Language-Dataset.text1K<n<10K0 likes138 downloads28d agoHugging Face13WillHeld /paloma_programming_languagestext10K<n<100K0 likes118 downloads1y agoHugging Face14dougiefresh /systems_programming_and_administrationtext10K<n<100K1 likes101 downloads1y agoHugging Face15Programming-Language /codeagent-pythontext100K<n<1M11 likes97 downloads3y agoHugging Face16Neura-parse /quantum-compilation-and-programming Neura Parse — Quantum Compilation & Programming A code-heavy vertical on the quantum software/compilation stack: turning abstract quantum circuits and unitaries into device-executable programs. Covers unitary decomposition and circuit synthesis (Euler/ZYZ, KAK/Cartan, Solovay-Kitaev, Ross-Selinger gridsynth, numerical synthesis with BQSKit), gate-set/basis transpilation to native gate sets, qubit layout/mapping and routing under connectivity constraints (SABRE, VF2, SWAP… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-compilation-and-programming.tabulartext-generation100K<n<1M0 likes92 downloads3mo agoHugging Face17ianncity /Hunter-Alpha-Programming-160000x Hunter-Alpha-Programming-160000x - just a filtered version of the original dataset with like 50k more programming questions, DO NOT FINETUNE ON BOTH ONLY USE ONE 160,000 programming reasoning traces distilled from Hunter Alpha on OpenRouter at high and xhigh reasoning Distribution: Includes: Webdev, C++, Java, JS, C, Ruby, Lua, Rust, and C# Token Count as of 3/17/2026: a lot idk probably 1 billion [!NOTE] This will be the last dataset update for hunter alpha I beleive… See the full description on the dataset page: https://huggingface.co/datasets/ianncity/Hunter-Alpha-Programming-160000x.text100K<n<1M20 likes90 downloads7mo agoHugging Face18jamesdborin /Nemotron-SFT-Competitive-Programming-v2-prompt-only Nemotron-SFT-Competitive-Programming-v2-prompt-only Prompt-only extraction from nvidia/Nemotron-SFT-Competitive-Programming-v2. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Competitive-Programming-v2-prompt-only.tabular100K<n<1M0 likes82 downloads3mo agoHugging Face19databounty-io /build-competitive-programming-problem-dataset-cmsokff2 Build Competitive Programming Problem Dataset Each item is a build competitive programming problem dataset example providing Problem statement, Constraints, Reference solution, Test cases, Time limit (ms). Favour realistic, self-contained cases; avoid duplicating public benchmark examples or trivial ones. About This dataset was produced by the DataBounty community and published here as part of an open, karma-only program. Accepted items: 1000 Language: Python… See the full description on the dataset page: https://huggingface.co/datasets/databounty-io/build-competitive-programming-problem-dataset-cmsokff2.text1K<n<10K0 likes81 downloads28d agoHugging Face20TaskPuppyAI /lunamax-multitask-programming-1000 LunaMax Multitask Programming 1000 A 1,000-record synthetic multitask programming dataset generated with ChatGPT LunaMax. The recovered dataset combines code review, implementation, bug and severity classification, and strict output-contract tasks across multiple programming languages. The historical source shards were reviewed with ChatGPT 5.6 Sol High according to dataset creator confirmation. During Hugging Face publication preparation, all 1,000 records received a new… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/lunamax-multitask-programming-1000.text1K<n<10K0 likes69 downloads1mo agoHugging Face21jamesdborin /Nemotron-Competitive-Programming-v1-prompt-only Nemotron-Competitive-Programming-v1-prompt-only Prompt-only extraction from nvidia/Nemotron-Competitive-Programming-v1. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row indexes where… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-Competitive-Programming-v1-prompt-only.tabular1M<n<10M0 likes67 downloads3mo agoHugging Face22NickyNicky /Code-290k-labels-programming_languages-NO_Chatgpt Para etiquetar los lenguajes de programación en un conjunto de datos extenso de fragmentos de código, se aplicaron técnicas automatizadas de procesamiento de texto y patrones específicos de cada lenguaje, sin recurrir al uso de modelos de lenguaje avanzados como ChatGPT o LLMs. Se inició con la extracción y preparación de datos usando pandas, una biblioteca de análisis de datos en Python, que facilitó la manipulación y el procesamiento del conjunto de datos obtenido de Hugging Face's… See the full description on the dataset page: https://huggingface.co/datasets/NickyNicky/Code-290k-labels-programming_languages-NO_Chatgpt.text100K<n<1M5 likes66 downloads3y agoHugging Face23theblackcat102 /multiround-programming-convo Multi-Round Programming Conversations Based on previous evol-codealpaca-v1 dataset with added sampled questions from stackoverflow, crossvalidated and make it multiround! It should be more suited to train a code assistant which works side by side. Tasks included in here: Data science, statistic, programming questions Code translation : translate a short function from Python, Golang, C++, Java, Javascript Code fixing : Fix randomly corrupts characters with no tab… See the full description on the dataset page: https://huggingface.co/datasets/theblackcat102/multiround-programming-convo.texttext-generation100K<n<1M9 likes61 downloads3y agoHugging Face24Tapos-Minmoy /python_programming_questionstext10K<n<100K2 likes61 downloads1y agoHugging Face25SciCode /SciCode-Programming-Problemsgated DATA3: Programming Problems Generation Dataset Dataset Overview DATA3 is a large-scale programming problems generation dataset that contains AI-generated programming problems inspired by real scientific computing code snippets. The dataset consists of 22,532 programming problems, each paired with a comprehensive solution. These problems focus on scientific computing concepts such as numerical algorithms, data analysis, mathematical modeling, and computational… See the full description on the dataset page: https://huggingface.co/datasets/SciCode/SciCode-Programming-Problems.texttext-generation10K<n<100K1 likes58 downloads8mo agoHugging Face26leo009 /python-programming-instructionstext100K<n<1M1 likes52 downloads2y agoHugging Face27abanm /arjoonn-codechef-competitive-programming-ChatGPT4otext1K<n<10K0 likes50 downloads2y agoHugging Face28dougiefresh /systems_programming_code_conversationstext1K<n<10K1 likes47 downloads1y agoHugging Face29StarsMakeGalaxy /competitive-programming-curated-600 🚀 Competitive Programming & Algorithmic Reasoning (Verbose CoT Reasoning) This dataset contains 600 curated training records with in-depth, verbose 4-phase <Thinking> Chain-of-Thought reasoning, 100 frozen evaluation benchmark samples, and 50 frozen regression verification samples formatted in standard ChatML (messages) and Prompt-Target pairs, strictly following the Pioneer / Prometheus research paper 3-slice curriculum design. 📊 Dataset Composition & 3-Slice… See the full description on the dataset page: https://huggingface.co/datasets/StarsMakeGalaxy/competitive-programming-curated-600.texttext-generationn<1K0 likes47 downloads2mo agoHugging Face30electricsheepafrica /africa-somalia-cash-based-programming-in-somalia-6c20447a Cash Based Programming in Somalia | Africa (OCHA Somalia) 7,141 rows - 1 Africa country/area - 2021 - source table - Engineered by Electric Sheep Africa TL;DR This dataset contains 7,141 rows from OCHA Somalia, covering Cash Based Programming in Somalia. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples. What This Dataset Measures Agriculture datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-somalia-cash-based-programming-in-somalia-6c20447a.tabulartabular-classification1K<n<10K0 likes46 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.