Team Ai
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Omarrran /StackPulse_778K_QnA_Code_dataset 💻 StackOverflow-778K: Multi-Year Developer Q&A Dataset Dataset Summary A large-scale Stack Overflow question dataset containing 778,929 unique questions sampled across 7 years (2015–2022). Each question includes the raw HTML body, plain-text version, tags, score, view count, answer count, and a rich set of derived features for immediate ML use. Collected across 8 sampling runs on Feb 27 2026, deduplicated to 778,929 unique questions with only 2 duplicates removed.… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/StackPulse_778K_QnA_Code_dataset.tabulartext-classification1M<n<10M0 likes187 downloads6mo agoHugging Face02RnniaSnow /st-code-datasetst_dataset_distillation_by_qwen2.5.jsonl 是由 stgenerate 蒸馏qwen2.514b生成的 /data/complier_improve 中的是在蒸馏的过程中增加编译器在环` st_dataset_local.jsonl 是编译验证通过 st_dpo_dataset.jsonl 是交给ai改正后编译验证通过形成DPO 其他的都是各种原因为通过编译 st_dataset_distillation_by_st_coder_clean 这个适合做反面教材 texttext-generation100K<n<1M0 likes178 downloads8mo agoHugging Face03Techta /backend-code-generator-dataset Backend Code Generation Dataset Dataset Description This dataset contains examples for training AI models to generate backend application code. It includes descriptions of backend requirements paired with complete, functional code implementations across multiple frameworks and programming languages. Dataset Summary The Backend Code Generation Dataset is designed to train models that can generate complete backend applications from natural language descriptions.… See the full description on the dataset page: https://huggingface.co/datasets/Techta/backend-code-generator-dataset.texttext-generationn<1K1 likes126 downloads1y agoHugging Face04Veri-Code /ReForm-Python2Dafny-Dataset Re:Form Datasets This repository contains the datasets associated with the paper "Re:Form -- Reducing Human Priors in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny". The paper introduces a framework that leverages Reinforcement Learning (RL) within Large Language Models (LLMs) to reduce reliance on human-annotated priors for formal software verification. The datasets provided are integral for training and evaluating models in formal language… See the full description on the dataset page: https://huggingface.co/datasets/Veri-Code/ReForm-Python2Dafny-Dataset.texttext-generation10K<n<100K2 likes105 downloads5mo agoHugging Face05suneeldk /arduino-code-dataset Arduino Programming Dataset Overview The Arduino Programming Dataset is a structured dataset designed for training and evaluating AI models that generate Arduino sketches for microcontroller projects. The dataset contains natural language instructions paired with complete Arduino programs. Each entry describes a hardware task and provides a working Arduino implementation along with additional metadata such as libraries used, wiring information, and difficulty level. This… See the full description on the dataset page: https://huggingface.co/datasets/suneeldk/arduino-code-dataset.texttext-generationn<1K1 likes78 downloads7mo agoHugging Face06ysn-rfd /FibonacciAi-CODE-Fine_Tuning-DatasetDocument Version: 1.0.2 | Last Updated: 07/17/2026 texttext-generation1K<n<10K1 likes63 downloads3mo agoHugging Face07Agnuxo /github-source-code-dataset Github Source Code Dataset Complete source code from Agnuxo projects. Part of the Agnuxo Ecosystem by Francisco Angulo de Lafuente. texttext-generation1K<n<10K0 likes54 downloads5mo agoHugging Face080-deng-0 /st-code-datasetst_dataset_distillation_by_qwen2.5.jsonl 是由 stgenerate 蒸馏qwen2.514b生成的 /data/complier_improve 中的是在蒸馏的过程中增加编译器在环` st_dataset_local.jsonl 是编译验证通过 st_dpo_dataset.jsonl 是交给ai改正后编译验证通过形成DPO 其他的都是各种原因为通过编译 st_dataset_distillation_by_st_coder_clean 这个适合做反面教材 texttext-generation100K<n<1M0 likes52 downloads2mo agoHugging Face09InfoBayAI /Legacy-Code-Datasetgated Legacy Codebase Dataset Dataset Description The Legacy Codebase Dataset is a large-scale collection of enterprise software repositories designed for training next-generation Large Language Models (LLMs), AI coding assistants, software engineering copilots, automated refactoring systems, repository understanding models, and intelligent program analysis pipelines. The complete collection contains 405 real-world legacy codebases spanning 23 major industries… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Legacy-Code-Dataset.texttext-generation10K<n<100K0 likes33 downloads2d agoHugging Face10InfoBayAI /Product-Source-Code-DatasetgatedDataset Description: This dataset is a large-scale collection of coding and data, designed to support the development of advanced AI systems for code generation, program understanding, software intelligence, debugging assistance, and next-generation developer AI applications. Additionally, this dataset can be integrated into pipelines for Supervised Fine-Tuning (SFT) and Reinforcement Learning with Human Feedback (RLHF) workflows, helping improve AI performance in code completion, automated… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Product-Source-Code-Dataset.texttext-generationn<1K0 likes25 downloads2d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.