Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jugalgajjar /MultiLang-Code-Parser-Dataset MultiLang Code Parser Dataset (MLCPD) MultiLang-Code-Parser-Dataset (MLCPD) provides a large-scale, unified dataset of parsed source code across 10 major programming languages, represented under a universal schema that captures syntax, semantics, and structure in a consistent format. Each entry corresponds to one parsed source file and includes: Language metadata Code-level statistics (lines, errors, AST nodes) Universal Schema JSON (normalized structural representation) MLCPD… See the full description on the dataset page: https://huggingface.co/datasets/jugalgajjar/MultiLang-Code-Parser-Dataset.tabular1M<n<10M2 likes345 downloads1y agoHugging Face02hieunguyenminh /code_contests_dp_datasettabular1K<n<10K0 likes310 downloads2y agoHugging Face03Omarrran /StackPulse_778K_QnA_Code_dataset 💻 StackOverflow-778K: Multi-Year Developer Q&A Dataset Dataset Summary A large-scale Stack Overflow question dataset containing 778,929 unique questions sampled across 7 years (2015–2022). Each question includes the raw HTML body, plain-text version, tags, score, view count, answer count, and a rich set of derived features for immediate ML use. Collected across 8 sampling runs on Feb 27 2026, deduplicated to 778,929 unique questions with only 2 duplicates removed.… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/StackPulse_778K_QnA_Code_dataset.tabulartext-classification1M<n<10M0 likes182 downloads6mo agoHugging Face04p-doom /crowd-code-dataset-0.1The crowd-code-dataset-0.1 is a raw, unfiltered dataset of fine-grained IDE interactions collected during the development of Jasmine using crowd-code, a VS Code/Cursor extension capturing large parts of the software engineering workflow. The dataset captures real research engineering workflows (character-level edits, navigation, terminal use, iterative debugging). The crowd-code-dataset-0.1 only includes data from the Jasmine authors. We are actively working on cleaning and curating the full… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/crowd-code-dataset-0.1.tabular100K<n<1M5 likes140 downloads9mo agoHugging Face05peiran20030408 /team1_dataset_play_diatonic_codeThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 15, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/peiran20030408/team1_dataset_play_diatonic_code.tabularrobotics10K<n<100K0 likes87 downloads2mo agoHugging Face06atul10 /reverse_engineering_code_dataset_O2_x64_O2tabular100K<n<1M2 likes77 downloads1y agoHugging Face07atul10 /reverse_engineering_code_dataset_O0_arm_O0tabular100K<n<1M1 likes64 downloads1y agoHugging Face08atul10 /reverse_engineering_code_dataset_O3_x64_O3tabular100K<n<1M1 likes62 downloads1y agoHugging Face09atul10 /reverse_engineering_code_dataset_O2_arm_O2tabular10K<n<100K0 likes61 downloads1y agoHugging Face10atul10 /reverse_engineering_code_dataset_O3_arm_O3tabular10K<n<100K0 likes60 downloads1y agoHugging Face11atlas-institute /code-trainer-offsec-datasettabular10K<n<100K0 likes60 downloads6mo agoHugging Face12atul10 /reverse_engineering_code_dataset_O2_mips_O2tabular10K<n<100K0 likes57 downloads1y agoHugging Face13atul10 /reverse_engineering_code_dataset_O0_x64_O0tabular100K<n<1M0 likes57 downloads1y agoHugging Face14atul10 /reverse_engineering_code_dataset_O1_mips_O1tabular10K<n<100K0 likes53 downloads1y agoHugging Face15BambusControl /AI-Code-Optimization-for-Sustainability-Dataset AI Code Optimization for Sustainability: Dataset Refactoring Python Code for Energy-Efficiency using Qwen3: Dataset based on HumanEval, MBPP, and Mercury 📄 Read the Paper | Zenodo Mirror | DOI: 10.5281/zenodo.18377893 | About the author This dataset is a part of a Master thesis research internship investigating the use of LLMs to optimize Python code for energy efficiency. The research was conducted as part of the Greenify My Code (GMC) project at the Netherlands Organisation for… See the full description on the dataset page: https://huggingface.co/datasets/BambusControl/AI-Code-Optimization-for-Sustainability-Dataset.tabular10K<n<100K0 likes53 downloads8mo agoHugging Face16atul10 /reverse_engineering_code_dataset_O1_arm_O1tabular10K<n<100K0 likes52 downloads1y agoHugging Face17atul10 /reverse_engineering_code_dataset_O1_x64_O1tabular100K<n<1M0 likes52 downloads1y agoHugging Face18ammarnasr /Python-React-Code-Datasettabular1K<n<10K2 likes51 downloads3y agoHugging Face19ammarnasr /Python-Security-Code-Datasettabular1K<n<10K3 likes49 downloads3y agoHugging Face20atul10 /reverse_engineering_code_dataset_O0_mips_O0tabular100K<n<1M1 likes43 downloads1y agoHugging Face21atul10 /recreated_reverse_engineering_code_dataset_O1_x86_O1tabular100K<n<1M0 likes43 downloads1y agoHugging Face22bermaneh /pde-llm-eval-code-perturbation-dataset pde-llm-eval-code-perturbation-dataset Code Perturbation Dataset (final). 256 PDE solver implementations: 64 base programs (32 physically valid, 32 with an injected bug that still runs to completion) expanded by 4 lexical perturbation conditions each -- original comments, comments removed, comments swapped in from a different implementation, and descriptive identifiers replaced by meaningless placeholders. The perturbations change the lexical surface only; executable behaviour… See the full description on the dataset page: https://huggingface.co/datasets/bermaneh/pde-llm-eval-code-perturbation-dataset.tabularn<1K0 likes43 downloads1mo agoHugging Face23ahmetggg /Dr-Zeon-Github-Python-Code-Dataset Luck Spark 1B - High Quality Code Dataset The first quality-scored, star-agnostic code dataset for training 1B MoE code models. Unlike The Stack / CodeParrot that filter by stars, this dataset scores every file by its content (0-10). A 2-star well-documented library scores higher than a 10k-star minified file. Continuously updated by an autonomous bot. Repo: ahmetggg/luck-spark-1b-code-dataset | Bot: github_to_hf_bot.py | License: Permissive only (MIT / Apache-2.0 / BSD /… See the full description on the dataset page: https://huggingface.co/datasets/ahmetggg/Dr-Zeon-Github-Python-Code-Dataset.tabulartext-generation10K<n<100K1 likes43 downloads1mo agoHugging Face24avhi-code /titanic_datasettabularn<1K0 likes41 downloads14d agoHugging Face25atul10 /reverse_engineering_code_dataset_O1_x86_O1tabular100K<n<1M0 likes38 downloads1y agoHugging Face26atul10 /recreated_reverse_engineering_code_dataset_O3_x86_O3tabular100K<n<1M0 likes38 downloads1y agoHugging Face27atul10 /reverse_engineering_code_dataset_O3_x86_O3tabular100K<n<1M0 likes34 downloads1y agoHugging Face28ChamaraVishwajithRajapaksha /Code_Vulnerability_Dataset 🔐 Code Vulnerability Dataset (CWE-Enriched) 📌 Overview This dataset is built from the bstee615/diversevul dataset and enhanced with structured vulnerability intelligence from the MITRE Common Weakness Enumeration (CWE) database. It provides a rich, machine-readable representation of software vulnerabilities, mapping raw vulnerable code samples to standardized CWE classifications. The dataset is designed for research and development in: Vulnerability detection models… See the full description on the dataset page: https://huggingface.co/datasets/ChamaraVishwajithRajapaksha/Code_Vulnerability_Dataset.tabular100K<n<1M0 likes32 downloads6mo agoHugging Face29jwang2373 /dataset_qwen2.5_code_1.5b_grpo_iter0_full_data_miao_0212_2_global_step_70filteredtabular10K<n<100K0 likes28 downloads2y agoHugging Face30atul10 /reverse_engineering_code_dataset_O3_mips_O3tabular10K<n<100K0 likes28 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.