Team Ai
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hschumann2 /TempleOS-Source-Codetext1K<n<10K0 likes832 downloads1y agoHugging Face02nisaefendioglu /synthetic-sensitive-data-in-source-code-n300 Synthetic Sensitive Data in Source Code (N=300) Synthetic dataset of 300 source-code / config snippets containing hardcoded secrets and PII.Every sample includes at least one sensitive finding (no clean negatives in the main file). Designed for local masking, secret detection, and OWASP LLM02 — Sensitive Information Disclosure workflows. Version 1.3.6: README Files table documents split extension (split_pattern_custom_n300) and clean CSV. v1.3.4: split files + Dataset Viewer… See the full description on the dataset page: https://huggingface.co/datasets/nisaefendioglu/synthetic-sensitive-data-in-source-code-n300.texttext-classificationn<1K1 likes172 downloads9d agoHugging Face03haowu89 /open_parallel_think_code_source open_parallel_think_code_source A large-scale code reasoning distillation dataset with 320,000 solution trajectories generated by 4 state-of-the-art thinking models across 10,000 unique coding problems. Source / raw pool. This is the per-trajectory dataset. The packed parallel-thinking datasets derived from it are haowu89/open_parallel_think_code_full (full reasoning + solution) and haowu89/open_parallel_think_code_cot (solution only). Each trajectory's metadata carries… See the full description on the dataset page: https://huggingface.co/datasets/haowu89/open_parallel_think_code_source.tabular100K<n<1M0 likes151 downloads4mo agoHugging Face04ananyarn /Algorithm_and_Python_Source_CodeAlgorithm_and_Python_Source_Code This dataset provides different algorithms and their corresponding source code in Python. credits: Source codes given here are taken from "iamtarun/python_code_instructions_18k_alpaca" dataset in Hugging Face. text10K<n<100K11 likes109 downloads3y agoHugging Face05shibing624 /source_code纯文本数据,内容:高质量编程源代码,包括Python,Java,CPP源代码texttext-generation10M<n<100M8 likes83 downloads4y agoHugging Face06codemaivanngu /dpca-b200-source dpca-b200-source Tooling that builds, qualifies, packages and publishes the portable B200 runtime for the DPCA reproduction of On-Policy Distillation (ivanniu/On-Policy-Distill). This repo is the harness, not the runtime. The runtime payload (venv, source tree, node helpers) and the model weights live in a separate dataset repo, because they are ~56 GB and the tooling is 61 KB. What is here File What it is dpca-b200-source-20261005.tar.gz The tooling… See the full description on the dataset page: https://huggingface.co/datasets/codemaivanngu/dpca-b200-source.textn<1K0 likes77 downloads2d agoHugging Face07Agnuxo /github-source-code-dataset Github Source Code Dataset Complete source code from Agnuxo projects. Part of the Agnuxo Ecosystem by Francisco Angulo de Lafuente. texttext-generation1K<n<10K0 likes58 downloads5mo agoHugging Face08param-bharat /sourcecode-detectiontext1M<n<10M0 likes53 downloads2y agoHugging Face09annawleo /python-algorithm-sourcecode Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description This dataset provides algorithms and corresponding Python source code which can be leveraged for any type of code conversion applications. Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/annawleo/python-algorithm-sourcecode.textn<1K0 likes52 downloads3y agoHugging Face10AmanPriyanshu /RLVR-Env-Retrieval-Source-code-search-net-javascript RLVR-Env-Retrieval-Source-code-search-net-javascript RLVR-ready retrieval environment derived from Nan-Do/code-search-net-javascript. Author: Aman Priyanshu What Is This A 100k-row retrieval QA dataset where each row contains a question, ground-truth chunks, and pre-mined distractor chunks (random + semantically similar). Designed for training and evaluating retrieval agents in an RLVR (Reinforcement Learning with Verifiable Rewards) setup — the agent searches through… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/RLVR-Env-Retrieval-Source-code-search-net-javascript.texttext-retrieval100K<n<1M0 likes30 downloads7mo agoHugging Face11AmanPriyanshu /RLVR-Env-Retrieval-Source-code-search-net-python RLVR-Env-Retrieval-Source-code-search-net-python RLVR-ready retrieval environment derived from Nan-Do/code-search-net-python. Author: Aman Priyanshu What Is This A 100k-row retrieval QA dataset where each row contains a question, ground-truth chunks, and pre-mined distractor chunks (random + semantically similar). Designed for training and evaluating retrieval agents in an RLVR (Reinforcement Learning with Verifiable Rewards) setup — the agent searches through distractors… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/RLVR-Env-Retrieval-Source-code-search-net-python.texttext-retrieval100K<n<1M0 likes26 downloads7mo agoHugging Face12InfoBayAI /Product-Source-Code-DatasetgatedDataset Description: This dataset is a large-scale collection of coding and data, designed to support the development of advanced AI systems for code generation, program understanding, software intelligence, debugging assistance, and next-generation developer AI applications. Additionally, this dataset can be integrated into pipelines for Supervised Fine-Tuning (SFT) and Reinforcement Learning with Human Feedback (RLHF) workflows, helping improve AI performance in code completion, automated… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Product-Source-Code-Dataset.texttext-generationn<1K0 likes17 downloads21d agoHugging Face13Mr-Vicky-01 /source-code-Review-vulntext10K<n<100K0 likes15 downloads1y agoHugging Face14Mr-Vicky-01 /vuln-with-source-codetext10K<n<100K1 likes14 downloads1y agoHugging Face15MegaBites-AI /Doom-source-codetext1 likes14 downloads6mo agoHugging Face16shidowake /llama-inst-filtered-1k-source-code Description Slightly modified and formatted version of a subset of the original dataset for my own purpose. Original Dataset togethercomputer/llama-instruct · Datasets at Hugging Face LICENSE LLAMA 2 COMMUNITY LICENSE Download Llama text1K<n<10K0 likes10 downloads3y agoHugging Face17LiteCoder /LiteCoder_SourceCode LiteCoder Experiment Reproducing package To run the pre-train objective use the following scripts: Reproduce LiteCoder with all objectives: Navigate the folder Pre-training containing the LiteCoder.py file Then, run Python LiteCoder.py --train-tt --train-cs --train-pd The pretrained model is released on hugging face, therefore it automatically loads. To run the ablation studies: Ablation 1: Python LiteCoder.py --train-tt Ablation 2: Python LiteCoder.py --train-tt… See the full description on the dataset page: https://huggingface.co/datasets/LiteCoder/LiteCoder_SourceCode.text0 likes9 downloads3y agoHugging Face18CodeT5SmallCAPS /SourceCodetext0 likes9 downloads3y agoHugging Face19christopherlancei /edusite-source-code ☀️ Sun Shine School — Narasapuram | Since 1982 A full-stack educational website for Sun Shine School, Narasapuram (West Godavari District, Andhra Pradesh), built with React (frontend) and Python/FastAPI (backend). Content migrated from www.sunshineschoolnsp.com. Pages 🏠 Home — About Sun Shine School, Education Excellence Award, academic programs overview 🌟 What We Offer — Smart Classes, Sports, Labs, Language Development, Hostel 🎨 Activities — Singing, Dancing… See the full description on the dataset page: https://huggingface.co/datasets/christopherlancei/edusite-source-code.textn<1K0 likes7 downloads6mo agoHugging Face20ParZiVal04 /Purr-Data_example_source_codesPurr-Data Patch Source Code Dataset: This dataset is designed for training language models to generate source code for Purr-Data patches. It focuses specifically on patches that output a particular message when a "bang" object is clicked. Dataset Creation: The dataset was created with the goal of evaluating the ability of large language models like Google's 2B GEMMA to be fine-tuned for Purr-Data source code generation. Dataset Characteristics: Content: Each data point consists of two… See the full description on the dataset page: https://huggingface.co/datasets/ParZiVal04/Purr-Data_example_source_codes.texttext-generationn<1K0 likes6 downloads3y agoHugging Face21lsok264257 /test-smells-with-source-codetabular1K<n<10K0 likes4 downloads1y agoHugging Face22Senthil-HIGS /FEEDBACK_BASED_SOURCE_CODE_GENERATIONtabular10K<n<100K0 likes4 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.