Team Ai
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ai-sec-lab /PatchBench PatchBench PatchBench is a benchmark for evaluating AI agents on realistic vulnerability patching tasks: 213 tasks drawn from 32 popular GitHub C/C++ projects. It selects vulnerabilities whose ground-truth fixes lie outside the crash stack, and uses vulnerability transplant plus code mutation to mitigate surface-level fixes and patch memorization. This repository holds the task metadata, one row per task to identify the project, the exact repository state, the crash, and the… See the full description on the dataset page: https://huggingface.co/datasets/ai-sec-lab/PatchBench.texttext-generationn<1K0 likes182 downloads20d agoHugging Face02michoo42 /Patchnoisseur Patchnoisseur A connoisseur's cellar of CVEs: every NVD CVE joined to its fixing-commit diff (when one could be found), its NVD description, and its associated CWE(s) (id, name, short description) — served as a single Parquet dataset. 351 884 CVEs · 25 015 with a real git diff attached · CVE-1999 → CVE-2026 · ~744 MB on disk (zstd-compressed Parquet, sharded ~300 MB each). What's in it One row per CVE in the NVD feed. CVEs without a retrievable patch are… See the full description on the dataset page: https://huggingface.co/datasets/michoo42/Patchnoisseur.tabulartext-classification100K<n<1M0 likes146 downloads5mo agoHugging Face03adastracomputing /nixpkgs-security-patches nixpkgs-security-patches Training dataset for fine-tuning LLMs on nixpkgs security patch generation. Derived from real merged security PRs in NixOS/nixpkgs. Dataset Details 588 training examples / 66 eval examples (654 total) Format: Multi-turn tool-calling conversations in ChatML JSONL Each example is a realistic agent session: the model reads the package file, finds the upstream fix, computes hashes via tools, and submits the fix for approval Hashes and URLs… See the full description on the dataset page: https://huggingface.co/datasets/adastracomputing/nixpkgs-security-patches.texttext-generationn<1K1 likes114 downloads7mo agoHugging Face04Transluce /act_patch_llama_3.1_8b_counterfact Training Language Models to Explain Their Own Computations Paper | Code This dataset contains activation patching results used for training explainer models to predict how internal interventions affect target model outputs. It was introduced in the paper "Training Language Models to Explain Their Own Computations". Dataset Summary The dataset covers the Activation Patching task for the Llama-3.1-8B target model, where explainer models learn to predict the effects of… See the full description on the dataset page: https://huggingface.co/datasets/Transluce/act_patch_llama_3.1_8b_counterfact.texttext-generation100K<n<1M0 likes107 downloads9mo agoHugging Face05ccerati /dx7-patches-and-prompts Yamaha DX7 Synthesizer Patches with AI-Generated Prompts Dataset Description This is a comprehensive, multi-task dataset designed for fine-tuning language models to understand and generate synthesizer patches for the Yamaha DX7. The dataset contains over 20,000 examples across three distinct but related tasks, making it ideal for creating models that can not only generate patches but also understand and reason about their structure and validity. How the Data Was… See the full description on the dataset page: https://huggingface.co/datasets/ccerati/dx7-patches-and-prompts.texttext-generation10K<n<100K0 likes89 downloads1y agoHugging Face06sunlab /patch_dbgated PatchDB: A Large-Scale Security Patch Dataset Description To foster large-scale research on vulnerability mitigation and to enable a comparison of different detection approaches, we make our dataset PatchDB from our DSN'21 paper publicly available. PatchDB is a large-scale security patch dataset that contains around 12,073 security patches and 23,742 non-security patches from the real world. You can find more details on the dataset in the paper "PatchDB: A Large-Scale… See the full description on the dataset page: https://huggingface.co/datasets/sunlab/patch_db.textfeature-extraction10K<n<100K25 likes70 downloads2y agoHugging Face07Tejaswiniprabhakaran19 /patchpilot-patchgen PatchPilot patch-generation dataset Supervised fine-tuning chats for PatchPilot's patch generator (A2). Each chat is exactly the prompt PatchPilot's agent sends to its model, followed by the developers' real fix written in the agent's SEARCH/REPLACE edit format. Source Built by scripts/build_patchgen_data.py (seed 42) from the SWE-bench training split (princeton-nlp/SWE-bench, train) and the gold files' contents from princeton-nlp/SWE-bench_oracle. The same seeded… See the full description on the dataset page: https://huggingface.co/datasets/Tejaswiniprabhakaran19/patchpilot-patchgen.texttext-generation1K<n<10K0 likes55 downloads6d agoHugging Face08Transluce /act_patch_qwen3_8b_counterfact Training Language Models to Explain Their Own Computations This dataset is part of the work presented in the paper "Training Language Models To Explain Their Own Computations". Code | Paper Introduction Language models (LMs) can learn to faithfully describe their internal computations. This dataset contains results from Activation Patching interventions on the Qwen3-8B target model using CounterFact data. Explainer models are trained on this data to predict how… See the full description on the dataset page: https://huggingface.co/datasets/Transluce/act_patch_qwen3_8b_counterfact.texttext-generation100K<n<1M0 likes41 downloads9mo agoHugging Face09SunBK201 /PatchBPtexttext-generation1K<n<10K0 likes8 downloads1y agoHugging Face10LIF1014 /ptdbench-llama-dapo-implementation-task-monkey-patch-011-dataset PTDBench dataset snapshot: task_monkey_patch_011 This repository stores the immutable runtime dataset snapshot for one materialized PTDBench task. It intentionally excludes model weights and training checkpoints. PTDBench family: llama_dapo_implementation Source evaluation metric: val-core/math_dapo/acc/mean@1 Provenance: Processed from BytedTsinghua-SIA/DAPO-Math-17k; task-specific bytes are pinned. License: Apache-2.0 The artifact manifest records every hydrated runtime path… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-llama-dapo-implementation-task-monkey-patch-011-dataset.texttext-generationn<1K0 likes7 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.