Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01laion /openswe-tasks-patched-v5text10K<n<100K0 likes1.6k downloads5mo agoHugging Face02patched-codes /static-analysis-evalA dataset of 76 Python programs taken from real Python open source projects (top 100 on GitHub), where each program is a file that has exactly 1 vulnerability as detected by a particular static analyzer (Semgrep), used in the paper Patched MOA: optimizing inference for diverse software development tasks. OpenAI used the synth-vuln-fixes and fine-tuned a new version of gpt-4o is now the SOTA on this benchmark. More details and code is available from their repo. More details on the benchmark… See the full description on the dataset page: https://huggingface.co/datasets/patched-codes/static-analysis-eval.textn<1K20 likes885 downloads1y agoHugging Face03CIRCL /vulnerability-cwe-patch Description This dataset, CIRCL/vulnerability-cwe-patch, provides structured, real-world vulnerabilities enriched with CWE identifiers and corresponding patches from platforms like GitHub and GitLab. It is designed to support the development of tools for vulnerability classification, triage, and automated remediation. Each entry includes metadata such as CVE/GHSA ID, a description, CWE categorization, and links to verified patch commits with associated diff content and commit… See the full description on the dataset page: https://huggingface.co/datasets/CIRCL/vulnerability-cwe-patch.text1K<n<10K5 likes542 downloads3mo agoHugging Face04anaumghori /patchlet-embed-preprocessedimage100K<n<1M0 likes540 downloads8mo agoHugging Face05patched-codes /generate-readme-eval Generate README Eval The generate-readme-eval is a dataset (train split) and benchmark (test split) to evaluate the effectiveness of LLMs when summarizing entire GitHub repos in form of a README.md file. The datset is curated from top 400 real Python repositories from GitHub with at least 1000 stars and 100 forks. The script used to generate the dataset can be found here. For the dataset we restrict ourselves to GH repositories that are less than 100k tokens in size to allow us to… See the full description on the dataset page: https://huggingface.co/datasets/patched-codes/generate-readme-eval.textsummarizationn<1K3 likes498 downloads2y agoHugging Face06closji /cc12m_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-13image10M<n<100M0 likes496 downloads4y agoHugging Face07zhcharyzhang /patchaudit-artifact PatchAudit Artifact PatchAudit audits security patches. You give it a CVE's initial fix — commit C1 — and a later commit Ci, and it tells you whether Ci is a future commit: a later commit that had to keep fixing the same problem because C1 was incomplete (it left the vulnerability reachable) or incorrect (its own change introduced a new defect). When such a future commit exists, C1 was a bad patch. When even the latest fix still leaves the hole open, the bug is a lingering… See the full description on the dataset page: https://huggingface.co/datasets/zhcharyzhang/patchaudit-artifact.text1K<n<10K0 likes449 downloads12d agoHugging Face08livebench /liveswebench-patchestextn<1K1 likes362 downloads2y agoHugging Face09DefendIntelligence /vessel-detection-labeled-patches Vessel Detection Labeled Patches Validated/confirmed satellite image patches exported from the military-boat-detection review workflow. Contents images/: patch images. metadata.csv: one row per patch, compatible with Hugging Face image-folder metadata. metadata.jsonl: rich patch metadata with nested objects. annotations.csv: one row per vessel annotation. annotations.jsonl: JSONL version of the object annotations. labels/: YOLO-format labels. Hard negatives have empty… See the full description on the dataset page: https://huggingface.co/datasets/DefendIntelligence/vessel-detection-labeled-patches.imageobject-detection1K<n<10K5 likes292 downloads5mo agoHugging Face10ByteDance /PatchEval 👋 Overview PatchEval-Verified is a benchmark for evaluating LLMs and coding agents on automated repair of real-world vulnerabilities. It contains 230 CVE cases with Docker-based evaluation environments, covering vulnerabilities reported between 2015 and 2025 across Go, JavaScript, and Python. PatchEval-Verified updates the evaluation environments from the original PatchEval release. In the original benchmark, some PoC tests were adapted from project regression tests and were… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance/PatchEval.textn<1K8 likes255 downloads2mo agoHugging Face11rasdani /github-patchestext10K<n<100K0 likes253 downloads1y agoHugging Face12Martingkc /LLaVa-CC3M-PostTraining-clip-vit-base-patch16text100K<n<1M0 likes235 downloads6mo agoHugging Face13prakanda /SynthMat_Patches_DStext100K<n<1M0 likes221 downloads2y agoHugging Face14Kushalkhemka /cybersec-chatml-vuln-patch-v1 Cybersecurity ChatML SFT Dataset (Detection + Patch + Multitask) This dataset contains ChatML records for 2 security tasks: Vulnerability detection (is_vulnerable, cwe, severity JSON output) Secure patch generation (assistant returns patched code only) Files chatml_detection_train.jsonl chatml_detection_val.jsonl chatml_patch_train.jsonl chatml_patch_val.jsonl chatml_multitask_train.jsonl chatml_multitask_val.jsonl chatml_build_manifest.json… See the full description on the dataset page: https://huggingface.co/datasets/Kushalkhemka/cybersec-chatml-vuln-patch-v1.text100K<n<1M1 likes220 downloads6mo agoHugging Face15DCAgent /rl__24GPU_base__swe_rebench_patched_oracle__r2egym-nl2bash-stacktext10K<n<100K0 likes218 downloads7mo agoHugging Face16open-athena /a3-rl-DCAgent_r2egym-patched-full-oracletext10K<n<100K0 likes216 downloads4mo agoHugging Face17DCAgent /swe_rebench_patchedtext1K<n<10K0 likes203 downloads7mo agoHugging Face18closji /cc12m_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-15image10M<n<100M0 likes196 downloads4y agoHugging Face19Martingkc /LLaVa-Instruct-150K-clip-vit-base-patch32text100K<n<1M0 likes180 downloads6mo agoHugging Face20ai-sec-lab /PatchBench PatchBench PatchBench is a benchmark for evaluating AI agents on realistic vulnerability patching tasks: 213 tasks drawn from 32 popular GitHub C/C++ projects. It selects vulnerabilities whose ground-truth fixes lie outside the crash stack, and uses vulnerability transplant plus code mutation to mitigate surface-level fixes and patch memorization. This repository holds the task metadata, one row per task to identify the project, the exact repository state, the crash, and the… See the full description on the dataset page: https://huggingface.co/datasets/ai-sec-lab/PatchBench.texttext-generationn<1K0 likes173 downloads18d agoHugging Face21michoo42 /Patchnoisseur Patchnoisseur A connoisseur's cellar of CVEs: every NVD CVE joined to its fixing-commit diff (when one could be found), its NVD description, and its associated CWE(s) (id, name, short description) — served as a single Parquet dataset. 351 884 CVEs · 25 015 with a real git diff attached · CVE-1999 → CVE-2026 · ~744 MB on disk (zstd-compressed Parquet, sharded ~300 MB each). What's in it One row per CVE in the NVD feed. CVEs without a retrievable patch are… See the full description on the dataset page: https://huggingface.co/datasets/michoo42/Patchnoisseur.tabulartext-classification100K<n<1M0 likes164 downloads5mo agoHugging Face22R2E-Gym /R2E-TestgenAgent-Patchestextn<1K1 likes163 downloads1y agoHugging Face23open-athena /rl__24GPU_shaped__swe_rebench_patched_oracle__r2egym-nl2bash-stacktext10K<n<100K0 likes157 downloads7mo agoHugging Face24DCAgent /Kimi-2.5-swe_rebench_patched-maxeps-32ktext10K<n<100K0 likes145 downloads5mo agoHugging Face25rasdani /github-patches-genesysimport re import json from datasets import load_dataset PROMPT_TEMPLATE = """\ We are currently solving the following issue within our repository. Here is the issue text: --- BEGIN ISSUE --- {issue} --- END ISSUE --- Below are some code segments, each from a relevant file. One or more of these files may contain bugs. --- BEGIN FILES --- {file_context} --- END FILES --- Please first localize the bug based on the issue statement, and then generate a patch according to the `git diff` format… See the full description on the dataset page: https://huggingface.co/datasets/rasdani/github-patches-genesys.text10K<n<100K0 likes144 downloads1y agoHugging Face26mlfoundations-dev /swe_gym_annotate_with_patchtext10K<n<100K0 likes136 downloads2y agoHugging Face27letitiaaa /patch_region_128image100K<n<1M0 likes133 downloads9mo agoHugging Face28TrevorJS /mtg-scryfall-cropped-art-embeddings-siglip-so400m-patch14-384image10K<n<100K0 likes127 downloads2y agoHugging Face29closji /cc12m_openai-clip-vit-patch32image1M<n<10M1 likes123 downloads4y agoHugging Face30closji /mscoco_train_2014_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-15tabular10M<n<100M0 likes122 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.