datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PatchCamelyon
PatchCamelyon (PCam)
Description
The PatchCamelyon benchmark is a new and challenging image classification dataset. It consists of 327.680 color images (96 x 96px) extracted from histopathologic scans of lymph node sections. Each image is annoted with a binary label indicating presence of metastatic tissue. PCam provides a new benchmark for machine learning models: bigger than CIFAR10, smaller than imagenet, trainable on a single GPU
Why PCam
Fundamental… See the full description on the dataset page: https://huggingface.co/datasets/1aurent/PatchCamelyon.ota-patchespatchrecoverygym-laguna
PatchRecoveryGym for Laguna
Submitted by: Kannappan Sirchabesan (@kannappans) · Poolside Research Hackathon (Foundations track)
A reproducible eval + RL environment that tests whether a coding agent can
recover from a wrong first attempt — a real, under-measured agentic-coding
weakness. Built for Poolside Laguna XS.2 on dependency-migration repair tasks.
📦 Installable Verifiers environment on the Prime Hub · 🎯 deterministic hidden-test reward · 🔁 144-candidate reranking… See the full description on the dataset page: https://huggingface.co/datasets/poolside-laguna-hackathon/patchrecoverygym-laguna.semantic_patch_cache
HeatTok Semantic Patch Cache
Precomputed .pt caches for HeatTok. Use with HEATTOK_SEMANTIC_CACHE_DIR=/path/to/cache.
Filename pattern: {image_hash}_g1_s28.pt or {image_hash}_g1_gd1_s28.pt
semantic_patch_cache_vrsbench
Dataset: VRSBench (512×512)
Caches do not store precomputed global tokens or patch orientations.
Gaussian parameters and patch metadata are stored.
No need to regenerate .pt files — HeatTok computes global tokens and orientations online at load… See the full description on the dataset page: https://huggingface.co/datasets/Yingying11/semantic_patch_cache.patch
Patch
Threads pulled from 2ch
To synchronize the threads:
python -m much sync -i assets/patch/index.tsv -p assets/patch/threads -t assets/orphan
python -m much sort -s assets/patch/index.tsv -d assets/patch/index.tsv -t assets/patch/threads
vlm_wsi_patch
VLM WSI Patch
Closed-answer WSI patch VQA dataset with 48,766 metadata-eligible rows and
48,609 valid 1344×1344 RGB top-4 mosaics, picked from of PathGen. Rows without an image retain an
explicit processing rejection status. Each patch uses a WSI level-0 top-left coordinate and
672×672 pixels; backfill and patch substitution are excluded.
The image field is a repository-relative PNG path. metadata/manifest.jsonl and
metadata/manifest.parquet contain QA labels, top-4 coordinates… See the full description on the dataset page: https://huggingface.co/datasets/HCOOH/vlm_wsi_patch.paper2-patches-period1-ARCHIVED-OLDCOVERAGE-20260710openswe-tasks-patched-v5imagenet1k-256x256-ztree-sdvae-patch2Dataset produced by https://github.com/theAdamColton/zero-tree-diffusion
patch size: 2, uses quantization, clip value 2.5, db3, level 4, imagenet images resized to 256x256, uses the stable diffusion vae
repo2rlenv-cve-patches
repo2rlenv-cve-patches
Generated by Repo2RLEnv — turning real GitHub repositories into verifiable RL environments.
💡 Browse this dataset in your browser — click the badge above or open
HuggingFaceH4/harbor-visualiser
to inspect every task's spec, instruction, oracle patch, test script, and Dockerfile.
Source repos (6):
Pylons/waitress
andialbrecht/sqlparse
lepture/mistune
pallets/flask
pallets/werkzeug
psf/requests
Pipeline: cve_patches
Tasks: 19
Visibility: public
Spec:… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/repo2rlenv-cve-patches.static-analysis-evalA dataset of 76 Python programs taken from real Python open source projects (top 100 on GitHub),
where each program is a file that has exactly 1 vulnerability as detected by a particular static analyzer (Semgrep), used in the paper Patched MOA: optimizing inference for diverse software development tasks.
OpenAI used the synth-vuln-fixes and fine-tuned
a new version of gpt-4o is now the SOTA on this benchmark. More details and code is available from their repo.
More details on the benchmark… See the full description on the dataset page: https://huggingface.co/datasets/patched-codes/static-analysis-eval.patch-aliasing-bayesian
Patch-Aliasing Bayesian Analysis Data
Raw and processed data from the Bayesian analysis of structural patch-aliasing in Chronos-Bolt Tiny.
Companion to the model weights at federicosabbadini/chronos-bolt-patch-aliasing-models.
Structure
clean_15model/ # 15 preregistered (P,S) configurations
full_22model/ # all 22 configurations (adds 7 robustness checks)
h1_fixed_offset/ # supplementary H1 analysis with fixed offset
Each run folder… See the full description on the dataset page: https://huggingface.co/datasets/federicosabbadini/patch-aliasing-bayesian.diffllama_patch_tokenizedvulnerability-cwe-patch
Description
This dataset, CIRCL/vulnerability-cwe-patch, provides structured, real-world vulnerabilities enriched with CWE identifiers and corresponding patches from platforms like GitHub and GitLab. It is designed to support the development of tools for vulnerability classification, triage, and automated remediation. Each entry includes metadata such as CVE/GHSA ID, a description, CWE categorization, and links to verified patch commits with associated diff content and commit… See the full description on the dataset page: https://huggingface.co/datasets/CIRCL/vulnerability-cwe-patch.patchlet-embed-preprocessedPatchMap_v1generate-readme-eval
Generate README Eval
The generate-readme-eval is a dataset (train split) and benchmark (test split) to evaluate the effectiveness of LLMs
when summarizing entire GitHub repos in form of a README.md file. The datset is curated from top 400 real Python repositories
from GitHub with at least 1000 stars and 100 forks. The script used to generate the dataset can be found here.
For the dataset we restrict ourselves to GH repositories that are less than 100k tokens in size to allow us to… See the full description on the dataset page: https://huggingface.co/datasets/patched-codes/generate-readme-eval.cc12m_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-13e621-tagger-patchpatchaudit-artifact
PatchAudit Artifact
PatchAudit audits security patches. You give it a CVE's initial fix — commit C1 — and a later commit Ci,
and it tells you whether Ci is a future commit: a later commit that had to keep fixing the same problem
because C1 was incomplete (it left the vulnerability reachable) or incorrect (its own change
introduced a new defect). When such a future commit exists, C1 was a bad patch. When even the latest fix
still leaves the hole open, the bug is a lingering… See the full description on the dataset page: https://huggingface.co/datasets/zhcharyzhang/patchaudit-artifact.liveswebench-patchestraining_03_05_patchpatch_camelyonvessel-detection-labeled-patches
Vessel Detection Labeled Patches
Validated/confirmed satellite image patches exported from the military-boat-detection review workflow.
Contents
images/: patch images.
metadata.csv: one row per patch, compatible with Hugging Face image-folder metadata.
metadata.jsonl: rich patch metadata with nested objects.
annotations.csv: one row per vessel annotation.
annotations.jsonl: JSONL version of the object annotations.
labels/: YOLO-format labels. Hard negatives have empty… See the full description on the dataset page: https://huggingface.co/datasets/DefendIntelligence/vessel-detection-labeled-patches.patchbench
PatchBench
A multi-language benchmark for evaluating whether an LLM can fix a bug correctly and without introducing new bugs. Each row presents a real issue report plus a large code context (~500–1500 lines); the model under test produces a patch; a per-row test suite then checks that (a) previously failing tests now pass and (b) the rest of the suite stays green.
Status: v0 spec. Rows are being built.
Row schema
Field
Type
In tasks config
Description… See the full description on the dataset page: https://huggingface.co/datasets/rootxhacker/patchbench.PatchCamelyon
PatchCamelyon (PCam)
This is a reupload of the PatchCamelyon (PCam) dataset to make it more readily usable instead of manipulating H5 files. The original can be found in the author's Github repo.
If you use this dataset, please cite the original publications:
@inproceedings{veeling2018rotation,
title={Rotation Equivariant CNNs for Digital Pathology},
author={Veeling, Bastiaan S and Linmans, Jasper and Winkens, Jim and Cohen, Taco and Welling, Max},
booktitle={Medical Image… See the full description on the dataset page: https://huggingface.co/datasets/zacharielegault/PatchCamelyon.PatchEval
👋 Overview
PatchEval-Verified is a benchmark for evaluating LLMs and coding agents on automated repair of real-world vulnerabilities. It contains 230 CVE cases with Docker-based evaluation environments, covering vulnerabilities reported between 2015 and 2025 across Go, JavaScript, and Python.
PatchEval-Verified updates the evaluation environments from the original PatchEval release. In the original benchmark, some PoC tests were adapted from project regression tests and were… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance/PatchEval.github-patchesLLaVa-CC3M-PostTraining-clip-vit-base-patch16patchvla-lehome-migration
PatchVLA / LeHome migration snapshot
This repository contains a Docker environment and runtime data snapshot exported from the owner's A6000 server for restoration on a PRO5000 server. It is an environment transfer bundle, not a Hugging Face model or a tabular dataset.
Contents
docker-images.tar.zst.part-*: a single Docker save archive containing lehome-lerobot:0.6.0-vla and openpi-pi05:ubuntu22, Zstandard-compressed and split into 4 GiB parts.… See the full description on the dataset page: https://huggingface.co/datasets/haxiaofeng/patchvla-lehome-migration.
