datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WavCaps
WavCaps
WavCaps is a ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language multimodal research, where the audio clips are sourced from three websites (FreeSound, BBC Sound Effects, and SoundBible) and a sound event detection dataset (AudioSet Strongly-labelled Subset).
Paper: https://arxiv.org/abs/2303.17395
Github: https://github.com/XinhaoMei/WavCaps
Statistics
Data Source
# audio
avg. audio duration (s)avg. text length
FreeSound… See the full description on the dataset page: https://huggingface.co/datasets/cvssp/WavCaps.cvdp-benchmark-datasetImportant please see "Files and versions" above for full list of files in the CVDP dataset.
Please see LICENSE and NOTICE for licensing information. See CHANGELOG for changes.
This is the Comprehensive Verilog Design Problems (CVDP) benchmark dataset to use with the CVDP infrastructure on GitHub.
C-VARCThis repository contains all the data associated with the paper "C-VARC: A Large-Scale Chinese Value Rule Corpus for Value Alignment of Large Language Models".
We propose a three-tier value classification framework based on core Chinese values, which includes three dimensions, twelve core values, and fifty derived values. With the assistance of large language models and manual verification, we constructed a large-scale, refined, and high-quality value corpus containing over 250,000 rules. We… See the full description on the dataset page: https://huggingface.co/datasets/Beijing-AISI/C-VARC.cve-reconstruction
CVE Reconstruction
This public dataset contains the complete case assets for 363 validated CVE
reconstruction cases. Each case has white-box and black-box variants, giving 726 tasks. The code-only evaluator is in the
residency-environments PR.
The public Prime environment
also provides the code-only evaluator.
The task index is data/tasks.jsonl. Each row identifies a real vulnerable and
fixed release and its public Prime sandbox image. cases/ contains every case
manifest and… See the full description on the dataset page: https://huggingface.co/datasets/wambosec/cve-reconstruction.SEED-Data-Edit-Part1-Openimages
SEED-Data-Edit
SEED-Data-Edit is a hybrid dataset for instruction-guided image editing with a total of 3.7 image editing pairs, which comprises three distinct types of data:
Part-1: Large-scale high-quality editing data produced by automated pipelines (3.5M editing pairs).
Part-2: Real-world scenario data collected from the internet (52K editing pairs).
Part-3: High-precision multi-turn editing data annotated by humans (95K editing pairs, 21K multi-turn rounds with a maximum of 5… See the full description on the dataset page: https://huggingface.co/datasets/AILab-CVC/SEED-Data-Edit-Part1-Openimages.All-CVE-Records-Training-Dataset
CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025)
1. Project Overview
This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/All-CVE-Records-Training-Dataset.SPADES-RGBcv-corpus-25.0-ja
Mozilla Common Voice 25.0 - Japanese Test Set (Complete)
Dataset Description
Complete Japanese test set from Mozilla Common Voice Corpus 25.0. This dataset contains all 9,019 validated test samples, compared to the partial 2,334-sample version previously available on HuggingFace.
Key Features
Size: 9,019 validated test utterances
Coverage: 100% of official Common Voice 25.0 Japanese test split
Multi-speaker: Diverse set of speakers with demographic metadata… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/cv-corpus-25.0-ja.MODG
MODG: Matched Outcomes, Divergent Gaze
Data for "Matched Outcomes, Divergent Gaze: How Foveated MLLMs Search Compared to Humans."
Paper: arXiv:2608.16514 · Code: github.com/kmamine/MODG
Three multimodal LLMs search COCO-Search18 scenes one fixation at a time, zero-shot:
Qwen3.5-35B-A3B
GLM-4.6V-Flash
Gemma-4-E4B
Each model sees the scene through a foveated renderer centred on its current gaze. The models are compared
with the ten human scanpaths recorded for each scene.… See the full description on the dataset page: https://huggingface.co/datasets/Amine-CV/MODG.mitre-stix-cve-exploitdb-dataset-alpaca-chatml-harmony
MITRE+NVD+ExploitDB Dataset (Alpaca/ChatML/Harmony)
A dataset for training AI assistants/agents on vulnerability analysis and pentesting Q&A. It is built by the pentestds pipeline, which fetches and merges data from MITRE CVE, NVD (CVSS enrichment), ExploitDB, and a small set of HuggingFace datasets. Provenance is recorded for every entry, and the pipeline emits Alpaca, ChatML, and Harmony JSONL files.
Dataset Summary
This dataset is designed for training AI agents to… See the full description on the dataset page: https://huggingface.co/datasets/jason-oneal/mitre-stix-cve-exploitdb-dataset-alpaca-chatml-harmony.cv22_azeros
FLEURS (Lhotse cuts)
Each language is a separate config. Load a single language's cuts as a HF Dataset of raw manifest records with, e.g.:
from datasets import load_dataset
ds = load_dataset("your-org/REPO_NAME", "bg_bg", split="train")
If audio shards (recording.NNNNN.tar) are present alongside the cuts, the LANG/SPLIT/ folder is a valid Lhotse Shar directory. Download it (e.g. via snapshot_download) and load with Lhotse directly:
from huggingface_hub import snapshot_download… See the full description on the dataset page: https://huggingface.co/datasets/sonalsannigrahi/cv22_azeros.HanDyVQA
HanDyVQA Dataset 👋
HanDyVQA (Hand-Object Dynamics Video Question Answering) Dataset is a new benchmark for evalutating abundant spatio-temporal dynamics, process, and effects contained in hand-object interactions. This dataset is built on top of Ego4D Dataset.
Get Started
0. Install LFS
If you haven’t already, install Git Large File Storage (LFS):
git lfs install
1. Clone Repository
git clone https://huggingface.co/datasets/aist-cvrt/HanDyVQA… See the full description on the dataset page: https://huggingface.co/datasets/aist-cvrt/HanDyVQA.All-CVE-Chat-MultiTurn-1999-2025-Dataset
CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025)
1. Project Overview
This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/All-CVE-Chat-MultiTurn-1999-2025-Dataset.cve_train
CVE-Factory Agent Traces
This dataset contains 4,078 distilled agent traces from 887 CVE reproduction tasks (the trainset/ split), generated using Claude Opus 4.5 with a Mini SWE-Agent harness. See CVE-Factory for the full pipeline.
Note: CVE-Factory also provides a trainset-2/ split with additional simpler tasks, which is not included in this training dataset.
🚀 Training Results
Fine-tuning on this dataset yields dramatic improvements across security benchmarks:… See the full description on the dataset page: https://huggingface.co/datasets/Luoberta/cve_train.CVPR_Papers
CVPR Papers
Since 2013, deep learning has revolutionized computer vision, with CVPR (IEEE Conference on Computer Vision and Pattern Recognition) serving as the premier venue documenting this transformative journey. From AlexNet's breakthrough to the rise of Transformers, CVPR papers chronicle the complete trajectory of computer vision advancement.
CVPR Papers is a comprehensive dataset containing all papers from CVPR 2013 to present, including metadata and… See the full description on the dataset page: https://huggingface.co/datasets/choucsan/CVPR_Papers.cve_train_v1.1
CVE-Factory Agent Traces v1.1
This dataset is an expanded version of cve_train, containing 18,783 distilled agent traces for CVE reproduction tasks. The traces were generated using Claude Opus 4.5 with a Mini SWE-Agent harness through the CVE-Factory pipeline.
What's New in v1.1
Compared to cve_train (v1.0):
18.8k total samples (up from ~4k in v1.0)
+3k agentic tasks from cve_tasks_3k_compressed
Additional traces from expanded CVE task coverage
Training… See the full description on the dataset page: https://huggingface.co/datasets/Luoberta/cve_train_v1.1.Cybersecurity-LLM-CVE2025.06.07 Updated data code :https://github.com/Bouquets-ai/Data-Processing/blob/main/CVE-Data.py
Change 121 lines of code (keyword="CVE-2025") to obtain the required CVE time
93 lines of code (json record=) to change the required format
Cybersecurity-LLM-CVE Dataset Introduction 🚀
Overview 🛡️
An open-source cybersecurity vulnerability dataset designed for training/evaluating Large Language Models (LLMs) in security domains. Covers all public CVE IDs from January 1, 2021 to April 9, 2025… See the full description on the dataset page: https://huggingface.co/datasets/Bouquets/Cybersecurity-LLM-CVE.ai-cve-2026
Ai Cve 2026
Part of the LEGION Intelligence dataset collection.
Provider: LEGION Systems
Access: Requires approval — submit request below
Usage
from datasets import load_dataset
dataset = load_dataset("gemmozero/ai-cve-2026")
API Access
Real-time access via LEGION API:
curl https://api.legion-api.com/incidents
API Docs · Pro Access €29/mo
License
CC BY-NC 4.0 — Research and non-commercial use only.
Commercial use requires API… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-cve-2026.cvalues_rlhfConverted from: https://modelscope.cn/datasets/damo/CValues-Comparison/summary. We obtained harmless set by selecting pos_type="拒绝为主" and neg_type="风险回复". We obtained helpful set by selecting pos_type="拒绝&正向建议" and neg_type="拒绝为主".
cve-to-metasploit-module
CVE To Metasploit Module Prompt
This dataset is a submodule to the overall project to create an LLM that can look at newly published CVE writeups and create metasploit modules. The main repo for the project can be found here.
Usage
TO-DO
References
TO-DO
Spatial-SSRL-81k
Spatial-SSRL-81k
📖Paper| 🏠Github |🤗Spatial-SSRL-7B Model |
🤗Spatial-SSRL-3B Model | 🤗Spatial-SSRL-Qwen3VL-4B Model |
🤗Spatial-SSRL-81k Dataset | 📰Daily Paper
Spatial-SSRL-81k is a training dataset for enhancing spatial understanding in large vision-language models. It contains 81,053 samples of five pretext tasks for self-supervised learning, offering simple, intrinsic supervision that scales RLVR efficiently.
📢 News
🚀 [2026/04/05] We have released… See the full description on the dataset page: https://huggingface.co/datasets/cvis-tmu/Spatial-SSRL-81k.kepler-arc-agi-3-traces
Kepler 1.0 ARC-AGI-3 trace corpus
Run artifacts from Kepler 1.0, an open-source agent harness for the 25 public
ARC-AGI-3 games. A stock CLI coding agent
encodes its theory of each game as an executable world_model.py and is
instructed to check it against recorded history before planning. Usable
predictions are compared with observations; a mismatch interrupts the plan.
Simulation errors and zero prediction coverage can permit execution. This is
conditional checking, not a… See the full description on the dataset page: https://huggingface.co/datasets/cveinnt/kepler-arc-agi-3-traces.cve-proof-corpus
CVE Proof Corpus
Six real vulnerability classes, each with a machine-checkable proof that the shipped fix
eliminates it — and a checker that shares no code with whatever produced the proof.
Every record carries the safety relation, the guard the upstream project shipped, the declared
attacker domain, and the nonnegative multipliers that prove the guard implies safety. All six verify.
pip install "certkit@git+https://github.com/nickharris808/certkit@main"
python verify.py… See the full description on the dataset page: https://huggingface.co/datasets/nickh007/cve-proof-corpus.jev-class-cve-bench
Jev-class: name the weakness class
You get the published description of one software vulnerability. You pick the
weakness class it belongs to, from a list of 30 CWE identifiers.
30 options is a lot for this kind of benchmark, and the number of options is
itself a limit. Some tools that do this job only accept short option lists — one
published implementation refuses any list longer than 16 before it even runs the
model. It does not score badly; it cannot answer at all.
In… See the full description on the dataset page: https://huggingface.co/datasets/trapstreet/jev-class-cve-bench.WeThink_Multimodal_Reasoning_120K
Dataset Card for WeThink
Repository: https://github.com/yangjie-cv/WeThink
Paper: https://arxiv.org/abs/2506.07905
Dataset Structure
Question-Answer Pairs
The WeThink_Multimodal_Reasoning_120K.jsonl file contains the question-answering data in the following format:
{
"problem": "QUESTION",
"answer": "ANSWER",
"category": "QUESTION TYPE",
"abilities": "QUESTION REQUIRED ABILITIES",
"refined_cot": "THINK PROCESS",
"image_path": "IMAGE PATH"… See the full description on the dataset page: https://huggingface.co/datasets/yangjie-cv/WeThink_Multimodal_Reasoning_120K.hardware-cvdp-problems
Hardware Design AI Training Dataset
This dataset contains processed hardware design problems and Verilog code for training AI models.
Contents
CVDP Problems: 160 evaluation problems organized by domain and complexity
Training Data: Instruction-code pairs for hardware design
Metadata: Rich annotations for each problem
Usage
from datasets import load_dataset
dataset = load_dataset("AbiralArch/hardware-cvdp-problems")
Categories
Module Generation… See the full description on the dataset page: https://huggingface.co/datasets/AbiralArch/hardware-cvdp-problems.C-VARCThis repository contains all the data associated with the paper "C-VARC: A Large-Scale Chinese Value Rule Corpus for Value Alignment of Large Language Models".
We propose a three-tier value classification framework based on core Chinese values, which includes three dimensions, twelve core values, and fifty derived values. With the assistance of large language models and manual verification, we constructed a large-scale, refined, and high-quality value corpus containing over 250,000 rules. We… See the full description on the dataset page: https://huggingface.co/datasets/stelvita/C-VARC.cv-pii-bench
CV-PII-Bench
200 fictional CVs in 12 languages with character-level annotations of personal data, built to evaluate PII masking in recruitment. It includes a 40-CV hold-out written independently of any masking system.
Code, pipeline, all experiment results and the paper: https://github.com/ammarisme/cv-pii-bench
Splits
Split
CVs
Required spans
Authorship
Intended use
probe20
20
107
Hand-written, one failure mode per CV
Smoke test only; the reference… See the full description on the dataset page: https://huggingface.co/datasets/abamerdeen/cv-pii-bench.cve_train
CVE-Factory Agent Traces
This dataset contains 4,078 distilled agent traces from 887 CVE reproduction tasks (the trainset/ split), generated using Claude Opus 4.5 with a Mini SWE-Agent harness. See CVE-Factory for the full pipeline.
Note: CVE-Factory also provides a trainset-2/ split with additional simpler tasks, which is not included in this training dataset.
🚀 Training Results
Fine-tuning on this dataset yields dramatic improvements across security… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/cve_train.cve-exploitation-signals
CVE exploitation signals
One row per CVE joining reference data (CVSS, CWE, affected vendors and products) with exploitation signals: CISA KEV listing and due date, whether a public exploit is known, and whether the vulnerability is used by ransomware operators.
Built from the ThreatCluster corpus. 60,879 rows, snapshot generated 2026-09-06.
Fields
Field
Description
cve_id
CVE identifier
description
Vulnerability description
published_date
CVE… See the full description on the dataset page: https://huggingface.co/datasets/threatcluster/cve-exploitation-signals.
