Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cvssp /WavCaps WavCaps WavCaps is a ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language multimodal research, where the audio clips are sourced from three websites (FreeSound, BBC Sound Effects, and SoundBible) and a sound event detection dataset (AudioSet Strongly-labelled Subset). Paper: https://arxiv.org/abs/2303.17395 Github: https://github.com/XinhaoMei/WavCaps Statistics Data Source # audio avg. audio duration (s)avg. text length FreeSound… See the full description on the dataset page: https://huggingface.co/datasets/cvssp/WavCaps.textn<1K56 likes22k downloads3y agoHugging Face02nvidia /cvdp-benchmark-datasetImportant please see "Files and versions" above for full list of files in the CVDP dataset. Please see LICENSE and NOTICE for licensing information. See CHANGELOG for changes. This is the Comprehensive Verilog Design Problems (CVDP) benchmark dataset to use with the CVDP infrastructure on GitHub. textn<1K40 likes4.9k downloads3mo agoHugging Face03Beijing-AISI /C-VARCThis repository contains all the data associated with the paper "C-VARC: A Large-Scale Chinese Value Rule Corpus for Value Alignment of Large Language Models". We propose a three-tier value classification framework based on core Chinese values, which includes three dimensions, twelve core values, and fifty derived values. With the assistance of large language models and manual verification, we constructed a large-scale, refined, and high-quality value corpus containing over 250,000 rules. We… See the full description on the dataset page: https://huggingface.co/datasets/Beijing-AISI/C-VARC.texttext-generation100K<n<1M3 likes3.1k downloads1y agoHugging Face04wambosec /cve-reconstruction CVE Reconstruction This public dataset contains the complete case assets for 363 validated CVE reconstruction cases. Each case has white-box and black-box variants, giving 726 tasks. The code-only evaluator is in the residency-environments PR. The public Prime environment also provides the code-only evaluator. The task index is data/tasks.jsonl. Each row identifies a real vulnerable and fixed release and its public Prime sandbox image. cases/ contains every case manifest and… See the full description on the dataset page: https://huggingface.co/datasets/wambosec/cve-reconstruction.texttext-generationn<1K0 likes2.4k downloads27m agoHugging Face05AILab-CVC /SEED-Data-Edit-Part1-Openimages SEED-Data-Edit SEED-Data-Edit is a hybrid dataset for instruction-guided image editing with a total of 3.7 image editing pairs, which comprises three distinct types of data: Part-1: Large-scale high-quality editing data produced by automated pipelines (3.5M editing pairs). Part-2: Real-world scenario data collected from the internet (52K editing pairs). Part-3: High-precision multi-turn editing data annotated by humans (95K editing pairs, 21K multi-turn rounds with a maximum of 5… See the full description on the dataset page: https://huggingface.co/datasets/AILab-CVC/SEED-Data-Edit-Part1-Openimages.tabulartext-to-image1M<n<10M10 likes1.1k downloads2y agoHugging Face06AlicanKiraz0 /All-CVE-Records-Training-Dataset CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025) 1. Project Overview This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/All-CVE-Records-Training-Dataset.texttext-generation100K<n<1M61 likes1k downloads1y agoHugging Face07CVI2-UniLU /SPADES-RGBimage100K<n<1M0 likes821 downloads28d agoHugging Face08FluidInference /cv-corpus-25.0-ja Mozilla Common Voice 25.0 - Japanese Test Set (Complete) Dataset Description Complete Japanese test set from Mozilla Common Voice Corpus 25.0. This dataset contains all 9,019 validated test samples, compared to the partial 2,334-sample version previously available on HuggingFace. Key Features Size: 9,019 validated test utterances Coverage: 100% of official Common Voice 25.0 Japanese test split Multi-speaker: Diverse set of speakers with demographic metadata… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/cv-corpus-25.0-ja.audioautomatic-speech-recognition1K<n<10K0 likes674 downloads6mo agoHugging Face09Amine-CV /MODG MODG: Matched Outcomes, Divergent Gaze Data for "Matched Outcomes, Divergent Gaze: How Foveated MLLMs Search Compared to Humans." Paper: arXiv:2608.16514 · Code: github.com/kmamine/MODG Three multimodal LLMs search COCO-Search18 scenes one fixation at a time, zero-shot: Qwen3.5-35B-A3B GLM-4.6V-Flash Gemma-4-E4B Each model sees the scene through a foveated renderer centred on its current gaze. The models are compared with the ten human scanpaths recorded for each scene.… See the full description on the dataset page: https://huggingface.co/datasets/Amine-CV/MODG.imagevisual-question-answering100K<n<1M0 likes672 downloads4d agoHugging Face10jason-oneal /mitre-stix-cve-exploitdb-dataset-alpaca-chatml-harmony MITRE+NVD+ExploitDB Dataset (Alpaca/ChatML/Harmony) A dataset for training AI assistants/agents on vulnerability analysis and pentesting Q&A. It is built by the pentestds pipeline, which fetches and merges data from MITRE CVE, NVD (CVSS enrichment), ExploitDB, and a small set of HuggingFace datasets. Provenance is recorded for every entry, and the pipeline emits Alpaca, ChatML, and Harmony JSONL files. Dataset Summary This dataset is designed for training AI agents to… See the full description on the dataset page: https://huggingface.co/datasets/jason-oneal/mitre-stix-cve-exploitdb-dataset-alpaca-chatml-harmony.text1M<n<10M15 likes543 downloads5mo agoHugging Face11sonalsannigrahi /cv22_azeros FLEURS (Lhotse cuts) Each language is a separate config. Load a single language's cuts as a HF Dataset of raw manifest records with, e.g.: from datasets import load_dataset ds = load_dataset("your-org/REPO_NAME", "bg_bg", split="train") If audio shards (recording.NNNNN.tar) are present alongside the cuts, the LANG/SPLIT/ folder is a valid Lhotse Shar directory. Download it (e.g. via snapshot_download) and load with Lhotse directly: from huggingface_hub import snapshot_download… See the full description on the dataset page: https://huggingface.co/datasets/sonalsannigrahi/cv22_azeros.tabular1M<n<10M0 likes269 downloads3mo agoHugging Face12aist-cvrt /HanDyVQA HanDyVQA Dataset 👋 HanDyVQA (Hand-Object Dynamics Video Question Answering) Dataset is a new benchmark for evalutating abundant spatio-temporal dynamics, process, and effects contained in hand-object interactions. This dataset is built on top of Ego4D Dataset. Get Started 0. Install LFS If you haven’t already, install Git Large File Storage (LFS): git lfs install 1. Clone Repository git clone https://huggingface.co/datasets/aist-cvrt/HanDyVQA… See the full description on the dataset page: https://huggingface.co/datasets/aist-cvrt/HanDyVQA.imagequestion-answering10K<n<100K0 likes239 downloads11mo agoHugging Face13Trendyol /All-CVE-Chat-MultiTurn-1999-2025-Dataset CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025) 1. Project Overview This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/All-CVE-Chat-MultiTurn-1999-2025-Dataset.texttext-generation100K<n<1M32 likes182 downloads1y agoHugging Face14Luoberta /cve_train CVE-Factory Agent Traces This dataset contains 4,078 distilled agent traces from 887 CVE reproduction tasks (the trainset/ split), generated using Claude Opus 4.5 with a Mini SWE-Agent harness. See CVE-Factory for the full pipeline. Note: CVE-Factory also provides a trainset-2/ split with additional simpler tasks, which is not included in this training dataset. 🚀 Training Results Fine-tuning on this dataset yields dramatic improvements across security benchmarks:… See the full description on the dataset page: https://huggingface.co/datasets/Luoberta/cve_train.texttext-generation1K<n<10K6 likes176 downloads8mo agoHugging Face15choucsan /CVPR_Papers CVPR Papers Since 2013, deep learning has revolutionized computer vision, with CVPR (IEEE Conference on Computer Vision and Pattern Recognition) serving as the premier venue documenting this transformative journey. From AlexNet's breakthrough to the rise of Transformers, CVPR papers chronicle the complete trajectory of computer vision advancement. CVPR Papers is a comprehensive dataset containing all papers from CVPR 2013 to present, including metadata and… See the full description on the dataset page: https://huggingface.co/datasets/choucsan/CVPR_Papers.texttext-classification10K<n<100K3 likes172 downloads2mo agoHugging Face16Luoberta /cve_train_v1.1 CVE-Factory Agent Traces v1.1 This dataset is an expanded version of cve_train, containing 18,783 distilled agent traces for CVE reproduction tasks. The traces were generated using Claude Opus 4.5 with a Mini SWE-Agent harness through the CVE-Factory pipeline. What's New in v1.1 Compared to cve_train (v1.0): 18.8k total samples (up from ~4k in v1.0) +3k agentic tasks from cve_tasks_3k_compressed Additional traces from expanded CVE task coverage Training… See the full description on the dataset page: https://huggingface.co/datasets/Luoberta/cve_train_v1.1.texttext-generation10K<n<100K5 likes155 downloads7mo agoHugging Face17Bouquets /Cybersecurity-LLM-CVE2025.06.07 Updated data code :https://github.com/Bouquets-ai/Data-Processing/blob/main/CVE-Data.py Change 121 lines of code (keyword="CVE-2025") to obtain the required CVE time 93 lines of code (json record=) to change the required format Cybersecurity-LLM-CVE Dataset Introduction 🚀 Overview 🛡️ An open-source cybersecurity vulnerability dataset designed for training/evaluating Large Language Models (LLMs) in security domains. Covers all public CVE IDs from January 1, 2021 to April 9, 2025… See the full description on the dataset page: https://huggingface.co/datasets/Bouquets/Cybersecurity-LLM-CVE.text100K<n<1M16 likes153 downloads1y agoHugging Face18gemmozero /ai-cve-2026gated Ai Cve 2026 Part of the LEGION Intelligence dataset collection. Provider: LEGION Systems Access: Requires approval — submit request below Usage from datasets import load_dataset dataset = load_dataset("gemmozero/ai-cve-2026") API Access Real-time access via LEGION API: curl https://api.legion-api.com/incidents API Docs · Pro Access €29/mo License CC BY-NC 4.0 — Research and non-commercial use only. Commercial use requires API… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-cve-2026.texttext-classificationn<1K0 likes148 downloads6d agoHugging Face19Skepsun /cvalues_rlhfConverted from: https://modelscope.cn/datasets/damo/CValues-Comparison/summary. We obtained harmless set by selecting pos_type="拒绝为主" and neg_type="风险回复". We obtained helpful set by selecting pos_type="拒绝&正向建议" and neg_type="拒绝为主". text10K<n<100K10 likes136 downloads3y agoHugging Face20icantiemyshoe /cve-to-metasploit-module CVE To Metasploit Module Prompt This dataset is a submodule to the overall project to create an LLM that can look at newly published CVE writeups and create metasploit modules. The main repo for the project can be found here. Usage TO-DO References TO-DO text1K<n<10K9 likes128 downloads3y agoHugging Face21cvis-tmu /Spatial-SSRL-81k Spatial-SSRL-81k 📖Paper| 🏠Github |🤗Spatial-SSRL-7B Model | 🤗Spatial-SSRL-3B Model | 🤗Spatial-SSRL-Qwen3VL-4B Model | 🤗Spatial-SSRL-81k Dataset | 📰Daily Paper Spatial-SSRL-81k is a training dataset for enhancing spatial understanding in large vision-language models. It contains 81,053 samples of five pretext tasks for self-supervised learning, offering simple, intrinsic supervision that scales RLVR efficiently. 📢 News 🚀 [2026/04/05] We have released… See the full description on the dataset page: https://huggingface.co/datasets/cvis-tmu/Spatial-SSRL-81k.imagevisual-question-answeringn<1K0 likes128 downloads27d agoHugging Face22cveinnt /kepler-arc-agi-3-traces Kepler 1.0 ARC-AGI-3 trace corpus Run artifacts from Kepler 1.0, an open-source agent harness for the 25 public ARC-AGI-3 games. A stock CLI coding agent encodes its theory of each game as an executable world_model.py and is instructed to check it against recorded history before planning. Usable predictions are compared with observations; a mismatch interrupts the plan. Simulation errors and zero prediction coverage can permit execution. This is conditional checking, not a… See the full description on the dataset page: https://huggingface.co/datasets/cveinnt/kepler-arc-agi-3-traces.tabularn<1K0 likes112 downloads10d agoHugging Face23nickh007 /cve-proof-corpus CVE Proof Corpus Six real vulnerability classes, each with a machine-checkable proof that the shipped fix eliminates it — and a checker that shares no code with whatever produced the proof. Every record carries the safety relation, the guard the upstream project shipped, the declared attacker domain, and the nonnegative multipliers that prove the guard implies safety. All six verify. pip install "certkit@git+https://github.com/nickharris808/certkit@main" python verify.py… See the full description on the dataset page: https://huggingface.co/datasets/nickh007/cve-proof-corpus.tabulartext-classificationn<1K0 likes99 downloads1mo agoHugging Face24trapstreet /jev-class-cve-bench Jev-class: name the weakness class You get the published description of one software vulnerability. You pick the weakness class it belongs to, from a list of 30 CWE identifiers. 30 options is a lot for this kind of benchmark, and the number of options is itself a limit. Some tools that do this job only accept short option lists — one published implementation refuses any list longer than 16 before it even runs the model. It does not score badly; it cannot answer at all. In… See the full description on the dataset page: https://huggingface.co/datasets/trapstreet/jev-class-cve-bench.texttext-classification1K<n<10K2 likes98 downloads7d agoHugging Face25yangjie-cv /WeThink_Multimodal_Reasoning_120K Dataset Card for WeThink Repository: https://github.com/yangjie-cv/WeThink Paper: https://arxiv.org/abs/2506.07905 Dataset Structure Question-Answer Pairs The WeThink_Multimodal_Reasoning_120K.jsonl file contains the question-answering data in the following format: { "problem": "QUESTION", "answer": "ANSWER", "category": "QUESTION TYPE", "abilities": "QUESTION REQUIRED ABILITIES", "refined_cot": "THINK PROCESS", "image_path": "IMAGE PATH"… See the full description on the dataset page: https://huggingface.co/datasets/yangjie-cv/WeThink_Multimodal_Reasoning_120K.text100K<n<1M12 likes86 downloads1y agoHugging Face26AbiralArch /hardware-cvdp-problems Hardware Design AI Training Dataset This dataset contains processed hardware design problems and Verilog code for training AI models. Contents CVDP Problems: 160 evaluation problems organized by domain and complexity Training Data: Instruction-code pairs for hardware design Metadata: Rich annotations for each problem Usage from datasets import load_dataset dataset = load_dataset("AbiralArch/hardware-cvdp-problems") Categories Module Generation… See the full description on the dataset page: https://huggingface.co/datasets/AbiralArch/hardware-cvdp-problems.texttext-generationn<1K0 likes79 downloads1y agoHugging Face27stelvita /C-VARCThis repository contains all the data associated with the paper "C-VARC: A Large-Scale Chinese Value Rule Corpus for Value Alignment of Large Language Models". We propose a three-tier value classification framework based on core Chinese values, which includes three dimensions, twelve core values, and fifty derived values. With the assistance of large language models and manual verification, we constructed a large-scale, refined, and high-quality value corpus containing over 250,000 rules. We… See the full description on the dataset page: https://huggingface.co/datasets/stelvita/C-VARC.texttext-generation100K<n<1M0 likes76 downloads7mo agoHugging Face28abamerdeen /cv-pii-bench CV-PII-Bench 200 fictional CVs in 12 languages with character-level annotations of personal data, built to evaluate PII masking in recruitment. It includes a 40-CV hold-out written independently of any masking system. Code, pipeline, all experiment results and the paper: https://github.com/ammarisme/cv-pii-bench Splits Split CVs Required spans Authorship Intended use probe20 20 107 Hand-written, one failure mode per CV Smoke test only; the reference… See the full description on the dataset page: https://huggingface.co/datasets/abamerdeen/cv-pii-bench.texttoken-classificationn<1K0 likes70 downloads14d agoHugging Face29ArkhAngelLifeJiggy /cve_train CVE-Factory Agent Traces This dataset contains 4,078 distilled agent traces from 887 CVE reproduction tasks (the trainset/ split), generated using Claude Opus 4.5 with a Mini SWE-Agent harness. See CVE-Factory for the full pipeline. Note: CVE-Factory also provides a trainset-2/ split with additional simpler tasks, which is not included in this training dataset. 🚀 Training Results Fine-tuning on this dataset yields dramatic improvements across security… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/cve_train.texttext-generation1K<n<10K0 likes70 downloads14d agoHugging Face30threatcluster /cve-exploitation-signals CVE exploitation signals One row per CVE joining reference data (CVSS, CWE, affected vendors and products) with exploitation signals: CISA KEV listing and due date, whether a public exploit is known, and whether the vulnerability is used by ransomware operators. Built from the ThreatCluster corpus. 60,879 rows, snapshot generated 2026-09-06. Fields Field Description cve_id CVE identifier description Vulnerability description published_date CVE… See the full description on the dataset page: https://huggingface.co/datasets/threatcluster/cve-exploitation-signals.texttabular-classification10K<n<100K0 likes67 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.