Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01heitorrosa /cvm-corpus CVM Filings Corpus (PT-BR) Brazilian CVM regulatory filings in Portuguese, cleaned and chunked for language-model pretraining. Built for DAPT on financial Portuguese. Contents Path What output/corpus.jsonl Full corpus (7.2 GB). Chunk schema: text, company, cnpj, category, subject, date, year, document_id, chunk_id, extraction_quality. output/corpus-250M.jsonl DSIR-selected 250M-token subset (token count by chars÷4 proxy ≈ 150M whitespace tokens): 187… See the full description on the dataset page: https://huggingface.co/datasets/heitorrosa/cvm-corpus.documenttext-generation100K<n<1M0 likes11k downloads13d agoHugging Face02Beijing-AISI /C-VARCThis repository contains all the data associated with the paper "C-VARC: A Large-Scale Chinese Value Rule Corpus for Value Alignment of Large Language Models". We propose a three-tier value classification framework based on core Chinese values, which includes three dimensions, twelve core values, and fifty derived values. With the assistance of large language models and manual verification, we constructed a large-scale, refined, and high-quality value corpus containing over 250,000 rules. We… See the full description on the dataset page: https://huggingface.co/datasets/Beijing-AISI/C-VARC.texttext-generation100K<n<1M3 likes3.1k downloads1y agoHugging Face03wambosec /cve-reconstruction CVE Reconstruction This public dataset contains the complete case assets for 365 validated CVE reconstruction cases. Each case has white-box and black-box variants, giving 730 tasks. The code-only evaluator is in the residency-environments PR. The public Prime environment also provides the code-only evaluator. The task index is data/tasks.jsonl. Each row identifies a real vulnerable and fixed release and its public Prime sandbox image. cases/ contains every case manifest and… See the full description on the dataset page: https://huggingface.co/datasets/wambosec/cve-reconstruction.texttext-generationn<1K0 likes2.4k downloads2h agoHugging Face04hitoshura25 /cvefixes CVEfixes Security Vulnerabilities Dataset Security vulnerability data from CVEfixes v1.0.8 with 12,987 vulnerability fix records across 11,726 unique CVEs and 4,205 repositories. Contains CVE metadata (descriptions, CVSS scores, CWE classifications), git commit data, and code diffs showing vulnerable vs fixed code. Usage from datasets import load_dataset dataset = load_dataset("hitoshura25/cvefixes") Citation If you use this dataset, please cite the original… See the full description on the dataset page: https://huggingface.co/datasets/hitoshura25/cvefixes.tabulartext-generation10K<n<100K5 likes1.8k downloads1y agoHugging Face05diyer22 /Panda-CVL-train Panda-CVL Training Split Overview Panda-CVL is a token-level correction dataset and benchmark annotated with the onPanda tool. Given a question-response pair, the model first judges whether the response is acceptable. If correction is needed, it must locate the first inappropriate token and replace it with an appropriate one, so generation can continue from the "correct prefix + corrected token" state and ultimately produce an acceptable response. Compared with… See the full description on the dataset page: https://huggingface.co/datasets/diyer22/Panda-CVL-train.text-generation1K<n<10K0 likes1.2k downloads18d agoHugging Face06AlicanKiraz0 /All-CVE-Records-Training-Dataset CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025) 1. Project Overview This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/All-CVE-Records-Training-Dataset.texttext-generation100K<n<1M61 likes1k downloads1y agoHugging Face07diyer22 /Panda-CVL-test Panda-CVL Test Split This dataset is based on the paper onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction. Code is available at GitHub. Overview Panda-CVL is a token-level correction dataset and benchmark annotated with the onPanda tool. Given a question-response pair, the model first judges whether the response is acceptable. If correction is needed, it must locate the first inappropriate token and replace it… See the full description on the dataset page: https://huggingface.co/datasets/diyer22/Panda-CVL-test.text-generationn<1K0 likes984 downloads18d agoHugging Face08AbiralArch /hardware-cvdp-complete CVDP - Comprehensive Verilog Design Problems (Complete Dataset) 🎯 782 out of 783 problems from the official CVDP benchmark by NVIDIA Research 🔥 Dataset Overview This is the most complete version of the Comprehensive Verilog Design Problems (CVDP) benchmark available, containing 782 problems across 13 task categories. CVDP is designed to evaluate Large Language Models and agents on RTL design and verification tasks. 📊 Dataset Statistics Total Problems: 772… See the full description on the dataset page: https://huggingface.co/datasets/AbiralArch/hardware-cvdp-complete.text-generation1K<n<10K1 likes193 downloads1y agoHugging Face09exploitintel /cve-cwe-consensus CVE-to-CWE Consensus Dataset A multi-label dataset mapping CVE vulnerability descriptions to their CWE weakness type(s), built for fine-tuning instruction-tuned LLMs (e.g. with Unsloth). Each label is a consensus assignment: a CWE is kept only when NVD and the CVE Numbering Authority (CNA) independently agree on it, after rolling both up to CWE View-1003 (the ~130-weakness "Weaknesses for Simplified Mapping of Published Vulnerabilities"). TL;DR Task: given a CVE… See the full description on the dataset page: https://huggingface.co/datasets/exploitintel/cve-cwe-consensus.texttext-classification10K<n<100K5 likes187 downloads4mo agoHugging Face10Trendyol /All-CVE-Chat-MultiTurn-1999-2025-Dataset CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025) 1. Project Overview This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/All-CVE-Chat-MultiTurn-1999-2025-Dataset.texttext-generation100K<n<1M32 likes182 downloads1y agoHugging Face11Luoberta /cve_train CVE-Factory Agent Traces This dataset contains 4,078 distilled agent traces from 887 CVE reproduction tasks (the trainset/ split), generated using Claude Opus 4.5 with a Mini SWE-Agent harness. See CVE-Factory for the full pipeline. Note: CVE-Factory also provides a trainset-2/ split with additional simpler tasks, which is not included in this training dataset. 🚀 Training Results Fine-tuning on this dataset yields dramatic improvements across security benchmarks:… See the full description on the dataset page: https://huggingface.co/datasets/Luoberta/cve_train.texttext-generation1K<n<10K6 likes176 downloads8mo agoHugging Face12Luoberta /cve_train_v1.1 CVE-Factory Agent Traces v1.1 This dataset is an expanded version of cve_train, containing 18,783 distilled agent traces for CVE reproduction tasks. The traces were generated using Claude Opus 4.5 with a Mini SWE-Agent harness through the CVE-Factory pipeline. What's New in v1.1 Compared to cve_train (v1.0): 18.8k total samples (up from ~4k in v1.0) +3k agentic tasks from cve_tasks_3k_compressed Additional traces from expanded CVE task coverage Training… See the full description on the dataset page: https://huggingface.co/datasets/Luoberta/cve_train_v1.1.texttext-generation10K<n<100K5 likes155 downloads7mo agoHugging Face13auren-research /cve-sft-v5 CVE SFT Dataset v5 CVE SFT Dataset v5 is a structured instruction-following dataset for fine-tuning language models on cybersecurity vulnerability analysis. Built by Auren Research, it combines authoritative vulnerability metadata from the NIST National Vulnerability Database (NVD) with five generated fields that teach models to explain, reason about, and remediate real-world CVEs — including side-by-side vulnerable vs. safe code examples.Unlike most security… See the full description on the dataset page: https://huggingface.co/datasets/auren-research/cve-sft-v5.texttext-generation10K<n<100K7 likes106 downloads5mo agoHugging Face14morpheuslord /cve-llm-training CVE-llm_dataset This dataset is intended to train an LLM model for an utterly CVE-focused input and output. Data extraction: For the data extraction, I first downloaded the CVE database from NVD lists and then loaded them using the cve_dataset_2.py and cve_dataset.py both have produce different datasets one is for llama and the other is for openai GPT. The CVE json files are mapped in this format: cves: | ├─1999 | ├─0xxx | | ├─CVE-1999-0001.json | | ├─.... | |… See the full description on the dataset page: https://huggingface.co/datasets/morpheuslord/cve-llm-training.text-generationn<1K21 likes101 downloads2y agoHugging Face15Publicus /cvefixes-security-autoformal-span-cache CVEfixes Security source spans This dataset contains exact, source-bound prose, code, and diff spans derived from the original-data configuration of Publicus/cvefixes-security-ir-graphrag at revision 6fd5918bed34f8851430e74a149502587a953fe2. The underlying source is hitoshura25/cvefixes at revision d4f5c4ea65329d9ccbb8a3b3149e5d06eda5edb2. The extraction considered all 12,987 original rows across all three shards. It retained 9,402 source rows and excluded 3,585. The release… See the full description on the dataset page: https://huggingface.co/datasets/Publicus/cvefixes-security-autoformal-span-cache.tabulartext-generation100K<n<1M0 likes83 downloads10d agoHugging Face16AbiralArch /hardware-cvdp-problems Hardware Design AI Training Dataset This dataset contains processed hardware design problems and Verilog code for training AI models. Contents CVDP Problems: 160 evaluation problems organized by domain and complexity Training Data: Instruction-code pairs for hardware design Metadata: Rich annotations for each problem Usage from datasets import load_dataset dataset = load_dataset("AbiralArch/hardware-cvdp-problems") Categories Module Generation… See the full description on the dataset page: https://huggingface.co/datasets/AbiralArch/hardware-cvdp-problems.texttext-generationn<1K0 likes79 downloads1y agoHugging Face17HPAI-BSC /NotSoTiny-25-12-CVDPThis is a version of NotSoTiny-25-12 benchmark modified to work with Nvidia's CVDP framework [!WARNING] If you plan to run this benchmark via CVDP framework, proceed with this version. Otherwise refer to the main dataset for additional info: HPAI-BSC/NotSoTiny-25-12 Subsets and Shuttles The default dataset contains all NotSoTiny-25-12 Tiny Tapeout shuttles combined onto a single dataset comprising 1114 total tasks. However, you can also access individual shuttles as subsets.… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/NotSoTiny-25-12-CVDP.text-generation1K<n<10K0 likes77 downloads5mo agoHugging Face18Shrutz72 /cvefixes CVEfixes Security Vulnerabilities Dataset Security vulnerability data from CVEfixes v1.0.8 with 12,987 vulnerability fix records across 11,726 unique CVEs and 4,205 repositories. Contains CVE metadata (descriptions, CVSS scores, CWE classifications), git commit data, and code diffs showing vulnerable vs fixed code. Usage from datasets import load_dataset dataset = load_dataset("hitoshura25/cvefixes") Citation If you use this dataset, please cite the original… See the full description on the dataset page: https://huggingface.co/datasets/Shrutz72/cvefixes.tabulartext-generation10K<n<100K0 likes77 downloads8mo agoHugging Face19stelvita /C-VARCThis repository contains all the data associated with the paper "C-VARC: A Large-Scale Chinese Value Rule Corpus for Value Alignment of Large Language Models". We propose a three-tier value classification framework based on core Chinese values, which includes three dimensions, twelve core values, and fifty derived values. With the assistance of large language models and manual verification, we constructed a large-scale, refined, and high-quality value corpus containing over 250,000 rules. We… See the full description on the dataset page: https://huggingface.co/datasets/stelvita/C-VARC.texttext-generation100K<n<1M0 likes76 downloads7mo agoHugging Face20ArkhAngelLifeJiggy /cve_train CVE-Factory Agent Traces This dataset contains 4,078 distilled agent traces from 887 CVE reproduction tasks (the trainset/ split), generated using Claude Opus 4.5 with a Mini SWE-Agent harness. See CVE-Factory for the full pipeline. Note: CVE-Factory also provides a trainset-2/ split with additional simpler tasks, which is not included in this training dataset. 🚀 Training Results Fine-tuning on this dataset yields dramatic improvements across security… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/cve_train.texttext-generation1K<n<10K0 likes70 downloads14d agoHugging Face21anicka /cve-backport-codegen-dataset CVE Backport Code Generation Dataset Per-hunk code generation dataset for CVE security patch backporting, derived from openSUSE Build Service maintenance patches. Task Given a region of vulnerable source code and a description of the upstream CVE fix, the model outputs the fixed version of the code. A programmatic diff then produces the final patch. This plays to LLM strengths in code completion and avoids format-sensitivity issues with direct diff generation.… See the full description on the dataset page: https://huggingface.co/datasets/anicka/cve-backport-codegen-dataset.text-generation10K<n<100K0 likes55 downloads6mo agoHugging Face22oi-uae /CVEsgated CVEs — a full-coverage CVE chat dataset 1,625,017 chat conversations covering all 361,190 usable CVEs (1999–2026), built for fine-tuning cybersecurity assistants. Every known CVE in the official CVE List with severity enrichment from NVD (via the fkie-cad community feeds), rendered as English user/assistant conversations with varied phrasings, honest handling of missing data, and a per-CVE 99/1 train/validation split with zero leakage. The schema matches oi-uae/cyber-security… See the full description on the dataset page: https://huggingface.co/datasets/oi-uae/CVEs.texttext-generation1M<n<10M3 likes48 downloads25d agoHugging Face23Amine-CV /thinkprob ThinkProbe: Thought Graphs of Open-Ended LLM Reasoning This is the companion dataset for ThinkProbe: Beyond Accuracy — Structural Profiling of Open-Ended LLM Reasoning Traces via Non-Generative Thought Graphs (Kerkouri et al., 2026). The code is at github.com/kmamine/ThinkProb. ThinkProbe turns each LLM reasoning trace into a Thought Graph: a directed graph, which can contain cycles, whose nodes are typed thought units (TUs) and whose edges are typed reasoning links. The… See the full description on the dataset page: https://huggingface.co/datasets/Amine-CV/thinkprob.tabulartext-generation100K<n<1M0 likes44 downloads4d agoHugging Face24llxtyin /FabricGen-CVPR26 Dataset of FabricGen: Microstructure-Aware Woven Fabric Generation (CVPR 26) The dataset has two components: Macro-scale texture dataset (img_gen): Dataset of microstructure-free fabric textures along with captions. The dataset is used to fine-tune an image diffusion model. Weaving pattern dataset (pattern_gen): Dataset of weaving drafts along with tags (sourced from Handweaving.net). The dataset is used to fine-tune a large language model. Detailed fine-tuning strategy and… See the full description on the dataset page: https://huggingface.co/datasets/llxtyin/FabricGen-CVPR26.image-to-text1K<n<10K2 likes43 downloads7mo agoHugging Face25ansulev /all-cve-chat-multiturn-1999-2025 CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025) 1. Project Overview This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/all-cve-chat-multiturn-1999-2025.texttext-generation100K<n<1M2 likes39 downloads7mo agoHugging Face26ChipHolmes /All-CVE-Records-Training-Dataset-archive CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025) 1. Project Overview This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/All-CVE-Records-Training-Dataset-archive.texttext-generation100K<n<1M1 likes39 downloads3mo agoHugging Face27sh111111111111111 /cve-analysis CVE & Vulnerability Analysis Dataset A comprehensive vulnerability analysis and CVE research dataset. Each row is a detailed security analysis covering root cause, exploitation methodology, detection rules (Sigma/Splunk/Suricata), CVSS v3.1 scoring, MITRE ATT&CK mapping, and remediation guidance — verified by the same model in an independent review pass. Overview This dataset contains 9,999 structured vulnerability analyses across 20 security domains. Unlike simple… See the full description on the dataset page: https://huggingface.co/datasets/sh111111111111111/cve-analysis.texttext-generation1K<n<10K1 likes36 downloads7mo agoHugging Face28lambdasec /cve-single-line-fixes Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/lambdasec/cve-single-line-fixes.texttext-generationn<1K3 likes31 downloads4y agoHugging Face29puwaer /cvalues_rlhf_en puwaer/cvalues_rlhf_jp This dataset is based on Skepsun/cvalues_rlhf and has been translated into appropriate English for DPO (Direct Preference Optimization). For the prompt and rejected (negative example) fields, outputs from huihui-ai/Huihui-gpt-oss-20b-mxfp4-abliterated-v2 were used. For the chosen (positive example) field, outputs from openai/gpt-oss-20b were used. このデータセットは、Skepsun/cvalues_rlhfをもとに、適切な英語に翻訳したdpo用のデータセットです。… See the full description on the dataset page: https://huggingface.co/datasets/puwaer/cvalues_rlhf_en.texttext-generation10K<n<100K1 likes29 downloads3mo agoHugging Face30puwaer /cvalues_rlhf_jp_cot puwaer/cvalues_rlhf_jp_cot This dataset is based on Skepsun/cvalues_rlhf and has been translated into appropriate Japanese for DPO (Direct Preference Optimization). For the prompt and rejected (negative example) fields, outputs from huihui-ai/Huihui-gpt-oss-20b-mxfp4-abliterated-v2 were used. For the chosen (positive example) field, outputs from openai/gpt-oss-20b were used. For the chosen chain of thought field, outputs from Qwen/Qwen3-235B-A22B-Instruct-2507 were used. For… See the full description on the dataset page: https://huggingface.co/datasets/puwaer/cvalues_rlhf_jp_cot.texttext-generation10K<n<100K0 likes26 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.