datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arxiv-complete
arXiv Complete Corpus
A snapshot of arXiv's metadata, version history, submission files and rendered
documents. It covers 3,148,796 papers and includes file contents, paths, sizes
and SHA-256 digests. Metadata comes from arXiv's OAI-PMH arXivRaw interface;
files come from the GCS mirror, S3 source archives and direct PDF fetches.
This release holds a PDF for 99.47% of papers and 99.54% of versions reported
with a non-zero submission size. It is a one-off snapshot; coverage gaps… See the full description on the dataset page: https://huggingface.co/datasets/secemp9/arxiv-complete.dojo_sector_precomputed
Languages: 简体中文 · English
dojo_sector_precomputed — Precomputed Sector Analytics
Overview
Derived sector analytics: L3 constituent snapshots, daily cap-weighted sector index levels, and per-constituent daily returns. Built offline from taxonomy, mappings, quotes, and stock K-lines.
Files
File
Description
manifest.json
Generation metadata: version, window start, row counts, latest trade dates
constituents.parquet
L3 constituent… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_sector_precomputed.security-auditsA collection of agent traces generated with Swival (not Claude Code, despite what the HF interface currently shows), an agent designed for open-source models.
These traces focus on security audits of opensource software.
Sharing traces with Swival
Swival can export full conversation traces with --trace-dir, which writes one <session_id>.jsonl file per session:
swival "Fix the login bug" --trace-dir traces/
Those JSONL files use Swival's Claude Code compatible trace export, and… See the full description on the dataset page: https://huggingface.co/datasets/jedisct1/security-audits.dojo_sector_info
Languages: 简体中文 · English
dojo_sector_info — Sector Taxonomy
Overview
Three-level sector taxonomy (L1 industry → L2 chain → L3 sub-sector) with bilingual Chinese/English names and definitions.
Files
File
Description
data.parquet
Taxonomy tree (one L1 row each; L2/L3 nested in children)
Key Fields
Field
Description
id
L1 sector ID
name / name_alias
L1 English name / Chinese alias
description /… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_sector_info.dojo_sector_symbol_relations
Languages: 简体中文 · English
dojo_sector_symbol_relations — Stock–Sector Mapping
Overview
Maps each stock to L1/L2/L3 sector paths with primary and secondary assignments. One row per (ticker, market) pair.
Files
File
Description
data.parquet
Full stock ↔ sector relations
Key Fields
Field
Description
ticker
Stock symbol
market
us, cn, or hk
primary
JSON object — primary sector path
secondary
JSON array —… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_sector_symbol_relations.SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset.
The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database.
The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/TeraflopAI/SEC-EDGAR.secqa
SecQA
SecQA is a specialized dataset created for the evaluation of Large Language Models (LLMs) in the domain of computer security.
It consists of multiple-choice questions, generated using GPT-4 and the
Computer Systems Security: Planning for Success textbook,
aimed at assessing the understanding and application of LLMs' knowledge in computer security.
Dataset Details
Dataset Description
SecQA is an innovative dataset designed to benchmark the… See the full description on the dataset page: https://huggingface.co/datasets/zefang-liu/secqa.cavaai-sec-mirrorSEC_Exhibit_10
SEC Exhibit 10 Material Contracts Dataset (2001-2024)
Overview
This dataset contains material contracts (Exhibit 10) from SEC filings between 2001 and 2024, representing the largest collection of commercial contracts to date. It includes both original filings and text-converted versions.
If you find this dataset useful, please cite:
Arbel, Yonathan A., The Readability of Contracts: Big Data Analysis, Forthcoming J. Empirical Legal Stud. 21:4 (Dec. 2024)
Methodological… See the full description on the dataset page: https://huggingface.co/datasets/yonathanarbel/SEC_Exhibit_10.pmc_open_access_sectionSecuTable
SecuTable: A Dataset for Semantic Table Interpretation in Security Domain
Dataset Overview
Security datasets are scattered on the Internet (CVE, CAPEC, CWE, etc.) and provided in CSV, JSON or XML formats. This makes it difficult to get a holistic view of the interconnectedness of information across different data sources. On the other hand, many datasets focus on specific attack vectors or limited environments, limiting generalisability. There is a lack of detailed… See the full description on the dataset page: https://huggingface.co/datasets/jiofidelus/SecuTable.SEC-bench
Data Instances
instance_id: (str) - A unique identifier for the instance
repo: (str) - The repository name including the owner
project_name: (str) - The name of the project without owner
lang: (str) - The programming language of the repository
work_dir: (str) - Working directory path
sanitizer: (str) - The type of sanitizer used for testing (e.g., Address, Memory, Undefined)
bug_description: (str) - Description of the vulnerability
base_commit: (str) - The base commit hash where the… See the full description on the dataset page: https://huggingface.co/datasets/SEC-bench/SEC-bench.pulsefeed-x402-security
PulseFeed — x402 Agent-Payment Security & Trust (open data)
Independent, daily-updated trust & safety data for the x402 agent-payment economy (HTTP 402 + stablecoins on Base) and the MCP server ecosystem — by PulseFeed.
AI agents increasingly pay for APIs autonomously over x402 and connect to MCP servers that can run code on install. But 20% of listed x402 endpoints are dead or invalid, and "live" is not the same claim as "payable": of 33724 endpoints that return a valid 402… See the full description on the dataset page: https://huggingface.co/datasets/Nikolife/pulsefeed-x402-security.DTap-Bench-Agent-Trajectories
DecodingTrust-Agent Platform
A Controllable and Interactive Red-Teaming Platform for AI Agents.
This is the full collection of the agent trajectories produced from evaluating the DTap-Bench from DecodingTrust-Agent Platform (DTAP),
spanning 14 real-world domains and 50+ simulation environments that replicate widely-used
systems such as Google Workspace, PayPal, Slack, Salesforce, Snowflake, and Databricks. Each task
ships the configuration the evaluator needs to spin up the… See the full description on the dataset page: https://huggingface.co/datasets/AI-Secure/DTap-Bench-Agent-Trajectories.SEC
SEC Annual Reports (Form 10-K) 1993-2024
Dataset Overview
This dataset comprises SEC annual reports (Form 10-K) for the years 1993 to 2024, providing comprehensive coverage of publicly traded companies' financial and business information. The reports are stored in Parquet format, ensuring efficient storage and quick access. This dataset was meticulously compiled using the EDGAR-Crawler toolkit, which facilitates the extraction and processing of SEC filings from the EDGAR… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/SEC.cyber-security
Cybersecurity AI Knowledge Base — PhD-Level Dataset
Overview
This is the most comprehensive cybersecurity knowledge base ever assembled for AI training. It covers all domains of cybersecurity at PhD-level depth — from offensive red teaming and bug bounty exploitation to defensive SOC operations, digital forensics, and cutting-edge AI/LLM security.
Size: 16 GB | Files: 507 | Domains: 30+ | Sources: 15+ platforms
Purpose
Train the world's most… See the full description on the dataset page: https://huggingface.co/datasets/Vyber07/cyber-security.security_instruct_mcq_2481White-Hat-Security-Agent-Prompts-600K
White Hat Security Agent Prompts 600K
Overview
The White-Hat-Security-Agent-Prompts-600K dataset is a practitioner-perspective security prompts corpus of 596,295 richly contextualized queries, designed to represent how real-world defensive security professionals communicate, interrogate, and reason through active threat scenarios.
Where most security datasets catalogue CVEs, malware signatures, or CTF write-ups, this collection teaches models to operate from inside the… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/White-Hat-Security-Agent-Prompts-600K.404mini
404-GEN Mini 3D
This dataset contains over 20,000 3D assets generated with text prompts using 3D Gaussian Splatting, designed for text-to-3D generation tasks. This is a sample of a much larger dataset comprised of 21.5M assets and 40TB in size, available by request at https://dataset.404.xyz
Dataset Description
Dataset Summary
404-GEN Mini 3D is a collection of over 20,000 3D assets generated from text prompts on Bittensor Subnet 17, providing mid-… See the full description on the dataset page: https://huggingface.co/datasets/Olague-Secret/404mini.domains
Internet Domains
Domains
HuggingFace Hub Mirror for https://github.com/pkgforge-security/domains
The Sync Workflow actions are at: https://github.com/pkgforge-security/domains
TOS & Abuse (To Hugging-Face's Staff)
Hi, if you are an offical from Hugging-Face here to investigate why this Repo is so Large and are considering deleting, & terminating our Account.
Please note that, this project benefits a lot of people (You can do a… See the full description on the dataset page: https://huggingface.co/datasets/pkgforge-security/domains.SecQue
SECQUE
Paper
SECQUE is a comprehensive benchmark for evaluating large language models (LLMs) in financial analysis tasks.
SECQUE comprises 565 expert-written questions covering SEC filings analysis across four key categories:
comparison analysis
ratio calculation
risk assessment
financial insight generation.
To assess model performance, we develop SECQUE-Judge, an evaluation mechanism leveraging multiple LLM-based judges, which demonstrates strong alignment with human… See the full description on the dataset page: https://huggingface.co/datasets/nogabenyoash/SecQue.sec-material-contracts
Material Contracts (Exhibit 10) from SEC/EDGAR
Because sometimes you need 1,141,632 examples of corporate legalese to train your next model ☕
Dataset Summary
Picture this: 1,141,632 material contracts (Exhibit 10) painstakingly collected from sec.gov's EDGAR database. We're talking about legal agreements spanning from 1994 to 2025 Q1, sourced from 10-K, 10-Q, and 8-K filings. Think of Exhibit 10 as the treasure trove where companies hide their most important legal… See the full description on the dataset page: https://huggingface.co/datasets/chenghao/sec-material-contracts.second-measurements
Second Measurements
Public claims about logs, datasets and software supply chains, recomputed by a different path than the one that produced them. One row per check: what was claimed, how we re-measured it, what we found, and what the other party said.
Papers and evidence: https://markovianprotocol.com/measurements/ · Scripts: https://github.com/MarkovianProtocol/second-measurements
What is in it
measurements.jsonl has one row per check.
field
meaning… See the full description on the dataset page: https://huggingface.co/datasets/MarkovianProtocol/second-measurements.generated-passport-faces-aditya-second-halfsec-8k-events
SEC Form 8-K Corporate Events
Every Form 8-K filed since the modern item taxonomy took effect — and, for each
one, the second the SEC accepted it, which is not the date printed on it.
1 761 353 filings · 3 676 835 item-level events · 23 August 2004 to today
The pipeline lives in recipe/ at the same revision as the data.
See PIPELINE.md for the method.
The problem this dataset exists to solve
Apple filed its June-quarter results on 30 July 2026. Here is the filing… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/sec-8k-events.DecodingTrust-Agent-Platform
DecodingTrust-Agent Platform
A Controllable and Interactive Red-Teaming Platform for AI Agents.
This is the per-task dataset for the DecodingTrust-Agent Platform (DTAP),
spanning 14 real-world domains and 50+ simulation environments that replicate widely-used
systems such as Google Workspace, PayPal, Slack, Salesforce, Snowflake, and Databricks. Each task
ships the configuration the evaluator needs to spin up the sandbox, run an agent, and verify the
outcome — config.yaml (task… See the full description on the dataset page: https://huggingface.co/datasets/AI-Secure/DecodingTrust-Agent-Platform.humanitys-second-last-exam
Humanity's Second Last Exam
Benchmark design, curation and release maintenance: Shashank Agnihotri.
Original questions retain their recorded authorship and source attribution.
This owner-reviewed retained release contains 365 target questions, 730
context examples, and 365 ordered target/A/B links: 1,095 question rows.
The owner review concluded on 16 September 2026. This is an owner-reviewed
release after suspected-AI-content exclusions, not a software-certified
guarantee of… See the full description on the dataset page: https://huggingface.co/datasets/shashankskagnihotri/humanitys-second-last-exam.secret-scan-remediation-trajectories
Secret Scan Remediation Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/secret-scan-remediation-trajectories.comptia_security_pluse_701financial-reports-secThe dataset contains the annual report of US public firms filing with the SEC EDGAR system.
Each annual report (10K filing) is broken into 20 sections. Each section is split into individual sentences.
Sentiment labels are provided on a per filing basis from the market reaction around the filing data.
Additional metadata for each filing is included in the dataset.
