Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01secemp9 /arxiv-complete arXiv Complete Corpus A snapshot of arXiv's metadata, version history, submission files and rendered documents. It covers 3,148,796 papers and includes file contents, paths, sizes and SHA-256 digests. Metadata comes from arXiv's OAI-PMH arXivRaw interface; files come from the GCS mirror, S3 source archives and direct PDF fetches. This release holds a PDF for 99.47% of papers and 99.54% of versions reported with a non-zero submission size. It is a one-off snapshot; coverage gaps… See the full description on the dataset page: https://huggingface.co/datasets/secemp9/arxiv-complete.tabulartext-generation100M<n<1B656 likes171k downloads21d agoHugging Face02AlphaDojo /dojo_sector_precomputed Languages: 简体中文 · English dojo_sector_precomputed — Precomputed Sector Analytics Overview Derived sector analytics: L3 constituent snapshots, daily cap-weighted sector index levels, and per-constituent daily returns. Built offline from taxonomy, mappings, quotes, and stock K-lines. Files File Description manifest.json Generation metadata: version, window start, row counts, latest trade dates constituents.parquet L3 constituent… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_sector_precomputed.tabular100K<n<1M0 likes25k downloads1d agoHugging Face03jedisct1 /security-auditsA collection of agent traces generated with Swival (not Claude Code, despite what the HF interface currently shows), an agent designed for open-source models. These traces focus on security audits of opensource software. Sharing traces with Swival Swival can export full conversation traces with --trace-dir, which writes one <session_id>.jsonl file per session: swival "Fix the login bug" --trace-dir traces/ Those JSONL files use Swival's Claude Code compatible trace export, and… See the full description on the dataset page: https://huggingface.co/datasets/jedisct1/security-audits.tabulartext-generation10K<n<100K18 likes17k downloads4mo agoHugging Face04AlphaDojo /dojo_sector_info Languages: 简体中文 · English dojo_sector_info — Sector Taxonomy Overview Three-level sector taxonomy (L1 industry → L2 chain → L3 sub-sector) with bilingual Chinese/English names and definitions. Files File Description data.parquet Taxonomy tree (one L1 row each; L2/L3 nested in children) Key Fields Field Description id L1 sector ID name / name_alias L1 English name / Chinese alias description /… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_sector_info.tabularn<1K0 likes16k downloads24d agoHugging Face05AlphaDojo /dojo_sector_symbol_relations Languages: 简体中文 · English dojo_sector_symbol_relations — Stock–Sector Mapping Overview Maps each stock to L1/L2/L3 sector paths with primary and secondary assignments. One row per (ticker, market) pair. Files File Description data.parquet Full stock ↔ sector relations Key Fields Field Description ticker Stock symbol market us, cn, or hk primary JSON object — primary sector path secondary JSON array —… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_sector_symbol_relations.text10K<n<100K0 likes16k downloads24d agoHugging Face06TeraflopAI /SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset. The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database. The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/TeraflopAI/SEC-EDGAR.texttext-generation1M<n<10M48 likes13k downloads6mo agoHugging Face07zefang-liu /secqa SecQA SecQA is a specialized dataset created for the evaluation of Large Language Models (LLMs) in the domain of computer security. It consists of multiple-choice questions, generated using GPT-4 and the Computer Systems Security: Planning for Success textbook, aimed at assessing the understanding and application of LLMs' knowledge in computer security. Dataset Details Dataset Description SecQA is an innovative dataset designed to benchmark the… See the full description on the dataset page: https://huggingface.co/datasets/zefang-liu/secqa.textmultiple-choicen<1K12 likes10k downloads2y agoHugging Face08Nicogarrr /cavaai-sec-mirror0 likes6.2k downloads5h agoHugging Face09yonathanarbel /SEC_Exhibit_10 SEC Exhibit 10 Material Contracts Dataset (2001-2024) Overview This dataset contains material contracts (Exhibit 10) from SEC filings between 2001 and 2024, representing the largest collection of commercial contracts to date. It includes both original filings and text-converted versions. If you find this dataset useful, please cite: Arbel, Yonathan A., The Readability of Contracts: Big Data Analysis, Forthcoming J. Empirical Legal Stud. 21:4 (Dec. 2024) Methodological… See the full description on the dataset page: https://huggingface.co/datasets/yonathanarbel/SEC_Exhibit_10.text1 likes5.6k downloads2y agoHugging Face10TomTBT /pmc_open_access_sectiontext1M<n<10M3 likes5.1k downloads2y agoHugging Face11jiofidelus /SecuTable SecuTable: A Dataset for Semantic Table Interpretation in Security Domain Dataset Overview Security datasets are scattered on the Internet (CVE, CAPEC, CWE, etc.) and provided in CSV, JSON or XML formats. This makes it difficult to get a holistic view of the interconnectedness of information across different data sources. On the other hand, many datasets focus on specific attack vectors or limited environments, limiting generalisability. There is a lack of detailed… See the full description on the dataset page: https://huggingface.co/datasets/jiofidelus/SecuTable.2 likes3.4k downloads11mo agoHugging Face12SEC-bench /SEC-bench Data Instances instance_id: (str) - A unique identifier for the instance repo: (str) - The repository name including the owner project_name: (str) - The name of the project without owner lang: (str) - The programming language of the repository work_dir: (str) - Working directory path sanitizer: (str) - The type of sanitizer used for testing (e.g., Address, Memory, Undefined) bug_description: (str) - Description of the vulnerability base_commit: (str) - The base commit hash where the… See the full description on the dataset page: https://huggingface.co/datasets/SEC-bench/SEC-bench.textn<1K10 likes3.1k downloads11mo agoHugging Face13Nikolife /pulsefeed-x402-security PulseFeed — x402 Agent-Payment Security & Trust (open data) Independent, daily-updated trust & safety data for the x402 agent-payment economy (HTTP 402 + stablecoins on Base) and the MCP server ecosystem — by PulseFeed. AI agents increasingly pay for APIs autonomously over x402 and connect to MCP servers that can run code on install. But 20% of listed x402 endpoints are dead or invalid, and "live" is not the same claim as "payable": of 33724 endpoints that return a valid 402… See the full description on the dataset page: https://huggingface.co/datasets/Nikolife/pulsefeed-x402-security.10K<n<100K0 likes2.8k downloads21h agoHugging Face14AI-Secure /DTap-Bench-Agent-Trajectories DecodingTrust-Agent Platform A Controllable and Interactive Red-Teaming Platform for AI Agents. This is the full collection of the agent trajectories produced from evaluating the DTap-Bench from DecodingTrust-Agent Platform (DTAP), spanning 14 real-world domains and 50+ simulation environments that replicate widely-used systems such as Google Workspace, PayPal, Slack, Salesforce, Snowflake, and Databricks. Each task ships the configuration the evaluator needs to spin up the… See the full description on the dataset page: https://huggingface.co/datasets/AI-Secure/DTap-Bench-Agent-Trajectories.text-generation1K<n<10K4 likes2.6k downloads3mo agoHugging Face15PleIAs /SEC SEC Annual Reports (Form 10-K) 1993-2024 Dataset Overview This dataset comprises SEC annual reports (Form 10-K) for the years 1993 to 2024, providing comprehensive coverage of publicly traded companies' financial and business information. The reports are stored in Parquet format, ensuring efficient storage and quick access. This dataset was meticulously compiled using the EDGAR-Crawler toolkit, which facilitates the extraction and processing of SEC filings from the EDGAR… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/SEC.tabulartext-generation100K<n<1M13 likes2.6k downloads2y agoHugging Face16Vyber07 /cyber-securitygated Cybersecurity AI Knowledge Base — PhD-Level Dataset Overview This is the most comprehensive cybersecurity knowledge base ever assembled for AI training. It covers all domains of cybersecurity at PhD-level depth — from offensive red teaming and bug bounty exploitation to defensive SOC operations, digital forensics, and cutting-edge AI/LLM security. Size: 16 GB | Files: 507 | Domains: 30+ | Sources: 15+ platforms Purpose Train the world's most… See the full description on the dataset page: https://huggingface.co/datasets/Vyber07/cyber-security.texttext-generationn<1K133 likes2.3k downloads1mo agoHugging Face17gussieIsASuccessfulWarlock /security_instruct_mcq_2481textn<1K0 likes2.2k downloads2y agoHugging Face18yatin-superintelligence /White-Hat-Security-Agent-Prompts-600K White Hat Security Agent Prompts 600K Overview The White-Hat-Security-Agent-Prompts-600K dataset is a practitioner-perspective security prompts corpus of 596,295 richly contextualized queries, designed to represent how real-world defensive security professionals communicate, interrogate, and reason through active threat scenarios. Where most security datasets catalogue CVEs, malware signatures, or CTF write-ups, this collection teaches models to operate from inside the… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/White-Hat-Security-Agent-Prompts-600K.texttext-generation100K<n<1M21 likes2.1k downloads7mo agoHugging Face19Olague-Secret /404mini 404-GEN Mini 3D This dataset contains over 20,000 3D assets generated with text prompts using 3D Gaussian Splatting, designed for text-to-3D generation tasks. This is a sample of a much larger dataset comprised of 21.5M assets and 40TB in size, available by request at https://dataset.404.xyz Dataset Description Dataset Summary 404-GEN Mini 3D is a collection of over 20,000 3D assets generated from text prompts on Bittensor Subnet 17, providing mid-… See the full description on the dataset page: https://huggingface.co/datasets/Olague-Secret/404mini.imagetext-to-3dn<1K0 likes2k downloads4mo agoHugging Face20pkgforge-security /domains Internet Domains Domains HuggingFace Hub Mirror for https://github.com/pkgforge-security/domains The Sync Workflow actions are at: https://github.com/pkgforge-security/domains TOS & Abuse (To Hugging-Face's Staff) Hi, if you are an offical from Hugging-Face here to investigate why this Repo is so Large and are considering deleting, & terminating our Account. Please note that, this project benefits a lot of people (You can do a… See the full description on the dataset page: https://huggingface.co/datasets/pkgforge-security/domains.text10B<n<100B5 likes2k downloads2mo agoHugging Face21nogabenyoash /SecQue SECQUE Paper SECQUE is a comprehensive benchmark for evaluating large language models (LLMs) in financial analysis tasks. SECQUE comprises 565 expert-written questions covering SEC filings analysis across four key categories: comparison analysis ratio calculation risk assessment financial insight generation. To assess model performance, we develop SECQUE-Judge, an evaluation mechanism leveraging multiple LLM-based judges, which demonstrates strong alignment with human… See the full description on the dataset page: https://huggingface.co/datasets/nogabenyoash/SecQue.textquestion-answeringn<1K4 likes1.8k downloads1y agoHugging Face22chenghao /sec-material-contracts Material Contracts (Exhibit 10) from SEC/EDGAR Because sometimes you need 1,141,632 examples of corporate legalese to train your next model ☕ Dataset Summary Picture this: 1,141,632 material contracts (Exhibit 10) painstakingly collected from sec.gov's EDGAR database. We're talking about legal agreements spanning from 1994 to 2025 Q1, sourced from 10-K, 10-Q, and 8-K filings. Think of Exhibit 10 as the treasure trove where companies hide their most important legal… See the full description on the dataset page: https://huggingface.co/datasets/chenghao/sec-material-contracts.texttext-generation1M<n<10M3 likes1.8k downloads1y agoHugging Face23MarkovianProtocol /second-measurements Second Measurements Public claims about logs, datasets and software supply chains, recomputed by a different path than the one that produced them. One row per check: what was claimed, how we re-measured it, what we found, and what the other party said. Papers and evidence: https://markovianprotocol.com/measurements/ · Scripts: https://github.com/MarkovianProtocol/second-measurements What is in it measurements.jsonl has one row per check. field meaning… See the full description on the dataset page: https://huggingface.co/datasets/MarkovianProtocol/second-measurements.textothern<1K0 likes1.8k downloads7d agoHugging Face24abhifdsdf /generated-passport-faces-aditya-second-halfimage1K<n<10K0 likes1.8k downloads10mo agoHugging Face25ZipLime /sec-8k-events SEC Form 8-K Corporate Events Every Form 8-K filed since the modern item taxonomy took effect — and, for each one, the second the SEC accepted it, which is not the date printed on it. 1 761 353 filings · 3 676 835 item-level events · 23 August 2004 to today The pipeline lives in recipe/ at the same revision as the data. See PIPELINE.md for the method. The problem this dataset exists to solve Apple filed its June-quarter results on 30 July 2026. Here is the filing… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/sec-8k-events.tabulartext-classification1M<n<10M1 likes1.6k downloads17h agoHugging Face26AI-Secure /DecodingTrust-Agent-Platform DecodingTrust-Agent Platform A Controllable and Interactive Red-Teaming Platform for AI Agents. This is the per-task dataset for the DecodingTrust-Agent Platform (DTAP), spanning 14 real-world domains and 50+ simulation environments that replicate widely-used systems such as Google Workspace, PayPal, Slack, Salesforce, Snowflake, and Databricks. Each task ships the configuration the evaluator needs to spin up the sandbox, run an agent, and verify the outcome — config.yaml (task… See the full description on the dataset page: https://huggingface.co/datasets/AI-Secure/DecodingTrust-Agent-Platform.text-generation1K<n<10K1 likes1.5k downloads4mo agoHugging Face27shashankskagnihotri /humanitys-second-last-exam Humanity's Second Last Exam Benchmark design, curation and release maintenance: Shashank Agnihotri. Original questions retain their recorded authorship and source attribution. This owner-reviewed retained release contains 365 target questions, 730 context examples, and 365 ordered target/A/B links: 1,095 question rows. The owner review concluded on 16 September 2026. This is an owner-reviewed release after suspected-AI-content exclusions, not a software-certified guarantee of… See the full description on the dataset page: https://huggingface.co/datasets/shashankskagnihotri/humanitys-second-last-exam.image1K<n<10K2 likes1.5k downloads22d agoHugging Face28rmems /secret-scan-remediation-trajectories Secret Scan Remediation Trajectories Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated payload is now published under data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/secret-scan-remediation-trajectories.0 likes1.4k downloads19d agoHugging Face29kala185 /comptia_security_pluse_701text1K<n<10K0 likes1.4k downloads1y agoHugging Face30JanosAudran /financial-reports-secThe dataset contains the annual report of US public firms filing with the SEC EDGAR system. Each annual report (10K filing) is broken into 20 sections. Each section is split into individual sentences. Sentiment labels are provided on a per filing basis from the market reaction around the filing data. Additional metadata for each filing is included in the dataset.tabularfill-mask10M<n<100M77 likes1.4k downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.