Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Lakera /b3-agent-security-benchmark-weak[paper] [blogpost] [game] b3 AI Security Benchmark: Breaking Agent Backbones Highly contextalized prompt injections crowd-sourced during the Agent Breaker Challenge. This is a low-quality version of the data behind Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents. The high quality dataset was used to evaluate the security of more than 30 LLMs. Dataset Summary Purpose: This dataset contains crowdsourced adversarial attacks… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/b3-agent-security-benchmark-weak.tabulartext-classificationn<1K9 likes1.3k downloads3d agoHugging Face02tickertruthorg /nse-india-security-master TickerTruth — NSE India Security Master (Explorer) A clean, normalized reference table of 2,389 NSE-listed equities — ISIN mappings, listing dates, company names, and active/delisted status — built from the TickerTruth India reference-data pipeline. Why this dataset exists India equity data is notoriously messy. NSE symbols change (renames, mergers, delistings), ISINs get reissued, and raw bhavcopy files carry no historical context. TickerTruth's pipeline… See the full description on the dataset page: https://huggingface.co/datasets/tickertruthorg/nse-india-security-master.textother1K<n<10K0 likes344 downloads3mo agoHugging Face03tumeteor /Security-TTP-Mapping The Security Attack Pattern (TTP) Recognition or Mapping Task We share in this repo the MITRE ATT&CK mapping datasets, with training, validation and test splits. The datasets can be considered as an emerging and challenging multilabel classification NLP task, with over 600 hierarchical classes. NOTE: due to their security nature, these datasets contain textual information about malware and other security aspects. Datasets TRAM This dataset belongs to CTID… See the full description on the dataset page: https://huggingface.co/datasets/tumeteor/Security-TTP-Mapping.texttext-classification10K<n<100K31 likes277 downloads3y agoHugging Face04tickertruth /nse-india-security-master TickerTruth — NSE India Security Master (Explorer) A clean, normalized reference table of 2,389 NSE-listed equities — ISIN mappings, listing dates, company names, and active/delisted status — built from the TickerTruth India reference-data pipeline. Why this dataset exists India equity data is notoriously messy. NSE symbols change (renames, mergers, delistings), ISINs get reissued, and raw bhavcopy files carry no historical context. TickerTruth's pipeline… See the full description on the dataset page: https://huggingface.co/datasets/tickertruth/nse-india-security-master.textother1K<n<10K0 likes167 downloads4mo agoHugging Face05ScortonAI /Cyber-Security-Breachestabular1K<n<10K12 likes85 downloads3y agoHugging Face06ArkhAngelLifeJiggy /Security-TTP-Mapping The Security Attack Pattern (TTP) Recognition or Mapping Task We share in this repo the MITRE ATT&CK mapping datasets, with training, validation and test splits. The datasets can be considered as an emerging and challenging multilabel classification NLP task, with over 600 hierarchical classes. NOTE: due to their security nature, these datasets contain textual information about malware and other security aspects. Datasets TRAM This dataset… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/Security-TTP-Mapping.texttext-classification10K<n<100K0 likes69 downloads10d agoHugging Face07tapsin /Securitytextn<1K0 likes55 downloads9d agoHugging Face08rogue-security /real-world-benign-use-cases Real-World Benign Use Cases A curated set of 178 real-world, 100%-benign examples (label == 0 for every row) pulled from production AI-coding-agent traffic — chat messages, tool output, shell commands, code snippets — built specifically to stress-test prompt-injection / jailbreak classifiers for false positives. Every row was independently judged benign with high confidence before inclusion. This is not a random sample of production traffic: rows were preferentially drawn from… See the full description on the dataset page: https://huggingface.co/datasets/rogue-security/real-world-benign-use-cases.tabulartext-classificationn<1K0 likes51 downloads2mo agoHugging Face09CodeNinjatools /ot-security-cip-evidence-us-ontology OT security and NERC CIP evidence ontology and model register for an electric utility The object model and the model and equipment register from Grid Context Watch: One system of context for OT Security and NERC CIP Evidence, an open reference architecture by CodeNinja for United States. Part of the Vertical-Driven Architectures series; every design in the series is also a row in the cumulative dataset https://huggingface.co/datasets/CodeNinjatools/vertical-driven-architectures.… See the full description on the dataset page: https://huggingface.co/datasets/CodeNinjatools/ot-security-cip-evidence-us-ontology.textn<1K0 likes49 downloads4d agoHugging Face10joylarkin /openclaw-security-newsSource Repo: https://github.com/joylarkin/openclaw-security-newsSource Feed: https://raw.githubusercontent.com/joylarkin/openclaw-security-news/main/feed.xml Description: OpenClaw Security News for AI builders, developers, and investors. Featuring global government warnings about OpenClaw, OpenClaw news headlines, and OpenClaw security vendor advisories. Last Update: 22 May 2026 textn<1K1 likes45 downloads5mo agoHugging Face11sublime-security /babbelphish BabbelPhish BabbelPhish is a dataset based on the Sublime Security Message Query Language (MQL) used for email security detection engineering. This dataset is specially created for the BabbelPhish project, which focuses on leveraging large language models to facilitate the work of detection engineers. This dataset comprises around 3,000 examples drawn from various sources. We've utilized the following: Sublime Security Documentation Message Data Model (Schema) Sublime Rules Repo… See the full description on the dataset page: https://huggingface.co/datasets/sublime-security/babbelphish.tabulartranslation1K<n<10K1 likes30 downloads3y agoHugging Face12patcon /japanchoice-2025-foreign-policy-and-securityFetched: 2026-02-16-2245 Data was gathered using the Polis software (see: compdemocracy.org/polis and github.com/compdemocracy/polis), and acquired from this URL: https://polis.japanchoice.jp/7cdcyjmsyh comments.csv was augmented with is-seed and is-meta columns fetched from https://polis.japanchoice.jp/api/v3/comments?conversation_id=7cdcyjmsyh&moderation=true&include_voting_patterns=true tabular10K<n<100K0 likes30 downloads8mo agoHugging Face13xtb20 /SecurityToolUsetext100K<n<1M0 likes30 downloads3d agoHugging Face14Shaharb /us-social-security-medicare-FAQs-testtabularn<1K0 likes21 downloads3y agoHugging Face15patcon /japanchoice-2026-foreign-policy-and-securityExported: 2026-02-20 This dataset covers the topic of Foreign Policy and Security (外交・安全保障) from the Japan Choice Polis platform. The conversation can be found at: https://polis.japanchoice.jp/3x3ancmrtc Data was exported from the database and provided by Nishio Hirokazu. It was gathered using the Polis software by The Computational Democracy Project (see: compdemocracy.org/polis and github.com/compdemocracy/polis). tabular10K<n<100K0 likes21 downloads8mo agoHugging Face16ClarusC64 /network-security-route-hijack-coherence-risk-v0.1What this repo is for Detect routing security incidents fast. Covers: unexpected origin ASN invalid ROAs AS-path anomalies suspicious more-specific prefixes observed traffic diversion whether mitigation happened This is high-impact because one leak can break many networks. texttext-classificationn<1K0 likes20 downloads8mo agoHugging Face17Kaballas /security_contenttext1K<n<10K1 likes19 downloads2y agoHugging Face18itayhf /security_steerability Security Steerability & the VeganRibs Benchmark Security steerability is defined as an LLM's ability to stick to the specific rules and boundaries set by a system prompt, particularly for content that isn't typically considered prohibited. To evaluate this, we developed the VeganRibs benchmark. The benchmark tests an LLM's skill at handling conflicts by seeing if it can follow system-level instructions even when a user's input tries to contradict them. VeganRibs works by presenting… See the full description on the dataset page: https://huggingface.co/datasets/itayhf/security_steerability.textn<1K3 likes18 downloads1y agoHugging Face19Shoriful025 /zero_trust_network_security_logstabularn<1K1 likes16 downloads9mo agoHugging Face20AlexSo79 /US_Social_Security_Medicare_FAQs_Sampletabularn<1K0 likes14 downloads3y agoHugging Face21loocorez /security_analysistabularn<1K0 likes14 downloads1y agoHugging Face22stu8king /securityincidentstabular1K<n<10K0 likes11 downloads2y agoHugging Face23fenar /iot-securityIoT Data Collection This document provides detailed information about the various data fields collected by the IoT Device with Multiple Sensors. Each field is described along with the type of data it includes. Data Fields ALS (Ambient Light Sensor) Explanation: Measures the ambient light levels around the thermostat. Data Type: Intensity of light in lux (lumens per square meter). PIR (Passive Infrared Sensor) Explanation: Detects motion by measuring the infrared (heat) emitted by objects in… See the full description on the dataset page: https://huggingface.co/datasets/fenar/iot-security.tabular100K<n<1M2 likes11 downloads2y agoHugging Face24boomer-rogers /social_security_embeddingstabularn<1K1 likes10 downloads3y agoHugging Face25rsh-raj /asss-securitytext1K<n<10K0 likes10 downloads1y agoHugging Face26Perfectyash /Human-Like-Reasoning-Security-Scenarios Human-Like Reasoning Security Scenarios A handcrafted dataset of realistic cybersecurity, system, and network scenarios focused on human-style reasoning, not just final answers. Why this dataset is different Includes human thought processes Covers common mistakes Explains correct decisions Designed for reasoning-based AI models Domains Cybersecurity Systems Networking Use Cases LLM fine-tuning AI agent training Cybersecurity education… See the full description on the dataset page: https://huggingface.co/datasets/Perfectyash/Human-Like-Reasoning-Security-Scenarios.textn<1K0 likes10 downloads9mo agoHugging Face27miomit /playwright_security_shell_generationtext1K<n<10K0 likes9 downloads7mo agoHugging Face28AidanFerrara /security_domain_knowlegetabular1K<n<10K1 likes8 downloads2y agoHugging Face29electricsheepafrica /africa-synth-urbanization-housing-tenure-security-africa-all Africa Synth Urbanization Housing Tenure Security Africa All | Africa (Electric Sheep Africa metadata inventory) Size category: 10K<n<100K - Formats: csv - Sector: governance_security - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-urbanization-housing-tenure-security-africa-all.tabulartabular-classification10K<n<100K0 likes8 downloads2mo agoHugging Face30AidanFerrara /security_finetunetext10K<n<100K0 likes7 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.