witfoo/precinct6-cybersecurity
WitFoo Precinct6 Cybersecurity Dataset Version 2.1.0 (built 2026-09-22). Regenerated to address feedback from the University of Canterbury PIDS evaluation: attacks and normal traffic now share a timeline, usernames are wired into the provenance graph, mis-parsed timestamps are repaired, and every number in this card is generated from the uploaded data. Overview A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC)… See the full description on the dataset page: https://huggingface.co/datasets/witfoo/precinct6-cybersecurity.
WitFoo Precinct6 Cybersecurity Dataset
Version 2.1.0 (built 2026-09-22). Regenerated to address feedback from the University of Canterbury PIDS evaluation: attacks and normal traffic now share a timeline, usernames are wired into the provenance graph, mis-parsed timestamps are repaired, and every number in this card is generated from the uploaded data.
Overview
A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC) data processed by WitFoo Precinct version 6.x. It contains 2,011,674 sanitized security events captured live from 4 organizations (2024-07-26 11:10:23 UTC to 2024-08-01 06:00:03 UTC), 51,371 incident provenance graphs with their 238,511 embedded triggering signals, per-incident GraphML files, natural-language attack reports, and a merged provenance graph (47,585 nodes, 1,740,438 edges).
Available in two sizes (same incidents, same methodology, same sanitization registry):
- `witfoo/precinct6-cybersecurity` — 2,011,674 live signals, incidents in context (this dataset)
- `witfoo/precinct6-cybersecurity-100m` — the full live capture
Generate your own: WitFoo Precinct 6.x customers can create datasets from their own data with the open-source pipeline `witfoo/dataset-from-precinct6`.
This dataset supports research in:
- Provenance graph-based intrusion detection (KnowHow, NodLink, and similar systems)
- AI-driven cyber defense simulation (CybORG and MARL-based defense policy training)
- Security alert classification (malicious vs. suspicious vs. benign event labeling)
- Attack lifecycle analysis using MITRE ATT&CK framework mappings
- Detection rule evaluation using WitFoo's 261 lead detection rules
What changed in v2
How this subset was drawn
This is the companion small dataset. It is a deterministic subset of the full live capture (`witfoo/precinct6-cybersecurity-100m`) built for incidents in context:
- every live row that is an incident lead (
malicious) — 7,728 rows; - every live row of the same organization within ±5 minutes of an in-window incident lead (the ordinary traffic around each incident, on the same hosts and network) — 1,606,249 rows;
- a uniform random fill across the whole capture window (fill probability 0.00371, seed 42) — 397,697 rows.
graph/incidents.jsonl, graph/incidents_graphml/ and graph/attack_reports.jsonl are identical to the full dataset. The signals table is subsampled, and the merged graph (graph/nodes.jsonl, graph/edges*.jsonl*, graph/graph.graphml) is rebuilt from the subsampled rows plus all incidents, so its live-derived nodes, edges and first_seen / last_seen / signal_count attributes describe the subset (counts in graph/metadata.json). Selection parameters are in build/subset_stats.json.
Versions
This is v2.1.0. It corrects v2.0.0, which stays available at the v2.0.0 tag:
Tokens are shared with v2.0.0 (this build reused its PII registry), apart from the incident org fields above and fewer than 200 values, mostly email addresses, whose registry entries were re-created and renumbered.
from datasets import load_dataset
signals = load_dataset("witfoo/precinct6-cybersecurity", "signals", split="train") # v2.1.0
signals_v2_0 = load_dataset("witfoo/precinct6-cybersecurity", "signals", split="train", revision="v2.0.0") # previous releaseThe 2026-05 release (v1) is not kept for download. Re-verifying it against the v2 tooling showed that some values had escaped sanitization — device and account names survived inside JSON-escaped Windows event text and in stream_name — so it was withdrawn rather than preserved at a tag. Its tokens are not comparable with v2: each used its own registry, so HOST-0042 there is a different machine from HOST-0042 here.
Quick Start
from datasets import load_dataset
REPO = "witfoo/precinct6-cybersecurity"
# Live capture, labeled in place (benign / suspicious / malicious share hosts and hours)
signals = load_dataset(REPO, "signals", split="train")
# In-window attacks with their surrounding traffic: filter on the incident ids
attacks = signals.filter(lambda x: x["label_binary"] == "malicious")
# Historical incident leads (2022–2024) — a separate timeline, same columns
incident_signals = load_dataset(REPO, "incident_signals", split="train")
# Provenance graph: hosts + credentials, host->host event edges, user->host edges, incident links
nodes = load_dataset(REPO, "graph_nodes", split="train")
edges = load_dataset(REPO, "graph_edges", split="train")
# Deterministic attack reports (one per incident)
reports = load_dataset(REPO, "attack_reports", split="train")
# Full incident graphs (nested dicts keyed by uuid; not a typed config)
import pandas as pd
incidents = pd.read_json("hf://datasets/" + REPO + "/graph/incidents.jsonl", lines=True)Join an incident lead to its live row: incident_signals.artifact_id == signals.artifact_id (both are the Precinct artifact timeuuid). Rows of signals that are leads carry the incident ids in incident_ids.
Temporal coverage
`signals` (live capture, `origin = live`) — 2,011,674 rows, 2024-07-26 11:10:23 UTC → 2024-08-01 06:00:03 UTC
`incident_signals` (embedded incident leads, `origin = incident_lead`) — 238,511 rows, 2022-05-30 14:43:45 UTC → 2024-07-28 16:39:47 UTC
The live capture is the complete artifact retention window of the archived Precinct cluster. Coverage is not uniform across organizations or days: check the per-organization table below, the per-label time ranges in signals/metadata.json, and the per-organization hourly histogram in build/label_stats.json before assuming a continuous capture.
Organizations
Incidents in context
1584 incidents have at least one triggering signal that was found among the live rows; those 7,728 lead artifacts are labeled malicious in place in signals (they are not duplicated in incident_signals). Leads of 49,797 incidents were not found among the live rows (their incidents pre-date the capture, or the live artifact was not retained) and live in incident_signals.
Label distribution
Across both tables (2,250,185 rows):
Disposition of malicious rows (raw Precinct incident status, see Ground Truth):
signals:
incident_signals:
Signal columns
Both signal tables share one schema (38 columns).
Graph data
- Node ids. Public IPs are global node ids; private IPs, hostnames and credentials are scoped by organization (
ORG-0004/10.44.0.7,ORG-0004/user:USER-0007) because the same private address or account name exists in several customer networks. Incident host and credential nodes are mapped onto the same ids (their Precinct uuids are kept inattrs.precinct_node_ids). Every node carries the same attribute keys;first_seen/last_seen/signal_countOther incident node types (SERVICE,FILE,ACTOR, ...) keep their Precinct uuid as node id. come from live signals only,incident_first_observed/incident_last_observedfrom incident membership. - Edges from signals carry the signal's labels,
attrs.origin,attrs.org_idandattrs.artifact_id; edgetypeis derived from the message type (NETWORK_FLOW,LOGON,AUDIT_EVENT, ...). A signal with a username adds aUSER_ACTIONedge from the credential node to the accessed host (attrs.host_rolesays whether that was the destination or the reporting host). - Edges from incidents are
INCIDENT_LINKwith the incident's labels;timestampis the edge's own start time. - Edge types:
USER_ACTION(692,375),AUDIT_EVENT(336,425),INCIDENT_LINK(335,645),NETWORK_FLOW(323,376),EVENT(45,063),DNS_RESOLVE(7,554). graph/graph.graphmlholds the whole merged graph (streaming GraphML).
Attack reports
graph/attack_reports.jsonl holds one natural-language threat-hunting report per incident (51,371), deterministically composed from the incident's structured metadata (modus operandi, set roles, lead descriptions, MITRE mappings, timestamps). Each report states that it reflects Precinct's automated correlation output, not an independent investigation. Derivation: `src/precinct6_dataset/attack_reports.py`.
Files
signals/signals.parquet— live signalssignals/incident_signals.parquet— embedded incident leadssignals/metadata.json— exact counts, per-label time ranges, org/stream/message-type distributionsgraph/nodes.jsonl,graph/edges.jsonl— merged provenance graph (NDJSON)graph/incidents.jsonl— full sanitized incident records with embeddednodes,edges,leads(dicts keyed by Precinct uuids, so this file is not exposed as aload_datasetconfig; read it withpandas.read_json(lines=True))graph/incidents_graphml/<x>/<incident_id>.graphml— one GraphML per incident (sharded by first hex character)graph/attack_reports.jsonl— attack reportsgraph/metadata.json— graph counts and node id schemereference/lead_rules_catalog.json— 261 lead detection rules, 158 products, 106 classification sets
Labeling methodology
Three-tier labels:
- `malicious` — the event is a triggering signal (lead) of a Precinct incident. In
signalsthese are live rows joined to their incident onartifact_id; inincident_signalsthey are the embedded copies of leads whose live rows fall outside the capture. - `suspicious` — the event matched one or more of WitFoo's 261 lead detection rules but is not a lead of any incident.
- `benign` — no rule matched and the event is not part of any incident.
A lead that belongs to several incidents is one row whose incident_ids lists them all; its mo_name, disposition and suspicion_score come from the highest-suspicion incident.
Ground truth and disposition
All labels derive from WitFoo Precinct's automated incident correlation engine — there is no independent, analyst-verified ground truth. Treat Precinct as a strong but imperfect oracle. disposition is the parent incident's Precinct status, and most statuses are set by the engine, not by an analyst:
Analyst-set statuses are rare, so disposition is not a usable ground-truth signal on its own; check the disposition_distribution in signal/metadata.json for how many rows carry each status.
Scoring
- `suspicion_score` — Precinct's proprietary score of the parent incident (0–1). Zero for benign and suspicious.
- `label_confidence` — how much corroborating evidence supports the tier (not a probability of maliciousness):
MITRE ATT&CK mappings
Tactics and techniques are derived from (1) WitFoo set role names on the incident, (2) the incident's modus operandi, and (3) per-product framework data embedded in incident.nodes.products.frameworks, deduplicated. They are priors, not analyst-confirmed per-event attributions. Mapping tables: `src/precinct6_dataset/mitre_mapping.py`.
Source products
The events in this build come from 19 security products across 12 vendors (exact counts in signals/metadata.json under product_distribution / vendor_distribution).
Most frequent products: AWS Instance Backup, Windows Active Directory, Windows Logs, VMWare VCenter, ASA Firewall, AWS VPC Security, Barracuda WAF, Linux PAM, ManageEngine ADManager, Graph, Barracuda ESS, Falcon, Cisco Network Operating System, Apache Web Server. Vendors: Microsoft, Amazon Web Services, VMWare, Cisco, Barracuda, Linux, ManageEngine, Crowdstrike, Apache, Symantec, SentinelOne, WitFoo.
The generator's rule catalog (reference/lead_rules_catalog.json) covers a much wider set — 158 products across firewalls, endpoint protection, network detection, identity, cloud, email security and infrastructure — because it is shared by every deployment; only the products above actually appear in this capture.
Top streams in this build: aws_cloudtrail_events (543,159), microsoft-windows-security-auditing (421,587), windows_security_audit (336,425), vcenter (266,201), java_stack_trace (154,465), cisco_asa (95,143), no_useful_info (48,246), aws_cloud_trail (34,484), dnsmasq (34,110), aws_vpc_flow_log (19,219).
Sanitization
All customer-identifying information was removed with the open-source four-layer pipeline (`witfoo/dataset-from-precinct6`):
- Structured field sanitization + Aho-Corasick multi-pattern sweep — deterministic tokens (public IPs → RFC 5737 TEST-NET, private IPs → HMAC-remapped RFC 1918, hostnames →
HOST-NNNN, accounts →USER-NNNN, organizations →ORG-NNNN, emails →user-NNNN@example.net, SIDs, AWS accounts/ARNs, machine accounts), then a sweep over every string field with the full registry. Record identifiers (artifact/incident uuids) are protected from the sweep. - Format-specific log parsing — Cisco ASA, Windows Security XML, WinLogBeat, AWS CloudTrail, Palo Alto, VMware vCenter, DNS, and a generic fallback.
- ML residual detection — Microsoft Presidio (spaCy) and BERT NER on a stratified sample; findings trigger full re-sanitization.
- LLM contextual review — sampled review by the local WitQ model (ran for this build: 1,500 records reviewed by
witq, 2 additional values registered). The ML layer (3) sampled 3,000 records and registered 470 additional values.
The same original value always maps to the same token across both signal tables, the incidents and the graph, so topology and identity are preserved. Both dataset sizes were produced from one registry, so tokens agree between them. The registry for this build holds 85,088 mappings.
Research context
Produced in collaboration with the University of Canterbury (New Zealand) Computer Science and Software Engineering department for two research projects: an AI cyber-security battle simulator (improving CybORG with realistic IDS observations and graph-based defense policies) and intrusion detection based on provenance graphs (evaluating KnowHow, NodLink and similar PIDS).
Limitations
- Label imbalance reflects production SOC reality; sample accordingly.
- Temporal scope: the live capture spans 2024-07-26 11:10:23 UTC to 2024-08-01 06:00:03 UTC with uneven per-organization coverage; incident leads span 2022-05-30 14:43:45 UTC to 2024-07-28 16:39:47 UTC.
- Ground truth: labels are Precinct's automated correlation; stratify with
disposition. - Sanitization trade-offs: some free-text detail is reduced by PII replacement.
- Tokens differ from v1: the v2 registry was rebuilt, so
USER-NNNN/HOST-NNNN/ORG-NNNNvalues do not correspond to v1 values.
Citation
@dataset{witfoo_precinct6_2026,
title={WitFoo Precinct6 Cybersecurity Dataset},
author={WitFoo, Inc.},
year={2026},
version={2.1.0},
url={https://huggingface.co/datasets/witfoo/precinct6-cybersecurity},
license={Apache-2.0}
}