The dataset matters because realistic, labeled SOC telemetry plus provenance graphs are rare: this release pairs in-place live signals with per-incident provenance so models see attacks in their operational context rather than isolated leads.
What Sets It Apart
- Combined live capture and incident provenance: a merged graph with ~47.6k nodes and 1.74M edges links hosts, credentials and events so graph-based detectors and GNNs can train on authentic topology and temporal context.
- Deterministic, reproducible subset design: every live incident lead in-window is included, surrounding ±5-minute traffic is retained, and a seeded random fill ties the small and large builds to the same sanitization registry for consistent tokens.
- Rich labels and metadata for ML workflows: three-tier labels (malicious/suspicious/benign) derived from 261 lead rules, MITRE ATT&CK tactic/technique fields, per-incident suspicion scores, matched rule lists, and machine-readable GraphML/NDJSON and Parquet exports.
- Transparent limitations documented: labels come from WitFoo Precinct’s automated correlation engine (not independent analyst verification), heavy class imbalance (malicious rows are a small fraction), and sanitization replaces PII which reduces some free-text fidelity.
Who it fits and trade-offs
Great fit if you need realistic SOC telemetry with linked provenance for training or evaluating graph-based intrusion detection systems, AI-driven cyber defense simulators (CybORG/MARL), or alert-classification models that leverage MITRE mappings. Look elsewhere if you require analyst-verified ground truth for every incident, continuous long-duration captures beyond the documented windows, or raw unsanitized logs (this release applies extensive PII removal).