AIAny
Icon for item

WitFoo Precinct6 Cybersecurity Dataset

Provides a sanitized, labeled SOC capture and merged provenance graph for intrusion-detection research, including 2,011,674 live signals, 51,371 incident graphs, MITRE ATT&CK mappings, and deterministic attack reports.

Introduction

The dataset matters because realistic, labeled SOC telemetry plus provenance graphs are rare: this release pairs in-place live signals with per-incident provenance so models see attacks in their operational context rather than isolated leads.

What Sets It Apart
  • Combined live capture and incident provenance: a merged graph with ~47.6k nodes and 1.74M edges links hosts, credentials and events so graph-based detectors and GNNs can train on authentic topology and temporal context.
  • Deterministic, reproducible subset design: every live incident lead in-window is included, surrounding ±5-minute traffic is retained, and a seeded random fill ties the small and large builds to the same sanitization registry for consistent tokens.
  • Rich labels and metadata for ML workflows: three-tier labels (malicious/suspicious/benign) derived from 261 lead rules, MITRE ATT&CK tactic/technique fields, per-incident suspicion scores, matched rule lists, and machine-readable GraphML/NDJSON and Parquet exports.
  • Transparent limitations documented: labels come from WitFoo Precinct’s automated correlation engine (not independent analyst verification), heavy class imbalance (malicious rows are a small fraction), and sanitization replaces PII which reduces some free-text fidelity.
Who it fits and trade-offs

Great fit if you need realistic SOC telemetry with linked provenance for training or evaluating graph-based intrusion detection systems, AI-driven cyber defense simulators (CybORG/MARL), or alert-classification models that leverage MITRE mappings. Look elsewhere if you require analyst-verified ground truth for every incident, continuous long-duration captures beyond the documented windows, or raw unsanitized logs (this release applies extensive PII removal).

Information

  • Websitehuggingface.co
  • OrganizationsWitFoo, Inc., University of Canterbury, Computer Science and Software Engineering
  • Published date2026/03/21

Categories

More Items

Hugging Face

Aggregated, screened corpus of 55,050 normalized Indian public-information text bodies and 65,209 source records for retrieval and question-answering. Exports include deduplicated CSV/Parquet with provenance, topic labels, extraction quality flags and a private SQLite backup.

Hugging Face

Provides 3,451 hours (2,051,810 clips) of AI‑generated 48 kHz Turkish speech with transcripts, spoken forms and per‑clip voice descriptions for TTS and ASR development. Includes 2,752 designed voices and is licensed CC BY 4.0 / CC BY‑SA 4.0 (attribution to PatientDesk AI required).

Hugging Face

Provides imagined interaction segments generated by world models for RoboTwin2.0 tasks, stored as fixed-length HDF5 chunks (21 observation frames, 20 actions, rewards and episode flags). Useful for training and evaluating world-model-based policies; currently limited to the RoboTwin2.0 subdataset.