AIAny
Icon for item

Vyber07/cyber-security

A 16 GB, 507-file PhD‑level cybersecurity knowledge base for training and evaluating security-focused LLMs and automation. Covers offensive/defensive/forensics/cloud/iot and AI-security across 30+ domains with real-world labs and framework mappings.

Introduction

Most AI assistants and detection models fail when confronted by realistic, cross-domain security content that mixes offensive techniques, detection rules, and forensic artifacts. This dataset assembles that full stack at academic depth so models can learn practical attack workflows, defensive playbooks, and AI-specific threats in one corpus.

What Sets It Apart
  • Broad cross-domain coverage: 30+ domains (red team, blue team, forensics, malware, cloud, IoT, web security, bug bounty, AI/ML security), so models trained on it encounter the end‑to‑end lifecycle of attacks and responses rather than isolated examples.
  • PhD‑level technical depth: 507 files (~16 GB) include playbooks, lab scenarios (HackTheBox/TryHackMe/PortSwigger-style), mappings to MITRE ATT&CK / OWASP / NIST, and detection artefacts (YARA, Sigma), which improves model fidelity on advanced tactics.
  • Practical, hands-on content: contains real-world labs, exploit chains, detection recipes and forensics artefacts — so the dataset is useful not only for text generation but for building detection rules, SOAR playbooks, and training red/blue exercises.
  • AI-security material included: sections on adversarial ML, prompt injection, model extraction and OWASP LLM/ML lists help bridge traditional cybersecurity with AI-specific threats.
Who It's For and Trade-offs

Great fit if you want to fine-tune or evaluate security-focused LLMs, build SIEM/SOAR automation, design red/blue training scenarios, or teach advanced cybersecurity courses. It is also useful as a reference corpus for detection engineering and threat modelling. Look elsewhere if you need curated, low-risk data for public demos—the corpus includes explicit offensive techniques and exploit recipes that require responsible handling and operational safeguards. The dataset is large and technical (16 GB), so expect nontrivial storage and preprocessing work. Licensed under Apache-2.0, suitable for research and tooling but demands ethical use and access controls.

Information

Categories

More Items

Hugging Face

Structured dataset for training and evaluating LLM agentic behavior: function-calling conversations, JSON-mode structured outputs, and extraction samples for teaching models to generate tool calls and strict structured responses. Includes single-turn and multi-turn scenarios across several configs.

Hugging Face

A multi-task English NLU benchmark for evaluating models across nine tasks (acceptability, sentiment, paraphrase, similarity, and various NLI setups), with a diagnostic evaluation set and an online leaderboard to compare generalization and transfer learning.

Hugging Face

Provides 1.3 billion platform-specific video URLs extracted from CommonCrawl along with crawl metadata (no media included), serving as the source corpus for the LAION-BVD multimodal video dataset; distributed on Hugging Face in Parquet format.