Most public smart-contract audit outputs are semi-structured and noisy; this dataset aggregates those raw findings into a single Parquet with consistent columns so researchers can train or evaluate vulnerability detection and report-generation models at scale. It explicitly trades immediate usability for breadth: you'll get many real audit write-ups, but you must filter and normalize before training.
What Sets It Apart
- Scale and scope: 23,625 individual audit findings collected from public audit/contest platforms, with fields for title, description, PoC, recommendation, and a normalized severity label.
- Raw provenance preserved: the original CSVs are kept under raw/ and the primary artifact is a single Parquet file for easy loading with pandas/datasets/polars.
- Write-up quality signal: each row includes a precomputed
bug_weight(0–1) that encodes write-up thoroughness via section lengths, PoC code-block counts, and a coarse severity multiplier — useful for ranking higher-quality examples. - Research-focused notes: the dataset is explicitly labeled “raw, semi-structured” and intended for defensive security research rather than out-of-the-box model training.
How bug_weight is computed (brief)
A length-based content score (description, recommendation, PoC) is combined with a code-block bonus and a severity multiplier, then log-normalized into [0,1]. Critical/High share the same multiplier; rows with nonstandard severity text are scored as Low. Treat bug_weight as a thoroughness proxy, not a definitive severity label.
Who It's For and Trade-offs
Great fit if you want a large corpus of real-world audit findings to build or evaluate vulnerability classifiers, PoC generators, or report summarizers and are prepared to run deduplication, severity-normalization, and PoC filtering. Look elsewhere if you need a small, expert-labeled, fully deduplicated dataset with guaranteed working exploit code — ~78.1% of rows have placeholder PoC values and ~12.2% have placeholder recommendations; a few exact duplicates and some inconsistent severity labels remain. License/provenance for underlying reports is unclear, so verify usage rights before redistribution or commercial use.