AIAny
Icon for item

Smart Contract Audit Findings

Provides 23,625 semi-structured smart-contract audit findings (title, description, PoC, recommendation, normalized severity) for defensive-security research; requires cleaning, deduplication, and PoC filtering before model training.

Introduction

Most public smart-contract audit outputs are semi-structured and noisy; this dataset aggregates those raw findings into a single Parquet with consistent columns so researchers can train or evaluate vulnerability detection and report-generation models at scale. It explicitly trades immediate usability for breadth: you'll get many real audit write-ups, but you must filter and normalize before training.

What Sets It Apart
  • Scale and scope: 23,625 individual audit findings collected from public audit/contest platforms, with fields for title, description, PoC, recommendation, and a normalized severity label.
  • Raw provenance preserved: the original CSVs are kept under raw/ and the primary artifact is a single Parquet file for easy loading with pandas/datasets/polars.
  • Write-up quality signal: each row includes a precomputed bug_weight (0–1) that encodes write-up thoroughness via section lengths, PoC code-block counts, and a coarse severity multiplier — useful for ranking higher-quality examples.
  • Research-focused notes: the dataset is explicitly labeled “raw, semi-structured” and intended for defensive security research rather than out-of-the-box model training.
How bug_weight is computed (brief)

A length-based content score (description, recommendation, PoC) is combined with a code-block bonus and a severity multiplier, then log-normalized into [0,1]. Critical/High share the same multiplier; rows with nonstandard severity text are scored as Low. Treat bug_weight as a thoroughness proxy, not a definitive severity label.

Who It's For and Trade-offs

Great fit if you want a large corpus of real-world audit findings to build or evaluate vulnerability classifiers, PoC generators, or report summarizers and are prepared to run deduplication, severity-normalization, and PoC filtering. Look elsewhere if you need a small, expert-labeled, fully deduplicated dataset with guaranteed working exploit code — ~78.1% of rows have placeholder PoC values and ~12.2% have placeholder recommendations; a few exact duplicates and some inconsistent severity labels remain. License/provenance for underlying reports is unclear, so verify usage rights before redistribution or commercial use.

Information

Categories

More Items

Hugging Face

Provides 7,366 recorded agent trajectories from H Company’s Holo4 benchmark runs, with step-level reasoning, actions, tool results, token usage and screenshots for replay and analysis. Bundled as JSON and image files for per-trajectory inspection and automated replay; released under Apache 2.0.

Hugging Face

Provides 150,000 source‑grounded decision examples for training models that pick options, judge yes/no propositions, or assign ordered scores. Each row pairs a 'state' with JEV-style typed questions (CHOICE/NOUL/SCORE); multiple configs and train/test splits are included, license mixed/unknown.

Hugging Face

Provides over 1.1M hours of high-bandwidth, multichannel multilingual speech with segment- and word-level timestamps, English translations, and per-file metadata for ASR, TTS and audio-representation research. Preserves original 48kHz multichannel OPUS audio and is released under CC BY 3.0.