RealVuln
Measures scanner performance on real-world vulnerable code with a standardized dataset and scoring methodology.
An open benchmark that measures how rule-based SAST, general-purpose LLM, and security-specialized scanners perform on real-world vulnerable code. It provides a standardized dataset of 1,903 vulnerabilities and 279 false-positive traps across 66 repositories, with all ground truth and scoring code released for independent reproduction. The benchmark targets developers, security teams, and IT teams evaluating security scanners. It is positioned as a living benchmark with a public leaderboard, allowing comparison of scanner performance on identical pinned commits.
Key features
- 1,903 vulnerabilities dataset
- 279 false-positive traps
- 66 repositories with pinned commits
- 26 scanners tested
- F2 and F3 scoring metrics
- Public leaderboard
- Open-source Apache 2.0 license
- Independent reproduction and audit
- No social media activity within the last 30 days
GTM channels
- Blog
ICP
- Security teams
- Software developers
- IT teams