The Overall tab spans the whole 140-repository corpus and ranks only scanners that covered every language in it. Runs that covered a single language are listed on that language's tab, so every ranking is like-for-like.
Leaderboard
ranked by active metric| # | Scanner ▼ | F3 ▼ | Python F3 ▼ | TS/JS F3 ▼ | Missed Vulns ▼ | Prec % ▼ | Repos ▼ | Cost $ ▼ |
|---|
Precision vs. recall
hover a pointPerformance vs. cost
F3 vs costRecall ranking
fraction of vulnerabilities foundPrecision ranking
fraction of flags that were realBy category
three-tier summaryDetection by vulnerability class
recall %, best by approach▸ LLM-based scanners dominate classes that need semantic data-flow understanding — SQL injection, command injection, insecure deserialization. ▸ Rule-based tools stay competitive only on syntactic patterns, and even there overall recall remains low.
Dataset composition
4,138 vulnerabilities · 280 FP traps · 140 repositoriesFindings
Languages (140 repos)
Frameworks
Scanner categories
All scanner results
A per-repository breakdown, detection profile and run conditions for each of the 32 scanners evaluated in v3.0.
Security-Specialized
General-Purpose LLM
All figures are live RealVuln results across 32 scanners and 140 repositories. F3 weights recall nine times over precision; strict mode counts unfinished repositories as misses. Cost is API spend normalized per 100,000 lines of code scanned (rule-based tools are free or variably priced). Metric definitions →