You are viewing an archived release — RealVuln v3.1.0. View the latest benchmark →
realvuln v3.1
Dashboard Methodology Dataset Findings Roadmap Journal Paper GitHub ↗

Benchmark dashboard

33 scanners · 140 repositories · ranked by F3 (strict)
Metric
33
Scanners
3 categories
140
Repositories
All apps
86.4
Best F3 (strict)
Kolega Enterprise
0.89
Highest recall %
Kolega Enterprise
741,034
Total LOC
across all repos
Language
Authorship

Leaderboard

ranked by active metric
# Scanner F3 Python F3 TS/JS F3 Missed Vulns Prec % Repos Cost $

Precision vs. recall

hover a point

Performance vs. cost

F3 vs cost

Recall ranking

fraction of vulnerabilities found

Precision ranking

fraction of flags that were real

By category

three-tier summary

Detection by vulnerability class

recall %, best by approach

LLM-based scanners dominate classes that need semantic data-flow understanding — SQL injection, command injection, insecure deserialization. Rule-based tools stay competitive only on syntactic patterns, and even there overall recall remains low.

Dataset composition

4,138 vulnerabilities · 280 FP traps · 140 repositories

Findings

4,138 vulnerabilities
280
Real vulnerabilities FP traps (6.3%)
18
CWE families
741,034
Lines of code

Languages (140 repos)

Python66
TypeScript53
JavaScript21

Frameworks

Express27
Django23
FastAPI23
Flask15
Next.js12
NestJS + Angular10
Remix10
Fastify + Vue10
React5
custom3
aiohttp1
Tornado1

Scanner categories

GP-LLM24
Rule SAST3
Sec.-spec.6
12
Frameworks
33
Scanners tested

All scanner results

A per-repository breakdown, detection profile and run conditions for each of the 33 scanners evaluated in v3.1.

All figures are live RealVuln results across 33 scanners and 140 repositories. F3 weights recall nine times over precision; strict mode counts unfinished repositories as misses. Cost is API spend normalized per 100,000 lines of code scanned (rule-based tools are free or variably priced). Metric definitions →