You are viewing an archived release — RealVuln v3.1.0. View the latest benchmark →
realvuln v3.1
Dashboard Methodology Dataset Findings Roadmap Journal Paper GitHub ↗
12 / journal

Journal

The narrative record of the benchmark — every scanner run, corpus change, metric decision, and release, in the order it happened. Each issue states what we did, what we measured, and why we think the numbers moved. Written by Faizan Raza, co-author of the RealVuln paper with John Pellew.

18 issues·2026-03-05 — 2026-09-10·Measurements only — no vendor endorsements
§

Issues, newest first

18

RealVuln 3.1.0: DeepSeek V4.1 Flash, and a harness bug that was costing every agentic scanner coverage

RealVuln 3.1.0 adds DeepSeek V4.1 Flash at full 140-repository coverage (F3 50.9, sixth Overall) and fixes a headless-agent permission bug that had been silently ending 5–15% of runs before they wrote any findings.

17

Roster update: GPT-6 Astra, Claude Sonnet 5 and GLM-5.3 join the Overall leaderboard

Three more scanners now cover the full 140-repository corpus. GLM-5.3 takes fourth Overall, GPT-6 Astra fifth and Claude Sonnet 5 sixth — and the leaderboard now shows each run's effort setting and cost per line of code.

16

Most scanners get worse on TypeScript, and the drop is mostly precision

Eight of nine scanners that cover both RealVuln corpora lose precision on TypeScript/JavaScript, by 15 to 36 points. Per line of code they are not noisier — the TS/JS corpus simply has four times fewer real vulnerabilities per line to find.

15

A second language: the TypeScript/JavaScript corpus and per-language leaderboards

How RealVuln added 74 TypeScript/JavaScript applications and 2,236 labeled vulnerabilities, why the leaderboard now has language tabs, and why the Overall tab only ranks scanners that saw every repository.

14

Non-scoring ground truth: what to do with a finding you cannot grade

RealVuln 3.0.0 adds a third ground-truth state. 176 reviewed locations a competent reviewer could call either way are kept, documented, and excluded from scoring in both directions — never a hit, never a miss.

13

RealVuln 3.0.0: TypeScript/JavaScript, non-scoring entries, and per-language leaderboards

RealVuln 3.0.0 doubles the corpus to 140 repositories by adding 74 TypeScript/JavaScript apps, introduces non-scoring ground-truth entries, and splits the leaderboard by language.

12

RealVuln 2.1.0: the current release and the dataset export

RealVuln 2.1.0 is the current frozen release: 66 repos, 30 scanners, 2,176 labeled findings — plus a generated Hugging Face dataset.

11

Frontier and community scanners: Claude Opus 5 and the first external contribution (Rowan)

RealVuln adds Claude Opus 5 and its first community-contributed scanner, Rowan — a milestone for an open, independent benchmark.

10

Local models and GPT-5.6: can open-weight scanners compete?

RealVuln adds open-weight local scanners (Qwen 3.6, Gemma 4, Ornith) and GPT-5.6 Codex CLI — testing whether self-hosted LLMs can scan code.

09

RealVuln 2.0.0: the full 66-repo, false-positive-aware release

RealVuln 2.0.0 freezes the 66-repo corpus with 280 false-positive traps — the first release to measure both recall and false positives openly.

08

RealVuln 1.0.0: freezing the first release and opening the public site

RealVuln 1.0.0 is frozen at realvuln.com/v/1.0.0/ — the first citable release, with a redesigned public site, sitemap, and analytics.

07

The v2 corpus: adding 40 LLM-generated applications and measuring authorship

RealVuln v2 adds 40 LLM-generated, company-style apps — letting us measure how scanners perform on the AI-written code now entering production.

06

Scaling the roster: Opus 4.7, GLM-5.1, Kimi K2.6, GPT-5.5 and DeepSeek-V4

RealVuln adds a wave of frontier LLM scanners — Opus 4.7, GLM-5.1, Kimi K2.6, GPT-5.5, DeepSeek-V4 — and what the expanding field shows.

05

The paper, and why we score F3 (recall weighted nine times over precision)

Why RealVuln's headline metric is F3, not F1: in security, a missed vulnerability costs far more than a false positive. The metric decision, explained.

04

Reproducible by design: prompt versioning and the manifest

How RealVuln makes every score reproducible: content-hash prompt versioning, a pinned manifest, and one-command scoring anyone can run.

03

First LLM scanner runs: the agentic models enter the leaderboard

Our first agentic LLM scanner F2 results: Gemini 3.1 Pro and Sonnet 4.6 lead, GLM-5 follows — and what the early data showed.

02

Introducing the RealVuln benchmark framework

The RealVuln benchmark framework: how we label real vulnerabilities, match scanner findings on three fields, and score scanners fairly.

01

Why we built RealVuln: the false-positive problem no benchmark measured

Why we started RealVuln: an open, independent benchmark measuring vulnerability-scanner accuracy and false positives that no existing benchmark covered.

Dates correspond to the repository's git history and release manifests. Version numbers reference the frozen releases at realvuln.com/v/<version>/.