realvuln v2.1
Dashboard Methodology Dataset Findings Roadmap Journal Paper GitHub ↗
12 / journal

Journal

The narrative record of the benchmark — every scanner run, corpus change, metric decision, and release, in the order it happened. Each issue states what we did, what we measured, and why we think the numbers moved. Written by Faizan Raza, co-author of the RealVuln paper with John Pellew.

12 issues·2026-03-05 — 2026-08-24·Measurements only — no vendor endorsements
§

Issues, newest first

12

RealVuln 2.1.0: the current release and the dataset export

RealVuln 2.1.0 is the current frozen release: 66 repos, 30 scanners, 2,176 labeled findings — plus a generated Hugging Face dataset.

11

Frontier and community scanners: Claude Opus 5 and the first external contribution (Rowan)

RealVuln adds Claude Opus 5 and its first community-contributed scanner, Rowan — a milestone for an open, independent benchmark.

10

Local models and GPT-5.6: can open-weight scanners compete?

RealVuln adds open-weight local scanners (Qwen 3.6, Gemma 4, Ornith) and GPT-5.6 Codex CLI — testing whether self-hosted LLMs can scan code.

09

RealVuln 2.0.0: the full 66-repo, false-positive-aware release

RealVuln 2.0.0 freezes the 66-repo corpus with 280 false-positive traps — the first release to measure both recall and false positives openly.

08

RealVuln 1.0.0: freezing the first release and opening the public site

RealVuln 1.0.0 is frozen at realvuln.com/v/1.0.0/ — the first citable release, with a redesigned public site, sitemap, and analytics.

07

The v2 corpus: adding 40 LLM-generated applications and measuring authorship

RealVuln v2 adds 40 LLM-generated, company-style apps — letting us measure how scanners perform on the AI-written code now entering production.

06

Scaling the roster: Opus 4.7, GLM-5.1, Kimi K2.6, GPT-5.5 and DeepSeek-V4

RealVuln adds a wave of frontier LLM scanners — Opus 4.7, GLM-5.1, Kimi K2.6, GPT-5.5, DeepSeek-V4 — and what the expanding field shows.

05

The paper, and why we score F3 (recall weighted nine times over precision)

Why RealVuln's headline metric is F3, not F1: in security, a missed vulnerability costs far more than a false positive. The metric decision, explained.

04

Reproducible by design: prompt versioning and the manifest

How RealVuln makes every score reproducible: content-hash prompt versioning, a pinned manifest, and one-command scoring anyone can run.

03

First LLM scanner runs: the agentic models enter the leaderboard

Our first agentic LLM scanner F2 results: Gemini 3.1 Pro and Sonnet 4.6 lead, GLM-5 follows — and what the early data showed.

02

Introducing the RealVuln benchmark framework

The RealVuln benchmark framework: how we label real vulnerabilities, match scanner findings on three fields, and score scanners fairly.

01

Why we built RealVuln: the false-positive problem no benchmark measured

Why we started RealVuln: an open, independent benchmark measuring vulnerability-scanner accuracy and false positives that no existing benchmark covered.

Dates correspond to the repository's git history and release manifests. Version numbers reference the frozen releases at realvuln.com/v/<version>/.