Issues, newest first
RealVuln 2.1.0: the current release and the dataset export
RealVuln 2.1.0 is the current frozen release: 66 repos, 30 scanners, 2,176 labeled findings — plus a generated Hugging Face dataset.
11Frontier and community scanners: Claude Opus 5 and the first external contribution (Rowan)
RealVuln adds Claude Opus 5 and its first community-contributed scanner, Rowan — a milestone for an open, independent benchmark.
10Local models and GPT-5.6: can open-weight scanners compete?
RealVuln adds open-weight local scanners (Qwen 3.6, Gemma 4, Ornith) and GPT-5.6 Codex CLI — testing whether self-hosted LLMs can scan code.
09RealVuln 2.0.0: the full 66-repo, false-positive-aware release
RealVuln 2.0.0 freezes the 66-repo corpus with 280 false-positive traps — the first release to measure both recall and false positives openly.
08RealVuln 1.0.0: freezing the first release and opening the public site
RealVuln 1.0.0 is frozen at realvuln.com/v/1.0.0/ — the first citable release, with a redesigned public site, sitemap, and analytics.
07The v2 corpus: adding 40 LLM-generated applications and measuring authorship
RealVuln v2 adds 40 LLM-generated, company-style apps — letting us measure how scanners perform on the AI-written code now entering production.
06Scaling the roster: Opus 4.7, GLM-5.1, Kimi K2.6, GPT-5.5 and DeepSeek-V4
RealVuln adds a wave of frontier LLM scanners — Opus 4.7, GLM-5.1, Kimi K2.6, GPT-5.5, DeepSeek-V4 — and what the expanding field shows.
05The paper, and why we score F3 (recall weighted nine times over precision)
Why RealVuln's headline metric is F3, not F1: in security, a missed vulnerability costs far more than a false positive. The metric decision, explained.
04Reproducible by design: prompt versioning and the manifest
How RealVuln makes every score reproducible: content-hash prompt versioning, a pinned manifest, and one-command scoring anyone can run.
03First LLM scanner runs: the agentic models enter the leaderboard
Our first agentic LLM scanner F2 results: Gemini 3.1 Pro and Sonnet 4.6 lead, GLM-5 follows — and what the early data showed.
02Introducing the RealVuln benchmark framework
The RealVuln benchmark framework: how we label real vulnerabilities, match scanner findings on three fields, and score scanners fairly.
01Why we built RealVuln: the false-positive problem no benchmark measured
Why we started RealVuln: an open, independent benchmark measuring vulnerability-scanner accuracy and false positives that no existing benchmark covered.
Dates correspond to the repository's git history and release manifests. Version numbers reference the frozen releases at realvuln.com/v/<version>/.