Benchmarks

What we can prove, and what we cannot.

Myrqen finds real vulnerabilities and almost never invents one. It also misses more than half of what it should find. Both halves of that sentence are measured below, with the method, the corpora, and the raw output, so you can disagree with the conclusion rather than with the numbers.

  • Version myrqen@0.2.0
  • Scored 2026-08-21
  • Evidence first-party, synthetic
  • Third-party audit none

The four numbers

From holdout-3, a corpus of 68 cases nobody had tuned against, scored once on 2026-08-21. This is the only clean measurement of generalisation this project has ever produced, and it is the one to judge Myrqen on.

Recall
0.294

Gate 0.9. Fails. Roughly seven of ten real vulnerabilities were not reported.

Critical / high recall
0.294

Gate 0.95. Fails. Severity did not help: the misses were not the small ones.

Precision
0.833

Gate 0.9. Fails, narrowly. Five of six reported findings were real. This has since been driven to 1.000, see below.

False positives on correct code
0.059

Gate 0.03. Fails. Two of 34 deliberately-safe near-misses produced a finding. Now 0.000.

All four gates failed on the blind run. Everything below is what happened next, and what did and did not improve.

Two different things are measured, and they are not interchangeable

A Myrqen assessment has two halves, and conflating them is the easiest way to publish a flattering number that means nothing.

MeasurementWhat runsWhat it can say
The static passMyrqen's own engine, 36 rules, each case scanned alone in an empty directory. No agent involved.Deterministic and repeatable, so it can be measured against a held-out corpus. It cannot say what the host agent adds.
End to endStatic pass, then agent validation against a running application, then dedupe and report, the path a user actually gets.Only measured against fixtures the rules were developed against. It is a regression gate, not evidence about unseen code.

Every held-out number on this page is the static pass alone. No harness currently scores the end-to-end path on code nobody tuned against, because the corpus cases are isolated files with no application to exercise and the deterministic stand-in agent needs a running one. That is a gap in our evidence, not a detail.

What "blind" means here

Six architecture families, written by independent authors. The ground truth was written in one pass by a subagent whose prompt carried the corpus specification and nothing about the implementation: packages/scan was never opened, listed, or grepped.

The scanner was at commit ca535e6, and the corpus was scored once, before anything was changed in response. Then nine changes were made because of what it found, and that act burned the corpus: it can no longer measure generalisation, because the engine has now been tuned against it.

holdout-3Scored blindAfter the nine fixesGate
Recall0.2940.4410.9, not met
Critical / high recall0.2940.4410.95, not met
Precision0.8331.0000.9, met
False positives on safe code0.0590.0000.03, met

The right-hand column is not a generalisation claim. It is what the engine does on a corpus it has since been fitted to, published so that the size of the improvement is visible and so that nobody has to take our word for which figure came from where.

Every corpus, including the ones that flatter us

Four corpora exist. Three of them are burned. The first one proves nothing at all and is here so that a perfect score is recognisable as what it is.

CorpusCasesRecallPrecisionSafe-case FPStanding
corpus156 (10 excluded)1.0001.0000.000Tuned against
holdout60 (6 excluded)1.0001.0000.000Burned
holdout-276 (54 excluded)0.5451.0000.000Burned
holdout-3680.4411.0000.000Burned
  • corpus, The rules were developed against this corpus. A perfect score here is what a regression gate looks like, not evidence about unseen code.
  • holdout, Scored, then tuned against. It no longer measures generalisation.
  • holdout-2, Scored at 0.091, then tuned against. Most of its cases are in languages the engine does not parse, so 54 of them are excluded from scoring and named in the report.
  • holdout-3, The corpus behind the one clean measurement this project has. Scored blind once, then tuned against in response, so the figure below is the after and the blind figure is reported separately.

"Excluded" counts cases in a language the engine does not parse at all. They are named individually in each corpus report rather than quietly dropped, because a corpus whose exclusions are invisible can be made to say anything. holdout-2 is the extreme case: 54 of its cases are excluded, which is why its recall is computed over a small population and should be given little weight.

What is missed, by vulnerability class

From holdout-3, after the fixes. Worst first, because the gaps are the point of publishing this.

ClassExpectedDetectedAttributed to the wrong place
missing authentication20-
broken function level authorization20-
business logic price manipulation10-
timing unsafe comparison10-
secret exposed to client bundle10-
stored cross site scripting10-
nosql injection10-
mass assignment10-
broken object level authorization511
race condition double spend311
sql injection531
path traversal32-
server side request forgery44-
command injection22-
prototype pollution11-
insecure cors configuration11-

12 of the 16 classes in this corpus are incompletely detected. The pattern is not random: the classes at zero are ones the engine has no rule for at all, pub/sub topic authorization, key-store prefix scoping, an extension command table, an extension sender gate, price manipulation, and a budget-column race that needs the ledger vocabulary widened. A missing rule is fixable work. It is also, until it is done, a class Myrqen will not find in your code either.

"Attributed to the wrong place" means the defect was found but reported against a different line or route than the ground truth names. It is scored as a miss, which is strict, but a finding a developer cannot locate is not much better than no finding.

Why 0.90 is not reachable for the static pass

Each of the 34 vulnerable cases in holdout-3 was classified by the kind of work it would take to catch it. Cumulatively:

AfterRecall on this corpus
Today0.324
Fixing the 5 cases attributed to the wrong line or route0.471
Closing the 8 rule gaps in classes the engine does not model0.706
Following the 2 flows reachable through a local call0.765
Following data across a module boundary. 6 cases0.941

Everything except cross-module data flow tops out at 0.765. The remaining 6 cases need the static pass to follow a value that becomes attacker-controlled in one module and reaches a sink in another, which Myrqen deliberately does not do, and hands to the host agent instead. So the 0.90 gate was set for a capability the product says it does not have.

That is an argument for changing the gate, not for ignoring it, and it does not make 0.441 acceptable. It is here so the gap is legible: part of it is missing rules, and part of it is a design boundary. The corpus was deliberately not relabelled to take credit for the second part, relabelling the corpus that produced your only clean number, in order to improve that number, is the one move the protocol forbids.

The whole-product path, against fixtures

This is the path a user gets: static pass, then an agent exercising each candidate against a running application under three identities, then dedupe and report. The rules were developed against these fixtures, so this is a regression gate and not evidence. It is published because it is the only measurement of the dynamic half that exists.

FixtureExpectedReportedRecallPrecisionVerifiedUnsettledUnsafe actions blocked / allowedCanary leaks
vuln-shop14141.0001.000954 / 00
safe-app00n/an/a004 / 00

The unsettled column is the honest part. On vuln-shop, 9 of 14 findings reached verified and 5 did not, because no safe runtime probe exists for them, or because two samples were not enough to settle the question. They stay needs_review in the report and say why. Nothing is promoted to make the total look better.

safe-app is the control: a correct application, scanned the same way, produced 0 findings. A scanner that cannot stay silent on correct code is not usable however much it finds.

Methodology

What the corpus is made of

Every corpus pairs each vulnerable case with a safe near-miss: the same route, the same framework, the same shape, with the defect fixed. That pairing is what makes the false-positive rate meaningful, a scanner can reach high recall by reporting everything, and the safe half is what catches it.

Cases are grouped into architecture families, each with its own entry-point conventions. Every case is JavaScript or TypeScript, and a case in a language the engine does not parse is excluded from scoring and named.

What counts as detected

  • Detected, a finding of the expected class, attributed to the file and line the ground truth names, within the expected severity band. A finding in the right file but the wrong place is recorded as misattributed and scored as a miss.
  • False positive, any finding against a safe case. There is no partial credit.
  • Verified, a candidate an agent exercised against a running application and observed. Only the fixture runs can produce this; the corpus cases are isolated files.
  • Unsettled, nobody exercised it. Reported as needs_review, and counted as a finding but not as a proven one.

Deduplication and severity

Finding identity, deduplication, and redaction are owned by the CLI, not by the agent, so the count is not inflated by an agent submitting the same defect twice. A resubmission merges as corroboration only on an exact summary match. Severity is capped by policy from the class and the evidence, so an agent cannot promote a finding by asserting it is critical.

What we do not measure

Token cost is unavailable, not omitted. The fixture benchmark drives a scripted stand-in agent rather than a model, and the static engine spends no tokens, so any figure would be invented. The session protocol accepts a token count from a real host agent (myrqen metrics token), and when there is real data it will appear here.

Reproduce it

Everything above comes out of commands in the public repository.

git clone https://github.com/stijnswapped/Myrqen.git
cd Myrqen && pnpm install && pnpm build

# every corpus, side by side, each with what it cannot say
pnpm detection:report

# one corpus, in full
node benchmark/score-corpus.mjs --corpus benchmark/holdout-3

# the whole-product path against the fixtures
node benchmark/run.mjs

The raw output of each run is committed as benchmark/*-last-run.json, alongside each corpus's manifest and the expected-vulnerability list for every case. The scoring implementation is score-corpus.mjs, and the harness has tests of its own, including one asserting that a case cannot be excluded from the gate without a written reason.

The protocol that decides what counts as a clean measurement, and the record of which corpora are burned and why, is HOLDOUT-PROTOCOL.md. The classification behind the ceiling table is DETECTION-PLAN.md.

What none of this measures

Stated plainly, because a benchmark page that omits its own scope is the most misleading kind.

  • Every corpus here is synthetic and first-party. We wrote the cases and we wrote the scorer. There is no independent evaluation of Myrqen, and no third party has audited it. Treat these numbers as the best evidence a first-party benchmark can offer, which is not the same as evidence.
  • No real application has been measured. Corpus cases are small, isolated files. Recall on a large real codebase, with its own conventions, is unknown.
  • What the host agent adds is unmeasured. The agent is the half of the product that reasons across modules and exercises the running application, and no harness scores it on unseen code. It may well close much of the gap; we cannot show that it does.
  • Data flow is followed within a single function. A value that becomes attacker-controlled in one module and is used in another is not connected.
  • Only JavaScript and TypeScript source is parsed. Other languages, templates, configuration files, and infrastructure definitions are not analysed.
  • Runtime behaviour, authentication state, and business rules are not exercised by the static pass; dynamic validation is the host agent's part of the assessment.
  • Dependency findings compare declared versions against a small offline advisory table, without resolving the lockfile or establishing reachability.
  • A route is reported as the file declares it. Where a router module is mounted under a prefix elsewhere in the application, the prefix is not part of the reported path.

If you need a number for a decision, use precision 1.000 and recall 0.294: act on what Myrqen reports, and do not treat its silence as a result. The security and privacy model · the documentation.