Measured against code nobody here wrote.

Every number on this page is produced by a script in the open-source repository, re-run for release 0.6.0, with the method beside it. The big ones are whole corpora, every item scanned: not a sample, and not a demo.

At scale

Every item, not a sample.

A number is only as good as what it was measured on. These runs scan every record and every file of their corpus, and each defect they found was fixed with a test before the number was written down.

249,646known-malicious records
caught, every one100%
known-malicious records, caught, every one: 100%

Every malicious-package record in the bundled intel, across npm, PyPI, NuGet, RubyGems, Cargo, Go, Maven, Composer and VS Code, planted in a lockfile and scanned: 287,899 checks, none missed.

bench/malicious_records.py
190,000packages in 1,687 real lockfiles
read the same as Trivy reads them99.6%
packages in 1,687 real lockfiles, read the same as Trivy reads them: 99.6%

The lockfiles of the most-downloaded projects on 14 registries, compared package by package. 98.7% of lockfiles agree on average, 12 of 14 registries at 98% or above, and every gap was read: each Cordon defect it found is fixed.

bench/parse_agreement.py
8,692Agent Threat Rules test cases
of the attacks detected97.7%
Agent Threat Rules test cases, of the attacks detected: 97.7%

3,941 of 4,034 attacks, each on its rule's scan path; 80.6% of the published evasions, read offline for intent; 92.4% of the catalogue's benign near-misses left clean.

bench/atr_bench.py
Head to head

The same samples, two tools.

GuardDog is the open-source malware scanner most teams reach for. Both ran over the same draw of real malware and the same popular packages, inside Docker with the network off.

Real malware detectedthe same 498 malicious samples, a fixed-seed draw. Higher is better.
0%25%50%75%100%Cordon95.2%GuardDog85.5%
Popular packages wrongly blockedthe top 1,000 PyPI and 1,000 npm packages. Lower is better.
0%5%10%15%20%Cordon1.6%GuardDog16.8%
Cordon 0.6.0 GuardDog 3.2.0The same samples for both, run in Docker with the network off.
100%of known-malicious package records caught, every one checked249,646 records
99.6%agreement with Trivy on what real lockfiles contain190,000 packages
94.2%of real malicious packages detected, 79.7% by reading the code alone39,328 samples
97.7%of AI-agent attacks in the Agent Threat Rules test cases detected4,034 attacks
Every corpus

Five questions, each against real code.

Malware from DataDog's public dataset and malregistry, every known-malicious record, the most-downloaded packages, real lockfiles and agent configs, and the infrastructure the cloud vendors publish as correct.

MalwareIs a malicious package caught?
  • Known-malicious records, by name and version100%
    Known-malicious records, by name and version: 100%249,646 records
  • Real malicious packages detected (79.7% by the code alone)94.2%
    Real malicious packages detected (79.7% by the code alone): 94.2%39,328 (DataDog, malregistry)
  • The same malware, Cordon against GuardDog95.2% vs 85.5%
    The same malware, Cordon against GuardDog: 95.2%498, a fixed-seed draw
False alarmsIs harmless code left alone?
  • Popular packages wrongly blocked1.6% vs GuardDog's 16.8%
    Popular packages wrongly blocked: 1.6%, lower is bettertop 1,000 PyPI + 1,000 npm
  • Widely used open-source repositories passing the gate (0.4.0 run)85.4%
    Widely used open-source repositories passing the gate (0.4.0 run): 85.4%1,427
ReadingIs what a project contains read correctly?
  • Packages in real lockfiles, against Trivy99.6%
    Packages in real lockfiles, against Trivy: 99.6%1,687 lockfiles, 14 registries
  • Operating-system packages in 175 container images, against Syft99.9%
    Operating-system packages in 175 container images, against Syft: 99.9%21,758 packages (deb, apk, rpm, pacman, portage)
  • Every package in those images, against Syft95.1% mean
    Every package in those images, against Syft: 95.1%175 images, after stated exclusions
  • CVEs, against Trivy and OSV-Scanner98.4%, every gap explained
    CVEs, against Trivy and OSV-Scanner: 98.4%100 lockfiles
AI agentsIs an attack on the agent caught?
  • Attacks, each on its rule's scan path97.7%
    Attacks, each on its rule's scan path: 97.7%4,034 cases
  • Published evasions, read offline for intent80.6%
    Published evasions, read offline for intent: 80.6%289 cases
  • Benign near-misses left clean92.4%
    Benign near-misses left clean: 92.4%4,369 cases
  • Agent configs in real repositories blocked1.1%, each correct
    Agent configs in real repositories blocked: 1.1%, lower is better372 repositories
InfrastructureIs a vendor's own published infrastructure judged fairly?
  • Reference infrastructure, as its vendors publish it (0.4.0 run)2,560 findings, 795 blocking
    13 repositories, 20,310 files
Three questions

Detection is not one number.

Whether a file attacks you, whether your graph resolves a named release, and whether an attack shape is covered at all need different evidence.

ContentDoes the file itself attack you?Real malicious releases, unpacked without executing, scanned and deleted, in batches inside a container with no network. No advisory lookup: what the code does, read cold.
AdvisoryDoes the graph resolve a release someone has named?Every known-malicious record in the intel, planted in a lockfile in its ecosystem's own shape and scanned. The path a lockfile scan depends on, and the one that reaches every ecosystem.
TechniqueIs every attack shape covered at all?The labelled corpus, one directory per named technique, each with an expectation file stating what has to be found. Every sample is a permanent regression test.
Content, by source39,328 real malicious packages: 94.2% detected, 79.7% by the code alone.
  • Malicious npm packages (DataDog)93.7%
    Malicious npm packages (DataDog): 93.7%25,766 samples · 75.8% by the code alone
  • Malicious PyPI packages (DataDog)93.4%
    Malicious PyPI packages (DataDog): 93.4%2,502 samples · 86.3% by the code alone
  • Malicious AI-agent skills and IDE extensions (DataDog)35.6%
    Malicious AI-agent skills and IDE extensions (DataDog): 35.6%326 samples · 35.6% by the code alone
  • Malicious packages (malregistry)97.4%
    Malicious packages (malregistry): 97.4%10,734 samples · 88.9% by the code alone
  • All94.2%
    All: 94.2%39,328 samples · 79.7% by the code alone

Public sample sets exist for npm and PyPI, so content is measured there; every ecosystem is measured by record, on the right.

Every known-malicious record249,646 records, 287,899 checks, none missed.
  • npm100%
    npm records caught: 100%260,391 checked
  • PyPI100%
    PyPI records caught: 100%17,130 checked
  • NuGet100%
    NuGet records caught: 100%5,215 checked
  • RubyGems100%
    RubyGems records caught: 100%5,047 checked
  • VS Code extensions100%
    VS Code extensions records caught: 100%69 checked
  • Cargo100%
    Cargo records caught: 100%22 checked
  • Go100%
    Go records caught: 100%20 checked
  • Maven100%
    Maven records caught: 100%3 checked
  • Composer100%
    Composer records caught: 100%1 checked
  • Git100%
    Git records caught: 100%1 checked

Each record is planted as a pinned dependency in its ecosystem's own lockfile and scanned. Records whose only listed version is npm's empty takedown placeholder are counted and set aside: there is nothing malicious left to install.

By technique

Every attack shape, as a regression test.

The typosquat miss is deliberate: one character appended to a short name is how ecosystems name companion packages (vuex, reacts), so that shape is excused by name and reported by a different route.

  • Install-hook attacks
    5 samples100%
    Install-hook attacks: 100%
  • Credential exfiltration
    9 samples100%
    Credential exfiltration: 100%
  • Obfuscated payloads
    7 samples100%
    Obfuscated payloads: 100%
  • Download-and-execute payloads
    11 samples100%
    Download-and-execute payloads: 100%
  • Persistence mechanisms
    1 sample100%
    Persistence mechanisms: 100%
  • Supply-chain integrity
    4 samples100%
    Supply-chain integrity: 100%
  • CI/CD pipeline attacks
    2 samples100%
    CI/CD pipeline attacks: 100%
  • Infrastructure misconfiguration
    4 samples100%
    Infrastructure misconfiguration: 100%
  • Dependency confusion
    3 samples100%
    Dependency confusion: 100%
  • Typosquatting
    20 samples95%
    Typosquatting: 95%
What these numbers are not

Where the 5.9% goes.

Content detection reads what a package does. A package that does nothing yet, or ships only a compiled file, has nothing to read.

About 2 in 5 content misses are almost-empty packages: a bare package.json, a placeholder, a researcher's proof of concept. There is no behaviour in them to read. They are caught by name: all 249,646 known records are.

About 1 in 5 are a prebuilt binary and nothing else. Also caught by name, and a binary under a source name is its own finding.

The rest is the long tail, fixed shape by shape, each fix measured against the malware and the false-alarm corpora first.

Runtime-only behaviour is out of reach by construction: a payload decoded from a network response does not exist until the code runs. Cordon never runs it; the separate, opt-in sandbox is what observes that, now with a second install under a clock moved 400 days ahead.

AI agents

Agent attacks, on the path each rule reads.

The Agent Threat Rules catalogue ships attack, evasion and benign cases for every rule. Cordon runs each on the scan path the catalogue gives it, then on 372 real repositories it was never tuned on.

97.7%Attacks, each on its rule's scan path: 97.7%Attacks, each on its rule's scan path3,941 of 4,034
80.6%Evasions, read offline for intent: 80.6%Evasions, read offline for intent233 of 289
92.4%Benign left clean (the catalogue's near-misses): 92.4%Benign left clean (the catalogue's near-misses)4,038 of 4,369
372 real repositories10.5% warned and 1.1% blocked. Every block was read by hand and is correct.
Offline first, then the judgeText is read for what it asks, in six languages with disguises folded away, with no model and no network. For what still gets past, a local 7B model on a laptop caught 70 of 100 new attack wordings and left 22 of 22 realistic configs clean: a floor, not a gate.
Noise: 1,427 maintained projectsSecurity tools included, because a security tool's own signature file is the canonical false positive. 85.4% pass the default gate, and both malware findings across the corpus are correct.
Infrastructure the vendors publishTerraform modules from AWS, Azure and Google, AWS's CloudFormation library, Kubernetes' own examples. All 79 blocking rule classes were read against the files they fired on.
20,000+ tests on every pushLinux, macOS and Windows, Python 3.11 to 3.13, fuzzing, latency budgets, byte-identical reproducibility, and Cordon scanning itself.
Run it yourselfThe harness ships in the repository. Samples are downloaded into a Docker volume and opened only inside a container with its network off.
Reproduce
docker build -f bench/Dockerfile -t cordon-bench:dev .docker volume create cordon-bench-datadocker run --rm -v cordon-bench-data:/data --entrypoint python cordon-bench:dev /bench/fetch.py all --count 15000docker run --rm --network none -v cordon-bench-data:/data:ro cordon-bench:dev malware benigndocker run --rm --entrypoint python cordon-bench:dev /bench/malicious_records.pydocker run --rm -v cordon-bench-data:/data --entrypoint python cordon-bench:dev /bench/fetch_lockfiles.pydocker run --rm --network none -v cordon-bench-data:/data:ro --entrypoint python cordon-bench:dev /bench/parse_agreement.py

The full method, per ecosystem and per technique: docs/08-ACCURACY.md at v0.6.0.

Point it at your own code.