If your team already pays for an AI coding agent, it's a fair question: why buy a security scanner when you can type “review this repository for security vulnerabilities” into the agent you already have? We wanted a real answer, so we ran exactly that through our internal security benchmark: Cursor's headless agent on Grok 4.5 against Patchlight's production scanner, over the same repositories with the same labelled vulnerabilities, scored the same way.
74.1
Patchlight F3
vs 51.5 for Cursor with Grok 4.5
+22.6
F3 points ahead
same repositories, same scoring
1.5×
More vulnerabilities found
76% recall vs 49%
<5¢
Per full-repo scan
average on our benchmark, pay as you go
The result
Patchlight scored an F3 of 74.1. Cursor with Grok 4.5 scored 51.5. Patchlight found 76% of the known vulnerabilities; the agent found 49%.
We rank on F3 because it weights recall over precision, and that's the right trade for a security scan. A false alarm costs a reviewer a few seconds to dismiss. A missed injection or a missing ownership check ships. Grok was careful with what it did report, and most of its findings were real. But it stopped long before it had covered the codebase, and left half of the known vulnerabilities on the table.
Ahead on every repository
Patchlight came out ahead on every codebase in the comparison, and the averages hide how far apart the two can land. On one, the agent caught 38% of the vulnerabilities and Patchlight caught 79%. On another, the agent caught under a third, and Patchlight more than doubled it.
Why a general-purpose agent falls short
Coding agents like Cursor, Claude Code and Codex are built to write and change code. When you ask one to audit a whole repository, it does what an agent does: it opens the files that look interesting, forms an opinion, and decides for itself when it's done. Coverage depends on that judgement call, and on a codebase of any size it's the wrong one to leave to chance.
Patchlight treats a scan as a coverage problem. It walks the whole repository in focused passes, checks every part against each vulnerability class, and reports only what it can point at: severity, CWE, file and line. Better prompting in the agent won't close that gap. The gap is the harness.
More than a better scan
A one-off chat with an agent produces a list you have to act on by hand. Patchlight turns the scan into a control that runs whether or not anyone remembers it:
- Scans on a schedule. Every repository gets its own schedule (daily, weekdays or the days you pick) at your time and in your timezone. Turn on scan-on-merge and the default branch is re-checked every time a pull request lands.
- Findings that file themselves. New findings above a severity you choose open GitHub issues automatically, filed once and never again after you close them.
- Every pull request reviewed. Inline comments on the diff, with fixes as GitHub suggestions you commit in one click.
- One place to see it all. Dashboards for findings, resolution rate and turnaround by severity, repository and contributor, plus weekly or monthly write-ups and an AI chat that answers from your own workspace data.
- Cents, not seats. A full-repository scan cost under 5 cents on our benchmark. Pricing is a prepaid balance with no subscription and no per-seat fees, and spend caps keep a busy month from becoming a surprise.
- Your code stays yours. It is never used to train a model, ours or a provider's, and model providers run under zero-retention terms.
Try it on your own code
Benchmarks are a fair fight on someone else's code. The comparison that matters is yours: run your agent of choice over a repository you know well, then run Patchlight over the same one and see what each of them catches.