A $25 Bug Worth $500,000: AI Just Broke VDP Economics

A researcher found a $500K-class WordPress RCE for $25 of AI compute. Here's what collapsing discovery costs do to your VDP, and how to re-tune it.

Ernest Bursa

Ernest Bursa

Founder · · 11 min read
A security engineer at a garage workbench comparing a $25 compute receipt against a $500,000 exploit-broker price sheet on a second monitor, morning light through the open garage door

The cost of finding a critical bug just collapsed. A researcher at Searchlight Cyber pointed a frontier model at WordPress core, spent about $25 in compute over roughly ten hours, and surfaced a pre-authentication SQL injection that escalates to remote code execution, a bug class that exploit brokers price at up to $500,000. That is not a story about AI spam. It is the opposite, and it is the failure mode almost no vulnerability disclosure program budgeted for: a surge of genuinely valid, genuinely critical findings arriving faster than any human triage rota can acknowledge, verify, or pay them.

If you run a VDP or a bug bounty program, your entire cost model rests on one quiet assumption: that finding a critical bug is expensive and rare. That assumption just broke. This is what it means for your SLAs, your bounty budget, and your triage queue, and what to change this quarter.

What actually happened, and what is verified

Adam Kues, a researcher at Searchlight Cyber, adapted an AI prompting recipe into a multi-agent vulnerability-hunting harness and aimed a frontier model at WordPress source code. The researcher reports it took ten-plus hours of model runtime, costing roughly $25, to discover a pre-auth SQL injection escalating to RCE. He then reproduced it against a stock WordPress install, asking the model to steal the admin’s email address, and it printed the address back within minutes. He held publication and prepared a disclosure report.

Two anchors matter here, and only one is independently verified. The $25 figure and the ten-hour runtime are the researcher’s own account of his own work, so treat them as reported, not established. The $500,000 valuation is the one you can check: Crowdfense’s public Exploit Acquisition Program price list quotes “WordPress (RCE): 500k USD,” the same tier as Apache HTTP Server and Microsoft IIS remote code execution, above Nginx at $350k and far above other content management systems like Joomla ($40k) or Drupal ($25k). The gray-market ceiling is real, not a rhetorical flourish. No CVE was assigned in the source, and the exact model naming is the article’s framing, not a shipping product spec.

The gap is the whole story. Cost to find it: about $25. Value to an attacker: up to half a million dollars. Those two numbers used to be roughly correlated, because scarce human expertise sat between them. AI just removed the human as the bottleneck.

“AI slop” was the easy problem

For two years the industry conversation about AI and bug bounties has been about noise. Fabricated CVEs, hallucinated function names, template reports, the kind of low-quality, high-volume garbage that made curl maintainer Daniel Stenberg publicly vent about AI-generated reports wasting his time. We wrote about that failure mode ourselves in AI slop is flooding bug bounty triage. Slop wastes triage hours. It is annoying, it is expensive at scale, and defenders have spent real effort tuning filters to reject it.

The WordPress finding is the inverse, and it is arguably worse. This is high-quality, high-severity signal, produced at machine speed and near-zero marginal cost. Slop threatens your triage time. A flood of real criticals threatens three things at once: your acknowledgment SLA, your bounty budget, and your patch pipeline. You tuned your intake to reject garbage. You did not budget for an abundance of gold.

Here is the uncomfortable operational reality: both failure modes now arrive together. The same technology minting valid criticals also produces convincing fakes, real-but-old CVEs resubmitted as new, plausible code references to functions that do not exist. Triage can no longer assume a report is either trustworthy or trash. It has to assume both modes are live in the same queue, every day.

Why cheap discovery breaks bug bounty economics

Bug bounty pricing encodes an implicit deal. Finding a critical bug is expensive and rare, so a few-thousand-dollar reward is a fair split of the value with a scarce expert who chose disclosure over the gray market. When discovery cost collapses to $25, that deal inverts. The broker still pays $500,000, because the offensive value of a pre-auth RCE across hundreds of millions of installs did not change. But the program’s payout, often one to three orders of magnitude lower, now looks like a rounding error against both the bug’s true worth and the volume of them about to arrive.

Put three numbers in one frame:

Amount Source
Cost to discover the bug ~$25 Researcher’s reported compute spend
Gray-market value up to $500,000 Crowdfense price list (verified)
What a typical VDP pays for the same class $0 to low five figures Program-dependent

The misalignment is not that programs are stingy. It is that the discovery-cost collapse widened an already-large gap between what a bug is worth to an attacker and what a defender can rationally pay, while simultaneously multiplying how many such bugs a program will receive. A researcher who can mint criticals for $25 each faces a sharp choice: submit to a program that pays $5,000 if anything and enforces disclosure terms, or sell to a broker at $500,000. Cheap discovery does not just strain triage capacity. It sharpens the pull toward the gray market for exactly the highest-severity findings your program most needs to receive.

No self-hosted program will ever match a broker. That is not the goal, and pretending otherwise wastes credibility. The realistic goal is capturing the large population of researchers who would rather disclose, provided the friction is low and the payout gap, though large, is not insulting. That is a program-design problem, and it has three levers.

The three things every VDP owner has to re-tune now

Cheap discovery does not require a new platform. It requires re-tuning three primitives your program already has, or should have: SLA tiers, reward economics, and triage capacity. Kit’s self-hosted CSIRT vertical encodes each of these as configuration, so I will use its defaults as a concrete reference point. The principles apply whether you run Kit, HackerOne, or a security.txt and a shared inbox.

Re-tuning SLA: separate “we heard you” from “we fixed it”

A flat SLA collapses under a surge of valid criticals. The fix is to split one promise into two. Acknowledgment is cheap and fast: it says “a human has your report and a clock is running.” Resolution is expensive and slow: it says “this is patched.” Conflating them means every spike in inbound threatens your public commitment to fix things on a deadline you no longer control.

Reserve your tightest resolution windows only for the top severity tier, so a wave of criticals cannot effectively deny-service your queue. Kit’s Csirt::SlaConfig ships this shape by default: a 72-hour acknowledgment promise, then severity-graded resolution targets of 24 hours for super-critical, 72 hours for critical, 7 days for high, and progressively longer for medium and low. The design point is that acknowledgment stays constant and cheap while resolution scales with severity. When a hundred valid reports land in a week, you can still honor “we heard you, within 72 hours” for all of them, and honor “we fixed it, within 24 hours” only for the genuine 10.0s.

Re-pricing rewards: build a defensible top tier and a dedup policy

Your severity vocabulary needs a band above “critical.” AI is about to surface more true CVSS 10.0s than programs historically saw in a year, and a pre-auth RCE across a huge install base is precisely that. Kit models this with a super_critical tier (CVSS 10.0) sitting above critical (9.0 to 9.9), and a graded BountyMatrixConfig that runs from $0 for informational up to $5,000 to $10,000 for super-critical.

Be honest about what that top tier is. Kit’s generous self-hosted default of $5,000 to $10,000 is still 50 to 100 times below the broker price for the exact bug Kues found. You will not close that gap, so stop trying to. What you can do is steepen the curve and widen the top so the highest-severity findings are clearly worth reporting rather than shelving. We walk through this in detail in how to structure bug bounty reward tiers.

Then add a deduplication policy, because it is suddenly load-bearing. When many researchers point the same model at the same popular target, collision on identical bugs is not a risk, it is the predictable outcome. Without a clear, published first-valid-report rule, you will spend your credibility arbitrating dead-heat disputes, the exact scenario we cover in resolving bug bounty payout disputes. Kit enables dedup by default in its TriageConfig for this reason.

Triage at machine speed: screen slop and fast-track the real criticals

Triage now has two jobs that pull in opposite directions. It must reject convincing fakes and rapidly validate genuine high-severity signal, often in the same queue on the same day. Human-paced triage cannot do both at volume.

Screening is the first half. Kit’s Csirt::AiScreening scores an ai_confidence_score (0.0 to 1.0, the likelihood a report is AI slop) and flags known fabrication signals like hallucinated functions, fabricated CVEs, a previous CVE cited as new, and reports with no specific proof of concept. That layer was built for the slop problem, and it still must run, because the slop is not going away.

But screening is only half the new problem. The frontier is fast-tracking the valid criticals. That means auto-escalation and on-call routing so a super-critical never sits in a queue waiting for someone to notice it. Kit’s TriageConfig supports escalation severities (critical and super_critical by default), retest requirements, and auto-assign-to-on-call, wired to an on-call rotation and PagerDuty. The takeaway for every program owner: slop-filtering is table stakes now; rapid validation and prioritization of genuine signal at volume is the capability you actually have to build next.

What to do this quarter

You cannot out-bid a broker, and you do not need to. You need a program that is fast, credible, and well-instrumented enough that a researcher who prefers disclosure has no friction-based excuse to broker instead. Concretely:

  1. Split your SLA. Publish a fast, constant acknowledgment window and severity-graded resolution windows. Reserve your tightest resolution target for a top tier above “critical.”
  2. Add a super-critical band and reprice the top. Widen and steepen the bounty curve so true 10.0s are worth reporting. Accept the gap to broker prices; close the gap to insulting.
  3. Publish a dedup rule before you need it. First valid, fully reproducible report wins. Say so in writing now, not during a dispute.
  4. Run two-sided triage. Score AI-likelihood to catch fakes, and auto-escalate valid criticals to on-call so nothing severe waits in a queue.
  5. Instrument acknowledgment time. If you cannot see your time-to-acknowledge per severity, you cannot defend it when the volume arrives.

The cost of finding a critical bug fell to $25. The value of that bug to an attacker did not move. Your program lives in the gap between those numbers, and the gap just got wider and more crowded. The programs that survive the next year will not be the ones with the biggest bounty pools. They will be the ones that acknowledge fast, price honestly, dedup fairly, and escalate the real criticals before anyone else even opens the ticket.

How Kit’s CSIRT handles the surge

Kit is an AI-native platform for startups, and its self-hosted CSIRT and VDP module exists precisely because founders end up owning security triage by default. It ships the levers this shift demands as configuration rather than a project: a super_critical severity band above critical, split acknowledgment and resolution SLAs, a graded bounty matrix you can reprice, dedup and auto-escalation to on-call, and AI screening that scores the likelihood a report is slop.

What it does not do is pretend to be a $500,000 broker, and that honesty is the point. A well-instrumented, fast, transparent program captures the researchers who would rather disclose than sell, as long as you do not make disclosure slow, opaque, or insulting. The cost curve for finding critical bugs collapsed. Your program’s design is the only lever you still fully control, so tune it before the flood, not during it.

Related articles

Ready to hire smarter?

Start free. No credit card required. Set up your first hiring pipeline in minutes.

Start hiring free