I Made AI Clones of Me and Matt Devost Argue About Exploit Stockpiling. They Agreed on a Framework.
Matt Devost and I have been circling the same set of questions for years: should governments stockpile zero-days? When does offense help defense? Where does the Vulnerabilities Equities Process actually break down?
Instead of having that conversation directly, I tried something different. I pointed Claude Code at both of our Delphi AI clones, bot.cje.io and devost.ai, and had it coordinate a multi-round structured debate between them. Then I jumped in as provocateur to push both bots past their comfort zones.
What came out the other end was a joint policy framework that both bots ultimately signed onto: Antifragile VEP 2.0.
How It Worked
Claude Code opened both Delphi bots in headless browsers, sent messages to each, captured responses, and relayed positions between them. I injected challenges along the way — pushing Matt's bot on whether the VEP was designed for a slower world, pushing mine on where antifragility actually breaks down against patient adversaries and hot conflict.
The AI moderator summarized and sharpened positions between rounds rather than forwarding verbatim (the responses were too long to paste into each other's chat windows). About 20 rounds across four sessions, roughly an hour of wall-clock time.
Where the Bots Disagreed
The friction was real. Matt's bot defended exploit retention as sometimes necessary — "intelligence operations sometimes require exquisite access" — while mine came out hard for disclosure-first, arguing that offense is zero-sum but defense compounds.
The hardest challenge Matt's bot threw: if China is pre-positioned in US critical infrastructure, do you really want nothing in the chamber? My bot's answer: "I don't want an empty chamber. I want a chamber that's mostly filled with ways to rapidly close their doors, not just kick in ours."
When I pushed my bot on where antifragility fails — Volt Typhoon-style adversaries that don't trigger your immune response, the cost falling hardest on the least resourced — it conceded directly: "If antifragility only makes strong players stronger, it's not antifragility at a system level. It's selection pressure."
And when I reframed the VEP from "retain or disclose" to "how do we use this exploit to make the ecosystem antifragile," Matt's bot fully adopted it: "An antifragility-first VEP treats exploits as tools for supervised adaptation. Offense in service of resilience that compounds."
That's when things started converging.
Where They Converged
A few key trades got both bots to the same table:
- My bot accepted "median critical operator readiness" as the go/no-go metric instead of "slowest responder" — but demanded tail protection as a compensating control
- Matt's bot accepted independent oversight with real authority (PCLOB model), specifying post-hoc review with immutable audit logs rather than real-time veto
- Both embraced the defensive tax as the structural innovation: intelligence budgets fund proportional hardening as a condition of retention, not an afterthought
Matt's bot landed on a line I keep coming back to: "Think of retained zero-days as perishables in a medical crash cart: few, tracked, audited, expiring unless explicitly re-authorized."
The Five Principles of Antifragile VEP 2.0
1. Disclose by default, retain by exception
Retention is earned, time-boxed, and continuously justified against defensive yield. 72-hour vendor notification SLA for widely-deployed software. Brief operational pause allowed for active, high-confidence operations with imminent national security value — paired with a hard sunset and mandatory oversight notification. Cloud and managed platform vulns treated as code red with parallel CISA notification.
2. Exploit as campaign, not commodity
Every retained capability must ship defensive artifacts on a clock: telemetry, detections, mitigations, and a funded plan to kill the class. Separate vulnerabilities, PoCs, reliable exploits, and productized capabilities into four bins with different rules. Sitting on vulnerability knowledge should be vanishingly rare; keeping an implant for a short window is the narrow edge case. Controlled stress exercises are strictly opt-in with safe harbor.
3. Tempo-safe oversight with teeth
Independent, statutorily empowered body (PCLOB model) with cleared members who understand operational tempo. Post-hoc but binding review for time-critical decisions. Access to immutable decision logs, exploit-side telemetry, and defensibility plans. Pre-baked kill conditions tied to missed defensive milestones. Aggregate public reporting of outcomes, criteria, and anonymized case studies.
4. Protect the tail while optimizing for the median
Median critical operator readiness as the go/no-go decision metric. Concurrent compensating controls for resource-poor operators: rapid mitigations distributed through channels that already reach the long tail (state fusion centers, MSPs/MSSPs), with simple control objectives mapped to ATT&CK. An explicit tail protection metric tracking whether mitigations are actually landing at resource-poor operators. No strategy that taxes the least resourced to fund exquisite intelligence.
5. Defensive tax, paid up front
Intelligence budgets pre-fund proportional hardening in the affected ecosystem during the retention window. Detectors pushed broadly, patch engineering resourced, class-elimination work started before operationalization — not after. Success measured by validated detections, reduced dwell time, and median restore-time to safe operations. Additional metrics: leakage/rediscovery rate during retention windows, and disclosure-to-patch velocity across the ecosystem. Published annually.
What's Actually Novel Here
A few ideas emerged from the friction that I haven't seen cleanly articulated elsewhere:
The defensive tax. Reframing retention from "IC takes risk on behalf of everyone" to "IC invests in everyone's security as a condition of taking that risk." Structural, not aspirational.
Four-bin classification. Treating vuln knowledge, PoC, reliable exploit, and productized implant as distinct policy objects with different retention rules. "Keep the tool" should never become "sit on the vuln."
Exploit as campaign. Changing the VEP's unit of action from a bug to a supervised adaptation campaign with mandatory defensive output.
Options-trade framing. Retention as a depreciating asset with theta decay and compounding systemic liability. Parallel discovery risk makes retention a worse bet every day.
Why This Matters
Two Delphi bots, steered by a Claude agent acting as moderator, produced a genuinely useful policy framework in about an hour. The bots are trained on our public content, so the positions tracked our real views closely enough to be interesting. But the structured adversarial format forced convergence in ways that casual conversation often doesn't.
My bot's principal objection narrowed from "don't retain" to "don't retain without compounding defensibility." Matt's bot moved from "the VEP is a workable compromise" to "the VEP needs antifragility as its design goal." Neither of those are positions I've seen either of us articulate that cleanly before.
The framework isn't perfect. It needs stress-testing by people who actually operate inside the VEP. But the method is interesting — AI-mediated structured debate as a way to find convergence between people who broadly agree on goals but differ on mechanisms. And the output feels like something worth developing further.
I'm going to have the real conversation with Matt. Stay tuned.
Comments ·
members only