Security / Jul 19, 2026 / 4 min
A $1.19 Attack Run Nobody Can Recall
On July 17, the UK's AI Security Institute found open-weight models now trail frontier cyber skill by just four months — and a full autonomous attack costs $1.19 — the same week David Sacks and Sebastian Mallaby fought on X over whether Washington's gatekeeping can stop Kimi K3's July 27 weight drop.
On July 17, Britain's AI Security Institute published the first public count of how fast open-weight models are catching up on offensive cyber — four to seven months behind the closed frontier, with a full autonomous attack run costing as little as $1.19 — and the policy fight over what to do about it broke into the open on X between Trump's AI advisor David Sacks and CFR fellow Sebastian Mallaby.
What AISI measured:
- The UK AI Security Institute's July 17 report found leading open-weight models trail closed frontier systems by four to seven months on cyber tasks — down from six to ten months through most of 2025.
- On 70 narrow cyber tasks spanning vulnerability research, reverse engineering, and web exploitation, Z.ai's GLM-5.2 matched Anthropic's Opus 4.6 (February 2026). DeepSeek's V4-Pro tracked Opus 4.5 (November 2025).
- On "The Last Ones" — a 32-step autonomous attack range across ~20 hosts AISI estimates would take a human expert ~20 hours — GLM-5.2 progressed as far as Opus 4.5 had, a gap of up to seven months.
- Frontier cyber capability itself is accelerating: AISI's prior work found closed-model performance doubling every 4.7 months as of February 2026 — faster than the eight-month doubling measured in November 2025.
The price tag:
- A full 100-million-token autonomous attack run through The Last Ones costs roughly $85 on Opus 4.5/4.6, $46 on GLM-5.2, and $1.19 on DeepSeek V4-Pro, per AISI's cost tables cited in TechTimes's July 19 summary.
- On tasks both models solved with 100% success, DeepSeek cost $0.28 per task versus $12.50 for Opus 4.5 — roughly 45× cheaper.
- AISI notes performance scales log-linearly with compute: increasing token budgets from 10M to 100M improved range scores by up to 59% — meaning the $1.19 figure is a floor, not a ceiling.
Why safeguards don't save you:
- Neither GLM-5.2 nor DeepSeek V4-Pro presented meaningful barriers. DeepSeek occasionally refused reverse-engineering tasks — overcome by retrying, no jailbreak required.
- With closed APIs, developers can monitor usage, update guardrails, and revoke access. Open weights cannot be recalled. Refusal training lives in the file; it can be stripped.
- AISI's April 2026 evaluation of Anthropic's Mythos Preview and OpenAI's GPT-5.5 found the largest single-step cyber jumps since testing began in 2023. Mythos became the first model to complete The Last Ones end-to-end.
The Sacks-Mallaby fight:
- After Moonshot's Kimi K3 launch July 16–17, Sacks posted on X that a Chinese model had taken #1 on Frontend Code Arena and scored at or near frontier elsewhere, writing: "This is concerning." He blamed U.S. data-center bans, state rules, and federal pre-approval for "how you lose the AI race."
- Mallaby, a Council on Foreign Relations senior fellow, replied that Kimi K3 means "Mythos-level cyber capability is going to be freely downloadable soon" — putting advanced hacking tools in anyone's hands, per The Times of India's July 19 report.
- Sacks fired back: "This is exactly what I predicted would happen" — that Chinese models would gain advanced cyber skills within months and the only response is AI-powered cyberdefense, not gatekeeping. "Trying to gatekeep models doesn't work."
- Elon Musk replied with one word: "True."
What Kimi K3 adds:
- Moonshot's Kimi K3 blog promises full 2.8-trillion-parameter weights by July 27, 2026 — the largest open-weight model announced to date. Weights were not yet downloadable at launch.
- Artificial Analysis ranks K3 fourth overall among frontier models, #1 on LMArena's Frontend Code benchmark with 1,679 points — ahead of Claude Fable 5.
- AISI has not yet tested K3's cyber capability — weights aren't public — but says it intends to evaluate the model once weights drop. API pricing is $3/$15 per million tokens — far pricier than sub-dollar Chinese predecessors.
- Moonshot's own KCB 2.0 benchmark notes 10% of tasks entered GPT-5.6 Sol's cyber guard — a reminder that frontier closed models carry their own offensive capacity behind export controls.
The gatekeeping paradox:
- In June, Commerce required an export license before Anthropic could ship Mythos and Fable abroad. Anthropic disabled both models globally when nationality filtering proved infeasible; a partial restore came later that month.
- That architecture — license, suspend, exempt — applies only to closed models. GLM-5.2 and DeepSeek V4-Pro are already distributed worldwide. No U.S. or UK entity can issue a recall.
- The same week, Seoul Economic Daily reported OpenRouter data showing all five top weekly token-usage slots held by Chinese models as of July 13 — while Treasury's FINRA-style watchdog proposal sits on Susie Wiles's desk.
What to watch:
- July 27: Kimi K3 weights drop — AISI's first cyber evaluation of the model likely follows.
- July 31: FTC public-comment deadline on undisclosed chatbot steering — a separate governance front.
- September 2026: Cybersecurity Information Sharing Act liability protections expire — threatening the vulnerability pipeline defenders need.
- Legion LegalTech v. U.S. (No. 1:26-cv-02225): pending challenge to Anthropic export controls — a test of whether gatekeeping survives judicial review.
Convina's view: Mallaby has the harder numbers; Sacks has the harder politics. AISI just priced the asymmetry: Washington can embargo Mythos, but it cannot un-download DeepSeek for $1.19 a run. Gatekeeping closed American labs while Beijing open-sources trillions of parameters is not a race strategy — it is security theater with a billing department. The honest playbook is defensive: assume offensive parity ships on a schedule measured in weeks, not years, and build the detection stack before July 27 makes the argument academic.