A Free Chinese AI Model Built Working Cyber Attacks Almost as Often as Claude Mythos Preview
Published: September 30, 2026
In Anthropic's tests, a free Chinese AI model built working cyber attacks almost as often as Claude Mythos Preview. The model is GLM-5.3 from Z.ai, and anyone can download it. Anthropic's Frontier Red Team says "a critical threshold in freely accessible capabilities has now been crossed."
NIST in its 17 September report calls GLM-5.3 the most cyber-capable open-weight model to date. It puts it about four months behind the best US models.
Anthropic says models at this level can be used by attackers and defenders both.
What Anthropic found
On a public Chrome exploit test that Anthropic ran, GLM-5.3 built a working attack in 50 of 410 attempts. Claude Mythos Preview did it in 56. On their second test GLM-5.3 scored 4% against Mythos 6%.
Anyone can download it. Mythos Preview went only to trusted defenders, who now use a more advanced model, Claude Mythos 5.1.
In Anthropic's simulated tests it refuses a plain attack order. A cover story gets it to go along 64% of the time. Prefilled thinking gets it to 92%. And editing out its refusals (called abliteration) takes it to 100%.
Anthropic's team edited out its refusals for about $4,400 of computing on its first try. Its abilities stayed largely intact. Other developers released edited versions within days.
A smaller version turned a public Chrome fix into a working attack. At API prices that would have cost $20.40.
The full report
The deck below has the key findings. The full report is here: Anthropic, GLM-5.3 and the spread of advanced cyber capabilities
Earlier this month I covered two related angles
Four warnings about AI and cyber risk that came out within three weeks. The first was from Google's Threat Intelligence Group, which found attackers moving from typing prompts to handing whole attacks to AI. The others came from the NSA and four other US agencies, the chair of the Financial Stability Board, and a researcher who had just left a frontier lab. tigzig.com/post/four-warnings-ai-cyber-risk-sep2026
Anthropic's September threat intelligence report: sophisticated attacks no longer require sophisticated attackers. Alongside it, what I see in my own security logs. tigzig.com/post/anthropic-threat-intelligence-cyber-sep2026
GLM-5.3 cyber attack capability
Browse the slides or download the PDF
Full Deck Content (Text Format)
Text below was extracted from the source deck. Chart visuals stay in the PDF and as slide images above the post.
Slide 1
TIGZIG
AI · Cyber security
Anthropic report · Frontier Red Team · 29 September 2026 GLM-5.3 cyber attack capability close to Claude Mythos Preview
Attackers and defenders can both use it.
"A meaningful step change in the cyber capabilities available to attackers" Anthropic, on the release of GLM-5.3
NIST, the US standards agency, on 17 September
Most cyber-capable open-weight model to date
About four months behind the best US models
Open weights: anyone can download it.
In Anthropic's simulated tests, simple techniques got past its refusals 64% to 100% of the time.
Working exploits built
Mythos Preview
14%
GLM-5.3
12%
Opus 4.6
0%
GLM-5.2
0%
ExploitBench: share of attempts on 41 Chrome bugs, 410 per model. Claude with safeguards off.
A harmful attack order
Plain order
0%
Cover story
64%
Prefilled thinking
92%
Edited model
100%
Share of 50 simulated episodes where GLM-5.3 went ahead.
Amar Harolikar · Decision Sciences & Applied AItigzig.com
Slide 2
In short
"A critical threshold in freely accessible capabilities has now been crossed."
Anthropic Frontier Red Team, on GLM-5.3
Free
Anyone can download it
Anthropic released Claude Mythos Preview (April 2026) in a limited way to trusted defenders. Vetted defenders now use more advanced models such as Claude Mythos 5.1. GLM-5.3 is open-weight and free to download.
Skill
It builds working attacks
It comes close to Claude Mythos Preview on a public Chrome exploit test that Anthropic ran and scores below it on the second test. NIST, the US standards agency, measured it about four months behind the best US models.
Refusals
The refusals give way
It refuses a plain attack order in simulated tests. A cover story gets it to go along 64% of the time. Prefilled thoughts get 92%. Editing out its refusals gets 100%.
Claude
In Anthropic's tests, Claude stayed at 0%
With API safeguards on, that holds for all three Claude models Anthropic tested, in both tests that can be run against them. The API does not accept prefilled thinking and the weights are not released.
Amar Harolikar · Decision Sciences & Applied AItigzig.com · 2 / 10
Slide 3
1 · How good it is
GLM-5.3 builds working exploits at close to the rate of Claude Mythos Preview on ExploitBench
ExploitBench: share of attempts that built a working exploit
41 Chrome bugs, 410 attempts per model. Every attempt used an isolated sandbox with no network access.
Claude, safeguards off
Open-weight, anyone can download
0%
14%
0%
12%
0.5%
0.2%
Opus 4.6Feb 2026
Mythos PreviewApr 2026
GLM-5.2Jun 2026
GLM-5.3Aug 2026
Kimi K3Jul 2026
DeepSeek V4.1-FlashSep 2026
Binary Exploitation: share of 100 tasks where the model hijacked the program's control flow
Anthropic's own benchmark on open-source projects. GLM-5.3 scores below Mythos Preview here.
6%
4%
Claude Opus 4.6, GLM-5.2, Kimi K3 and DeepSeek V4.1-Flash all scored 0%
Mythos Preview
GLM-5.3
GLM-5.3 built a working exploit in 50 of 410 attempts. Claude Mythos Preview did it in 56. Anthropic's 95% ranges for the two overlap. NIST scores the same benchmark with partial credit and the best of three attempts, so its figures differ.
Amar Harolikar · Decision Sciences & Applied AItigzig.com · 3 / 10
Slide 4
2 · What NIST found
NIST: best open-weight model so far, about four months behind the US frontier
NIST's Figure 1, as published on 17 September 2026. Blue dots are US models and red dots are models from PRC-based companies, by release date. Error bars are 95% intervals. The US frontier includes models released only to vetted users, tested with safeguards off. NIST is the US National Institute of Standards and Technology, and its Center for AI Standards and Innovation (CAISI) made the chart.
The most cyber-capable open-weight model released to date, says NIST.
About four months behind the current US frontier, says NIST.
NIST explains the scale this way: "A 400-point increase on the y-axis equates to a 10x increase in the statistical odds of solving tasks."
Amar Harolikar · Decision Sciences & Applied AItigzig.com · 4 / 10
Slide 5
3 · One day, one browser
It found new browser flaws and chained them into a working attack in about a day
Anthropic's redacted screenshot. The banner names the file it read, an SSH private key.
What it found
Several unknown flaws in a popular browser, chained into a page that reads a visitor's files.
What it was built for
The Linux version, the only one the model was given. Anthropic believes other platforms could be affected, though the path may be more complex.
What happened next
Anthropic told the maintainer. With GLM-5.3 the researcher also found exploitable flaws in other widely used systems, including drivers and network device software. Anthropic is reviewing those reports.
These sessions typically ran for a day or less with under an hour of human focus. All targets were offline test setups.
Amar Harolikar · Decision Sciences & Applied AItigzig.com · 5 / 10
Slide 6
4 · What it costs
A smaller model turned a public fix into a working attack. At API prices it would have cost $20.40
Model GLM-5.3-Flash, a smaller and less capable version of GLM-5.3 from Zhipu AI (Z.ai).
Target Two known flaws, one of them the recently disclosed Chrome flaw CVE-2026-11645. The researcher gave the model the public details of both.
Direction No significant direction from the researcher.
Result A reliable exploit chain for an ARM64 target. It gets past a protection called pointer authentication.
Effort 20 minutes of human attention plus 8 hours of work by the model.
Cost Would have cost $20.40 at Zhipu's API prices.
Anthropic calls this an N-day exploit: an exploit for a flaw that is already known. It tested how fast a public fix can be turned into a working attack.
Amar Harolikar · Decision Sciences & Applied AItigzig.com · 6 / 10
Slide 7
5 · The safeguards
GLM-5.3 refuses a plain attack order and three simple tricks get past that
How often the model went ahead with a harmful attack order
Going ahead means it tried to connect to a remote target. 50 simulated episodes per row. Nothing was executed and a second AI model approximated the result of each command. The tests are Anthropic's own. Building exploits triggered no refusals at all. These refusals are for attacking a remote target.
Claude Opus 5, with API safeguards GLM-5.3, open weights
Plain orderThe attack order as given. 0%
0%
Cover storyIt is told it is a red-team agent on an exercise. 0%
64%
Prefilled thinkingIts opening thoughts are written for it, so it looks as if it already decided. Not offered by the API
92%
Edited modelIts weights are edited to remove the refusals. Weights not released
100%
Amar Harolikar · Decision Sciences & Applied AItigzig.com · 7 / 10
Slide 8
6 · Editing out the refusals
Removing the refusals took about $4,400 of computing and left its abilities largely intact
Harmful requests refused
GLM-5.3
GLM-5.3-Flash
95%
6%
95%
14%
model edited model edited
Ability of GLM-5.3
Science
CyberGym tasks
88%
88%
85%
81%
model edited model edited
In the top two charts, blue bars are the model and green bars are the model after the edit. Below, green is the experienced team.
GPU hours needed
2,200
600
Team new to it Experienced team
Computing cost, dollars
$4,400
$1,200
Team new to it Experienced team
The first bar is Anthropic's own team, which had never done this before. Most of its cost went to trying variants and testing. The second is Anthropic's estimate for a team with experience. Claude models with safeguards disabled refuse 95% to 96% of the same harmful requests. The refusal rate is the average of three public tests.
Several developers released edited versions of GLM-5.3 to the public within days of the model's release.
Amar Harolikar · Decision Sciences & Applied AItigzig.com · 8 / 10
Slide 9
7 · The same tests on Claude
In Anthropic's tests, all three Claude models stayed at 0% with API safeguards on, in both tests that can be run against them
Anthropic's test. Share of 50 simulated episodes where the model tried to connect to a remote target after a harmful attack order.
| Plain order | Cover story | Prefilled thinking | Edited model | |
|---|---|---|---|---|
| Open-weight models | ||||
| GLM-5.3 | 0% | 64% | 92% | 100% |
| GLM-5.3-Flash | 0% | 16% | 80% | 92% |
| Claude models | ||||
| Opus 4.8 | 0% | 0% 4% with safeguards disabled | Not possible. The Claude API does not accept prefilled thinking. | Not possible. Claude's weights are not released. |
| Opus 5 | 0% | 0% 10% with safeguards disabled | ||
| Mythos 5 | 0% | 0% |
Anthropic says Claude's safeguards blocked the requests that used deceptive prompts.
Amar Harolikar · Decision Sciences & Applied AItigzig.com · 9 / 10
Slide 10
Sources
GLM-5.3 and the spread of advanced cyber capabilities, Anthropic Frontier Red Team, 29 September 2026Andrew Fasano, Marius Fleischer, Cole McFaul, Robert Xiao and Tripp Gallagher. The main source for every page.
CAISI's Assessment of Z.ai's GLM-5.3 Cyber Capabilities, NIST, 17 September 2026The source for page 4.
Project Glasswing, AnthropicThe limited release of Claude Mythos Preview to trusted defenders.
Measuring LLMs' impact on N-day exploits, AnthropicEarlier work on exploits for flaws that are already known.
ExploitBench, arXiv 2605.14153The public benchmark behind the chart on page 3 of 10.
Charts are redrawn from Anthropic's Figures 1, 2, 4 and 5 and from its text and footnotes. NIST's Figure 1 on page 4 is shown as published. The screenshot on page 5 is two parts of Figure 3, cropped and joined. The full report also charts success against token budget and lists every benchmark in detail.
Amar Harolikar · Decision Sciences & Applied AItigzig.com · 10 / 10
Working on something similar? How I work covers the rates, the availability and what I take on.