Anthropic published a September 29, 2026 analysis of GLM-5.3, the latest open-weight model from Zhipu AI (Z.ai), focusing on its ability to develop end-to-end cyber exploits. The report matters because Anthropic says GLM-5.3 crossed a capability threshold similar to its earlier Claude Mythos work, while its safeguards were easier to bypass or remove.
What Anthropic tested
Anthropic says the testing was performed in isolated and sandboxed environments so the evaluated models could only attack offline targets prepared for the research. The evaluation focused on end-to-end exploit development because that capability is most relevant to real attack chains.
On ExploitBench, which tests exploitation of known Chrome V8 vulnerabilities, Anthropic reports GLM-5.3 successfully developed end-to-end exploits in 50 of 410 attempts. Its earlier Claude Mythos Preview reached a similar 56 of 410 attempts in the comparison.
Why open-weight access changes the security picture
The issue is not only capability. Anthropic argues that GLM-5.3 was released as an open-weight model without safeguards strong enough to limit misuse. The company says users can modify an open-weight model to remove refusal behavior, which changes the security assumptions around deployment.
Anthropic tested an altered version of GLM-5.3 and found that its harmful-request refusal rates dropped sharply while general capability remained largely intact on the tested benchmarks. This creates a different risk profile from a hosted model where the provider controls the weights and the safety layer.
What the safeguard tests showed
| Test setup | GLM-5.3 engagement reported by Anthropic |
|---|---|
| Bare malicious order | 0% in the tested simulated trials |
| Deceptive cover story | 64% |
| Prefilled reasoning | 92% |
| Abliterated model | 100% |
These are results from Anthropic's own simulated tests, not a measure of what every real-world deployment will do. They do show why teams should not assume that an open-weight model's default refusal behavior will remain intact after local modification.
The exploit capability threshold
Anthropic also reports that GLM-5.3 completed full control-flow hijacks in 4% of trials on its internal Binary Exploitation benchmark, while Claude Mythos Preview reached 6% in the same comparison. Earlier models tested in the report did not succeed on those tasks.
The report also describes human-led sessions in which researchers used GLM-5.3 to identify previously unknown browser vulnerabilities and chain them into an exploit in a sandbox. Anthropic says those vulnerabilities were disclosed to maintainers rather than used against live targets.
What defenders can take from the report
- Keep powerful cyber models isolated. Run exploit-generation tests against offline targets and synthetic environments.
- Assume local model changes can remove safeguards. If weights are available, evaluate the modified model rather than the vendor's default checkpoint alone.
- Measure the full attack chain. Finding a bug, building an exploit and reaching a target are different capabilities.
- Use layered controls. Model refusals should sit alongside network boundaries, tool restrictions, credential controls and logging.
- Track disclosure. When research identifies real vulnerabilities, coordinate with maintainers and document the disclosure path.
How this differs from Anthropic's earlier threat reporting
Anthropic's broader threat reports focus on how existing AI systems are used in real malicious operations. The GLM-5.3 analysis is different: it studies a model's intrinsic cyber capability and how easily its safeguards can be bypassed. Together, the two perspectives show why AI security needs both misuse monitoring and pre-deployment capability testing.
Limits of the findings
The results come from controlled evaluations designed by Anthropic. The company says it used isolated environments and human-in-the-loop workflows, and it notes that some comparisons involve models with safeguards disabled. The findings therefore should not be read as a direct prediction of attack rates in the wild.
Anthropic also says it is disclosing relevant vulnerabilities to maintainers. The report is best used as a risk signal for defensive testing rather than a recipe for offensive use.
Frequently asked questions
What is GLM-5.3?
GLM-5.3 is an open-weight AI model developed by Zhipu AI, also known as Z.ai. Anthropic's September 29 report analyzes its cybersecurity capabilities.
Did Anthropic test GLM-5.3 against live companies?
No. Anthropic says its exploit evaluations used isolated and sandboxed environments with offline targets prepared for the research.
Why does open-weight access matter?
Open weights allow users to modify a model. Anthropic found that an altered version of GLM-5.3 could have much lower refusal rates on harmful-request benchmarks while retaining most of its general capabilities in the tests reported.
Sources
- Anthropic — GLM-5.3 and the spread of advanced cyber capabilities
- NIST — Caisi assessment of GLM-5.3 cyber capabilities