Anthropic and the US NIST find that Z.ai's open-weight model GLM-5.3 built working exploits in some tests and that simple techniques bypassed the model's refusals. These were simulated tests without access to real systems; the lesson is that such downloadable models require capability-specific tests, strict isolation and monitoring that keep working after modification.
On 29 September 2026, Anthropic's Frontier Red Team published an assessment of GLM-5.3 examining its cyber capabilities and the robustness of the model's refusals. In it, the company states that the model produced complete, working exploits in a material but minority share of the ExploitBench tasks, close to Anthropic's own Claude Mythos Preview. In addition, Anthropic reports that its simulated attempts were often able to bypass GLM-5.3's refusal mechanisms: 64% with misleading prompts and 92% with pre-filled reasoning tokens; for a modified, ablated version Anthropic reported 100%.
Important for interpretation: Anthropic is a competitor of Z.ai, and the figures on bypassing safeguards are simulated results reported by the company itself. The US NIST independently assessed GLM-5.3 as the most cyber-capable open-weight model released to date, although according to NIST it still trails current US frontier models in total capability.
What exactly did Anthropic and NIST find about GLM-5.3?
Anthropic reports that in its ExploitBench test GLM-5.3 built working end-to-end exploits in 50 of 410 attempts, compared with 56 of 410 for its own Claude Mythos Preview. It also reports that misleading prompts, pre-filled reasoning tokens and a modified version of the model bypassed the built-in refusals to a high degree; Anthropic states that its own Claude safeguards did not fail under comparable conditions.
- Exploit capability: working exploits in 50 of 410 ExploitBench attempts, a minority but relevant share of the tasks.
- Weak safeguards: bypass with simple techniques, rising to a fully removed refusal in a modified model version.
- Independent confirmation: NIST calls the model the most cyber-capable open-weight model to date, but still behind the top.
Z.ai itself documents in its official model repository the release and the downloadable distribution of GLM-5.3, including model weights and the technical context. That the model is freely downloadable is not a detail but the heart of the matter.
Why is 'exploit capability in a test' different from a real breach?
A model that writes an exploit in a controlled test has compromised nothing by doing so. Anthropic ran the tests in simulation, without giving the model direct access to external systems. Capability is therefore not the same as a real incident, and a reported test result is no proof of a public breach.
In our assessment, this very distinction is the reason to act now rather than wait: the risk is demonstrably present in the laboratory, and the step from laboratory to production is determined by how an organisation deploys the model, not by the model alone. A second distinction that, in our analysis, follows from the news: a model's default refusal is different from a protection that keeps working under adversarial prompts or after the weights are modified. Those who mistake the first for the second overestimate, in our assessment, their control.
What changes because GLM-5.3 is an open-weight model that can be modified?
With a hosted model, some safety measures are enforced by the provider. With an open-weight model you download, the operator must take responsibility for enforcing deployment controls, including controls that may otherwise be provided by the service. The reported bypasses, up to and including an ablated version for which Anthropic reported a bypass rate of 100% in simulation, underline, in our analysis, that a built-in refusal cannot be relied on as the only security layer after modification.
The sources leave one question open: which measures remain effective when the model is modified by third parties after download? Our answer, as editorial analysis: not the model's refusal layer, but the environment around it. In our assessment, that means isolation from production systems, least-privilege access, outbound network traffic blocked by default, and monitoring for exploit-like output and tool calls. These controls are intended to remain enforceable outside the model’s refusal layer even if the weights are changed, although their effectiveness still depends on correct implementation, configuration and monitoring. This aligns with the broader point that a high benchmark score is no proof for safe deployment.
What does this mean for directors, lawyers and CISOs working with sensitive information?
Our analysis: because GLM-5.3 can demonstrably produce working exploits, a generic AI approval is, in our assessment, insufficient; we therefore advise the responsible CISO to require a model- and version-specific cyber evaluation before any production deployment, and to record whether testing was done with hosted safeguards, modified weights, tools or network access. Because the refusal layer can fall away after modification, relying on the default refusal is, in our assessment, risky; the infrastructure administrator therefore preferably places the model in an isolated environment with default-deny outbound traffic and least-privilege credentials, to reduce the likelihood that a bypass can reach real systems Because the tests were simulated and a real incident cannot be ruled out, we advise the organisation to set up monitoring for exploit-oriented output and tool calls plus a stop, escalation and reporting procedure, so that intervention does not begin only once damage occurs. And because this concerns a downloadable model whose provenance and version matter, the lawyer or compliance officer, in our assessment, records model identity, weight provenance, evaluation conditions and approvals, so that it can later be reconstructed under which conditions the model was admitted.
Summarised as a checklist (in our analysis):
- Require a model- and version-specific cyber evaluation before approval.
- Note whether testing was done with hosted safeguards, modified weights, tools or network access.
- Isolate the model from production systems and sensitive data.
- Apply least-privilege credentials and default-deny outbound traffic.
- Monitor for exploit-oriented output and tool calls; retain logs and model provenance.
- Record stop, escalation and reporting procedures.
We saw the same logic in earlier incidents: treat the case where an AI test environment that hits production as a chain-wide incident response matter, and build secure private LLM inference in layers. More background on this kind of trade-off is in our hub on AI security and model protection.
Sources and references
Sources: The article relies on Anthropic's assessment of 29 September 2026, the independent assessment by NIST (CAISI) and Z.ai's official release documentation.