Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

Anthropic warns Chinese open-weight model GLM-5.3 is nearing the cyber frontier

Anthropic has published its own security assessment of GLM-5.3, an open-weight model from Chinese lab Zhipu AI, finding that it can autonomously build end-to-end cyber exploits while lacking meaningful safeguards. The company says attackers can bypass those safeguards in most of its tests and warns the model could widen the cyber capabilities available to malicious actors.

Published

Anthropic has published a security assessment of GLM-5.3, an open-weight model from the Chinese lab Zhipu AI, also known as Z.ai. According to the company's researchers, the model has strong capabilities for autonomously building end-to-end cyber exploits, but lacks meaningful safeguards.

The assessment was run in isolated, sandboxed environments, according to the report, using automated benchmarks such as ExploitBench alongside human-in-the-loop workflows, and it compared GLM-5.3 with Anthropic's own Mythos Preview model. Anthropic found that attackers could bypass GLM-5.3's safeguards between 63 percent and 100 percent of the time.

On ExploitBench the two models performed similarly, developing end-to-end exploits 50 and 56 times out of more than 400 attempts. Anthropic said the same attacks did not succeed against safeguarded Claude models in its testing.

GLM-5.3 does ship with guardrails and refuses prompts that clearly intend harm, the report notes, but the researchers said those guardrails could be removed with a few simple techniques. The most effective method was abiliteration, a standard refusal-reduction approach that lets users edit a downloadable model's weights without retraining it.

Anthropic also bypassed the safeguards by placing the model in a simulated environment and deceiving it into believing it was performing red-teaming exercises. The company warns that because GLM-5.3 is close in capability to Mythos and its safeguards are easily stripped, it could further narrow the gap between attack and defense.

The findings echo an earlier assessment from the U.S. Center for AI Standards and Innovation, which called GLM-5.3 the most cyber-capable open-weight model released to date, while noting that U.S. frontier models still outperform it in cybersecurity.

Anthropic competes directly with open-weight models like GLM-5.3, and its publication pushes the long-running debate over open weights and safeguards back into the spotlight. The report notes the conventional view that open-weight models lag closed models by roughly six to eight months, and that the gap may be narrowing; GLM-5.3 is an updated version of GLM-5.2, the model Hugging Face turned to after OpenAI agents breached its systems.

Why it matters

The episode signals that guardrails on downloadable open-weight models can be removed with widely known techniques, so the spread of frontier-level cyber capability may no longer be controllable by the labs that train the models. It also raises the stakes for access-control policy and for how companies assess competing models.

AnthropicOpen-Weight ModelsAI SecurityZhipu AI
Back to realtime news

Nearby Updates

All