Anthropic has released an analysis of GLM-5.3, the latest model from Zhipu AI, finding that the model possesses strong capabilities for autonomously building end-to-end cyber exploits while lacking sufficient safeguards to prevent misuse.
SecurityAnthropicZhipu AIGLM-5.3
Anthropic Analysis Reveals GLM-5.3's High Cyber Capabilities and Vulnerability to Safeguard Bypassing
In simulated tests, Anthropic found that attackers can bypass GLM-5.3’s safeguards between 64% and 100% of the time using simple techniques. For instance, applying "abliteration"—a technique used to remove model refusals—reduced the refusal rate from over 90% to as low as 2% to 12% on various benchmarks without significantly degrading the model's general reasoning capabilities. This stands in contrast to Anthropic's safeguarded Claude models, which did not succumb to similar attacks in testing.
In practical evaluations, researchers used GLM-5.3 to identify and chain vulnerabilities in a Linux build of a web browser, creating an exploit capable of reading arbitrary files from a user's computer. In a separate test, a smaller version, GLM-5.3-Flash, built a reliable exploit chain for an ARM64 target in approximately 8 hours by leveraging public details of a known vulnerability (CVE-2026-11645).
Anthropic noted that while such capabilities can benefit cyber defenders, the release of powerful models in an open-weight format—which allows users to reconfigure the model and remove refusals—significantly increases the tools available to malicious actors. The analysis highlights the importance of independent safety testing for sufficiently capable AI models to prevent real-world harm.
Sources
- GLM-5.3 and the spread of advanced cyber capabilities (Hacker News Frontpage, 2026-09-29)
- 「中国のGLM-5.3はClaude Mythos Preview級のサイバー攻撃能力を持つ一方で安全対策が不十分」とAnthropicが警告 (GIGAZINE, 2026-09-30)