Answer in brief
Two assessments describe notable cyber capability in an open-weight model, but they test different questions. Their results need careful reading before claims about real-world attacks.
A new test, an earlier government baseline
Anthropic published an assessment of Z.ai's GLM-5.3 on 29 September, two weeks after the US National Institute of Standards and Technology released its own cyber evaluation. NIST described GLM-5.3 as the most cyber-capable open-weight model it had assessed, yet below the current US frontier on its aggregate measure. Anthropic focused more narrowly on whether a model could build working exploits and whether simple simulated prompts could get past safeguards. The studies therefore inform one another without producing a single ranking for every use.
Read the denominators before the headline
In Anthropic's controlled ExploitBench trials on known V8 vulnerabilities, GLM-5.3 produced an end-to-end exploit in 50 of 410 attempts; its restricted Claude Mythos Preview did so in 56 of 410. Anthropic also reported a 4% result for GLM-5.3 on a selected 100-task internal binary-exploitation benchmark. These figures refer to configured tests and specific success criteria. They do not estimate the chance that an arbitrary website will be compromised or tell a defender how many new flaws exist in a particular product.
Access and safeguards are separate questions
NIST's comparison included frontier US models whose cyber safeguards were disabled for testing and, in some cases, whose releases were restricted to vetted users. Anthropic argues that freely downloadable GLM-5.3 changes the practical risk because its own simulated safeguard-bypass trials succeeded at rates from 64% to 100%, depending on the method. That is a claim about the evaluated configuration and attack prompts, not proof that all deployments behave alike. It also leaves room for defensive uses of the same capability when researchers have authorization.
What evidence would change the assessment
Security teams should take the results as a reason to review exposure, logging and patch response rather than as evidence of a specific incoming campaign. Future independent replications, documentation of model versions and real incident analysis would help show how benchmark ability transfers outside a lab. The two sources also have different interests: NIST sets a comparative public baseline, while Anthropic evaluates a competing model and its own safeguards. As of the 30 September cutoff, the measured test results are public; the scale of any real-world effect remains uncertain.
Questions and answers
Did Anthropic report attacks against live users?
No. The published comparison describes isolated benchmark and researcher-led tests against controlled targets. It should not be read as a count of incidents in the wild.
Why do the NIST and Anthropic findings differ?
They emphasize different measurements. NIST compares aggregate cyber capability across four benchmarks; Anthropic adds end-to-end exploit tests and simulated attempts to bypass model safeguards.
