September 21, 2026
GLM-5.3 trails US models in finding vulnerabilities by 4 months: assessment by a NIST-affiliated institute
On SEC-Bench Pro, GLM-5.3 scored 40,4%, while the strongest American model scored 90,2%.

On cyber tasks, GLM-5.3 is stronger than any other open model. Yet it still trails the strongest US models by four months.
CAISI, an institute at NIST, ran the model through four vulnerability-finding benchmarks and published the figures on 17 September. Open weights are now measured not by a vendor's claims, but by an external run on identical tasks.
Where the gap is sharper. On SEC-Bench Pro, the model solved 74 of 183 tasks, or 40,4%. The strongest American model scored 90,2%. On ExploitGym (Userspace), the score drops to 9,4% versus 44,4%.
In August, Z.ai gave its own figures: ExploitBench at 54,4% versus 24,4% for the previous GLM-5.2. CAISI gave GLM-5.3 even more on the same benchmark, 61,1%, but the strongest US model scored 100%.
How to get it. The weights are open: the zai-org/GLM-5.3 repository on Hugging Face, 753B MoE parameters, launched through vllm or sglang. You cannot run this at home. The FP8 checkpoint weighs about 753 GB, BF16 about 1,51 TB; it requires a server with several accelerators.
That leaves the API: model id glm-5.3, a 1 million-token window, and reasoning can no longer be disabled. On Z.ai's price list, input costs $1,4 per million tokens and output $4,4; Claude Opus 5 output costs $25. The GLM Coding Plan subscription starts at $18 per month.
This is CAISI's second look at Z.ai models: in July, the institute examined GLM-5.2 across 21 pages.
