Skip to content
← Back to Insights

Z.ai Launches GLM-5.3: Cyber Benchmark Reaches 84.5%, Open Weights Still Two Weeks Away

AI Z.ai GLM-5.3 Open Weights Cybersecurity AI Agents News

TL;DR

Z.ai released GLM-5.3 on August 14, 2026, claiming coding and exploitation gains from post-training alone; public benchmarks improved, but weights remain unavailable and harder cyber tests still trail closed models.

Z.ai Launches GLM-5.3: Cyber Benchmark Reaches 84.5%, Open Weights Still Two Weeks Away

The claim that GLM-5.3 has approached closed-model cyber capability will become testable in two weeks. Z.ai must release the promised weights, and independent teams must rerun the model with the same tools, time budgets, and task sets. If the CyberGym result of 84.5% cannot be reproduced, or if the model continues to trail badly on harder exploitation tasks, the release will show that post-training improved vendor-run evaluations—not that operational offensive capability has caught up with the frontier.

Z.ai released GLM-5.3 on August 14, 2026. The company says it retained the same base model as GLM-5.2 and obtained every gain by scaling post-training environments, task diversity, and compute. On public evaluations, Terminal-Bench 3.0 rose from 4.6 to 28.3, while DeepSWE v1.1 improved from 46.2 to 66.9. Z.ai also reports a 50% gain on its internal Z.ai Code Bench. That private benchmark may reduce contamination from public test sets, but outsiders cannot inspect its full task collection, failure cases, or scoring process.

The cybersecurity results tell two different stories. GLM-5.3 scored 84.5% on CyberGym, above GLM-5.2 at 77.2%, Mythos 5 at 83.8%, and GPT-5.6 Sol at 83.6%. Z.ai ran this as single-run Pass@1 over 1,507 tasks in Claude Code 2.1.207, without a total timeout per task. On ExploitBench, GLM-5.3 more than doubled its predecessor’s result, moving from 24.4% to 54.4%. Yet Mythos 5 reached 78.0% and GPT-5.6 Sol reached 76.5%. The model is competitive at finding and validating isolated flaws, while a substantial gap remains when a task requires deeper reasoning through an exploitation chain.

Z.ai further says that security teams working with the company found 2,436 vulnerabilities across 269 real projects after expert review, screening, and deduplication. The release does not provide the project list, a control group, or a complete false-positive rate, so that total cannot be converted directly into deployment value. Reuters independently reported the launch on the same day and focused on the comparison with Anthropic’s Mythos 5, but the benchmark figures themselves still originate with Z.ai.

API Access Comes Before the Weights

GLM-5.3 is available through Z.ai’s API and Coding Plan. Users can select low, high, or max reasoning effort, while disabling reasoning is no longer supported. Z.ai says the model weights will become publicly available within two weeks. Until those files appear, “open weights” describes a delivery commitment rather than something researchers can download and audit today.

The release also identifies a deployment constraint. Generating and verifying the long-horizon training environments still requires meaningful human involvement, so the post-training pipeline is not fully automated. Evaluation settings vary by benchmark as well: context limits, generation lengths, time budgets, and tool access differ. Comparisons are useful only when those conditions travel with the headline scores.

Three observable results now matter: whether the weights arrive by the promised deadline, whether outside teams reproduce the 84.5% CyberGym score, and whether later versions narrow the gaps on ExploitBench and ExploitGym. Those checks will determine whether GLM-5.3 represents auditable open-model progress or mainly a strong set of vendor benchmarks awaiting external verification.

Sources:

Get the latest insights

Join the newsletter to receive my latest articles on GenAI, AI Agents, and architecture.

No spam. Unsubscribe anytime.