HomeZ.ai GLM-5.3 tops CyberGym cybersecurity AI model benchmarkUncategorizedZ.ai GLM-5.3 tops CyberGym cybersecurity AI model benchmark

Z.ai GLM-5.3 tops CyberGym cybersecurity AI model benchmark

Z.ai has released GLM-5.3 with a CyberGym benchmark score of 84.5 percent, placing it ahead of rival AI models used for cybersecurity.

The previous leader, Anthropic’s Mythos 5, scored 83.8 percent on the benchmark. OpenAI’s GPT-5.6 Sol was only marginally behind Anthropic with a 83.6 percent score.

The company evaluated GLM-5.3 using the Claude Code 2.1.207 harness. It configured the model for maximum reasoning effort, with no web tools, a temperature of 1.0, top_p of 1.0 and a maximum output length of 128,000 tokens.

CyberGym contains 1,507 tasks. Z.ai reports a single-run Pass@1 result with no time limit per task.

The agent ran inside each task container. Z.ai says it removed Git-related information and applied a domain whitelist intended to prevent cheating. The permitted domains included pypi.org and deb.debian.org for basic tool installation.

That configuration affects how developers should read the result. The score reflects a constrained environment with source-code access, task containers and a defined verification process. Z.ai does not provide live vulnerability discovery rates, false-positive rates, or remediation outcomes for GLM-5.3.

Z.ai reports larger gains on exploitation benchmarks

CyberGym assesses vulnerability discovery and validation, according to Z.ai. The company also tested GLM-5.3 on ExploitBench and ExploitGym, which it presents as evaluations of later stages in vulnerability analysis and exploitation.

On ExploitBench, Z.ai reports a score of 54.4 percent for GLM-5.3. GLM-5.2 scored 24.4 percent in the same comparison, showing the apparent leap in capability. However, on this benchmark, the rival models continued to hold a significant lead with a 78 percent score for Mythos 5 and 76.5 percent for GPT-5.6 Sol.

ExploitBench used 41 tasks across three revisions. Z.ai limited each agent to 300 interaction rounds with its environment. The reported result is an average coverage score. Z.ai says each task’s coverage came from the combined capabilities achieved across the three revisions.

ExploitGym produced a separate time-normalised task-completion measure. GLM-5.3 completed 105 tasks within two hours and 130 tasks within six hours, according to Z.ai. GLM-5.2 completed 29 and 39 tasks under the same reported budgets.

Z.ai states that ExploitGym covers 869 tasks. It used Claude Code 2.1.207, maximum reasoning effort, no web tools, a 128,000-token output limit, temperature 1.0 and top_p 1.0.

The time budgets are not simple wall-clock limits. Z.ai says they combine non-API overhead with API inference time rescaled through model-specific tokens-per-second figures from Artificial Analysis. The company used 115 tokens per second for GLM-5.3, 40 for Kimi K3, and 47 for Qwen3.8 Max.

Mythos 5 was the only model that remained ahead of GLM-5.3 on the ExploitGym results. Z.ai lists 181 completed tasks in two hours and 247 in six hours for that model. For comparison, the third place model – Kimi K3 – only managed 36 tasks in two hours and 70 in six hours.

Results for various cybersecurity AI model benchmarks across GLM-5.3, GPT-5.6 Sol, Mythos 5, and Kimi K3.

Post-training added vulnerability environments

GLM-5.3 uses the same base model as GLM-5.2, according to Z.ai. The company attributes the stated coding and cyber gains to post-training.

Vulnerability-discovery data and task environments entered the training mix as part of that process. Z.ai says it expected gains in finding and reasoning about vulnerabilities, then observed further capability growth as post-training scaled.

The company says GLM-5.3 began to reason across multiple stages of exploitation and produce plans for complete exploitation chains. Its benchmark results provide the reported evidence for that statement, with the highest rank among its cited comparison models on CyberGym and lower scores than Mythos 5 and GPT-5.6 Sol on ExploitBench.

Z.ai says its wider post-training environments are executable and verifiable. Some represent long-horizon engineering and research tasks with multi-step dependencies and hidden state.

Research agents collect task patterns and generate runnable environments, according to Z.ai. A judge agent attempts tasks to establish whether they can be solved. Z.ai says its verifiers are synthesised without access to the reference solution and tested with oracle, no-op, and unsolved-state checks.

Security disclosure ledger records 2,436 findings

Z.ai says it has worked with several security teams in China since GLM-5.2 to run its models against real-world codebases. After expert review, screening, and deduplication, the company reports 2,436 vulnerabilities across 269 projects.

Its Security Disclosure Ledger lists 1,097 critical and high-severity findings. The ledger records 107 critical issues and 990 high-severity issues. Z.ai also lists 1,286 medium-severity findings and 53 low-severity findings.

The company says 53 findings have been publicly-disclosed. It lists 2,383 as under embargo.

Findings cover system kernels and operating systems. Z.ai also names browser engines, open-source infrastructure, web applications, and network protocols. The supplied material does not identify individual affected projects in the article text.

Z.ai says the oldest identified flaw was introduced in 1981. It reports that the average vulnerability remained in code for 26.6 years before discovery.

The disclosure ledger is intended to track findings through the disclosure process. For publicly disclosed issues, Z.ai says it records the affected project, severity, CVE where available, and the period the flaw remained in the codebase.

Developers face an API migration alongside the release

The release changes GLM-5.3’s thinking configuration. The model accepts low, high, and max values for reasoning_effort, with max as the default and Z.ai’s recommended setting for coding tasks.

Disabled thinking is no longer supported. Applications that currently send thinking.type: “disabled” must change that value to enabled and select reasoning_effort: “low” before updating the model ID to glm-5.3.

Z.ai states that requests will fail if developers update the model ID without making the configuration change. A compatible request uses enabled thinking:

{

 “model”: “glm-5.3”,

 “thinking”: { “type”: “enabled” },

 “reasoning_effort”: “max”

}

Z.ai says it will release GLM-5.3 weights two weeks after launch, following safety evaluation and hardening. Until then, the company’s 84.5 percent CyberGym result remains its reported basis for calling GLM-5.3 the leading model in that evaluation.

See also: OpenAI Daybreak adds GPT-5.6-Cyber for defensive security work

Banner for Cyber Security Expo by TechEx events.

Want to learn more about cybersecurity from industry leaders? Check out Cyber Security & Cloud Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the AI & Big Data Expo. Click here for more information.

Developer is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

Home
Services
Careers
Call Us
Contact