Z.ai Unveils GLM-5.3 with Major Enhancements for Coding and Cybersecurity

Blog WriterCybersecurity News - Original News Source is cybersecuritynews.com

Spread the love

Z.ai has released GLM-5.3, a new AI model built to handle complex coding, long-running agent tasks, and cybersecurity analysis. The company says the model uses the same base model as GLM-5.2, with improvements coming entirely from larger-scale post-training.

The release focuses on training AI agents in realistic task environments rather than short coding exercises. These environments can include codebases, documentation, storage systems, experiment results, testing tools, and multi-step workflows.

In one example, an agent may need to identify a bottleneck in a machine-learning training stack, implement an optimization, run tests, and demonstrate improved performance without breaking functionality.

LLM Performance Evaluation (source : z.ai )
LLM Performance Evaluation (source : z.ai )

Z.ai said GLM-5.3 delivered major coding gains on several benchmarks. It scored 28.3 on Terminal-Bench 3.0, compared with 4.6 for GLM-5.2. On DeepSWE v1.1, the new model reached 66.9, up from 46.2.

The company also reported a 50% improvement on its internal Z.ai Code Bench, which measures coding agents in more realistic local development environments. The model includes three reasoning settings: low, high, and max.

GLM-5.3 Major Enhancements

Z.ai recommends the max setting for coding tasks because it allows the model to spend more effort planning, implementing, testing, and verifying work. Unlike older versions, GLM-5.3 does not support fully disabling reasoning.

Cybersecurity is one of the most notable areas of improvement. Z.ai said it added vulnerability-discovery data and security-focused task environments during post-training.

The company expected better bug finding, but said the model also improved at connecting several stages of an attack path, including vulnerability analysis and exploitation reasoning.

 Agent coding performance by effort level (source : z.ai )
Agent coding performance by effort level (source : z.ai )

On CyberGym, a benchmark that tests white-box vulnerability discovery in source code, GLM-5.3 scored 84.5%, compared with 77.2% for GLM-5.2.

On ExploitBench, its score rose from 24.4% to 54.4%. In ExploitGym, GLM-5.3 completed 105 exploitation tasks within two hours and 130 within six hours, while GLM-5.2 completed 29 and 39 tasks under the same time budgets.

Z.ai emphasized that the strongest gains appeared further along the exploitation chain. This could help defenders identify complex weaknesses that involve multiple components, unsafe assumptions, and chained flaws.

Cybersecurity Evaluation (source : z.ai )

However, the company also acknowledged that leading closed models still scored higher on some exploitation benchmarks. The company said GLM-5.3 has already been tested with security teams against real-world codebases.

After expert review and duplicate removal, the model reportedly identified 2,436 vulnerabilities across 269 projects. Of these findings, 1,097 were rated medium to high severity.

The affected software reportedly included kernels, operating systems, browser engines, web applications, network protocols, and open-source infrastructure.

Z.ai has created a public Security Disclosure Ledger to track findings through coordinated disclosure. At launch, 53 findings had been publicly disclosed, while 2,383 remained under embargo.

The oldest reported flaw was introduced in 1981, showing how long vulnerabilities can remain hidden in widely used code. Model weights are expected to be released two weeks after launch, following safety evaluation and hardening.

 Strengthen Your SOC by Accelerating Threat Detection & Rapid Investigations. -> Integrate ANY.RUN With Your SOC Now.

The post Z.ai Unveils GLM-5.3 with Major Enhancements for Coding and Cybersecurity appeared first on Cyber Security News.