Interestana
Home/News/Z.ai Ships GLM-5.3 With Scaled Post-Training Gains
MarkTechPost4 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Z.ai Ships GLM-5.3 With Scaled Post-Training Gains

Z.ai has released GLM-5.3, a new iteration of its language model that demonstrates substantial performance gains, particularly in complex coding and long-horizon tasks. Notably, GLM-5.3 operates on the same 743 billion parameter base model as its predecessor, GLM-5.2. The reported advancements are attributed entirely to scaled post-training methodologies, which involved expanding the variety and complexity of task environments and increasing the duration of training. This approach has yielded significant improvements across several key benchmarks.

In coding, GLM-5.3 shows a dramatic increase on the Terminal-Bench 3.0 benchmark, moving from a score of 4.6 with GLM-5.2 to 28.3. This benchmark is designed to evaluate performance on long-horizon coding tasks. Further coding improvements are evident on the DeepSWE v1.1 benchmark, where GLM-5.3 achieved a score of 66.9, up from 46.2. The Agents' Last Exam (CLI) benchmark also saw an increase, from 23.8 to 28.5. On the GDPval-AA v2 benchmark, which assesses performance across 44 different occupations, GLM-5.3 scored 1,769. Z.ai's internal evaluation, the Z.ai Code Bench, reported a 50% improvement over GLM-5.2, with GLM-5.3 achieving 31.4% accuracy at approximately 50,000 output tokens per task. For comparison, Claude Opus 4.8 scored 29.5% at 120,000 tokens, and Claude Fable 5 achieved 39.5% at its maximum effort, though Z.ai argues its private benchmark reduces contamination risk.

Beyond coding, GLM-5.3 has made significant strides in cybersecurity. The CyberGym benchmark reached a score of 84.5%, a performance level that Z.ai states exceeded their expectations. The company has made GLM-5.3 partially deployable, with access available through the Z.ai API, the GLM Coding Plan, and ZCode. However, the model weights are not yet publicly available. Z.ai plans to release the weights approximately two weeks after the launch, following the completion of safety evaluations and hardening processes. This phased release strategy is intended to ensure model security and stability before wider distribution.

Z.ai suggests that startups and mid-market engineering organizations can begin adopting GLM-5.3 immediately through the Coding Plan or API. Enterprises with strict data-residency requirements or extensive vendor review processes are advised to wait for the public release of the weights. The model is expected to provide significant value to security vendors and Managed Security Service Providers (MSSPs), offering enhanced signal detection and policy exposure capabilities. Industries that stand to benefit include developer tooling, cloud infrastructure, application security, fintech and e-commerce engineering, as well as vendors developing core software components like kernels, browser engines, and network stacks. Potential applications range from large-scale repository refactoring and long-horizon CLI agents to CI failure triage, white-box vulnerability discovery, crash analysis, and secure code review.

Original source — read the full reporting at the publisher:

Read on MarkTechPost

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next