GLM-5.3 Scales Coding and Cyber Capability

Z.ai kept the GLM-5.2 base model and scaled post-training, reporting large coding gains and unexpectedly fast growth in exploitation capability; weights will follow after two weeks of safety work.

✓ Verified Source Z.ai official launch post and security disclosure ledger; benchmarks and vulnerability counts are vendor-reported ⚑ Model release

The 60-second version

GLM-5.3 keeps the GLM-5.2 base model but scales long-horizon post-training, producing reported gains in coding and cyber tasks.

Key points

  • Z.ai attributes the upgrade to more environments, more diverse tasks, and more reinforcement-learning compute rather than new pretraining.
  • Vendor-reported coding scores rise sharply, although GLM-5.3 does not lead every benchmark or every closed model.
  • The largest relative gains appear in multi-stage vulnerability exploitation, making the release materially dual-use.
  • Z.ai will publish weights after a two-week safety evaluation and hardening period; the hosted model is available now.
  • Thinking cannot be disabled, so production migrations must account for reasoning effort, latency, tokens, and tool controls.

Verdict. The technical story is credible as a first-party launch, but the performance and security numbers remain vendor evidence until independently reproduced; the weight release and safety documentation are the next test.

What changedThe base stayed; post-training scaled

GLM-5.3 uses the same base model as GLM-5.2. Z.ai says every gain came from another month of post-training: more executable environments, more varied long-horizon tasks, and more reinforcement-learning compute.

The tasks are designed as units of professional work rather than isolated exercises. An agent may need to inspect code, infrastructure, documentation, and experiment results, then implement and verify an improvement end to end. Environment synthesis, judge agents, and hidden-state verifiers help scale that training, although Z.ai says meaningful human involvement remains.

CodingThe reported gains are large, not universal

28.3Terminal-Bench 3.0, up from 4.6
66.9DeepSWE v1.1, up from 46.2
28.5Agents' Last Exam, up from 23.8

Z.ai calls GLM-5.3 the strongest open-weights coding model, even though its weights are not available at launch. Its own table shows a more nuanced result: the model leads some open-model tests but trails certain closed systems and does not top every benchmark.

Z.ai's comparison table covers coding, cyber, and agentic benchmarks. The values should be read as vendor-reported results under the footnoted evaluation settings.
Z.ai's comparison table covers coding, cyber, and agentic benchmarks. The values should be read as vendor-reported results under the footnoted evaluation settings. · Z.ai

Dual useCyber capability grew fastest downstream

Z.ai added vulnerability-discovery environments on purpose. The unexpected part, it says, was the speed at which the model moved from finding isolated flaws toward planning multi-stage exploitation chains.

CyberGym84.5% for GLM-5.3 versus 77.2% for GLM-5.2, according to Z.ai.
ExploitBench54.4% versus 24.4%; some closed models remain substantially ahead.
ExploitGym105 tasks in two hours and 130 in six, versus 29 and 39 for GLM-5.2 under Z.ai's normalized budgets.

The company reports 2,436 findings across 269 open-source projects, including 1,097 critical and high-severity issues. Its public disclosure ledger repeats those totals and distinguishes 53 disclosed findings from 2,383 still under embargo. The ledger improves traceability, but it is still maintained by Z.ai rather than an independent auditor.

Release controlWeights wait while the service ships

GLM-5.3 is available through Z.ai's Coding Plan and ZCode, but the company says the weights will be published two weeks after launch, after safety evaluation and hardening. The eventual safety report, license, and weight safeguards will determine how meaningful that delay is.

Bottom lineEvaluate the agent, not only the base model

Teams considering GLM-5.3 should retest coding quality, latency, token use, tool permissions, sandbox boundaries, and security review as one package. The release shows that post-training can materially change both usefulness and dual-use risk without changing the underlying base model.