FrontierThe story, in brief

Z.ai Ships GLM-5.3 Without Retraining the Base Model: Better at Complex Coding and Long-Horizon Tasks

Z.ai just shipped a 743B model that punches way harder on coding and long-horizon tasks—without touching the base weights. Post-training scaling alone moved Terminal-Bench from 4.6 to 28.3.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

GLM-5.3 demonstrates that frontier capability gains no longer require expensive base-model retraining; scaled post-training on task-specific environments delivers outsized gains on reasoning and coding—a shift in where labs are investing.

The key facts

9 to know
  1. GLM-5.3 released August 14, 2026

  2. Reuses 743B GLM-5.2 base unchanged

  3. Terminal-Bench 3.0: 4.6 → 28.3 (+515%)

  4. DeepSWE v1.1: 46.2 → 66.9 (+45%)

  5. CyberGym: 84.5%

  6. ExploitBench: doubled to 54.4%

  7. Gains from scaled post-training only (longer training, more environments, more task types)

  8. Weights available in ~2 weeks

  9. Cybersecurity gains exceeded Z.ai's reported plans

Go to the source

MarkTechPostmarktechpost.com

Publisher excerpt: Z.ai released GLM-5.3 on August 14, 2026. The model reuses the 743B GLM-5.2 base unchanged. Every reported gain comes from scaled post-training: more long-horizon task environments, more environment types, longer training. Terminal-Bench 3.0 moves from 4.6 to 28.3, and DeepSWE v1.1 from 46.2 to…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier