Z.ai Ships GLM-5.3 Without Retraining the Base Model: Better at Complex Coding and Long-Horizon Tasks
- Z.ai just dropped GLM-5.3 and here's the kicker — they didn't touch the base model. Same 743B GLM-5.2 engine, just more post-training bolted on. Terminal-Bench 3.0 jumped from 4.6 to 28.3, which is either impressive or a sign the old number was embarrassing. DeepSWE v1.1 went 46.2 to 66.9. CyberGym hit 84.5%, edging Mythos 5 and GPT-5.6 Sol. ExploitBench more than doubled to 54.4% — still trailing Mythos 5 at 78%, but the trajectory is the story. Z.ai admits the cybersecurity gains were unplanned. They fed it vuln-discovery data expecting incremental improvement and got coherent exploitation chains instead. That's not a feature, that's a liability with a release date. Weights ship in about two weeks pending safety hardening. API and Coding Plan are live now. Startups can move today; enterprises with vendor-review rules should wait. The lesson: post-training scaling works, but the deeper you go into exploitation chains, the bigger the gains — and the bigger the questions about who gets access. Aggressive passive income for whoever ships the first red-team agent on this thing.