GLM-5.3

GLM-5.3 brings stronger coding and a cybersecurity edge to open-source AI.

GLM-5.3 is here, and it's not just another incremental step. Z.ai's latest model, built on the same base as GLM-5.2, shows what happens when you pour everything into post-training: coding scores jump from 4.6 to 28.3 on…

4 min readVentureBeat
GLM-5.3 brings stronger coding and a cybersecurity edge to open-source AI.

**The post-training ceiling keeps moving. That's the story of GLM-5.3, and it's a meaningful one for anyone building on these models.** Z.ai didn't scale up a new base model this time. Instead, it took the same 743-billion-parameter foundation from GLM-5.2 and pushed it entirely through post-training, longer-horizon environments, and more reinforcement-learning compute. The result is a jump on Terminal-Bench from 4.6 to 28.3, and a 66.9 on DeepSWE. That's not just incremental. It's evidence that the frontier isn't only about pretraining dollars anymore. Explore Xiaomi’s MiMo-V2.6: AI Model Training Achieves $3.5M Benchmark shows a similar logic: the cost of training is dropping, and the leverage is shifting to how you run the loop, not just how big the model is. But GLM-5.3 also brings a complication that's harder to price: its cyber capabilities advanced faster than Z.ai expected, and it reportedly found a serious vulnerability in Cursor. That's not a bug. It's the point.

**Here's the practical tension for developers.** The same long-horizon agentic training that makes GLM-5.3 better at fixing your codebase makes it better at breaking into one. Z.ai says it introduced vulnerability-discovery environments expecting gains in finding flaws. Instead, capability progressed further down the exploitation chain. That's why they're holding back open weights for about two weeks and gating access behind a "trusted access" approach. For you, that means two things. First, if you're deploying coding agents, the line between "helpful" and "offensive" is now blurring inside the model itself. Second, the safety controls aren't just a PR gesture. They're a direct response to the model's own trajectory. The company reported 2,436 vulnerability findings across 269 projects, with 1,097 rated critical or high severity. That's not a hypothetical risk. It's a real capability, and it's moving faster than the governance around it.

**Here's the good news: you don't need to wait for the open weights to benefit.** GLM-5.3 is available now through the GLM Coding Plan and ZCode, and it's already showing a practical edge beyond benchmarks. It's more token-efficient than GLM-5.2, scoring 34.5% on Z.ai's private Code Bench at Max effort while using about 75,000 output tokens per task, compared to 96,000 for the previous version at a lower score. That matters in production, where long-running agent loops make inference cost and latency compound quickly. AI Made Me 5x Faster. It Also Made Me 5x Worse at My Job. is a cautionary tale about speed without supervision. GLM-5.3 doesn't solve that, but it does reduce the cost of letting the agent run longer, which is a prerequisite for building the kind of verification loops you'll need. Accelerating Research: Can Review Systems Handle AI-Driven Productivity? points to the same bottleneck: the model is only as good as your ability to review its work.

**The detail worth watching is the API migration.** GLM-5.3 requires you to change `thinking.type` from `"disabled"` to `"enabled"` and specify a reasoning effort. If you don't, requests fail. That's a small technical note, but it signals a larger shift: reasoning is no longer optional. For enterprise teams, that means planning for a real migration, not just a model swap. Our take? If you're building on GLM, budget for the two-week window between release and open weights, and treat the cyber capability as a feature to manage, not a headline to ignore. The question isn't whether Z.ai can keep pushing post-training. It's whether they can keep the access controls ahead of the capability curve. Watch the disclosure ledger. That's where the real signal will be.

From VentureBeat

Chinese AI startup Z.ai, known internationally for its growing lineup of powerful, largely open source GLM series of language models, today released GLM-5.3 with substantial gains in long-horizon coding and a more consequential — and potentially sensitive — jump in cybersecurity capabilities.

Already, GLM-5.3's cyber capabilities have found a "potentially serious vulnerability in Cursor," the AI coding startup recently acquired by SpaceX, according to z.ai developer advocate Lou, posting on X. VentureBeat also tagged Cursor for confirmation on X and is awaiting response.

Read the original at VentureBeat