Saturday, 15 August 2026 Login

Spreadsheets. Data. Now.

BREAKING
Data Workflow

New AI model uncovers major Cursor flaw

New AI model uncovers major Cursor flaw - ai model
New AI model uncovers major Cursor flaw

Chinese AI startup Z.ai has launched GLM-5.3, the newest version of its open-source language model series. The update introduces notable improvements in long-horizon coding and cybersecurity without retraining the underlying model.

Post-training scaling delivers benchmark improvements

GLM-5.3 retains the same 743-billion-parameter base as GLM-5.2 but achieves gains through expanded post-training. Z.ai scaled reinforcement learning across diverse tasks, including simulated engineering workflows that mimic multi-day projects rather than isolated exercises.

Performance metrics reflect this progress. On Z.ai’s Terminal-Bench 3.0, GLM-5.3 scores 28.3, a significant increase from the previous version’s 4.6. DeepSWE v1.1 performance rises to 66.9 from 46.2, while AutomationBench improves to 48.2 from 26.2. The model also advances on Agents’ Last Exam CLI, reaching 28.5 from 23.8.

Efficiency has improved as well. On Z.ai’s private Code Bench evaluation, the new version achieves a 34.5% task completion rate at its highest reasoning setting while using roughly 75,000 output tokens per task. The earlier version required 96,000 tokens for a 23.4% completion rate. At a lower effort level, GLM-5.3 reaches 31.4% with about 50,000 tokens—matching Claude Opus 4.8’s 29.5% but with less than half the token usage.

These results come from Z.ai’s internal benchmarks. The trend indicates post-training can yield meaningful improvements without the expense of retraining a new base model.

Related: DeepSeek Harness Debuts Open-Source Rival to Claude Code

Cybersecurity capabilities exceed initial expectations

Z.ai incorporated vulnerability-discovery environments during post-training, anticipating gradual improvements in identifying software flaws. The model’s abilities expanded further than expected, advancing along the exploitation chain.

The dual-use nature of these capabilities presents challenges. The same features that enhance software engineering also improve security research—and potentially offensive operations. Z.ai has introduced controls for the model’s advanced functions, including a “trusted access” system for sensitive tasks.

This caution may explain the staged rollout. GLM-5.3 is currently available only through Z.ai’s GLM Coding Plan and ZCode environment. API access and open weights will follow once safety evaluations are complete. The company plans to release the weights about two weeks after launch.

Breaking changes and enterprise adoption

Developers upgrading to GLM-5.3 must adjust their workflows. The new version introduces three reasoning-effort levels—low, high, and max—with max as the default. Thinking cannot be disabled, unlike in previous releases. Applications using thinking.type: “disabled” must update to enabled and specify a reasoning effort before switching, or requests will fail.

The change aligns with Z.ai’s focus on agentic engineering. GLM-5.3 is optimized for long-running autonomous tasks, including planning, implementation, testing, and verification. ZCode supports these workflows with remote task control and is available on macOS, Windows, and Linux.

Pricing remains competitive. Individual plans start at $12.60 per month for the Lite tier, which includes 10,000 credits weekly. Pro and Max tiers cost $56 and $117.60 monthly, respectively, with higher usage limits. Team plans range from $88 to $188 per user per month. A points-based quota system reduces costs for off-peak usage.

Related: Why Qwen and Opus 5 scores miss the cost mark

API pricing for GLM-5.3 hasn’t been announced, making direct comparisons with earlier versions difficult. The gradual release suggests Z.ai is prioritizing safety before broader availability.

Early vulnerability discovery demonstrates real-world impact

A Z.ai developer advocate reported on X that GLM-5.3 identified a “potentially serious vulnerability” in Cursor, the AI coding startup recently acquired by SpaceX. Cursor has not confirmed the claim, despite being tagged for comment.

For enterprise users, this incident illustrates both the potential and risks of advanced coding agents. A model capable of autonomously discovering vulnerabilities could speed up security audits—or create new attack surfaces if misused. Z.ai’s decision to delay open weights and API access may reflect efforts to balance these risks while delivering performance gains.

Post-training has proven effective in evolving model capabilities. GLM-5.3’s improvements stem from refining interactions with complex environments rather than scaling the base model. This approach may offer a more cost-effective path forward than continuous pretraining.

The cybersecurity results highlight a key question. How far can post-training push a model before safety controls become limiting? Z.ai’s next steps—particularly around API access—will likely influence enterprise adoption in the coming months.

Tags:

Leave a Reply

Your email address will not be published. Required fields are marked *