
Image: Flickr / Wikimedia Commons / Unsplash
GLM Coding Plan Pricing: Current Lite, Pro, and Max Limits
Current Z.ai plan prices, five-hour and weekly credit limits, supported models, and the practical cost tradeoffs for coding-agent users.
This article contains an affiliate link. AI Eating The World may earn a commission on qualifying purchases, at no extra cost to you.
Z.ai's current GLM Coding Plan starts at $18 per month. This update separates standard monthly prices from promotional prices and documents the current credit limits and model routing.
A model built for long jobs, not quick replies
Z.ai quietly rolled out GLM-5.2 to subscribers on the GLM Coding Plan on June 13, 2026, three days before releasing the model's weights publicly under an MIT license. For US engineering teams watching the fallout from Anthropic's export-control disruption to Claude Fable 5 and Mythos earlier that month, the timing read as more than coincidence. Z.ai used the moment to position GLM-5.2 as a credible option for developers who want frontier-level coding performance without depending on a single US vendor.
The model is built specifically for long-horizon work: multi-step engineering tasks that require an AI system to hold context across an entire codebase rather than a single file or function. Its headline spec is a 1 million-token context window, paired with a mixture-of-experts architecture and a sparse attention design that Z.ai calls IndexShare, which reuses the same indexing layer across every four attention layers to cut per-token compute at long context lengths.
In practice, that translates to a model meant to track module boundaries, API contracts, and architectural decisions across a project without losing the thread partway through a task. Z.ai benchmarked GLM-5.2 on long-horizon evaluations like FrontierSWE, PostTrainBench, and SWE-Marathon, all designed to test whether a model can carry a large engineering job from start to finish rather than answer isolated prompts.
Current GLM Coding Plan prices and limits

As checked on August 24, 2026, Z.ai lists standard monthly prices of $18 for Lite, $80 for Pro, and $168 for Max. The purchase page was also showing promotional prices of $12.60, $56, and $117.60. Treat promotional prices as temporary and verify the checkout total before subscribing.
Z.ai's current documentation lists 2,000 five-hour and 10,000 weekly credits for Lite, 12,000 and 60,000 for Pro, and 28,000 and 140,000 for Max. Credits reset on separate rolling schedules, and actual task usage depends on input, cached input, output, model, and tool multipliers.
All current plans support GLM-5.3, GLM-5-Turbo, and GLM-4.7. Z.ai says requests for GLM-5.2 or GLM-5.1 are automatically routed to GLM-5.3, so a subscription is no longer a way to pin the older model.
Sources for this section
Where GLM-5.2 lands next to Claude and GPT-5.5

On raw benchmark scores, GLM-5.2 doesn't claim the top spot. Independent testing puts it around 62 on SWE-bench Pro, compared to roughly 69 for Claude Opus 4.8. But the gap narrows on tasks built specifically for long-horizon agentic work, and it's the first open-weights model reported to cross 80% on Terminal-Bench, ahead of every other open model currently available and competitive with Gemini's closed-weight results.
The reaction from coding-tool builders was immediate. The team behind Kilo Code confirmed day-one integration on launch, and Cline's developers highlighted the combination of the 1M-token context window and the model's Max effort mode as the differentiator, rather than any single benchmark number. An independent Semgrep evaluation also found GLM-5.2 matching or slightly outperforming Claude on certain security-focused coding tasks.
The more relevant comparison for most teams is cost-adjusted performance. GLM-5.2's output tokens run roughly five to seven times cheaper than Claude Opus 4.8 or GPT-5.5, while landing within a few points of both on the benchmarks that matter for day-to-day engineering work. For US startups and mid-size engineering teams running high query volumes, that ratio changes the math on which model becomes the default rather than the fallback.
Sources for this section
Who this actually makes sense for
The MIT license is the part worth pausing on. Unlike a subscription to a closed model, GLM-5.2's weights can be downloaded and run entirely on a team's own infrastructure. For US companies in regulated industries or with strict data-residency requirements, that means frontier-level coding assistance without routing proprietary code through an external API at all. It also removes the vendor lock-in risk that's become more visible to operators since the recent Claude export-control episode.
Beyond self-hosting, GLM-5.2 fits teams doing sustained agentic coding work: refactors across large repositories, compiler or kernel-level development, and production-grade service builds where a model needs to hold architectural context for hours rather than minutes. Early testers on the GLM Coding Plan also reported stronger client-side and mobile engineering support, including a more complete on-device debugging loop than prior GLM releases offered.
It's a weaker fit for teams that mostly need short, single-file completions or occasional AI assistance. The Coding Plan's five-hour prompt cycles and per-tool quota consumption reward developers who work in extended sessions, and the pricing advantage shrinks if usage is light and sporadic.
Sources for this section
The tradeoffs worth knowing before you switch
GLM-5.2 consumes quota faster than Z.ai's other Coding Plan models, roughly three times the rate during peak hours and twice the rate off-peak compared to GLM-4.7 or GLM-4.5-Air, so the effective throughput on Lite and Pro tiers is lower than the headline prompt counts suggest. Some Max-tier users have also reported throttling during peak windows, which matters for teams that need guaranteed availability during business hours rather than flexible scheduling.
It's also still behind Claude Opus 4.8 on raw SWE-bench Pro performance, and pricing pages across third-party trackers don't always agree on exact figures, since Z.ai has run multiple promotional periods since launch. Anyone evaluating the plan should confirm current tier pricing and quota limits directly on Z.ai's site before committing to annual billing, since promo rates aren't guaranteed to hold.
Sources for this section
Getting started with the GLM Coding Plan
For US developers deciding whether to add a second model to their coding stack, compare the current credit limits against the kinds of tasks you run. Lite is positioned for one lightweight project, Pro for regular work across one or two projects, and Max for heavier concurrent use.
The current plan routes older GLM-5.2 and GLM-5.1 requests to GLM-5.3. Test the model and tool integration on work you can evaluate, then compare output quality and credits consumed against your existing stack before committing to a longer billing cycle.
Sources
Brian Weerasinghe is the Founder and Editor of AI Eating The World. AI Eating The World is the independent AI publication for builders, operators, and leaders navigating how AI is changing work and the world.
Brian Weerasinghe is the founder and editor of AI Eating The World, where he covers artificial intelligence, tech companies, layoffs, startups, and the future of work. His reporting focuses on how AI is transforming businesses, products, and the global workforce. He writes about major developments across the AI industry, from enterprise adoption and funding trends to the real-world impact of automation and emerging technologies.


