Z.ai icon

Z.ai

GLM-5.3-Flash

Builds websites, creates presentation slides, and analyzes data using GLM-5 and GLM-4.7 models to provide chat-based assistance and information.

Content updated today·GLM-5.3-Flash released yesterday

Pricing:Free + from $18/mo
Categories:
Try for Free
Jump to section
Z.ai interface for bilingual chat, agent workflows, ZCode, and GLM-5.3 coding model access

More tools to compare

Omniwork icon

Omniwork

Coasty icon

Coasty

MakersClaw icon

MakersClaw

Freebuff icon

Freebuff

AutoClaw icon

AutoClaw

Blocks.ai icon

Blocks.ai

Pros & Cons

Pros

  • Multiple access paths for chat, coding agents, APIs, and self-hosting
  • Coding plans integrate with more than 20 supported developer tools
  • MIT-licensed checkpoints such as GLM-5.2 allow self-hosting and fine-tuning
  • GLM-5.3 posts strong vendor-reported gains in coding and long-horizon tasks
  • GLM-5.3-Flash adds native multimodality and a higher-quota, lower-cost route for frequent work
  • Native bilingual support for English and Chinese without plugins
  • Points discounts and cached-input accounting can improve off-peak quota efficiency

Cons

  • Smaller plugin and integration ecosystem compared to ChatGPT
  • Some features lack the polish of more established platforms
  • Coding plans now start at $18 per month, well above earlier introductory pricing
  • Community and third-party tooling still maturing outside China
  • Coding Plan quota cannot be used as an unrestricted general API
  • Large open-weight models still require substantial deployment infrastructure
  • Model-specific reasoning, multimodal inputs, and quota behavior require integration testing before production routing

Overview

Z.ai is a bilingual AI chat, agent, and developer platform built around Zhipu AI's GLM model family. It combines a consumer web app with ZCode, GLM Coding Plan subscriptions, supported coding-tool integrations, and separately metered API services.

The platform serves developers, researchers, and everyday users who need an AI assistant for coding, writing, research, and document generation. Its August 2026 model line now pairs GLM-5.3, the heavier flagship for complex engineering, with GLM-5.3-Flash, a faster native-multimodal option for frequent coding, visual, and professional workflows.

Z.ai continues to publish major checkpoints as open weights, although availability is version-specific. GLM-5.2 and GLM-5.3-Flash have MIT-licensed weights; hosted access, model identifiers, and local deployment requirements still differ by release.

Key Features

  • GLM-5.3-Flash multimodal model — Combines a 1M-token context window with native image, video, chart, interface, and document understanding while using 18B active parameters for lower-cost inference.

  • GLM-5.3 flagship model — Builds on GLM-5.2's 1M-context base with expanded post-training for the most demanding coding, long-horizon agent, and authorized security-research workloads.

  • Agent mode — Converts natural language prompts into multi-step workflows, generating production-ready .docx, .pdf, and .xlsx documents directly from text instructions.

  • AI coding assistant — Works with Claude Code, OpenCode, ZCode, and other supported clients through the GLM Coding Plan, with model routing for flagship or higher-quota Flash workloads.

  • Long-horizon execution — Goal-oriented workflows can plan, edit code, run tests, inspect results, and revise strategy across extended tasks rather than stopping after a first attempt.

  • English and Chinese support — Official Z.ai model documentation lists both languages for the core GLM models.

  • Open-weight GLM ecosystem — GLM-5.2 and GLM-5.3-Flash support MIT-licensed self-hosting through current inference frameworks. Check each version separately because architecture and hardware needs vary substantially.

How Access Works

Z.ai spans several products that have different pricing and usage rules. The web app is the simplest entry point, while the Coding Plan is limited to supported coding tools and uses its own endpoint and quota system.

Access path Best suited to Important limitation
Z.ai web app Chat, research, writing, and agent tasks Model availability and usage limits can change by account
GLM Coding Plan Claude Code, OpenCode, ZCode, and supported development tools Subscription quota cannot be treated as a general-purpose API allowance
General API Custom applications and production integrations Separate pay-as-you-go billing and model availability apply
Open weights Private deployment, research, and customization Availability and hardware requirements vary by model version

Compared with ChatGPT, Claude, and Gemini, Z.ai combines English-Chinese model support, an open-weight model line, and Coding Plan subscriptions for officially supported third-party agent tools. GLM-5.3-Flash adds a lower-cost multimodal route, but benchmark leadership still varies by task and should be judged with matching harnesses and token budgets.

Pricing & Plans

Z.ai offers free web access, paid GLM Coding Plan subscriptions, and separate pay-as-you-go API billing. Current GLM Coding Plan monthly prices are:

Plan Monthly price Positioning
Lite $18 Lightweight work on one small repository
Pro $80 Daily development on one or two mid-sized projects; 6× Lite usage
Max $168 Advanced work across larger or concurrent projects; 14× Lite usage

Annual billing currently lists effective monthly rates of $12.60 for Lite, $56 for Pro, and $117.60 for Max. The Coding Plan uses credits-based quota accounting for input, cached input, and output tokens. Current rolling limits are 2,000 credits per 5 hours and 10,000 per week for Lite, 12,000 and 60,000 for Pro, and 28,000 and 140,000 for Max; 5-hour credits reset dynamically five hours after consumption, while weekly credits reset every seven days from subscription activation. Calls outside 14:00–18:00 UTC+8 on weekdays consume 50% of the standard credits. Z.ai says GLM-5.3-Flash provides three times the usable quota of GLM-5.3 within these plans.

The Coding Plan is restricted to officially supported tools and uses a dedicated coding endpoint. General API calls use separate pay-as-you-go billing; check the live model pricing page because rates and available model IDs can change independently of subscription prices.

Best For

  • Developers seeking a coding assistant that can route demanding work to GLM-5.3 and frequent visual or coding tasks to GLM-5.3-Flash
  • Teams needing bilingual English-Chinese AI support for cross-border collaboration
  • Open-source advocates willing to use a released checkpoint such as GLM-5.2 for AI chatbots
  • Researchers and students looking for a free, capable AI chatbot for daily use
  • Teams processing screenshots, charts, documents, or interfaces alongside code in multimodal agent workflows
  • Teams running long-horizon coding or authorized security-review workflows with explicit reasoning controls

Avoid If

  • You need a single subscription that also covers unrestricted production API traffic; Coding Plan quota and general API billing are separate.
  • Your procurement requires fixed model identifiers, quotas, or prices without checking the live plan pages.
  • You depend on integrations outside Z.ai's officially supported Coding Plan tool list, or you do not value its open-weight and multimodal options.

Sources & Verification

FAQ

Is Z.ai free to use?

Z.ai provides free web access, while higher-volume coding workflows use paid GLM Coding Plans. Model availability, limits, and free allowances can change, so check the live account page before relying on a specific quota.

How does Z.ai compare to ChatGPT?

Z.ai emphasizes English-Chinese model support, coding-tool integrations, and open-weight checkpoints. The better coding choice depends on the model, harness, task length, and subscription limits rather than a single headline benchmark.

What models does Z.ai support?

Z.ai runs the GLM family, including GLM-5.3 for flagship coding and agent work, GLM-5.3-Flash for lower-cost native-multimodal workflows, GLM-5.2 as an earlier downloadable 1M-context checkpoint, and specialized variants. Availability differs between chat, Coding Plan, general API, and open weights.

Can I use Z.ai for coding?

Yes. The GLM Coding Plan works with Claude Code, OpenCode, ZCode, Cline, and other supported clients. GLM-5.3 handles the heaviest long-horizon tasks, while GLM-5.3-Flash adds native visual input and three times the usable plan quota for frequent workflows. Confirm current model identifiers and reasoning settings before switching production clients.

Is Z.ai open source?

Some major GLM checkpoints are open weight, but do not assume every hosted model is downloadable. GLM-5.2 and GLM-5.3-Flash are available under the MIT License; deployment cost and runtime support vary by architecture and precision.

Does Z.ai support languages other than English?

Z.ai's official model documentation lists English and Chinese support for its core GLM models. Availability and quality can still vary by model and task.

What is the API pricing for Z.ai?

API pricing is separate from the Coding Plan. Through September 9, 2026 (24:00 UTC+8), GLM-5.3-Flash costs $0.075 per 1M input tokens, $0.015 per 1M cached-input tokens, and $0.25 per 1M output tokens; its list prices are $0.15, $0.03, and $0.50 respectively.

Is my data secure on Z.ai?

For users requiring additional data control, released open-weight GLM checkpoints can be self-hosted on private infrastructure, giving teams more control over where inference runs and data is processed; actual data-residency and compliance guarantees depend on the deployment. Check Z.ai's official documentation for current security and compliance details.

Version History

GLM-5.3-Flash

Current Version

Released on August 26, 2026

View Update
+What's new
3 updates
  • Outperform GLM-5.2 across coding and agent benchmarks with a 320B multimodal model that activates 18B parameters for lower-cost inference
  • Process images, video, charts, and interfaces natively, adding visual feedback to coding, browser-use, and professional document workflows
  • Stretch GLM Coding Plan capacity to three times GLM-5.3's usable quota, or deploy MIT-licensed weights with supported inference frameworks

GLM-5.3

Released on August 14, 2026

View Update
+What's new
3 updates
  • Complete complex coding and long-horizon agent tasks with a 50% gain over GLM-5.2 on Z.ai Code Bench while using fewer output tokens
  • Discover and validate software vulnerabilities with vendor-reported gains from 77.2% to 84.5% on CyberGym and more than 2× on ExploitBench
  • Control GLM-5.3 reasoning with low, high, or max effort, and update direct API integrations because `thinking.type: disabled` is no longer supported for this model

GLM-5.2

Released on June 16, 2026

View Update
+What's new
3 updates
  • Sustain repository-scale coding and research workflows with a 1M-token context window trained for reliable long-horizon execution
  • Balance capability, latency, and compute with selectable High and Max thinking effort, reaching 81.0 on Terminal-Bench 2.1 in Z.AI's published evaluation
  • Reduce long-context indexer FLOPs by 2.9× with IndexShare and raise speculative-decoding acceptance length by up to 20% overall

GLM-5.1

Released on April 7, 2026

View Update
+What's new
3 updates
  • Run long-horizon coding and engineering tasks autonomously for up to 8 hours, carrying work from planning and implementation through testing, refinement, and final delivery
  • Take on complex projects with a 200K context window and 128K maximum output, with stronger planning, debugging, tool use, and sustained execution across multi-step workflows
  • Build production-grade software and office deliverables with improved coding, frontend, PowerPoint, Word, PDF, and Excel capabilities aligned by Z.AI with Claude Opus 4.6

GLM-5V-Turbo

Released on April 1, 2026

View Update
+What's new
3 updates
  • Process images, videos, design drafts, and document layouts natively as a multimodal vision coding model with 200K context window and 128K max output tokens for long-horizon agentic tasks
  • Execute perception, planning, and action in GUI workflows, with strong results on AndroidWorld, WebVoyager, and ZClawBench while integrating with agents such as OpenClaw
  • Fuse visual understanding and code generation through CogViT vision encoder and 30+ task joint reinforcement learning across STEM, grounding, video, and coding domains

GLM-5.1

Released on March 27, 2026

View Update
+What's new
3 updates
  • Score 45.3 on Claude Code coding benchmark—94.6% of Claude Opus 4.6 performance—with 28% improvement over GLM-5, establishing a new frontier in cost-efficient agentic coding
  • Generate code at 55+ tokens/sec with estimated 200K context window, enabling long-horizon multi-file refactoring and distributed system architecture design
  • Access frontier-level coding intelligence from $3/month via Coding Plan with native compatibility for Claude Code, Cline, and Roo Code MCP tool integrations

GLM-5-Turbo

Released on March 15, 2026

+What's new
3 updates
  • Execute complex OpenClaw agent workflows with superior tool invocation reliability, scheduled task continuity, and high-throughput long-chain execution optimized since training phase
  • Decompose and follow multi-layered complex instructions with enhanced comprehension, supporting collaborative task division among multiple agents and MCP tool integrations
  • Outperform GLM-5 across multiple ZClawBench task categories while supporting 200K context input with multiple thinking modes for dynamic, long-running agent tasks

GLM-5

Released on February 12, 2026

View Update
+What's new
3 updates
  • Handle complex systems engineering and long-horizon agentic tasks with 744B parameters (40B active) and DeepSeek Sparse Attention integration
  • Generate production-ready documents (.docx, .pdf, .xlsx) directly from text with built-in Agent mode and multi-turn collaboration
  • Execute code with best-in-class open-source performance on reasoning benchmarks, approaching frontier model capabilities

GLM-4.7-Flash

Released on January 19, 2026

+What's new
3 updates
  • Get lightweight version of GLM-4.7 with faster response times and high throughput optimized for real-time coding, writing, and translation tasks
  • Deploy efficiently with competitive performance at smaller scale while maintaining strong general capabilities across reasoning and content generation
  • Access free-tier model designed for high-frequency use cases with best-in-class aesthetic outputs, low latency, and simplified deployment

GLM-4.7-Flash

Released on January 19, 2026

+What's new
3 updates
  • Get lightweight version of GLM-4.7 with faster response times and high throughput optimized for real-time coding, writing, and translation tasks
  • Deploy efficiently with competitive performance at smaller scale while maintaining strong general capabilities across reasoning and content generation
  • Access free-tier model designed for high-frequency use cases with best-in-class aesthetic outputs, low latency, and simplified deployment

GLM-4.7

Released on December 22, 2025

+What's new
3 updates
  • Build cleaner modern webpages and professional slides with major improvements in UI aesthetics, visual quality, and accurate layout sizing for frontend development
  • Solve multilingual coding tasks faster with 73.8% on SWE-bench and 41% on Terminal Bench 2.0, delivering stronger performance across agent frameworks
  • Reason through complex mathematical and logical problems with 42.8% on HLE benchmark while enhancing tool-using and web browsing capabilities

GLM-4.6

Released on September 30, 2025

+What's new
3 updates
  • Handle longer conversations and complex multi-file codebases with expanded 200K context window, enabling more sophisticated agentic task execution
  • Code more efficiently in Claude Code, Cline, and Roo Code with superior benchmark performance and improved real-world coding accuracy
  • Leverage enhanced reasoning capabilities with native tool use support during inference, delivering stronger results in search-based agent workflows

GLM-4.5

Released on July 28, 2025

+What's new
3 updates
  • Unify reasoning, coding, and agentic capabilities in a single model delivering balanced performance across complex problem-solving and rapid content generation
  • Switch between thinking mode for deep analysis and non-thinking mode for instant responses, adapting intelligence level to task complexity on demand
  • Build full-stack web applications with stronger frontend quality, and integrate the model into Claude Code, Roo Code, or custom agent workflows through tool APIs

ChatGLM3-6B

Released on October 27, 2023

+What's new
3 updates
  • Execute code directly and invoke external tools with new Code Interpreter and Function Call capabilities, enabling autonomous agent-style task completion
  • Handle semantics, mathematics, reasoning, code, and knowledge tasks more strongly, with ChatGLM3-6B-Base evaluated across eight representative Chinese and English benchmarks
  • Deploy locally on consumer hardware with open-source 6B-parameter model supporting both academic research and free commercial use after registration

ChatGLM2-6B

Released on June 25, 2023

+What's new
3 updates
  • Handle longer conversations with expanded 32K context window using FlashAttention technology, enabling deeper multi-turn dialogue understanding
  • Get responses 42% faster with improved inference speed and INT4 quantization, supporting extended dialogues on consumer GPUs with only 6GB VRAM
  • Achieve stronger performance across reasoning and knowledge benchmarks with enhanced training on 1.4T bilingual tokens covering diverse domains

ChatGLM-6B

Released on March 14, 2023

+What's new
3 updates
  • Deploy locally on consumer-grade graphics cards with lightweight 6.2B-parameter bilingual model, enabling private ChatGPT-style conversations
  • Chat naturally in Chinese and English with open-source conversational AI trained on approximately 1 trillion tokens of diverse text data
  • Use the model for academic research or registered free commercial deployment, with open model weights and code plus P-Tuning v2 support for downstream adaptation

Top alternatives

Related categories

From the blog

View all →

Track Z.ai in ToolWorthy Weekly

Important tool updates, better alternatives, and selected AI signals in one weekly brief.

Weekly only. Unsubscribe anytime.