Skip to navigation

GLM 5.3

GLM 5.3 is Z.ai’s model for complex coding and long-horizon work. It supports tool-driven workflows with always-on reasoning and native Low, High, and Max effort levels.

Model details

  • Identifier: z-ai/glm-5.3
  • Context window: 256K tokens
  • Capabilities: Text · Reasoning · Tools
  • Reasoning levels: Low · High · Max

Best for

  • Working across large codebases
  • Long-running coding and migration tasks
  • Multi-step tool use and automation
  • Sustained agent workflows that need adjustable reasoning effort

Platform defaults

When sampling parameters are omitted, Mixlayer applies:

{
"temperature": 1.0,
"top_p": 0.95
}

Reasoning behavior

GLM 5.3 defaults to native Max effort. Thinking cannot be disabled.

Requested effortGLM 5.3 behavior
Omitted or defaultNative Max
noneRejected
minimal or lowNative Low
medium or highNative High
xhigh or maxNative Max

Set the value through reasoning_effort in Chat Completions or reasoning.effort in Responses.

Requests with none return an invalid_request_error with code unsupported_parameter.

Budget enough output tokens for both reasoning and the final answer. Different effort values select the model’s native behavior, not a fixed reasoning-token budget.

Working with long context

Provide relevant repository context, identify authoritative files, and state the expected outcome. Break long tasks into verifiable milestones so tool results and intermediate changes can be checked before continuing.

When preparing a new turn, Mixlayer clears reasoning from assistant messages before the latest user message while preserving their visible answers and tool calls. Reasoning from assistant messages after the latest user message is retained, including during tool-call continuations.

Use the Mixlayer console or GET /v1/models to confirm the model identifier, context limit, and availability for your API key.

See the official GLM 5.3 model card for upstream architecture and capabilities.