Keep the exact editor and CLI you already use. Just point the engine at a cheaper, open model. We start in Cursor - where most people want this first - then make the identical move in code: Claude Code, Codex, and a raw API call. Every command below was verified against the official Z.ai, OpenRouter, and claude-code-router docs (June 2026).

The engine: GLM-5.2, Z.ai's open-weight coding model (MIT license, 1M-token context, released June 16, 2026). It's the highest-scoring open-weights model on the Artificial Analysis Intelligence Index, scores 81.0 on Terminal-Bench 2.1 (up from 63.5 for GLM-5.1; Claude Opus 4.8 sits at 85.0), and 62.1 on SWE-bench Pro. Near-frontier coding at a fraction of the price.

Meet the family (you'll use more than one):

  • glm-5.2 - the coding default, 1M context. Your main driver.

  • glm-5v-turbo (GLM-5V-Turbo) - the vision / multimodal model. Feed it screenshots, design mockups, and UI states and code against the image instead of describing it. This is the family's "VL." Note: there is no vision variant of 5.2 itself - GLM-5V-Turbo is the separate model for that.

  • glm-4.5-air - small, fast, free on Z.ai. Good for the cheap background calls.

Two ways in:

  • Z.ai direct - the first-party route. Most reliable, official, fewest moving parts.

  • OpenRouter / claude-code-router - the "router" route. One key, many models (GLM, DeepSeek, Qwen, and more) from a single config. More flexible, slightly more setup, and unofficial inside Cursor.

The one thing to know up front: there are two API dialects. Cursor and most editors speak OpenAI's format; Claude Code speaks Anthropic's. Z.ai exposes both, so the only thing that changes per tool is which base URL you paste:

  • Anthropic Messages (Claude Code): https://api.z.ai/api/anthropic

  • OpenAI coding (Cursor, Cline, Coding Plan keys): https://api.z.ai/api/coding/paas/v4

  • OpenAI general (pay-as-you-go keys, vision models): https://api.z.ai/api/paas/v4

That's the whole trick.

Part 1: Run GLM-5.2 in Cursor (start here)

Cursor + GLM-5.2 via Z.ai (the reliable path)

Cursor takes an OpenAI-compatible endpoint, so you point it at Z.ai's coding endpoint and add GLM-5.2 as a custom model. You need Cursor Pro or higher - custom models aren't on the free tier.

1) Open Cursor Settings then Models (Cmd+Shift+J on Mac, Ctrl+Shift+J on Windows).

2) Find OpenAI API Key, turn it ON, and paste your Z.ai API key.

3) Turn ON Override OpenAI Base URL and enter exactly:

https://api.z.ai/api/coding/paas/v4

Use the coding endpoint above - it's what a GLM Coding Plan key requires. A pay-as-you-go Z.ai key uses the general endpoint (https://api.z.ai/api/paas/v4) instead. Using the wrong one is the usual cause of failed requests or odd billing.

4) Click Add Custom Model, choose OpenAI Protocol, and enter the name in uppercase: GLM-5.2. (The API model id is glm-5.2; Cursor wants the uppercase label.) Save.

5) Open Chat, Agent, or Composer, pick GLM-5.2 in the model picker, and send a test message to confirm it answers.

Add the vision model: GLM-5V-Turbo

Want GLM to see? Add a second custom model the same way and name it GLM-5V-TURBO (API id glm-5v-turbo). It reads screenshots, design mockups, charts, and UI states, so you can hand it an image and code against it - useful for "make it look like this" and for debugging from a screenshot.

One wiring note: multimodal models run on Z.ai's general API, not the coding endpoint. If the coding base URL rejects glm-5v-turbo, set that model's base URL to https://api.z.ai/api/paas/v4 with a standard Z.ai API key.

Swap any model: Cursor via OpenRouter (the router path)

Prefer one key for many models? Route Cursor through OpenRouter:

1) Create a key at openrouter.ai/settings/keys.

2) Settings then Models: turn OpenAI API Key ON and paste the OpenRouter key.

3) Turn ON Override OpenAI Base URL and enter:

https://openrouter.ai/api/v1

4) Add Custom Model and enter the slug z-ai/glm-5.2. Swap the slug for any model OpenRouter carries.

The catch, and it's a real one: Cursor doesn't officially support OpenRouter. The supported bring-your-own-key providers are OpenAI, Anthropic, Google, Azure, and Bedrock; an OpenRouter base-URL override is known to be partially broken - you can hit request-format and tool-use errors. For Cursor specifically, a direct Z.ai key is the more reliable path. Use the router only when you genuinely need to hot-swap models.

Two Cursor quirks worth knowing: enabling the base-URL override can disrupt Cursor's built-in models - toggle the OpenAI API Key OFF when you want them back, ON again for GLM (the base URL stays saved). And Cursor may not pass Z.ai-specific params like reasoning_effort or expose the full 1M context window in its UI. Note too that Cursor's separate CLI doesn't support custom models or your own key yet - this is an in-editor setup.

Part 2: The same trick in code

Cursor is the easy on-ramp, but the move - keep the tool, swap the engine - works anywhere. Here's each harness, start to finish.

Claude Code via Z.ai (easiest, no proxy)

Claude Code speaks Anthropic's format and Z.ai exposes an Anthropic-compatible endpoint, so this is zero-proxy. The GLM Coding Plan starts at ~$18/mo (Lite), with Pro and Max tiers above (Max ~$80/mo). To try GLM first, Z.ai's standalone API ships free trial credits and a free GLM-4.5-Air model.

1) Install Claude Code:

npm install -g @anthropic-ai/claude-code

2) Get a Z.ai API key: z.ai then GLM Coding Plan.

3) Configure, automated:

npx @z_ai/coding-helper

or manual, in ~/.claude/settings.json:

"env": {
  "ANTHROPIC_BASE_URL": "https://api.z.ai/api/anthropic",
  "ANTHROPIC_AUTH_TOKEN": "your_zai_api_key",
  "ANTHROPIC_API_KEY": "",
  "ANTHROPIC_DEFAULT_OPUS_MODEL": "glm-5.2[1m]",
  "ANTHROPIC_DEFAULT_SONNET_MODEL": "glm-5.2[1m]",
  "ANTHROPIC_DEFAULT_HAIKU_MODEL": "glm-4.5-air"
}

(glm-5.2[1m] is the full 1M-context model. glm-4.5-air handles the fast/cheap calls.)

4) Open a NEW terminal, run claude, check /status. Base URL should be the Z.ai endpoint.

Claude Code via claude-code-router + OpenRouter (swap any model)

Use this when you want to switch between models, not just GLM. OpenRouter does expose an Anthropic-compatible endpoint, but it only guarantees it for Anthropic's own models - for GLM-5.2 it's best-effort and can break on tool-use or streaming. So for GLM you route Claude Code through claude-code-router (a tiny local translator).

1) Install both:

npm install -g @anthropic-ai/claude-code @musistudio/claude-code-router

2) Create ~/.claude-code-router/config.json:

{
  "Providers": [
    {
      "name": "openrouter",
      "api_base_url": "https://openrouter.ai/api/v1/chat/completions",
      "api_key": "YOUR_OPENROUTER_KEY",
      "models": ["z-ai/glm-5.2"],
      "transformer": { "use": ["openrouter"] }
    }
  ],
  "Router": {
    "default": "openrouter,z-ai/glm-5.2"
  }
}

(z-ai/glm-5.2 is the exact OpenRouter slug: openrouter.ai/z-ai/glm-5.2. Swap in any other slug to change engines.)

3) Launch Claude Code through the router:

ccr code

Switch models anytime with ccr model; open the dashboard with ccr ui.

Claude Code via Ollama (truly local)

Same router, but the provider is your own machine. Fully private, no per-token cost.

1) Install Ollama and pull a coding model:

ollama pull qwen2.5-coder:14b

2) Add an Ollama provider to ~/.claude-code-router/config.json:

{
  "Providers": [
    {
      "name": "ollama",
      "api_base_url": "http://localhost:11434/v1/chat/completions",
      "api_key": "ollama",
      "models": ["qwen2.5-coder:14b"]
    }
  ],
  "Router": {
    "default": "ollama,qwen2.5-coder:14b"
  }
}

3) Run it:

ccr code

The honest catch on truly-local GLM: running GLM-5.2 itself on your own machine needs serious hardware (roughly 48 to 80GB VRAM, quantized). Most people don't have that. So for "cheap," the Z.ai or OpenRouter routes are the realistic move; local with a smaller model (Qwen Coder, DeepSeek Coder) is the "fully private, runs on what I've got" move.

Codex CLI via OpenRouter

Codex reads its model from a profile and authenticates through a provider you define. Point that provider at OpenRouter and you get GLM-5.2 in Codex.

1) Create an OpenRouter API key at openrouter.ai/settings/keys.

2) Store the key in the macOS Keychain, so it isn't sitting in a plaintext config:

security add-generic-password \
  -a "$USER" \
  -s codex-openrouter-api-key \
  -w "sk-or-v1-YOUR_KEY_HERE" \
  -U

3) Add the OpenRouter provider to ~/.codex/config.toml (replace YOUR_MAC_USERNAME with the output of whoami):

[model_providers.openrouter]
name = "OpenRouter"
base_url = "https://openrouter.ai/api/v1"
wire_api = "responses"

[model_providers.openrouter.auth]
command = "/usr/bin/security"
args = ["find-generic-password", "-a", "YOUR_MAC_USERNAME", "-s", "codex-openrouter-api-key", "-w"]
timeout_ms = 5000
refresh_interval_ms = 0

4) Create a Codex profile at ~/.codex/openrouter-glm.config.toml:

model = "z-ai/glm-5.2"
model_provider = "openrouter"
model_context_window = 1000000
model_auto_compact_token_limit = 750000
model_reasoning_effort = "medium"

5) Run Codex with GLM-5.2:

codex --profile openrouter-glm

Outside a git repo, add --skip-git-repo-check. Quick test:

codex --profile openrouter-glm exec --skip-git-repo-check "Reply with exactly OK."

It should print OK.

Bonus: Z.ai's own agent (Z Code)

If you'd rather not reroute anything, Z.ai ships its own Codex-style coding agent, Z Code, built around the GLM models. The free tier gives 5M tokens/day on GLM-5.2 - a no-cost way to feel out the model in a real agent before you wire it into Cursor or Claude Code. Install it and sign in with your Z.ai account per Z.ai's docs.

Raw API: GLM-5.2 in your own scripts

Past the editors, GLM-5.2 is a normal OpenAI-compatible chat endpoint, so any script or app that already talks to OpenAI just needs the base URL, key, and model id swapped:

curl https://api.z.ai/api/paas/v4/chat/completions \
  -H "Authorization: Bearer $ZAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "glm-5.2",
    "messages": [{"role": "user", "content": "Reply with exactly OK."}]
  }'

Point the same call at glm-5v-turbo with an image in the message to use vision. With an OpenAI SDK, set base_url to the endpoint above and your model id - nothing else changes.

Gotchas (the stuff that wastes an hour)

  • Cursor needs the right endpoint AND uppercase model name. Coding Plan key uses https://api.z.ai/api/coding/paas/v4; pay-as-you-go key and vision models use https://api.z.ai/api/paas/v4. The model label in Cursor must be uppercase (GLM-5.2), even though the API id is lowercase.

  • OpenRouter inside Cursor is unofficial. It can work, but it's known to break on request format and tool use. A direct Z.ai key is the reliable Cursor path; save OpenRouter for the CLIs or for genuine model-swapping.

  • ANTHROPIC_API_KEY must be "" (empty string) for Claude Code, not unset, or it falls back to Anthropic's servers. A real key in your shell profile overrides everything; set it to "" and restart the terminal.

  • Previously logged into Claude Code? Run /logout once, then quit and relaunch. A stale Anthropic session causes confusing "model not found" errors.

  • There's no GLM-5.2 vision model. Vision is a separate model - glm-5v-turbo. Don't wait for a "GLM-5.2V" that doesn't exist; add GLM-5V-Turbo as its own model.

  • Always verify. In Claude Code, /status should show ANTHROPIC_AUTH_TOKEN and your provider's base URL. In Cursor, use the built-in API-key test before you start.

Which to pick

  • Coding inside Cursor, most reliable: Cursor + Z.ai direct (GLM-5.2, plus GLM-5V-Turbo for visual work).

  • Cheapest CLI setup, runs today: Claude Code via Z.ai (GLM Coding Plan).

  • Want to swap between many models freely: the router path (OpenRouter via claude-code-router, or Codex via OpenRouter).

  • Fully private / local, and you've got the hardware: Claude Code via Ollama.

  • Just want to try GLM free first: Z Code (5M tokens/day on GLM-5.2).

The model is becoming a commodity. The real skill is knowing which engine to point your harness at - and now you can point all of them.