glm-4.6

文本

Text-only model in the GLM-4 family with improved coding benchmarks and tool calls during reasoning, suited for software development and agent workflows.

Input price

$0.487915% off

Market: $0.574

/ M

Output price

$1.949915% off

Market: $2.294

/ M

Cache read

$0.097715% off

Market: $0.1149

/ M

Cache write

$015% off

/ M

Input

Text

Output

Text

API

Quick Start

Drop-in code to call this model with EasyAPI's OpenAI-compatible API.

1

Get your API key

Create an API key from your dashboard and set it as an environment variable:

Replace sk-xxxxxxxx with your API token from the console (Bearer token).

Create API token
shell
export EASYAPI_API_KEY=sk-xxxxxxxx
2

Make your first request

Use zai/glm-4.6 with the EasyAPI API:

cURL
curl https://token.easyapi.com/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $EASYAPI_API_KEY" \
  -d '{
  "model": "zai/glm-4.6",
  "messages": [
    {
      "role": "user",
      "content": "Hello!"
    }
  ]
}'
3

Enable streaming

Add "stream": true to your request body to receive responses as server-sent events:

curl
curl -N https://token.easyapi.com/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $EASYAPI_API_KEY" \
  -d '{"model":"zai/glm-4.6","messages":[{"role":"user","content":"Hello!"}],"stream":true}'

Endpoint

Sends a chat completion request. Supports streaming and non-streaming modes.

POSThttps://token.easyapi.com/v1/chat/completions
Authorization
Bearer $EASYAPI_API_KEY
Content-Type
application/json
Model
zai/glm-4.6

Send a chat completion request, with streaming or non-streaming responses. Docs

GEThttps://token.easyapi.com/v1/models
Authorization
Bearer $EASYAPI_API_KEY

List the models available to the current account. Docs

Parameters

NameTypeDefaultDescription
modelstringModel ID, e.g. openai/gpt-4o or anthropic/claude-3-5-sonnet
messagesarrayList of conversation messages, each with role and content
streambooleanfalseWhen true, the response is streamed via SSE
temperaturefloat1Sampling temperature; higher values are more random
max_tokensintegerMaximum number of tokens generated in a single reply
top_pfloat1Nucleus sampling; limits the cumulative probability of candidate tokens
frequency_penaltyfloat0Frequency penalty; reduces repeated wording
presence_penaltyfloat0Presence penalty; encourages new topics

Using third-party SDKs

Official OpenAI and other SDKs work by swapping the base URL. See integration guide