Agents and app development
Build agents on the OpenAI-compatible API
Every agent step is a call, and the model must return tool calls, streams and JSON reliably. EasyAPI passes these capabilities through without rewriting requests. Change base_url and your code runs.
Recommended models available now
Chosen for agents: reliable tool calls, structured output, context that holds many turns. Multi-step loops cost grows linearly with steps, so get the flow working on a cheap model, then upgrade to see the difference.
Platform prices per 1M tokens, billed on actual usage. The list follows the live pricing API; retired models never show up here. Pick a model and the config and cost below update to match.
Exact configuration
The API base and the selected model qwen3.8-max are already filled in. Replace the key with your own.
from openai import OpenAI
client = OpenAI(base_url="https://token.easyapi.com/v1", api_key="sk-your-key")
tools = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get current weather for a city",
"parameters": {"type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"]},
},
}]
resp = client.chat.completions.create(
model="qwen3.8-max",
messages=[{"role": "user", "content": "What is the weather in Shanghai today?"}],
tools=tools,
tool_choice="auto",
)
call = resp.choices[0].message.tool_calls[0]
print(call.function.name, call.function.arguments)Get one request working first
Swap in your key and run it as is. A 200 means the base URL, key and model are all right; otherwise match the error against the list below.
curl https://token.easyapi.com/v1/chat/completions \
-H "Authorization: Bearer sk-your-key" \
-H "Content-Type: application/json" \
-d '{"model":"qwen3.8-max","messages":[{"role":"user","content":"Hello"}]}'What to know
- tools, tool_choice, stream and response_format are passed to the upstream as is. Support varies by model; the parameter table on each model's detail page says what takes effect.
- Multi-turn agents accumulate messages. The context window is a hard limit: a request over it fails instead of being truncated.
- Watch for 429 under high concurrency on one key. Both platform rate limiting and an exhausted balance return 429; the error message says which.
- The models fallback list and provider.sort let a request declare backup models and routing preference, switching automatically on upstream failure. See the API docs.
Common errors
- tool_calls is empty and the model answered in prose
- Usually the model does not support tool calls or tool_choice is unset. Pick a model from the recommended list and set tool_choice to auto or a specific function.
- No deltas arrive when streaming
- Confirm stream is true and the client is not buffering. curl with -N shows the raw SSE and pinpoints the issue fast.
- Returned JSON fails to parse
- Use response_format with json_object or json_schema and say 'output JSON only' in the system prompt. Models without response_format rely on prompting alone and are not guaranteed.
- 429 Too Many Requests
- Read the message: insufficient quota means top up; rate limit means too much concurrency, so slow down or spread across keys.
Cost example
At the current price of qwen3.8-max, one typical request with 6,000 input + 800 output tokens, 2,000 times a day:
You pay only for tokens actually used. No platform fee, routing fee or monthly minimum. Real usage is listed request by request in the console logs.
Connect now
After signing up, the console home shows an onboarding guide that generates copy-ready config for your key and model, and stays until your first successful call.
Other solutions: Connect Claude Code, Cursor and Codex to the API · Use models in Dify, n8n and FastGPT