Model list & per-request billing
GET /v1/models returns a structured model list with pricing, context length and capabilities; GET /v1/generation looks up the token counts and actual cost of a single request by its request ID.
List models
https://token.easyapi.com/v1/modelsReturns the models available to the current key. Besides the native OpenAI fields (id / object / created / owned_by), each model carries pricing, context length, input/output modalities, supported parameters, capability flags and callable endpoints, so clients can show prices and context limits without a separate pricing lookup.
Request parameters
| Name | In | Type | Required | Description |
|---|---|---|---|---|
Authorization | header | string | Yes | EasyAPI API Key. Format: Bearer `API_KEY` |
Response fields
| Name | Type | Description |
|---|---|---|
id | string | Model ID; use it as the model field when calling |
name | string | Display name |
category | string | Category: text / image / audio / video / embedding / rerank |
context_length | integer | Context window in tokens; omitted when unknown |
architecture.modality | string | Input/output modalities, e.g. text+image->text |
top_provider.max_completion_tokens | integer | Maximum output tokens per request |
pricing.prompt | string | Price per input token (USD, string), already converted for your key group |
pricing.completion | string | Price per output token (USD) |
pricing.input_cache_read | string | Price per cached input token (USD); omitted when the model has no cache pricing |
pricing.request | string | Per-request price (USD); "0" for token-billed models |
billing.group | string | Group used for pricing; combine with model_ratio / group_ratio to verify the price |
supported_parameters | array<string> | Request parameters accepted by this model |
capabilities | object | Capability flags: stream / tools / vision / json_mode / parallel_tool_calls (text models only) |
endpoints | array<object> | Callable endpoints: type / path / method |
deprecation | object | Deprecation info: status (scheduled / offline), offline_at, suggested_model; omitted when not scheduled |
Response example
{
"object": "list",
"data": [
{
"id": "gpt-4o",
"object": "model",
"created": 1759190400,
"owned_by": "openai",
"name": "gpt-4o",
"category": "text",
"context_length": 128000,
"architecture": {
"modality": "text+image->text",
"input_modalities": ["text", "image"],
"output_modalities": ["text"]
},
"top_provider": {
"context_length": 128000,
"max_completion_tokens": 16384,
"is_moderated": false
},
"pricing": {
"prompt": "0.0000025",
"completion": "0.00001",
"input_cache_read": "0.00000125",
"request": "0"
},
"billing": {
"mode": "token",
"group": "default",
"group_ratio": 1,
"model_ratio": 1.25,
"completion_ratio": 4
},
"supported_parameters": ["temperature", "top_p", "max_tokens", "stream", "tools", "tool_choice", "response_format"],
"capabilities": {
"stream": true,
"tools": true,
"vision": true,
"json_mode": true,
"parallel_tool_calls": true
},
"supported_endpoint_types": ["openai", "openai-response"],
"endpoints": [
{ "type": "openai", "path": "/v1/chat/completions", "method": "POST" },
{ "type": "openai-response", "path": "/v1/responses", "method": "POST" }
]
}
]
}Get generation
https://token.easyapi.com/v1/generationExample request: https://token.easyapi.com/v1/generation?id=2026100812000000a1b2c3
Returns token counts, cached tokens, the final charge (USD), the ratio snapshot and timings of a single request by request ID. Only requests of the account that owns the key can be queried. To get the cost inline, send usage.include = true in the request body and the usage object gains a cost field; that value is an estimate made before the response is written, this endpoint is authoritative.
Request parameters
| Name | In | Type | Required | Description |
|---|---|---|---|---|
Authorization | header | string | Yes | EasyAPI API Key. Format: Bearer `API_KEY` |
id | query | string | Yes | Request ID from the X-EasyAPI-Request-Id response header |
Response fields
| Name | Type | Description |
|---|---|---|
id | string | Request ID |
model | string | Model that was called |
created_at | integer | Record time (unix seconds) |
is_stream | boolean | Whether the request was streamed |
tokens_prompt | integer | Input tokens |
tokens_completion | integer | Output tokens |
native_tokens_cached | integer | Input tokens served from cache |
native_tokens_cache_write | integer | Tokens written to cache |
total_cost | number | Final charge (USD) |
quota | integer | Internal quota units; 500000 = 1 USD |
generation_time | integer | Total request time (ms, second precision) |
latency | integer | Time to first token (ms); 0 for non-streaming requests |
billing | object | Ratio snapshot used for settlement: model_ratio / completion_ratio / group_ratio / cache_ratio |
Response example
{
"data": {
"id": "2026100812000000a1b2c3",
"model": "gpt-4o",
"created_at": 1759905600,
"is_stream": true,
"token_name": "my-key",
"group": "default",
"is_byok": false,
"tokens_prompt": 1000,
"tokens_completion": 200,
"total_tokens": 1200,
"native_tokens_cached": 80,
"native_tokens_cache_write": 0,
"total_cost": 0.0045,
"quota": 2250,
"generation_time": 3000,
"latency": 350,
"billing": {
"model_ratio": 1.25,
"completion_ratio": 4,
"group_ratio": 1,
"cache_ratio": 0.5
}
}
}Status codes
| Status code | Meaning | Description | Data model |
|---|---|---|---|
200 | OK | Billing record returned | GenerationResponse |
400 | Bad Request | Missing id | ErrorResponse |
401 | Unauthorized | Invalid API key | ErrorResponse |
404 | Not Found | Record not written yet, or the request belongs to another account; retry shortly | ErrorResponse |