Model list & per-request billing

Model list & per-request billing

GET /v1/models returns a structured model list with pricing, context length and capabilities; GET /v1/generation looks up the token counts and actual cost of a single request by its request ID.

FieldValue
Base URLhttps://token.easyapi.com/v1
AuthenticationAuthorization: Bearer `API_KEY`
FormatOpenAI-compatible; extra fields follow OpenRouter
BillingNeither endpoint is billed
Note: Prices in /v1/models are converted for the group of the calling API key and returned as USD-per-token strings. Every response carries an X-EasyAPI-Request-Id header; pass it to /v1/generation. The usage record is written after the request completes, so an immediate lookup may return 404; retry shortly.

List models

GEThttps://token.easyapi.com/v1/models

Returns the models available to the current key. Besides the native OpenAI fields (id / object / created / owned_by), each model carries pricing, context length, input/output modalities, supported parameters, capability flags and callable endpoints, so clients can show prices and context limits without a separate pricing lookup.

Request parameters

NameInTypeRequiredDescription
AuthorizationheaderstringYesEasyAPI API Key. Format: Bearer `API_KEY`

Response fields

NameTypeDescription
idstringModel ID; use it as the model field when calling
namestringDisplay name
categorystringCategory: text / image / audio / video / embedding / rerank
context_lengthintegerContext window in tokens; omitted when unknown
architecture.modalitystringInput/output modalities, e.g. text+image->text
top_provider.max_completion_tokensintegerMaximum output tokens per request
pricing.promptstringPrice per input token (USD, string), already converted for your key group
pricing.completionstringPrice per output token (USD)
pricing.input_cache_readstringPrice per cached input token (USD); omitted when the model has no cache pricing
pricing.requeststringPer-request price (USD); "0" for token-billed models
billing.groupstringGroup used for pricing; combine with model_ratio / group_ratio to verify the price
supported_parametersarray<string>Request parameters accepted by this model
capabilitiesobjectCapability flags: stream / tools / vision / json_mode / parallel_tool_calls (text models only)
endpointsarray<object>Callable endpoints: type / path / method
deprecationobjectDeprecation info: status (scheduled / offline), offline_at, suggested_model; omitted when not scheduled

Response example

{
  "object": "list",
  "data": [
    {
      "id": "gpt-4o",
      "object": "model",
      "created": 1759190400,
      "owned_by": "openai",
      "name": "gpt-4o",
      "category": "text",
      "context_length": 128000,
      "architecture": {
        "modality": "text+image->text",
        "input_modalities": ["text", "image"],
        "output_modalities": ["text"]
      },
      "top_provider": {
        "context_length": 128000,
        "max_completion_tokens": 16384,
        "is_moderated": false
      },
      "pricing": {
        "prompt": "0.0000025",
        "completion": "0.00001",
        "input_cache_read": "0.00000125",
        "request": "0"
      },
      "billing": {
        "mode": "token",
        "group": "default",
        "group_ratio": 1,
        "model_ratio": 1.25,
        "completion_ratio": 4
      },
      "supported_parameters": ["temperature", "top_p", "max_tokens", "stream", "tools", "tool_choice", "response_format"],
      "capabilities": {
        "stream": true,
        "tools": true,
        "vision": true,
        "json_mode": true,
        "parallel_tool_calls": true
      },
      "supported_endpoint_types": ["openai", "openai-response"],
      "endpoints": [
        { "type": "openai", "path": "/v1/chat/completions", "method": "POST" },
        { "type": "openai-response", "path": "/v1/responses", "method": "POST" }
      ]
    }
  ]
}

Get generation

GEThttps://token.easyapi.com/v1/generation

Example request: https://token.easyapi.com/v1/generation?id=2026100812000000a1b2c3

Returns token counts, cached tokens, the final charge (USD), the ratio snapshot and timings of a single request by request ID. Only requests of the account that owns the key can be queried. To get the cost inline, send usage.include = true in the request body and the usage object gains a cost field; that value is an estimate made before the response is written, this endpoint is authoritative.

Request parameters

NameInTypeRequiredDescription
AuthorizationheaderstringYesEasyAPI API Key. Format: Bearer `API_KEY`
idquerystringYesRequest ID from the X-EasyAPI-Request-Id response header

Response fields

NameTypeDescription
idstringRequest ID
modelstringModel that was called
created_atintegerRecord time (unix seconds)
is_streambooleanWhether the request was streamed
tokens_promptintegerInput tokens
tokens_completionintegerOutput tokens
native_tokens_cachedintegerInput tokens served from cache
native_tokens_cache_writeintegerTokens written to cache
total_costnumberFinal charge (USD)
quotaintegerInternal quota units; 500000 = 1 USD
generation_timeintegerTotal request time (ms, second precision)
latencyintegerTime to first token (ms); 0 for non-streaming requests
billingobjectRatio snapshot used for settlement: model_ratio / completion_ratio / group_ratio / cache_ratio

Response example

{
  "data": {
    "id": "2026100812000000a1b2c3",
    "model": "gpt-4o",
    "created_at": 1759905600,
    "is_stream": true,
    "token_name": "my-key",
    "group": "default",
    "is_byok": false,
    "tokens_prompt": 1000,
    "tokens_completion": 200,
    "total_tokens": 1200,
    "native_tokens_cached": 80,
    "native_tokens_cache_write": 0,
    "total_cost": 0.0045,
    "quota": 2250,
    "generation_time": 3000,
    "latency": 350,
    "billing": {
      "model_ratio": 1.25,
      "completion_ratio": 4,
      "group_ratio": 1,
      "cache_ratio": 0.5
    }
  }
}

Status codes

Status codeMeaningDescriptionData model
200OKBilling record returnedGenerationResponse
400Bad RequestMissing idErrorResponse
401UnauthorizedInvalid API keyErrorResponse
404Not FoundRecord not written yet, or the request belongs to another account; retry shortlyErrorResponse