Run your first completion

This page sends one chat completion through the gateway, walks through what came back, and points at the things you’ll come back to.

Send the request

Set your key in an environment variable so it doesn’t end up in shell history:

Bash
export LOWROUTER_API_KEY="sk-lr-..."

Then call the gateway:

Bash
curl https://api.lowrouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $LOWROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "auto/mistralai/mistral-large-2512",
    "messages": [
      {"role": "user", "content": "In one sentence, what is a vector database?"}
    ]
  }'

An auto/<creator>/<model> ID names the model you want and leaves the route to us: among the providers and regions serving it, LowRouter prefers EU-sovereign routes, then the lowest-carbon one (modeled from annual grid averages; see the methodology). You get the model you named either way; LowRouter only decides where it runs.

The model field is required. You can name the route yourself instead, for example openai/openai/gpt-4.1/global or anthropic/anthropic/claude-haiku-4.5/global, pinning a region with a UN/LOCODE as the fourth segment (e.g. aws-bedrock/mistralai/ministral-3-3b-instruct/br-gru), or call an alias you can repoint later without changing code. See routing. Custom auto-routing policies (a carbon or price ceiling, a provider or region allowlist, a preference order) are set on the Auto-Routing page and apply account-wide or per API key.

What comes back

The response is OpenAI-shaped. The fields you’ll use most are:

JSON
{
  "id": "chatcmpl-01J9...",
  "object": "chat.completion",
  "created": 1714150000,
  "model": "openai/openai/gpt-4.1",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "A vector database stores …"
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 18,
    "completion_tokens": 26,
    "total_tokens": 44,
    "cost": 0.000034,
    "currency": "EUR",
    "remaining_balance": 9.87
  },
  "lowrouter_metadata": {
    "provider": "openai",
    "region": "eu-west",
    "energy_wh": 0.0021,
    "carbon_gco2e": 0.00057,
    "carbon_intensity_gco2_per_kwh": 45.0,
    "estimation_methodology": "ecologits_formula",
    "routing_mode": "auto",
    "routing_reason": "lowest_carbon_intensity",
    "fallback_occurred": false,
    "providers_attempted": ["openai"]
  }
}

The OpenAI-compatible parts (id, choices, usage, …) follow the standard OpenAI Chat Completions shape. LowRouter adds three billing fields to usage:

  • cost: what this request cost, in currency. Because it is a per-request figure, you can attribute it to a user, tenant, or workflow without scraping a dashboard or maintaining your own price table.
  • currency: ISO-4217 code for cost and remaining_balance (EUR).
  • remaining_balance: your credit balance after this request.

Costs are billed in integer micros (millionths of a euro) and rounded up to the next micro, so a request never bills below what it cost us. On very small requests that rounding is visible: the amount you are charged can exceed the exact computed cost by up to €0.000001.

The lowrouter_metadata block carries the LowRouter-specific fields:

  • provider: which upstream actually served the request.
  • region: the region the upstream served from.
  • energy_wh and carbon_gco2e: the energy and carbon estimate for this request. Read methodology before quoting these numbers anywhere.
  • routing_mode, routing_reason, fallback_occurred, and providers_attempted: the routing trace, showing how the model was picked and whether a fallback was used.

To look up the full record for a request later, open its transaction on the dashboard.

Stream the response

For interactive UIs, set "stream": true and read Server-Sent Events:

Bash
curl https://api.lowrouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $LOWROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -N \
  -d '{
    "model": "auto/mistralai/mistral-large-2512",
    "stream": true,
    "messages": [{"role": "user", "content": "Count to 5 slowly"}]
  }'

The stream is a Server-Sent Events stream in the same format the OpenAI streaming API uses, so any OpenAI-compatible SDK consumes it unchanged. See integrations for per-SDK snippets.

Streaming keeps cost attribution. The final chunk before data: [DONE] carries lowrouter_metadata and one complete usage object with token counts and money together, in the same shape the non-streaming response returns:

JSON
{
  "object": "chat.completion.chunk",
  "choices": [],
  "usage": {
    "prompt_tokens": 12,
    "completion_tokens": 40,
    "total_tokens": 52,
    "cost": 0.000034,
    "currency": "EUR",
    "remaining_balance": 9.87
  },
  "lowrouter_metadata": { "provider": "openai", "...": "..." }
}

usage arrives in exactly one frame, so a client that reads “the usage chunk” gets everything from it. Cost is only known after the generation finishes and is settled, which is why it can only appear on this final frame. Token counts are held back to accompany it instead of being sent earlier on their own.

This frame has an empty choices array, and it is sent whether or not you set stream_options.include_usage, because it is the request’s receipt. A loop that indexes choices[0] on every chunk must skip it; the Python and TypeScript pages show the one-line guard.

This holds for every provider. Some attach their usage to a chunk that also carries generated text or the end of the turn; that chunk is passed on with its usage removed, and the counts arrive on the final frame with the cost.

If a stream is cut short (you cancel it, your client times out, or the connection drops), you are still billed for the tokens the provider generated before the cut. Those requests are marked interrupted on the dashboard, in the CSV export, and on the generation record, so you can reconcile them against your own logs.

Things that might surprise you

  • The model in the response is the resolved four-segment ID, not necessarily the string you sent. If you sent an auto/ or alias/ ID, the response tells you which route actually ran.
  • usage reflects upstream tokens, which may include caching discounts (for providers that support them). The credits charged on your dashboard match this usage.
  • The carbon estimate is absent on requests we couldn’t classify, e.g. a model whose parameter count is unknown. The dashboard shows the same record without an eco number rather than a fabricated one.

If it didn’t work

A 401, 402, or 404 here almost always means a key, balance, or model-id problem. The errors reference lists every payload the gateway returns, with the one-line fix for each.

Next

Tour the dashboard →