Usage accounting

Every request through the gateway produces a record. This page describes the record’s fields and the ways to read them.

Per-request record

The fields stored for each request:

generation_id
Opaque, globally unique. Returned in the response and used to look up the record later.
created_at
UTC timestamp when the gateway accepted the request.
api_key_id
The API key used. Names are joined in for display.
user_identifier
Identifier of the account the key belongs to.
model_id
The model that served the request. When the caller sent an auto-routed or aliased ID (e.g. auto/mistralai/mistral-large-2512), this is the resolved route.
provider_id
Upstream provider that served the request.
region
Region the upstream served from.
prompt_tokens
Token count of the input.
completion_tokens
Token count of the output.
total_tokens
Sum.
latency_ms
First-byte latency for streaming, end-to-end for non-streaming.
request_duration_ms
Total request duration measured at the gateway.
cost
Amount debited for this request, in EUR, quantized to €0.000001 (see Cost resolution).
energy_wh
Estimated energy for the inference, in watt-hours.
carbon_gco2e
Estimated CO₂e for the request, in grams.
routing_mode
How the model was selected (e.g. explicit vs. auto).
routing_reason
Why the router picked this model/provider.
fallback_occurred
Whether a fallback provider served the request after a primary failed.
status_code
HTTP status returned to the caller.

Prompt and response content are not stored.

Where to read it

  • API lookup: a completion response carries an id (the same chatcmpl-…/cmpl-… value OpenAI clients already read). Pass it to GET /v1/generation/{id} for the full record, or to GET /v1/metrics/{id} for just the token/energy/carbon figures. Both require the same API key and only return records billed to that key’s account. See Looking up a generation below.
  • The Activity page: the last ~50 requests, with filtering by date range, model, provider, region, key.
  • Transaction detail page: the full record for a single request, accessed by clicking a row in the transactions table (the detail path is keyed on the transaction’s referenceId).
  • Export: Export CSV on the Activity page produces a CSV of the filtered range, capped at 50,000 rows per export. The same file is served to scripts by GET /v1/usage/export?format=csv with your API key (see scripted export below). Use either for reconciliation or piping into your own data warehouse.

CSV columns

In order: timestamp, provider, model, tokens, cost, energy_wh, carbon_gco2e, savings_gco2e, routing_mode, interrupted, generation_id, api_key, region, eu_sovereign, prompt_tokens, completion_tokens, cache_read_tokens, cache_creation_tokens, reasoning_tokens, grid_basis.

New columns are always appended last, so a pipeline reading by position keeps working across releases. Read by header name if you can.

Eight of these exist so a single row stands on its own as evidence:

  • region: the serving region. Carbon intensity varies roughly fivefold between one provider’s regions, so a carbon_gco2e without it cannot be checked against published grid intensities.
  • eu_sovereign: whether the route was served by an EU/EEA-controlled, CLOUD-Act-free provider from an EU/EEA region. false also covers routes whose region could not be resolved: it means “not attested sovereign”, not “attested non-sovereign”.
  • prompt_tokens / completion_tokens: input and output are priced separately per 1M tokens, so the summed tokens alone cannot reconcile against cost.
  • cache_read_tokens / cache_creation_tokens: the part of prompt_tokens served from, or written to, the provider’s prompt cache. Cached reads bill at the read rate (typically a tenth of the input rate) and cache writes at the write rate, so a prompt_tokens × input price check on a cached session comes out high (60% high on a 344k-token Claude Code session) unless these are subtracted first.
  • reasoning_tokens: the part of completion_tokens the model spent on hidden reasoning, where the provider reports it. Billed at the output rate; a row with an empty answer and a full completion_tokens is usually this.
  • grid_basis: the grid carbon_gco2e was computed against: the serving region’s locode when its own grid intensity was used, or <CC>-AVERAGE / GLOBAL-AVERAGE when an average stood in for a region without one. A Scope-3 report should disclose averaged rows as such. Empty when no carbon was computed, and on requests recorded before the column existed.

The Activity page filters on region and on EU-sovereign routes, and the export honours whichever filters are active. The same cache and reasoning counts are on each row’s token tooltip there.

Scripted export

The Activity page’s file is also served on the API, so a FinOps pipeline can fetch it without a browser:

Bash
curl "https://api.lowrouter.ai/v1/usage/export?format=csv&date_from=2026-09-01T00:00:00Z&date_to=2026-10-01T00:00:00Z" \
  -H "Authorization: Bearer $LOWROUTER_API_KEY" \
  -o usage-september.csv

The key’s account is the scope: every key of the account is included, whichever key made the call. format is required and csv is the only value; anything else is a 400 naming it, so a client asking for JSON never receives CSV it then mis-parses. The response is text/csv with a Content-Disposition filename of lowrouter-usage-<date>.csv.

Optional query parameters, the same ones the Activity page sends:

date_from, date_to
RFC 3339 bounds on the request time. Omit both for the whole history.
provider, model
Catalogue UUIDs, repeatable — the ids the Activity page’s provider and model filters send.
key
An API key id (lrk_…) to narrow to one key.
region
Serving-region locodes, repeatable.
eu_sovereign=true
Only rows served by an EU-sovereign route.

Rows are most recent first, and the export is capped at 50,000 rows; narrow the date range and page by month when an account exceeds it. The columns are exactly the CSV columns above, in that order, and new columns are appended last there too.

Looking up a generation

Every completion response includes a top-level id. That value is the generation’s lookup key; there is no separate generation_id field in the response body.

Bash
# The id from a prior chat/completions response
GENERATION_ID="chatcmpl-abc123"

curl https://api.lowrouter.ai/v1/generation/$GENERATION_ID \
  -H "Authorization: Bearer $LOWROUTER_API_KEY"

The record carries the billed amount (cost.total_cost, in cost.currency), the route that served the request as the same {provider}/{creator}/{model}/{locode} id the response’s model field and the Activity export use, and routing_info.routing_reason in the same vocabulary as the response’s lowrouter_metadata.routing_reason, so a row can be reconciled against tokens, prices and the Activity page without a second lookup. carbon_metrics.grid_basis names the grid the carbon figure was computed against, as the CSV column above does.

The record is scoped to the calling key’s account: a key can only read generations billed to its own customer. An unknown id, or one that belongs to another account, returns 404, so the id space can’t be probed across accounts.

For just the metrics (tokens, energy, carbon, duration) without the routing detail, use GET /v1/metrics/{id} with the same id.

Aggregates

The dashboard pre-computes a small set of aggregates and updates them on each request:

  • Tokens per day, per model.
  • Cost per day, per model.
  • Energy and carbon per day, per model.

These aggregates power the charts. They are derived from the per-request records and are reproducible from a CSV export.

Retention

DataRetention
Per-request records90 days from created_at, then rolled into daily aggregates.
Daily aggregates36 months.
Account deletion audit trail12 months after deletion, then anonymised.
Account profileFor the lifetime of the account.

After retention expires the per-request rows are deleted and replaced by anonymised aggregates. Aggregates are kept for sustainability reporting and platform analytics; they cannot be used to reconstruct individual requests.

You can request earlier deletion of all your usage records via the support email; the deletion is irreversible and may affect your ability to dispute past invoices.

Cost resolution and rounding

cost is quoted in EUR at a resolution of €0.000001 (one micro-euro), the same resolution the ledger stores. Each request’s cost is computed exactly from integer token counts and the per-1M EUR rate, then rounded up to the next micro-euro. Two consequences:

  • A single request over-counts by strictly less than €0.000001, so for any individual call the rounding is negligible.
  • The rounding is systematic, not symmetric: at very high volumes of very small calls, the summed cost exceeds the exact price by up to one micro-euro per request. A million minimal calls can carry up to €1 of accumulated round-up.

We round up rather than to-nearest so that billing can never undercharge below the provider price. The bound above is the whole cost of that choice.

Reconciliation tips

  • The sum of cost over a day should equal the daily cost on the dashboard within a rounding tolerance.
  • The sum of total_tokens over a day grouped by model is what the upstream provider’s usage report (if you have one) should show.
  • The carbon estimate is reproducible: given the same model_id, region, and total_tokens, recomputing with the formula in methodology should yield the same gram count.

If the numbers diverge more than rounding allows, that is a bug: open an issue.