Usage accounting
Every request through the gateway produces a record. This page describes the record’s fields and the ways to read them.
Per-request record
The fields stored for each request:
generation_id- Opaque, globally unique. Returned in the response and used to look up the record later.
created_at- UTC timestamp when the gateway accepted the request.
api_key_id- The API key used. Names are joined in for display.
user_identifier- Identifier of the account the key belongs to.
model_id- The model that served the request. When the caller sent an auto-routed or aliased ID (e.g.
auto/mistralai/mistral-large-2512), this is the resolved route. provider_id- Upstream provider that served the request.
region- Region the upstream served from.
prompt_tokens- Token count of the input.
completion_tokens- Token count of the output.
total_tokens- Sum.
latency_ms- First-byte latency for streaming, end-to-end for non-streaming.
request_duration_ms- Total request duration measured at the gateway.
cost- Amount debited for this request, in EUR, quantized to €0.000001 (see Cost resolution).
energy_wh- Estimated energy for the inference, in watt-hours.
carbon_gco2e- Estimated CO₂e for the request, in grams.
routing_mode- How the model was selected (e.g. explicit vs. auto).
routing_reason- Why the router picked this model/provider.
fallback_occurred- Whether a fallback provider served the request after a primary failed.
status_code- HTTP status returned to the caller.
Prompt and response content are not stored.
Where to read it
- API lookup: a completion response carries an
id(the samechatcmpl-…/cmpl-…value OpenAI clients already read). Pass it toGET /v1/generation/{id}for the full record, or toGET /v1/metrics/{id}for just the token/energy/carbon figures. Both require the same API key and only return records billed to that key’s account. See Looking up a generation below. - The Activity page: the last ~50 requests, with filtering by date range, model, provider, region, key.
- Transaction detail page: the full record for a single request,
accessed by clicking a row in the transactions table (the detail
path is keyed on the transaction’s
referenceId). - Export: Export CSV on the Activity page produces a
CSV of the filtered range, capped at 50,000 rows per export. The same
file is served to scripts by
GET /v1/usage/export?format=csvwith your API key (see scripted export below). Use either for reconciliation or piping into your own data warehouse.
CSV columns
In order: timestamp, provider, model, tokens, cost,
energy_wh, carbon_gco2e, savings_gco2e, routing_mode,
interrupted, generation_id, api_key, region, eu_sovereign,
prompt_tokens, completion_tokens, cache_read_tokens,
cache_creation_tokens, reasoning_tokens, grid_basis.
New columns are always appended last, so a pipeline reading by position keeps working across releases. Read by header name if you can.
Eight of these exist so a single row stands on its own as evidence:
region: the serving region. Carbon intensity varies roughly fivefold between one provider’s regions, so acarbon_gco2ewithout it cannot be checked against published grid intensities.eu_sovereign: whether the route was served by an EU/EEA-controlled, CLOUD-Act-free provider from an EU/EEA region.falsealso covers routes whose region could not be resolved: it means “not attested sovereign”, not “attested non-sovereign”.prompt_tokens/completion_tokens: input and output are priced separately per 1M tokens, so the summedtokensalone cannot reconcile againstcost.cache_read_tokens/cache_creation_tokens: the part ofprompt_tokensserved from, or written to, the provider’s prompt cache. Cached reads bill at the read rate (typically a tenth of the input rate) and cache writes at the write rate, so aprompt_tokens × input pricecheck on a cached session comes out high (60% high on a 344k-token Claude Code session) unless these are subtracted first.reasoning_tokens: the part ofcompletion_tokensthe model spent on hidden reasoning, where the provider reports it. Billed at the output rate; a row with an empty answer and a fullcompletion_tokensis usually this.grid_basis: the gridcarbon_gco2ewas computed against: the serving region’s locode when its own grid intensity was used, or<CC>-AVERAGE/GLOBAL-AVERAGEwhen an average stood in for a region without one. A Scope-3 report should disclose averaged rows as such. Empty when no carbon was computed, and on requests recorded before the column existed.
The Activity page filters on region and on EU-sovereign routes, and
the export honours whichever filters are active. The same cache and
reasoning counts are on each row’s token tooltip there.
Scripted export
The Activity page’s file is also served on the API, so a FinOps pipeline can fetch it without a browser:
curl "https://api.lowrouter.ai/v1/usage/export?format=csv&date_from=2026-09-01T00:00:00Z&date_to=2026-10-01T00:00:00Z" \
-H "Authorization: Bearer $LOWROUTER_API_KEY" \
-o usage-september.csvThe key’s account is the scope: every key of the account is included,
whichever key made the call. format is required and csv is the only
value; anything else is a 400 naming it, so a client asking for JSON
never receives CSV it then mis-parses. The response is
text/csv with a Content-Disposition filename of
lowrouter-usage-<date>.csv.
Optional query parameters, the same ones the Activity page sends:
date_from,date_to- RFC 3339 bounds on the request time. Omit both for the whole history.
provider,model- Catalogue UUIDs, repeatable — the ids the Activity page’s provider and model filters send.
key- An API key id (
lrk_…) to narrow to one key. region- Serving-region locodes, repeatable.
eu_sovereign=true- Only rows served by an EU-sovereign route.
Rows are most recent first, and the export is capped at 50,000 rows; narrow the date range and page by month when an account exceeds it. The columns are exactly the CSV columns above, in that order, and new columns are appended last there too.
Looking up a generation
Every completion response includes a top-level id. That value is the
generation’s lookup key; there is no separate generation_id field in
the response body.
# The id from a prior chat/completions response
GENERATION_ID="chatcmpl-abc123"
curl https://api.lowrouter.ai/v1/generation/$GENERATION_ID \
-H "Authorization: Bearer $LOWROUTER_API_KEY"The record carries the billed amount (cost.total_cost, in
cost.currency), the route that served the request as the same
{provider}/{creator}/{model}/{locode} id the response’s model field
and the Activity export use, and routing_info.routing_reason in the
same vocabulary as the response’s lowrouter_metadata.routing_reason,
so a row can be reconciled against tokens, prices and the Activity page
without a second lookup. carbon_metrics.grid_basis names the grid the
carbon figure was computed against, as the CSV column above does.
The record is scoped to the calling key’s account: a key can only read
generations billed to its own customer. An unknown id, or one that
belongs to another account, returns 404, so the id space can’t be
probed across accounts.
For just the metrics (tokens, energy, carbon, duration) without the
routing detail, use GET /v1/metrics/{id} with the same id.
Aggregates
The dashboard pre-computes a small set of aggregates and updates them on each request:
- Tokens per day, per model.
- Cost per day, per model.
- Energy and carbon per day, per model.
These aggregates power the charts. They are derived from the per-request records and are reproducible from a CSV export.
Retention
| Data | Retention |
|---|---|
| Per-request records | 90 days from created_at, then rolled into daily aggregates. |
| Daily aggregates | 36 months. |
| Account deletion audit trail | 12 months after deletion, then anonymised. |
| Account profile | For the lifetime of the account. |
After retention expires the per-request rows are deleted and replaced by anonymised aggregates. Aggregates are kept for sustainability reporting and platform analytics; they cannot be used to reconstruct individual requests.
You can request earlier deletion of all your usage records via the support email; the deletion is irreversible and may affect your ability to dispute past invoices.
Cost resolution and rounding
cost is quoted in EUR at a resolution of €0.000001 (one
micro-euro), the same resolution the ledger stores. Each request’s
cost is computed exactly from integer token counts and the per-1M EUR
rate, then rounded up to the next micro-euro. Two consequences:
- A single request over-counts by strictly less than €0.000001, so for any individual call the rounding is negligible.
- The rounding is systematic, not symmetric: at very high volumes of
very small calls, the summed
costexceeds the exact price by up to one micro-euro per request. A million minimal calls can carry up to €1 of accumulated round-up.
We round up rather than to-nearest so that billing can never undercharge below the provider price. The bound above is the whole cost of that choice.
Reconciliation tips
- The sum of
costover a day should equal the daily cost on the dashboard within a rounding tolerance. - The sum of
total_tokensover a day grouped by model is what the upstream provider’s usage report (if you have one) should show. - The carbon estimate is reproducible: given the same
model_id,region, andtotal_tokens, recomputing with the formula in methodology should yield the same gram count.
If the numbers diverge more than rounding allows, that is a bug: open an issue.
