Routing

Every request goes through the router. For an explicit model the router only resolves it to a healthy upstream that serves it. For an auto-routed model it also chooses which provider and region serve it. This page describes both.

Three modes

  • Explicit model. You send a full model ID (see available models); the router resolves it to the provider and region encoded in that ID and sends it upstream.
  • Auto-routed model (auto/<creator>/<model>). You name the model; the router picks which provider serves it and from where. See auto-routing a model below.
  • Alias (alias/<name>). You send the name of one of your account’s model aliases; the router resolves it to the stored route and continues exactly like an explicit model. The alias/ prefix (like auto/) is a reserved namespace, so no provider can ever be named that.

The model field is required. There is no per-request route object and no separate routing parameters: everything the router needs comes from the model ID you send and the model catalogue. To constrain provider or region, encode it in the model ID (below).

lowrouter/auto has been removed. It used to let the router pick the model as well as the route, ranked on carbon per token, which meant the greenest answer was always the smallest model in the catalogue: a capability downgrade you had not asked for and could not see. You choose the model, and the router chooses the route. Requests naming it now return 400 and point here.

Auto-routing a model

The same model is often served by several providers, in several regions, at different prices and on very different electricity grids. auto/<creator>/<model> lets you name the model and leave that choice to the router:

Text
auto/mistralai/mistral-large-2512

You get the same canonical model, with the same weights and output; only the route is decided for you. Auto-routed IDs appear in GET /v1/models alongside the explicit ones (see available models) for every model served by more than one route, so they are selectable in any OpenAI-compatible client. They are omitted when you filter that listing with ?jurisdiction=, because auto ranking is global and cannot promise to stay inside a facet, so a filtered catalogue lists only explicit IDs, which name their provider and region outright.

The priority order

This is the built-in standard routing profile: what every account starts with, and what a request runs under until you configure a routing policy. Each step only breaks ties left by the previous one:

  1. EU-sovereign routes first. A route counts as sovereign only when both halves hold: the provider is an EU-sovereign company and the serving region is in the EU/EEA. An EU-sovereign company serving from a US datacenter does not qualify, and neither does a US-controlled provider serving from Frankfurt.
  2. Lowest grid carbon intensity. A region whose grid intensity we have no data for ranks below every region we do. An unknown is treated as unknown, never as the greenest.
  3. Lowest price. Because carbon is decided first, two identically-priced routes can never resolve to the dirtier grid. A route with no published price ranks below every priced one.
  4. Round-robin across whatever is still tied.

Routes a provider has taken out of service are excluded outright, at every step.

Sovereignty here is a preference, not a guarantee

If a model has no sovereign route at all, an auto request does not fail. It falls through to the greenest, then cheapest route available anywhere. auto/anthropic/claude-sonnet-5 routes, to a non-EU provider.

The response tells you which case you are in. lowrouter_metadata.eu_sovereign tells you whether the route you got was sovereign, alongside provider and region, and the model field carries the fully resolved four-segment ID. If you need sovereignty as a hard constraint rather than a preference, use an explicit ID. It fails loudly instead of falling back outside your constraint, and it carries the same eu_sovereign attestation, so the pin is checkable on every response rather than only at the moment you chose it.

lowrouter_metadata.routing_reason names the step that decided: eu_sovereign_preferred, lowest_carbon_intensity, lowest_cost, round_robin, or only_route when the model has just one route, plus eu_hosted_preferred, cloud_act_free_preferred and europe_preferred when a policy lists those criteria. Beside it, lowrouter_metadata.routing_profile names the profile the decision was made under (standard unless you configured one).

Routing policy

You can change the order above, and add constraints it never had, from /app/routing. A routing profile has two halves that work differently:

  • Preference order: the criteria the router ranks by, highest priority first. Drag them into any order. Available: eu-sovereign (the pair above), eu-hosted (EU/EEA region only; says nothing about who controls the provider), cloud-act-free (provider entity outside US legal reach; says nothing about where it serves from), europe (geographic Europe, wider than EU/EEA: the UK, Switzerland, Norway and Iceland qualify), lowest-carbon and lowest-price. Criteria you leave unlisted fall to the end in default order, and round-robin is always the final tiebreak. Reordering changes which route wins, never which routes are candidates. An unknown carbon or price figure still ranks last, never first, whatever the order; that rule is not configurable.
  • Hard gates: absolute constraints. A route a gate excludes is not a candidate: it is never ranked, and never used as a fallback either. Gates: deny_cloud_act (no CLOUD-Act-exposed provider, including any whose exposure is not attested), a carbon ceiling (max_carbon_gco2e_per_1k_tokens), a price ceiling (max_price_per_1k_tokens, in your billing currency), and allow/deny lists of providers, model creators and regions (locodes, or 2-letter country codes matching every region in that country).

Ceilings fail closed on unknown values. While a carbon or price ceiling is set, a route with no published figure is excluded, because an unknown value cannot be shown to satisfy a ceiling. This is different from the preference keys, where an unknown merely ranks last. A generous ceiling you expect to barely bind can therefore drop every route with a data gap; the dashboard’s tester shows you exactly which, before you save.

If your gates leave no eligible route for a model, the request is refused with 403 routing_policy_exhausted, naming the gate. It is never silently downgraded to a route you excluded, and never a 404 (the model exists) or a 503 (nothing is down). See the errors reference.

A few rules of scope:

  • Policy shapes auto/ selection only. An explicit four-segment ID or an alias is you choosing the route; no gate applies to it and no profile is reported. Pinning remains the escape hatch when a policy refuses a request.
  • A routing policy is set per API key: assign a profile to a key from the API keys page. Exactly one profile is tagged default (the built-in standard until you mark another one), and it applies to every key you have not assigned a profile to, new keys included. The key’s own profile wins over the default. A profile change takes effect on the very next request.
  • standard is always there. Making another profile the default demotes it, never removes it: it stays listed, stays editable, and can be made the default again at any time. It cannot be deleted or renamed.
  • Changing the default asks you to confirm the blast radius (how many of your keys will switch) before anything is sent. None of the routing actions asks you to re-enter your password: a routing change can only choose among LowRouter’s own providers, every change is recorded with who made it and the policy before and after, and every change is reversible. (Creating an API key does ask, because a key can spend your balance from outside the dashboard.)
  • A profile cannot be deleted while it is the default (make another profile the default first) or while any key is assigned to it (unassign those keys on the API keys page first). The dashboard says which keys. Deleting a profile never silently changes the policy a key runs under.
  • An order that undercuts itself (lowest-price above eu-sovereign, say, which lets a cheaper non-sovereign route beat a sovereign one) is accepted, not rejected. The tester makes the effect visible.
  • Every change to a profile is recorded: who, when, and the policy before and after. “Never CLOUD Act” is a compliance claim, and when it is relaxed someone will need to prove when and by whom.

The routing tester at the bottom of the same page takes an auto/ ID and shows the full ranked ladder under the profile you are editing, saved or not: which routes form the round-robin pool at the top, each route’s sovereignty verdicts, carbon and price, and any route a gate excluded together with the gate that did it. It runs the same ranking and gate evaluation as a live request and sends nothing to a provider.

Budgeting an auto ID before you send

An auto ID has no single price or carbon figure, because the route is chosen per request, so its catalogue entry carries the envelope instead:

JSON
{
  "id": "auto/openai/gpt-oss-120b",
  "auto_routing": {
    "candidates": [
      "scaleway/openai/gpt-oss-120b/fr-par",
      "nebius/openai/gpt-oss-120b/fi-mnt",
      "aws-bedrock/openai/gpt-oss-120b/br-gru"
    ],
    "pricing_range": {
      "prompt_per_1m_tokens":     { "min": 0.033, "max": 1.00 },
      "completion_per_1m_tokens": { "min": 0.150, "max": 4.20 },
      "currency": "EUR"
    },
    "carbon_range": {
      "carbon_per_token_gco2e":             { "min": 1.8e-5, "max": 1.4e-3 },
      "grid_carbon_intensity_gco2_per_kwh": { "min": 8, "max": 640 }
    }
  }
}
  • candidates is every concrete route the ID can resolve to, in the order a request would try them: the route auto would pick right now first, then that provider’s other regions, then the remaining providers in priority order. Each is a four-segment ID listed elsewhere in the same catalogue, so you can join it against that entry’s regions[] for exact figures, or pin it outright. (The sample above is abridged to three; a widely-served model really returns a few dozen.)
  • pricing_range and carbon_range are the bounds across those candidates, including the ones a failover may fall through to. The response’s usage.cost and lowrouter_metadata.carbon for any request sent to this ID fall inside them.
  • A range is published only when every candidate carries that figure. If one route has no published price, or no energy data, the corresponding range is omitted rather than narrowed. candidates still names the route, so you can see which one the gap is.

The envelope is derived from the same route set the router ranks, at the moment the catalogue is built, under the default profile. For an account that has not configured a routing policy it cannot drift from what a request will actually do.

With a configured policy, the envelope is a default-profile prediction, not a per-account guarantee. The catalogue is global, unfiltered and cacheable, and no policy enters it, so an account whose order or gates differ from the default can resolve to a route outside the published candidates[] (its order ranked a different rung first), see a price or carbon figure outside the published range (its gates removed the rung that set the bound), or be refused with 403 routing_policy_exhausted for an ID the catalogue lists candidates for. None of those are catalogue errors. The accurate per-policy view is the routing tester on /app/routing, which ranks the same route set under your profile.

Grid intensity is resolved the same way in the listing and the response, by the most specific figure available: the serving region’s own, else its country or bloc average (an eu region is accounted at the EU average), else the worldwide average. A region’s carbon_metrics key tells you which: the locode for a regional figure, <CC>-AVERAGE or GLOBAL-AVERAGE for an averaged one.

What auto-routing does not do

  • It never changes which model you get. That is your choice, made by naming it.
  • It does not score latency or answer quality.
  • It does not consult your aliases: auto/ is its own namespace and always applies your routing profile.
  • A version shorthand still works (auto/mistralai/mistral-large resolves to the latest version), but a region does not: there is no 4th segment to write, because choosing the region is the point.

Pinning a provider or region

Region and provider are part of the model ID, not a separate field. The public ID has the form {provider}/{creator}/{model}[/{locode}]:

GoalHow
Pin the providerUse an explicit model ID; the first segment is the provider, e.g. aws-bedrock/anthropic/claude-sonnet-5.
Pin the regionAppend a UN/LOCODE as the 4th segment, e.g. aws-bedrock/mistralai/ministral-3-3b-instruct/br-gru.
Default regionOmit the 4th segment; see how IDs resolve below.

If you request a region a model isn’t served in, the request is rejected rather than silently served elsewhere, so the region you pin is the region you get.

How IDs resolve

You do not have to write the full four-segment ID. What you leave out is filled in from the catalogue, most-specific first, and the response always tells you what you actually got.

You sendWhat happens
mistral/mistralai/mistral-large-2512/euUsed exactly as written.
mistral/mistralai/mistral-large-2512Region defaulted: resolves to /global.
mistral/mistralai/mistral-largeVersion and region defaulted: resolves to mistral-large-2512/global.
anthropic/anthropic/claude-sonnetResolves to the latest sonnet at /global.

Missing version resolves to the most recent version of that same model. The version is the only thing that moves: claude-sonnet resolves to the newest sonnet, never to an opus, and mistral-large never to a mistral-small. A shorthand that would be ambiguous between genuinely different models, such as openai/openai/gpt-oss (where gpt-oss-20b and gpt-oss-120b are separate models rather than two versions of one), is rejected, not guessed.

Missing region resolves to global when the model has a global endpoint. When it doesn’t, it resolves to the model’s only region, or, where several exist, to the lowest-carbon one. A region whose grid intensity we don’t have data for is never chosen over one we do: an unknown is treated as unknown, not as zero.

A global endpoint makes no promise about where the request runs. Mistral’s, for example, commits to no inference location, while its /eu endpoint processes in the EU at a 10% premium. If the location matters to you, pin the region rather than relying on the default.

The response’s model field always carries the fully resolved four-segment ID, so a request is reproducible from its own response: copy it back and you pin exactly what ran. lowrouter_metadata.region reports the region independently.

Pinning still fails loudly

Resolution only ever fills in what you left out. It never overrides what you wrote:

  • A version that doesn’t exist is an error, not a nudge to the nearest one. mistral/mistralai/mistral-large-9999 → 404.
  • A region the model isn’t served in is an error, not a reroute. mistral/mistralai/mistral-large-2512/us-iad → 404.
  • An explicit /global on a model with no global endpoint is an error, not a silent switch to a regional row.

The asymmetry is deliberate. Defaults exist so the IDs in our own docs and model list are callable as printed; pins exist so that when you name a version or a jurisdiction, the thing you named is the thing that ran. Aliases are stricter still: an alias target must be a full, concrete {provider}/{creator}/{model}/{locode} triple, so a stored route can never drift under you.

Worked scenario: EU-only first, lowest carbon second

Say a data-residency commitment requires EU inference, and within that constraint you want the lowest carbon available. Sovereignty here is a constraint and carbon is an optimisation, and the two are handled by different mechanisms.

  1. Encode the constraint. Filter the providers page to EU providers, then compare per-token carbon across the models that remain in the catalogue.
  2. Pin the winner with an explicit ID, e.g. mistral/mistralai/mistral-large-2512/eu, or better, point an alias at it, so you can re-run the comparison next quarter and repoint without touching client code.
  3. Verify. Check lowrouter_metadata.region on the responses; the pin is proven on every call.

This scenario does not ask the router for “EU only”, because auto-routing does not treat a jurisdiction as a hard constraint. auto/<creator>/<model> does prefer EU-sovereign routes above everything else, and reports whether it got one, but when a model has no sovereign route it falls back rather than failing, so it is an optimisation, not a commitment. Per-key jurisdiction restrictions are coming; until then the right tool for a hard constraint is the explicit ID, which fails loudly rather than falling back outside it.

What happens on failure

The two ID forms behave differently on failure, and that difference is the main reason to choose one over the other.

An explicit ID is a pin, and it fails loudly. If the route you named is down, the request returns 503 (or 504 if the provider accepted the connection and never answered) naming that provider. It is never quietly served by a different one. The four-segment form exists to make that guarantee: substituting a provider would bill against another pricing row and serve from another jurisdiction than the ID promised, so when you have pinned a route, we would rather fail than move you.

An auto/ ID may fall through its published candidates, in order. You delegated the route choice, so trying the next eligible route is what you asked for. The walk stops at the first success, and while eu-sovereign leads your preference order (it does under standard) it never leaves the sovereign set: a request that resolved to an EU-sovereign route is never rescued by a non-sovereign one. Under a routing policy that ranks price or carbon above sovereignty, a sovereign winner is incidental and every route the policy admits is a fallback. Every response reports what happened: lowrouter_metadata.providers_attempted lists the routes tried in order, fallback_occurred says whether the first choice served it, fallback_reason says why it did not, and eu_sovereign states whether the route that answered satisfies the sovereign pair.

Be aware of the consequence: a request that falls through is billed against the pricing row of the route that actually served it, and may run in a different jurisdiction than the first-ranked route would have. auto/ makes that trade on your behalf. The candidates list and the pricing_range / carbon_range envelope published with every auto ID show the full set of outcomes before you send anything. If you need a single row, pin it.

The walk is bounded, so a model whose every route is slow cannot make you wait indefinitely: it stops after a fixed number of attempts, or once the remaining time budget cannot fund another one. When it stops with routes untried, the failure body says so, so “everything is down” is distinguishable from “we ran out of budget”:

JSON
{
  "error": {
    "type": "provider_unavailable",
    "code": "provider_timeout",
    "message": "... (tried 2 of 5 candidate routes within the routing budget; 2 remaining — pin one of the model's published candidates to reach it directly)",
    "providers_attempted": ["berget", "tensorx"],
    "breaker_skipped": ["nebius/us-mci"],
    "candidates_remaining": 2
  }
}

candidates_remaining: 0 means the ladder was walked to its end and the model is genuinely unavailable right now. Anything above zero means routes were left untried, and pinning one of them from the candidates list may succeed where the auto/ ID did not. breaker_skipped names routes we declined to try because recent failures had opened their circuit breaker. They are published candidates, so they are named to make the arithmetic add up.

Fallback models

Everything above moves a request between routes of the model you asked for. A fallback chain is the other half: an ordered list of other models to serve the request when every route of the one you named is down, or when your prompt does not fit its context window.

You have exactly three choices, set on a routing policy or on an alias:

fallback.modeWhat happens
noneNothing. A request the model cannot serve fails, with the receipt above explaining what was tried. This is the default.
defaultA chain we curate per model, published in our repository and reviewed like any other change. The dashboard lists the published chains and only offers this choice once at least one exists; a model with no published chain behaves exactly as none. No chains are published yet.
customYour own list, one to five entries, tried in the order you give.
JSON
{
  "order": ["eu-sovereign", "lowest-carbon", "lowest-price"],
  "fallback": {
    "mode": "custom",
    "models": ["auto/anthropic/claude-sonnet-5", "scaleway/z-ai/glm-5.3"]
  }
}

An entry can be any ID form the gateway accepts. Prefer auto/{creator}/{model}: it picks a route at the moment the fallback is needed, so the entry is not tied to a region or provider chosen in advance. Each entry is dispatched once, to the best-ranked route for that model at that moment; if that route fails too, the walk moves to the next entry rather than through that model’s other providers. A four-segment ID names one route and gets no ladder of its own, and alias/{name} resolves to that alias’s target route only, never to the alias’s own chain, so a chain always reads top to bottom and cannot loop.

When a chain is consulted

  • Every route of your model failed: connection refused, timeout, 429, 5xx. The chain is tried after the last of them.
  • Your prompt does not fit the model’s window, or it rejected an input or parameter the model does not support. Here the chain is tried immediately: every route of one model rejects a given payload identically, so there is nothing to gain from trying them, but a model with a bigger window may take it. If every fallback model also rejects the request, you get your model’s error back, not the last one’s, because the window you need to know about is the one you asked for.
  • Never mid-stream. Once the first byte of a streamed answer has reached you, the request is committed. A second model’s output appended to a partial answer would be worse than an error.
  • Never past a gate. A model your hard gates exclude is not used as a fallback, and a request that started on an EU-sovereign route is never rescued by a non-sovereign model.

A chain is per request and temporary. The request that hit the failure is served by the entry that answered; the next request starts from your model again, and the circuit breaker is what keeps traffic off a dead route in between. There is no stickiness and no flag to clear.

Seeing that it happened

A response served by a fallback model carries lowrouter_metadata.requested_model (the route you would have been served by) beside model, which names what actually answered, and fallback_reason, which is context_length when the chain ran because your prompt did not fit. Your activity page marks those requests, and because the tokens and cost on that row belong to the model that answered, requested_model is what a cost or carbon comparison is computed against.

Entries are checked when you save them, not only when they run, because a chain is used during outages and a typo discovered then means the chain fails when it is needed. A model withdrawn after you save is skipped at request time rather than failing the request.

What the router does not do

  • It does not choose the model; you do, by naming it.
  • It does not benchmark output quality. It optimises for sovereignty, then carbon, then cost.
  • It does not accept a per-request route object or a prefer_low_carbon flag. Low carbon is already ranked ahead of price, so there is nothing to opt into.
  • It does not substitute a model unless you asked it to. Without a fallback chain, a request is served by the model you named or it fails.