Routing
Every request goes through the router. For an explicit model the router only resolves it to a healthy upstream that serves it. For an auto-routed model it also chooses which provider and region serve it. This page describes both.
Three modes
- Explicit model. You send a full model ID (see available models); the router resolves it to the provider and region encoded in that ID and sends it upstream.
- Auto-routed model (
auto/<creator>/<model>). You name the model; the router picks which provider serves it and from where. See auto-routing a model below. - Alias (
alias/<name>). You send the name of one of your account’s model aliases; the router resolves it to the stored route and continues exactly like an explicit model. Thealias/prefix (likeauto/) is a reserved namespace, so no provider can ever be named that.
The model field is required. There is no per-request route object
and no separate routing parameters: everything the router needs comes
from the model ID you send and the model catalogue. To constrain
provider or region, encode it in the model ID (below).
lowrouter/auto has been removed. It used to let the router pick
the model as well as the route, ranked on carbon per token, which
meant the greenest answer was always the smallest model in the
catalogue: a capability downgrade you had not asked for and could not
see. You choose the model, and the router chooses the route.
Requests naming it now return 400 and point here.
Auto-routing a model
The same model is often served by several providers, in several
regions, at different prices and on very different electricity grids.
auto/<creator>/<model> lets you name the model and leave that choice
to the router:
auto/mistralai/mistral-large-2512You get the same canonical model, with the same weights and output;
only the route is decided for you. Auto-routed IDs appear in GET /v1/models alongside the explicit ones
(see available models) for every model served by more
than one route, so they are selectable in any OpenAI-compatible
client. They are omitted when you filter that listing with
?jurisdiction=, because auto ranking is global and cannot promise to
stay inside a facet, so a filtered catalogue lists only explicit IDs, which
name their provider and region outright.
The priority order
This is the built-in standard routing profile: what every account
starts with, and what a request runs under until you configure a
routing policy. Each step only breaks ties left by
the previous one:
- EU-sovereign routes first. A route counts as sovereign only when both halves hold: the provider is an EU-sovereign company and the serving region is in the EU/EEA. An EU-sovereign company serving from a US datacenter does not qualify, and neither does a US-controlled provider serving from Frankfurt.
- Lowest grid carbon intensity. A region whose grid intensity we have no data for ranks below every region we do. An unknown is treated as unknown, never as the greenest.
- Lowest price. Because carbon is decided first, two identically-priced routes can never resolve to the dirtier grid. A route with no published price ranks below every priced one.
- Round-robin across whatever is still tied.
Routes a provider has taken out of service are excluded outright, at every step.
Sovereignty here is a preference, not a guarantee
If a model has no sovereign route at all, an auto request does not
fail. It falls through to the greenest, then cheapest route
available anywhere. auto/anthropic/claude-sonnet-5 routes, to a
non-EU provider.
The response tells you which case you are in.
lowrouter_metadata.eu_sovereign tells you whether the route you got
was sovereign, alongside provider and region, and the model
field carries the fully resolved four-segment ID. If you need
sovereignty as a hard constraint rather than a preference, use an
explicit ID. It fails loudly instead of falling back outside your
constraint, and it carries the same eu_sovereign attestation, so the
pin is checkable on every response rather than only at the moment you
chose it.
lowrouter_metadata.routing_reason names the step that decided:
eu_sovereign_preferred, lowest_carbon_intensity, lowest_cost,
round_robin, or only_route when the model has just one route,
plus eu_hosted_preferred, cloud_act_free_preferred and
europe_preferred when a policy lists those criteria. Beside it,
lowrouter_metadata.routing_profile names the profile the decision
was made under (standard unless you configured one).
Routing policy
You can change the order above, and add constraints it never had, from
/app/routing. A routing
profile has two halves that work differently:
- Preference order: the criteria the router ranks by, highest
priority first. Drag them into any order. Available:
eu-sovereign(the pair above),eu-hosted(EU/EEA region only; says nothing about who controls the provider),cloud-act-free(provider entity outside US legal reach; says nothing about where it serves from),europe(geographic Europe, wider than EU/EEA: the UK, Switzerland, Norway and Iceland qualify),lowest-carbonandlowest-price. Criteria you leave unlisted fall to the end in default order, and round-robin is always the final tiebreak. Reordering changes which route wins, never which routes are candidates. An unknown carbon or price figure still ranks last, never first, whatever the order; that rule is not configurable. - Hard gates: absolute constraints. A route a gate excludes is not
a candidate: it is never ranked, and never used as a fallback either.
Gates:
deny_cloud_act(no CLOUD-Act-exposed provider, including any whose exposure is not attested), a carbon ceiling (max_carbon_gco2e_per_1k_tokens), a price ceiling (max_price_per_1k_tokens, in your billing currency), and allow/deny lists of providers, model creators and regions (locodes, or 2-letter country codes matching every region in that country).
Ceilings fail closed on unknown values. While a carbon or price ceiling is set, a route with no published figure is excluded, because an unknown value cannot be shown to satisfy a ceiling. This is different from the preference keys, where an unknown merely ranks last. A generous ceiling you expect to barely bind can therefore drop every route with a data gap; the dashboard’s tester shows you exactly which, before you save.
If your gates leave no eligible route for a model, the request is
refused with 403 routing_policy_exhausted, naming the gate. It is never
silently downgraded to a route you excluded, and never a 404 (the model
exists) or a 503 (nothing is down). See the
errors reference.
A few rules of scope:
- Policy shapes
auto/selection only. An explicit four-segment ID or an alias is you choosing the route; no gate applies to it and no profile is reported. Pinning remains the escape hatch when a policy refuses a request. - A routing policy is set per API key: assign a profile to a key
from the API keys page. Exactly one profile is tagged default
(the built-in
standarduntil you mark another one), and it applies to every key you have not assigned a profile to, new keys included. The key’s own profile wins over the default. A profile change takes effect on the very next request. standardis always there. Making another profile the default demotes it, never removes it: it stays listed, stays editable, and can be made the default again at any time. It cannot be deleted or renamed.- Changing the default asks you to confirm the blast radius (how many of your keys will switch) before anything is sent. None of the routing actions asks you to re-enter your password: a routing change can only choose among LowRouter’s own providers, every change is recorded with who made it and the policy before and after, and every change is reversible. (Creating an API key does ask, because a key can spend your balance from outside the dashboard.)
- A profile cannot be deleted while it is the default (make another profile the default first) or while any key is assigned to it (unassign those keys on the API keys page first). The dashboard says which keys. Deleting a profile never silently changes the policy a key runs under.
- An order that undercuts itself (
lowest-priceaboveeu-sovereign, say, which lets a cheaper non-sovereign route beat a sovereign one) is accepted, not rejected. The tester makes the effect visible. - Every change to a profile is recorded: who, when, and the policy before and after. “Never CLOUD Act” is a compliance claim, and when it is relaxed someone will need to prove when and by whom.
The routing tester at the bottom of the same page takes an auto/
ID and shows the full ranked ladder under the profile you are editing,
saved or not: which routes form the round-robin pool at the top,
each route’s sovereignty verdicts, carbon and price, and any route a
gate excluded together with the gate that did it. It runs the same
ranking and gate evaluation as a live request and sends nothing to a
provider.
Budgeting an auto ID before you send
An auto ID has no single price or carbon figure, because the route is chosen per request, so its catalogue entry carries the envelope instead:
{
"id": "auto/openai/gpt-oss-120b",
"auto_routing": {
"candidates": [
"scaleway/openai/gpt-oss-120b/fr-par",
"nebius/openai/gpt-oss-120b/fi-mnt",
"aws-bedrock/openai/gpt-oss-120b/br-gru"
],
"pricing_range": {
"prompt_per_1m_tokens": { "min": 0.033, "max": 1.00 },
"completion_per_1m_tokens": { "min": 0.150, "max": 4.20 },
"currency": "EUR"
},
"carbon_range": {
"carbon_per_token_gco2e": { "min": 1.8e-5, "max": 1.4e-3 },
"grid_carbon_intensity_gco2_per_kwh": { "min": 8, "max": 640 }
}
}
}candidatesis every concrete route the ID can resolve to, in the order a request would try them: the route auto would pick right now first, then that provider’s other regions, then the remaining providers in priority order. Each is a four-segment ID listed elsewhere in the same catalogue, so you can join it against that entry’sregions[]for exact figures, or pin it outright. (The sample above is abridged to three; a widely-served model really returns a few dozen.)pricing_rangeandcarbon_rangeare the bounds across those candidates, including the ones a failover may fall through to. The response’susage.costandlowrouter_metadata.carbonfor any request sent to this ID fall inside them.- A range is published only when every candidate carries that
figure. If one route has no published price, or no energy data, the
corresponding range is omitted rather than narrowed.
candidatesstill names the route, so you can see which one the gap is.
The envelope is derived from the same route set the router ranks, at the moment the catalogue is built, under the default profile. For an account that has not configured a routing policy it cannot drift from what a request will actually do.
With a configured policy, the envelope is a default-profile
prediction, not a per-account guarantee. The catalogue is global,
unfiltered and cacheable, and no policy enters it, so an account
whose order or gates differ from the default can resolve to a route
outside the published candidates[] (its order ranked a different
rung first), see a price or carbon figure outside the published range
(its gates removed the rung that set the bound), or be refused with
403 routing_policy_exhausted for an ID the catalogue lists candidates
for. None of those are catalogue errors. The accurate per-policy view
is the routing tester on /app/routing, which ranks the
same route set under your profile.
Grid intensity is resolved the same way in the listing and the
response, by the most specific figure available: the serving
region’s own, else its country or bloc average (an eu region is
accounted at the EU average), else the worldwide average. A region’s
carbon_metrics key tells you which: the locode for a regional figure,
<CC>-AVERAGE or GLOBAL-AVERAGE for an averaged one.
What auto-routing does not do
- It never changes which model you get. That is your choice, made by naming it.
- It does not score latency or answer quality.
- It does not consult your aliases:
auto/is its own namespace and always applies your routing profile. - A version shorthand still works (
auto/mistralai/mistral-largeresolves to the latest version), but a region does not: there is no 4th segment to write, because choosing the region is the point.
Pinning a provider or region
Region and provider are part of the model ID, not a separate
field. The public ID has the form
{provider}/{creator}/{model}[/{locode}]:
| Goal | How |
|---|---|
| Pin the provider | Use an explicit model ID; the first segment is the provider, e.g. aws-bedrock/anthropic/claude-sonnet-5. |
| Pin the region | Append a UN/LOCODE as the 4th segment, e.g. aws-bedrock/mistralai/ministral-3-3b-instruct/br-gru. |
| Default region | Omit the 4th segment; see how IDs resolve below. |
If you request a region a model isn’t served in, the request is rejected rather than silently served elsewhere, so the region you pin is the region you get.
How IDs resolve
You do not have to write the full four-segment ID. What you leave out is filled in from the catalogue, most-specific first, and the response always tells you what you actually got.
| You send | What happens |
|---|---|
mistral/mistralai/mistral-large-2512/eu | Used exactly as written. |
mistral/mistralai/mistral-large-2512 | Region defaulted: resolves to /global. |
mistral/mistralai/mistral-large | Version and region defaulted: resolves to mistral-large-2512/global. |
anthropic/anthropic/claude-sonnet | Resolves to the latest sonnet at /global. |
Missing version resolves to the most recent version of that same
model. The version is the only thing that moves: claude-sonnet
resolves to the newest sonnet, never to an opus, and
mistral-large never to a mistral-small. A shorthand that would be
ambiguous between genuinely different models, such as
openai/openai/gpt-oss (where gpt-oss-20b and gpt-oss-120b are
separate models rather than two versions of one), is rejected, not
guessed.
Missing region resolves to global when the model has a global
endpoint. When it doesn’t, it resolves to the model’s only region, or,
where several exist, to the lowest-carbon one. A region whose grid
intensity we don’t have data for is never chosen over one we do: an
unknown is treated as unknown, not as zero.
A global endpoint makes no promise about where the request runs. Mistral’s,
for example, commits to no inference location, while its /eu endpoint
processes in the EU at a 10% premium. If the location matters to you, pin
the region rather than relying on the default.
The response’s model field always carries the fully resolved
four-segment ID, so a request is reproducible from its own response:
copy it back and you pin exactly what ran. lowrouter_metadata.region
reports the region independently.
Pinning still fails loudly
Resolution only ever fills in what you left out. It never overrides what you wrote:
- A version that doesn’t exist is an error, not a nudge to the nearest
one.
mistral/mistralai/mistral-large-9999→404. - A region the model isn’t served in is an error, not a reroute.
mistral/mistralai/mistral-large-2512/us-iad→404. - An explicit
/globalon a model with no global endpoint is an error, not a silent switch to a regional row.
The asymmetry is deliberate. Defaults exist so the IDs in our own
docs and model list are callable as printed; pins exist so that when
you name a version or a jurisdiction, the thing you named is the thing
that ran. Aliases are stricter still: an alias target must
be a full, concrete {provider}/{creator}/{model}/{locode} triple, so
a stored route can never drift under you.
Worked scenario: EU-only first, lowest carbon second
Say a data-residency commitment requires EU inference, and within that constraint you want the lowest carbon available. Sovereignty here is a constraint and carbon is an optimisation, and the two are handled by different mechanisms.
- Encode the constraint. Filter the providers page to EU providers, then compare per-token carbon across the models that remain in the catalogue.
- Pin the winner with an explicit ID, e.g.
mistral/mistralai/mistral-large-2512/eu, or better, point an alias at it, so you can re-run the comparison next quarter and repoint without touching client code. - Verify. Check
lowrouter_metadata.regionon the responses; the pin is proven on every call.
This scenario does not ask the router for “EU only”, because
auto-routing does not treat a jurisdiction as a hard constraint.
auto/<creator>/<model> does prefer EU-sovereign routes above
everything else, and reports whether it got one, but when a model has
no sovereign route it falls back rather than failing, so it is an
optimisation, not a commitment. Per-key jurisdiction restrictions are
coming; until then the right tool for a hard constraint is the
explicit ID, which fails loudly rather than falling back outside it.
What happens on failure
The two ID forms behave differently on failure, and that difference is the main reason to choose one over the other.
An explicit ID is a pin, and it fails loudly. If the route you
named is down, the request returns 503 (or 504 if the provider
accepted the connection and never answered) naming that provider. It
is never quietly served by a different one. The four-segment form
exists to make that guarantee: substituting a provider would bill
against another pricing row and serve from another jurisdiction than
the ID promised, so when you have pinned a route, we would rather fail
than move you.
An auto/ ID may fall through its published candidates, in
order. You delegated the route choice, so trying the next eligible
route is what you asked for. The walk stops at the first success, and
while eu-sovereign leads your preference order (it does under
standard) it never leaves the sovereign set: a request that resolved
to an EU-sovereign route is never rescued by a non-sovereign one. Under
a routing policy that ranks price or carbon above
sovereignty, a sovereign winner is incidental and every route the
policy admits is a fallback. Every response reports what happened:
lowrouter_metadata.providers_attempted lists the routes tried in
order, fallback_occurred says whether the first choice served it,
fallback_reason says why it did not, and eu_sovereign states
whether the route that answered satisfies the sovereign pair.
Be aware of the consequence: a request that falls through is billed
against the pricing row of the route that actually served it, and may
run in a different jurisdiction than the first-ranked route would
have. auto/ makes that trade on your behalf. The candidates list
and the pricing_range / carbon_range envelope published with every
auto ID show the full set of outcomes before you send anything. If you
need a single row, pin it.
The walk is bounded, so a model whose every route is slow cannot make you wait indefinitely: it stops after a fixed number of attempts, or once the remaining time budget cannot fund another one. When it stops with routes untried, the failure body says so, so “everything is down” is distinguishable from “we ran out of budget”:
{
"error": {
"type": "provider_unavailable",
"code": "provider_timeout",
"message": "... (tried 2 of 5 candidate routes within the routing budget; 2 remaining — pin one of the model's published candidates to reach it directly)",
"providers_attempted": ["berget", "tensorx"],
"breaker_skipped": ["nebius/us-mci"],
"candidates_remaining": 2
}
}candidates_remaining: 0 means the ladder was walked to its end and
the model is genuinely unavailable right now. Anything above zero
means routes were left untried, and pinning one of them from the
candidates list may succeed where the auto/ ID did not.
breaker_skipped names routes we declined to try because recent
failures had opened their circuit breaker. They are published
candidates, so they are named to make the arithmetic add up.
Fallback models
Everything above moves a request between routes of the model you asked for. A fallback chain is the other half: an ordered list of other models to serve the request when every route of the one you named is down, or when your prompt does not fit its context window.
You have exactly three choices, set on a routing policy or on an alias:
fallback.mode | What happens |
|---|---|
none | Nothing. A request the model cannot serve fails, with the receipt above explaining what was tried. This is the default. |
default | A chain we curate per model, published in our repository and reviewed like any other change. The dashboard lists the published chains and only offers this choice once at least one exists; a model with no published chain behaves exactly as none. No chains are published yet. |
custom | Your own list, one to five entries, tried in the order you give. |
{
"order": ["eu-sovereign", "lowest-carbon", "lowest-price"],
"fallback": {
"mode": "custom",
"models": ["auto/anthropic/claude-sonnet-5", "scaleway/z-ai/glm-5.3"]
}
}An entry can be any ID form the gateway accepts. Prefer
auto/{creator}/{model}: it picks a route at the moment the fallback
is needed, so the entry is not tied to a region or provider chosen in
advance. Each entry is dispatched once, to the best-ranked route for
that model at that moment; if that route fails too, the walk moves to
the next entry rather than through that model’s other providers.
A four-segment ID names one route and gets no ladder of its own, and
alias/{name} resolves to that alias’s target route only, never
to the alias’s own chain, so a chain always reads top to bottom and
cannot loop.
When a chain is consulted
- Every route of your model failed: connection refused, timeout, 429, 5xx. The chain is tried after the last of them.
- Your prompt does not fit the model’s window, or it rejected an input or parameter the model does not support. Here the chain is tried immediately: every route of one model rejects a given payload identically, so there is nothing to gain from trying them, but a model with a bigger window may take it. If every fallback model also rejects the request, you get your model’s error back, not the last one’s, because the window you need to know about is the one you asked for.
- Never mid-stream. Once the first byte of a streamed answer has reached you, the request is committed. A second model’s output appended to a partial answer would be worse than an error.
- Never past a gate. A model your hard gates exclude is not used as a fallback, and a request that started on an EU-sovereign route is never rescued by a non-sovereign model.
A chain is per request and temporary. The request that hit the failure is served by the entry that answered; the next request starts from your model again, and the circuit breaker is what keeps traffic off a dead route in between. There is no stickiness and no flag to clear.
Seeing that it happened
A response served by a fallback model carries
lowrouter_metadata.requested_model (the route you would have been
served by) beside model, which names what actually answered, and
fallback_reason, which is context_length when the chain ran
because your prompt did not fit. Your activity page marks those
requests, and because the tokens and cost on that row belong to the
model that answered, requested_model is what a cost or carbon
comparison is computed against.
Entries are checked when you save them, not only when they run, because a chain is used during outages and a typo discovered then means the chain fails when it is needed. A model withdrawn after you save is skipped at request time rather than failing the request.
What the router does not do
- It does not choose the model; you do, by naming it.
- It does not benchmark output quality. It optimises for sovereignty, then carbon, then cost.
- It does not accept a per-request
routeobject or aprefer_low_carbonflag. Low carbon is already ranked ahead of price, so there is nothing to opt into. - It does not substitute a model unless you asked it to. Without a fallback chain, a request is served by the model you named or it fails.
