Claude Code

Claude Code speaks the Anthropic Messages API rather than the OpenAI-compatible one. LowRouter serves that shape at /v1/messages, so it can be pointed at the gateway directly.

Configure

Verified against Claude Code 2.1.x. The config below is re-run against a pinned release every week.

Claude Code reads its endpoint and key from the environment:

Bash
export ANTHROPIC_BASE_URL="https://api.lowrouter.ai"
export ANTHROPIC_AUTH_TOKEN="sk-lr-..."

Then run claude as usual.

ANTHROPIC_API_KEY works too, as does the Anthropic SDK directly: /v1/messages accepts your key on either Authorization: Bearer or the x-api-key header the SDK sends by default. This is specific to this endpoint: the OpenAI-compatible surface takes Authorization: Bearer, which is what every OpenAI SDK sends. Requests are routed, billed and carbon-accounted exactly like any other request on your key, with the same balance, usage reporting and routing constraints.

Two settings you will probably need

Claude Code assumes an Anthropic model on the other end, so two of its defaults need adjusting when you point it elsewhere.

Bash
# Claude Code only knows its own model ids, and warns about ours. This
# tells it to trust the context window the API reports.
export CLAUDE_CODE_DISABLE_UNKNOWN_MODEL_WINDOW_ENFORCEMENT=1

# It requests 32000 output tokens by default, which exceeds what some
# models allow (gpt-4o-mini caps at 16384). Lower it to fit your model.
export CLAUDE_CODE_MAX_OUTPUT_TOKENS=8000

Without the second one you will see a provider error naming the real ceiling (max_tokens is too large: 32000). That is the model’s limit, not ours.

Two notices you can ignore

Both are Claude Code talking about its own assumptions, and neither means anything is wrong with the routing:

  • [claude-code:unrecognized_model] is logged because Claude Code only knows Anthropic’s own model IDs and ours are three-segment. The request goes through regardless. The first setting above stops it from also clamping the context window on that basis.
  • The auto-mode classifier notice (“this session isn’t eligible for no-charge classifier requests… your requests go through api.lowrouter.ai”) appears the first time auto mode wants to run a safety check. Claude Code names the gateway it has detected and runs its classifier through the session’s model instead, billed like any other request on your key. Auto mode keeps working. See Anthropic’s auto-mode classifier page for what it is checking.

Choosing a model

/v1/messages accepts any model id in the catalogue, not only Anthropic ones. The id is the same three-segment form the rest of the API uses:

Bash
export ANTHROPIC_MODEL="anthropic/anthropic/claude-sonnet-5"

Pointing it at a non-Anthropic model works, and is the more interesting configuration if you are routing for cost, sovereignty or carbon:

Bash
export ANTHROPIC_MODEL="auto/qwen/qwen-3-coder"

Tool-calling behaviour is a property of the model, not of the protocol. Models differ in how reliably they emit well-formed tool calls, and an agentic client depends on that heavily. If a model produces malformed tool arguments, the model is at fault, not the translation.

The routes below are verified with the real CLI and the configuration on this page, tool loop included. Every week each one is run against production: Claude Code has to read a file with its Read tool and write what it holds to a new file with its Write tool. The check looks at the file on disk. A run that ends cleanly without writing it counts as a failure.

The check is a text tool loop. Each route also says which inputs it takes beyond text, from its input_modalities in /v1/models, and the weekly run has Claude Code read those too: a screenshot on a route that takes images, a PDF on one that takes documents. A route that does not take an input refuses it with a 400 that ends the session, so if you paste screenshots or open PDFs, pick a route whose inputs include them.

  • anthropic/anthropic/claude-sonnet-5, and Anthropic’s other models. Inputs: text, images and documents.
  • openai/openai/gpt-5-mini. Inputs: text and images.
  • openai/openai/gpt-4o-mini. This one needs CLAUDE_CODE_MAX_OUTPUT_TOKENS at 16384 or below, as set above, because the model caps its output there. Inputs: text, images and documents.
  • deepinfra/z-ai/glm-5.3-flash. Inputs: text and images.
  • mistral/mistralai/mistral-medium-3.5, EU-sovereign at /eu (see below). This is the configuration to reach for when routing for sovereignty. Inputs: text and images.
  • mistral/mistralai/codestral-2508, EU-sovereign at /eu. Inputs: text only.
  • scaleway/openai/gpt-oss-120b, EU-sovereign. Inputs: text only.

The Mistral IDs above name no region, so they resolve to Mistral’s global endpoint, which commits to no inference location. For EU-sovereign processing, append /eu (for example mistral/mistralai/mistral-medium-3.5/eu), which runs on Mistral’s EU endpoint at a 10% premium.

The EU-sovereign routes currently serve without prompt caching, so the same session reads its whole context at full price on every turn and costs several times more than on a route that caches. See the usage block’s cache_read_input_tokens.

The list follows the weekly check, in both directions. A route above that fails two weeks running moves to the list below, and a route below that passes three weeks running moves back. Each move arrives as a pull request carrying every route’s recent pass rate.

The routes below failed the check and are for plain prompting only until they pass it again. They answer single requests with well-formed tool calls, but under Claude Code’s full agent prompt they have stopped using the tools: writing tool names and fragments of arguments as text and ending the session with no error and no work done, while every turn is still billed, or running out of turns without writing anything. Their results have varied from week to week, which is why they are still measured.

  • mistral/mistralai/mistral-large-2512, EU-sovereign at /eu. Inputs: text and images.
  • mistral/mistralai/ministral-14b-2512, EU-sovereign at /eu. Inputs: text and images.

Mistral refuses the end-user field that Claude Code’s metadata becomes. On a Mistral route it is removed before the request leaves, and the response lists metadata in lowrouter_unapplied on every turn. Nothing about the answer changes. See request parameters.

Gemini models are not usable for tool calling here: the Gemini API requires a thought_signature on every function call it previously emitted, and that value has no equivalent in the Anthropic request shape, so we cannot reconstruct it. Gemini works fine for plain prompting.

What the endpoint supports

  • Content blocks: text, image and document, with base64, url and plain-text sources. See Input modalities for which providers accept which.
  • system as a top-level field, string or array of blocks.
  • Tools: tools, tool_use, tool_result and tool_choice, translated onto whatever dialect the selected provider speaks. A tool_result may carry image and document blocks, which is how Claude Code delivers every screenshot and PDF its Read tool opens. See Reading files for what each route does with them.
  • metadata: its user_id goes to the provider as the end-user identifier where the route accepts one. Where it does not (Mistral, OVHcloud, Gemini and non-Claude Bedrock models), the response says so in lowrouter_unapplied as metadata.
  • cache_control, including the TTL variant, preserved exactly as sent on routes that implement it (Anthropic, Claude on Bedrock and Vertex, and Bedrock’s own cache points). The 5-minute and 1-hour TTLs bill at different rates, so we do not normalise one into the other. On a route with no such breakpoint (every OpenAI-compatible provider, such as Mistral, OpenAI, Scaleway and DeepInfra, and Gemini), the marker is removed before the request leaves, and the response says so in lowrouter_unapplied (cache_control, unsupported_on_this_path). Claude Code marks its system prompt on every turn, so on those routes you will see that entry on every turn and get no cache discount, because the provider offers none to mark. Read the usage block’s cache counters, not the marker, to know what was cached.
  • Streaming: the Anthropic SSE event sequence (message_start, content_block_start, content_block_delta, content_block_stop, message_delta, message_stop). A stream that is cut off before the model finishes ends with an error event and nothing after it (no content_block_stop and no message_stop), as Anthropic’s own streams do when they fail. What was received so far is partial.
  • usage in Anthropic’s field names, including cache_creation_input_tokens, cache_read_input_tokens and the cache_creation TTL breakdown.
  • lowrouter_metadata, the per-request receipt: the route that served the request, the carbon estimate and its methodology, the cache counters and, because Anthropic’s usage object has no room for them, the money fields cost, currency and remaining_balance that the OpenAI-compatible surface keeps on usage. On a non-streaming response it is a top-level field. On a stream it rides message_delta, next to usage. See per-request metadata.

input_tokens excludes cached tokens, following Anthropic’s convention, so the three usage fields sum to your total prompt size rather than overlapping.

Reading files

When Claude Code reads a PNG or a PDF, the file does not arrive as a user message: the CLI puts it inside the tool_result block it constructs for its own Read call. That block goes through on every route. The weekly verification drives the real CLI through this case (a screenshot and a PDF read in one session) against the model configured above:

Bash
claude -p "Use the Read tool on ./screenshot.png, then on ./report.pdf. Reply with only the word printed in the PDF." \
  --output-format text --max-turns 6

What happens to the image or document depends on the route:

  • Anthropic’s own models, Claude on Bedrock and Vertex, and Bedrock’s other models take images and documents inside a tool result natively. The block is forwarded as sent.
  • Every OpenAI-compatible provider, and Gemini have a text-only tool result. The text stays in the tool result, and the image or document is moved into a user message placed immediately after it, so the model still sees the file, as a user turn rather than as the tool’s own output. The response says so in lowrouter_unapplied (tool_result, unsupported_on_this_path, with the number of blocks moved), so nothing is dropped silently.

Whether the model can read the file is the model’s own capability: check input_modalities on the route in /v1/models, and see Input modalities.

Extended thinking

thinking is applied. It becomes a reasoning level that each provider receives in its own form. On Claude models it becomes extended thinking again: adaptive thinking on Claude 4.6 and later, a token budget on earlier models. Other providers get their own reasoning control. A token budget maps to the nearest level: up to 2048 tokens is low, up to 8192 medium, up to 16384 high, and anything larger xhigh. With adaptive thinking, output_config.effort sets the level and defaults to high.

Sometimes a route cannot apply it. The model may not support thinking, the request may set a non-default temperature, or an older Claude model may be in the middle of a tool loop. The answer still comes back, and lowrouter_unapplied on the message says so. When streaming, it is on message_delta. See request parameters.

output_config.format (structured output) and top_k are carried too.

A thinking (or redacted_thinking) content block replayed in a later turn is dropped rather than forwarded. It is the model’s own scratchpad from a previous turn, no provider we route to can accept it in another model’s request, and its signature is bound to whoever minted it, so replaying it is meaningless even where a block type exists. Unlike a dropped image, a reasoning trace the model never sees does not change the answer, which is why this one is not disclosed per-request.

What we don’t claim

  • usage on message_start is a zeroed placeholder. Anthropic sends usage on that event too, but nothing is known yet when we emit it. The real counts (input_tokens, output_tokens and the cache counters) arrive on message_delta at the end of the stream, together with the receipt. They are the billed counts, so there is no need to poll /v1/generation/{id} for them afterwards.
  • We do not implement Anthropic’s provider-hosted tools. Server-side tools we cannot meter are refused rather than silently forwarded.
  • This is a translation, not an emulation. Behaviour that comes from the model (reasoning quality, tool-call reliability, instruction following) is the model’s, and changes when you change the routed model.