The Responses API

POST /v1/responses accepts requests in OpenAI’s Responses format. Use it with clients that prefer that format over chat-completions, such as Codex and the OpenAI Agents SDK.

It is the same service behind a different request shape. Any model in the catalogue works, auto/ routing and failover included, and a request costs what it would cost on /v1/chat/completions.

Bash
curl https://api.lowrouter.ai/v1/responses \
  -H "Authorization: Bearer YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "auto/mistralai/mistral-large-2512",
    "instructions": "Answer in one sentence.",
    "input": "Why is the sky blue?"
  }'

With the OpenAI SDK, set the base URL and call responses.create as usual:

Python
from openai import OpenAI

client = OpenAI(base_url="https://api.lowrouter.ai/v1", api_key="YOUR_KEY")
response = client.responses.create(
    model="auto/mistralai/mistral-large-2512",
    input="Why is the sky blue?",
)
print(response.output_text)

LowRouter stores nothing

On OpenAI, a response is stored, and a later request can continue it by id. LowRouter does not store responses. That is the one real difference, and it decides what happens to every stateful field.

Send the whole conversation in input on every request. Pass back the output items from the previous response, followed by your new items. This is what clients do when they set store: false.

Fields that depend on stored state are refused with a 400 that names the field. Answering without them would give you a reply to a conversation the model never saw.

previous_response_id
400
conversation
400
prompt (a stored prompt template)
400
background: true
400
An item_reference input item
400
file_id on an image or file part
400. Send image_url, file_data or file_url instead.

GET, DELETE and the other routes under /v1/responses/{id} return a 404, since there is no stored response to address. Like every other route, they return a 401 without a valid API key. To look up what a request used and cost, call GET /v1/generation/{id} with the response’s id.

store: true is accepted, because it does not change the answer. The response says "store": false, and store is listed in lowrouter_unapplied.

What is supported

input as a string or an array of items
applied
instructions
applied, as the system message
message items, with input_text, output_text, input_image and input_file parts
applied
function_call and function_call_output items
applied
reasoning items
removed. They are the model’s notes from an earlier turn and don’t affect the answer.
tools of type function, tool_choice, parallel_tool_calls
applied
max_output_tokens, temperature, top_p
applied
reasoning.effort
applied, as reasoning_effort
text.format (text, json_object, json_schema)
applied, as response_format
metadata, user, service_tier
as on chat-completions: see request parameters
stream
applied
Provider-hosted tools (web_search, file_search, code_interpreter, mcp, …)
400. Only function tools are accepted.
top_logprobs
400. Use /v1/chat/completions with logprobs.
store, include, truncation: "auto", reasoning.summary, text.verbosity, strict on a tool, max_tool_calls, safety_identifier, prompt_cache_key, prompt_cache_retention, stream_options
accepted and disclosed in lowrouter_unapplied

A parameter the serving route could not apply is handled as on chat-completions: a 400 if it changes the answer, a disclosure if it does not. See request parameters.

The response

The response is a Responses API response object. output holds a message item for the text and one function_call item per tool call.

JSON
{
  "id": "chatcmpl-01J9...",
  "object": "response",
  "status": "completed",
  "model": "mistral/mistralai/mistral-large-2512",
  "output": [
    {
      "type": "message",
      "id": "msg_chatcmpl-01J9..._0",
      "status": "completed",
      "role": "assistant",
      "content": [
        { "type": "output_text", "text": "Sunlight scatters...", "annotations": [] }
      ]
    }
  ],
  "usage": {
    "input_tokens": 21,
    "input_tokens_details": { "cached_tokens": 0 },
    "output_tokens": 34,
    "total_tokens": 55,
    "cost": 0.000071,
    "currency": "EUR",
    "remaining_balance": 9.87
  },
  "store": false,
  "lowrouter_metadata": { "provider": "mistral", "routing_mode": "auto" }
}
  • status is incomplete when the answer was cut short, with the reason in incomplete_details: max_output_tokens or content_filter.
  • usage has the Responses API’s field names, plus cost, currency and remaining_balance. input_tokens includes the cached tokens. output_tokens_details is present only when the provider reported a reasoning count: an unknown count is omitted, not stated as 0.
  • lowrouter_metadata is the same block as on chat-completions.
  • lowrouter_unapplied lists what was not applied, using the names you sent: reasoning.effort, not reasoning_effort. It is absent when everything was applied.

Errors have the same shape as on chat-completions. See errors.

Streaming

With stream: true the response is a stream of typed events, each with a sequence_number:

Text
response.created
response.in_progress
response.output_item.added
response.content_part.added
response.output_text.delta          (one per fragment)
response.output_text.done
response.content_part.done
response.output_item.done
response.completed

A tool call is an output item of its own. It streams as response.function_call_arguments.delta events, followed by response.function_call_arguments.done.

The last event carries the whole response: output, usage with the cost, lowrouter_metadata and lowrouter_unapplied. It is one of:

EventMeaning
response.completedThe answer is whole.
response.incompleteThe answer was cut short. See incomplete_details.
response.failedThe request failed after the stream started. See error. Nothing follows this event.

A stream that is cut off before the model finishes also ends with response.failed, with the error code server_error. No output item is closed and the event carries no output and no usage: what you received as deltas is partial, and a tool call that was still streaming must not be run. You are billed for the tokens generated before the cut; the request is marked interrupted on the dashboard and in the CSV export.