The Responses API
POST /v1/responses accepts requests in OpenAI’s Responses format.
Use it with clients that prefer that format over chat-completions, such
as Codex and the OpenAI Agents SDK.
It is the same service behind a different request shape. Any model in
the catalogue works, auto/ routing and failover included, and a
request costs what it would cost on /v1/chat/completions.
curl https://api.lowrouter.ai/v1/responses \
-H "Authorization: Bearer YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "auto/mistralai/mistral-large-2512",
"instructions": "Answer in one sentence.",
"input": "Why is the sky blue?"
}'With the OpenAI SDK, set the base URL and call responses.create as
usual:
from openai import OpenAI
client = OpenAI(base_url="https://api.lowrouter.ai/v1", api_key="YOUR_KEY")
response = client.responses.create(
model="auto/mistralai/mistral-large-2512",
input="Why is the sky blue?",
)
print(response.output_text)LowRouter stores nothing
On OpenAI, a response is stored, and a later request can continue it by id. LowRouter does not store responses. That is the one real difference, and it decides what happens to every stateful field.
Send the whole conversation in input on every request. Pass back
the output items from the previous response, followed by your new
items. This is what clients do when they set store: false.
Fields that depend on stored state are refused with a 400 that names the field. Answering without them would give you a reply to a conversation the model never saw.
previous_response_id- 400
conversation- 400
prompt(a stored prompt template)- 400
background: true- 400
- An
item_referenceinput item - 400
file_idon an image or file part- 400. Send
image_url,file_dataorfile_urlinstead.
GET, DELETE and the other routes under /v1/responses/{id} return a
404, since there is no stored response to address. Like every other
route, they return a 401 without a valid API key. To look up what a
request used and cost, call GET /v1/generation/{id} with the
response’s id.
store: true is accepted, because it does not change the answer. The
response says "store": false, and store is listed in
lowrouter_unapplied.
What is supported
inputas a string or an array of items- applied
instructions- applied, as the system message
messageitems, withinput_text,output_text,input_imageandinput_fileparts- applied
function_callandfunction_call_outputitems- applied
reasoningitems- removed. They are the model’s notes from an earlier turn and don’t affect the answer.
toolsof typefunction,tool_choice,parallel_tool_calls- applied
max_output_tokens,temperature,top_p- applied
reasoning.effort- applied, as
reasoning_effort text.format(text,json_object,json_schema)- applied, as
response_format metadata,user,service_tier- as on chat-completions: see request parameters
stream- applied
- Provider-hosted tools (
web_search,file_search,code_interpreter,mcp, …) - 400. Only
functiontools are accepted. top_logprobs- 400. Use
/v1/chat/completionswithlogprobs. store,include,truncation: "auto",reasoning.summary,text.verbosity,stricton a tool,max_tool_calls,safety_identifier,prompt_cache_key,prompt_cache_retention,stream_options- accepted and disclosed in
lowrouter_unapplied
A parameter the serving route could not apply is handled as on chat-completions: a 400 if it changes the answer, a disclosure if it does not. See request parameters.
The response
The response is a Responses API response object. output holds a
message item for the text and one function_call item per tool call.
{
"id": "chatcmpl-01J9...",
"object": "response",
"status": "completed",
"model": "mistral/mistralai/mistral-large-2512",
"output": [
{
"type": "message",
"id": "msg_chatcmpl-01J9..._0",
"status": "completed",
"role": "assistant",
"content": [
{ "type": "output_text", "text": "Sunlight scatters...", "annotations": [] }
]
}
],
"usage": {
"input_tokens": 21,
"input_tokens_details": { "cached_tokens": 0 },
"output_tokens": 34,
"total_tokens": 55,
"cost": 0.000071,
"currency": "EUR",
"remaining_balance": 9.87
},
"store": false,
"lowrouter_metadata": { "provider": "mistral", "routing_mode": "auto" }
}statusisincompletewhen the answer was cut short, with the reason inincomplete_details:max_output_tokensorcontent_filter.usagehas the Responses API’s field names, pluscost,currencyandremaining_balance.input_tokensincludes the cached tokens.output_tokens_detailsis present only when the provider reported a reasoning count: an unknown count is omitted, not stated as 0.lowrouter_metadatais the same block as on chat-completions.lowrouter_unappliedlists what was not applied, using the names you sent:reasoning.effort, notreasoning_effort. It is absent when everything was applied.
Errors have the same shape as on chat-completions. See errors.
Streaming
With stream: true the response is a stream of typed events, each with
a sequence_number:
response.created
response.in_progress
response.output_item.added
response.content_part.added
response.output_text.delta (one per fragment)
response.output_text.done
response.content_part.done
response.output_item.done
response.completedA tool call is an output item of its own. It streams as
response.function_call_arguments.delta events, followed by
response.function_call_arguments.done.
The last event carries the whole response: output, usage with the
cost, lowrouter_metadata and lowrouter_unapplied. It is one of:
| Event | Meaning |
|---|---|
response.completed | The answer is whole. |
response.incomplete | The answer was cut short. See incomplete_details. |
response.failed | The request failed after the stream started. See error. Nothing follows this event. |
A stream that is cut off before the model finishes also ends with
response.failed, with the error code server_error. No output item
is closed and the event carries no output and no usage: what you
received as deltas is partial, and a tool call that was still streaming
must not be run. You are billed for the tokens generated before the
cut; the request is marked interrupted on the dashboard and in the
CSV export.
