
# The Responses API

`POST /v1/responses` accepts requests in OpenAI's Responses format.
Use it with clients that prefer that format over chat-completions, such
as Codex and the OpenAI Agents SDK.

It is the same service behind a different request shape. Any model in
the catalogue works, `auto/` routing and failover included, and a
request costs what it would cost on `/v1/chat/completions`.

```bash
curl https://api.lowrouter.ai/v1/responses \
  -H "Authorization: Bearer YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "auto/mistralai/mistral-large-2512",
    "instructions": "Answer in one sentence.",
    "input": "Why is the sky blue?"
  }'
```

With the OpenAI SDK, set the base URL and call `responses.create` as
usual:

```python
from openai import OpenAI

client = OpenAI(base_url="https://api.lowrouter.ai/v1", api_key="YOUR_KEY")
response = client.responses.create(
    model="auto/mistralai/mistral-large-2512",
    input="Why is the sky blue?",
)
print(response.output_text)
```

## LowRouter stores nothing

On OpenAI, a response is stored, and a later request can continue it by
id. LowRouter does not store responses. That is the one real difference,
and it decides what happens to every stateful field.

**Send the whole conversation in `input` on every request.** Pass back
the `output` items from the previous response, followed by your new
items. This is what clients do when they set `store: false`.

Fields that depend on stored state are refused with a 400 that names
the field. Answering without them would give you a reply to a
conversation the model never saw.

| Field | Result |
|---|---|
| `previous_response_id` | **400** |
| `conversation` | **400** |
| `prompt` (a stored prompt template) | **400** |
| `background: true` | **400** |
| An `item_reference` input item | **400** |
| `file_id` on an image or file part | **400**. Send `image_url`, `file_data` or `file_url` instead. |

`GET`, `DELETE` and the other routes under `/v1/responses/{id}` return a
404, since there is no stored response to address. Like every other
route, they return a 401 without a valid API key. To look up what a
request used and cost, call `GET /v1/generation/{id}` with the
response's `id`.

`store: true` is accepted, because it does not change the answer. The
response says `"store": false`, and `store` is listed in
`lowrouter_unapplied`.

## What is supported

| Field | Result |
|---|---|
| `input` as a string or an array of items | applied |
| `instructions` | applied, as the system message |
| `message` items, with `input_text`, `output_text`, `input_image` and `input_file` parts | applied |
| `function_call` and `function_call_output` items | applied |
| `reasoning` items | removed. They are the model's notes from an earlier turn and don't affect the answer. |
| `tools` of type `function`, `tool_choice`, `parallel_tool_calls` | applied |
| `max_output_tokens`, `temperature`, `top_p` | applied |
| `reasoning.effort` | applied, as [`reasoning_effort`](request-parameters#reasoning_effort) |
| `text.format` (`text`, `json_object`, `json_schema`) | applied, as [`response_format`](request-parameters#response_format) |
| `metadata`, `user`, `service_tier` | as on chat-completions: see [request parameters](request-parameters) |
| `stream` | applied |
| Provider-hosted tools (`web_search`, `file_search`, `code_interpreter`, `mcp`, …) | **400**. Only `function` tools are accepted. |
| `top_logprobs` | **400**. Use `/v1/chat/completions` with `logprobs`. |
| `store`, `include`, `truncation: "auto"`, `reasoning.summary`, `text.verbosity`, `strict` on a tool, `max_tool_calls`, `safety_identifier`, `prompt_cache_key`, `prompt_cache_retention`, `stream_options` | accepted and disclosed in `lowrouter_unapplied` |

A parameter the serving route could not apply is handled as on
chat-completions: a 400 if it changes the answer, a disclosure if it
does not. See [request parameters](request-parameters).

## The response

The response is a Responses API `response` object. `output` holds a
`message` item for the text and one `function_call` item per tool call.

```json
{
  "id": "chatcmpl-01J9...",
  "object": "response",
  "status": "completed",
  "model": "mistral/mistralai/mistral-large-2512",
  "output": [
    {
      "type": "message",
      "id": "msg_chatcmpl-01J9..._0",
      "status": "completed",
      "role": "assistant",
      "content": [
        { "type": "output_text", "text": "Sunlight scatters...", "annotations": [] }
      ]
    }
  ],
  "usage": {
    "input_tokens": 21,
    "input_tokens_details": { "cached_tokens": 0 },
    "output_tokens": 34,
    "total_tokens": 55,
    "cost": 0.000071,
    "currency": "EUR",
    "remaining_balance": 9.87
  },
  "store": false,
  "lowrouter_metadata": { "provider": "mistral", "routing_mode": "auto" }
}
```

- `status` is `incomplete` when the answer was cut short, with the
  reason in `incomplete_details`: `max_output_tokens` or
  `content_filter`.
- `usage` has the Responses API's field names, plus `cost`, `currency`
  and `remaining_balance`. `input_tokens` includes the cached tokens.
  `output_tokens_details` is present only when the provider reported a
  reasoning count: an unknown count is omitted, not stated as 0.
- [`lowrouter_metadata`](../models/per-request-metadata) is the same
  block as on chat-completions.
- `lowrouter_unapplied` lists what was not applied, using the names you
  sent: `reasoning.effort`, not `reasoning_effort`. It is absent when
  everything was applied.

Errors have the same shape as on chat-completions. See
[errors](errors).

## Streaming

With `stream: true` the response is a stream of typed events, each with
a `sequence_number`:

```text
response.created
response.in_progress
response.output_item.added
response.content_part.added
response.output_text.delta          (one per fragment)
response.output_text.done
response.content_part.done
response.output_item.done
response.completed
```

A tool call is an output item of its own. It streams as
`response.function_call_arguments.delta` events, followed by
`response.function_call_arguments.done`.

The last event carries the whole response: `output`, `usage` with the
cost, `lowrouter_metadata` and `lowrouter_unapplied`. It is one of:

| Event | Meaning |
|---|---|
| `response.completed` | The answer is whole. |
| `response.incomplete` | The answer was cut short. See `incomplete_details`. |
| `response.failed` | The request failed after the stream started. See `error`. Nothing follows this event. |

A stream that is cut off before the model finishes also ends with
`response.failed`, with the error code `server_error`. No output item
is closed and the event carries no `output` and no `usage`: what you
received as deltas is partial, and a tool call that was still streaming
must not be run. You are billed for the tokens generated before the
cut; the request is marked `interrupted` on the dashboard and in the
CSV export.
