
# Input modalities

Text-generation models on LowRouter accept more than text. Images, documents
and audio can be sent as content blocks in a message, using the same
OpenAI-compatible request shape as everything else.

What varies is what each upstream provider accepts. The answer is not uniform,
so this page states exactly where the limits fall.

## The short version

**Send binaries inline as base64 data URIs.** That form works on every provider
we route to, for every modality that provider supports.

**Remote `http(s)` URLs are not portable.** Some providers fetch them
server-side, some refuse them, and some fetch them and fail on their own
network policy. If you control the payload, inline it.

## Images

```json
{
  "role": "user",
  "content": [
    { "type": "text", "text": "What is in this image?" },
    {
      "type": "image_url",
      "image_url": { "url": "data:image/png;base64,iVBORw0KGgo..." }
    }
  ]
}
```

Base64 data URIs work across every provider that serves a vision-capable model.

Remote URLs are per-provider. Where a provider fetches the URL itself, the fetch
happens from that provider's network, not ours, so its success depends on their
egress and on whether your host serves them. Where a provider does not support
URL sources at all, you get a `400` naming that limitation rather than a silent
drop.

Give a remote URL a real file extension (`.png`, `.jpg`, `.pdf`). Some upstream
APIs require an explicit media type alongside the URL, and the extension is what
we derive it from. What happens to a URL we cannot type depends on the route:

- **Images on Anthropic, and on the providers we reach through their
  OpenAI-compatible API**, are forwarded as you sent them. The provider fetches
  the image and works out its type itself, so an extensionless CDN or signed
  URL works (Anthropic and OpenAI both accept one).
- **Documents on Anthropic, and any file on Gemini**, are refused with a message
  saying so: those APIs need the media type with the URL, and we will not guess
  one.
- **Claude on Vertex AI and on Bedrock, and Bedrock's other models**, take no
  remote URLs: those APIs accept base64 only. LowRouter fetches the image for
  you and sends it inline, within the limits Anthropic sets there: `https`
  URLs only, on a public host, up to 5 MB each, a JPEG, PNG, GIF or WebP, and
  at most 100 images and 32 MB in all per request; the extension is not
  needed. An image you already sent as a data URI is passed on untouched.
  Documents and audio on those routes still have to be data URIs.

So an untyped URL can succeed on one route and be refused on another; the
extension is what makes it portable.

## Documents

Two spellings are accepted. The OpenAI `file` block:

```json
{
  "type": "file",
  "file": {
    "file_data": "data:application/pdf;base64,JVBERi0xLjcK...",
    "filename": "report.pdf"
  }
}
```

and the Anthropic `document` block, which also works on `/v1/messages`:

```json
{
  "type": "document",
  "source": {
    "type": "base64",
    "media_type": "application/pdf",
    "data": "JVBERi0xLjcK..."
  }
}
```

PDF is the widely-supported case, and the only document type some providers
accept at all. Plain text, Markdown and CSV reach fewer; office formats
(`.docx`, `.xlsx`) reach fewer still. On the providers we reach through their
native SDKs, a media type the provider cannot express is refused with a `400`
naming both the type and the provider. Documents are otherwise always
forwarded, and never refused locally on the strength of the route's published
`input_modalities`, because a published list can be narrower than what the
route accepts (see [Discovering what a route accepts](#discovering-what-a-route-accepts)).

There is one exception to the rule above. On the
providers we reach through their OpenAI-compatible API, the content block is
forwarded as you sent it. Most of those providers reject a block they do not
understand, which gives you a clear `400`. A small number instead return `200`
and answer as though no document were attached. If a reply looks like the model
never saw your file, ask it to quote the document back before trusting the
answer, and prefer a provider that fails loudly for anything that matters.

## Audio

```json
{
  "type": "input_audio",
  "input_audio": { "data": "UklGRi...", "format": "wav" }
}
```

`format` is the bare container name (`wav`, `mp3`, `flac`, `ogg`, …), not a
media type, and `data` is raw base64 rather than a data URI. This block follows
the OpenAI spelling exactly.

Audio reaches fewer providers than documents do. Some providers have no audio
input block in their API at all; a request for one of those returns a `400`
naming the provider, so you can route the request elsewhere.

## Tool results

A tool result can carry an image or a document too: `image_url` / `file`
parts in a `tool`-role message on this surface, `image` / `document` blocks
inside a `tool_result` on `/v1/messages`. This is how Claude Code delivers
every screenshot and PDF its `Read` tool opens. Anthropic and Bedrock routes
take those blocks inside the tool result natively. Every OpenAI-compatible
provider and Gemini have a text-only tool result, so on those routes the
image or document is moved into a user message placed immediately after the
result, where the model still sees it, and the response says so in
`lowrouter_unapplied` (`tool_result`, `unsupported_on_this_path`). See
[request parameters](request-parameters#tool_result-with-an-image-or-document).

## When a provider cannot accept a modality

You get a `400` with `invalid_request_error`, naming the modality and the
provider, in two situations.

**Before the request leaves.** When the route's provider publishes
`input_modalities` for the model and that list excludes an image, audio or
video block you sent, the request is refused locally:

```json
{
  "error": {
    "type": "invalid_request_error",
    "message": "mistral does not accept image input on mistral/mistralai/ministral-3b-2410/global (published input modalities: text); send it to a route whose model lists image among its input modalities"
  }
}
```

This check acts only on a statement the provider made. A route with no
published list is forwarded and the provider answers for itself, and
documents are forwarded regardless (a published list can be narrower than
what a route accepts; see below). For an `auto/` id, routes whose provider
states the modality away rank last and are skipped by failover, so the
request is refused only when no capable or undocumented route exists.

**At the provider.** When the request is forwarded and the native API has
no block for the modality at all, the adapter refuses it by name:

```json
{
  "error": {
    "type": "invalid_request_error",
    "message": "anthropic does not accept audio input; send audio to a provider whose model lists audio among its input modalities"
  }
}
```

On the OpenAI-compatible providers the block reaches the upstream as sent,
and a rejection is the upstream's own message relayed back, which can read
as a complaint about your request's shape rather than about the modality.
The local check above exists to catch the common case before that happens.
Documents skip that check, so when a provider rejects a request carrying one
and its published list leaves `document` out, the `400` names that as the
likely cause and quotes the provider after it:

```json
{
  "error": {
    "type": "invalid_request_error",
    "message": "Provider mistral rejected the request: mistral/mistralai/mistral-large-2512/global may not accept document input (published input modalities: image, text); the provider said: …"
  }
}
```

Both refusals are deliberate. A content block that was quietly dropped would
produce a confident answer about a document the model never saw, and you would
have no way to tell that from a real answer, so the request fails instead.

## Billing

Multimodal input is billed from the token counts the provider itself reports,
not from an estimate of ours. Providers count images very differently, and the
spread is far larger than rounding: the same 640×200 PNG measured at 208 to 550
prompt tokens on Mistral, Scaleway and Bedrock and at **14,188 on
`openai/openai/gpt-4o-mini`** (roughly 50×), because that model's image
tokeniser scales small images up before counting them. Your cost for an
identical request therefore varies by route by more than the per-token price
alone suggests; check `usage.prompt_tokens` on the first response before
committing a batch of images to a route. The per-request cost is on the
response (`usage.cost`) and on `/v1/generation/{id}` (`cost.total_cost`).

## Discovering what a route accepts

`/v1/models` reports input modalities in the `capabilities` block, so you can
check before you send:

```json
"capabilities": {
  "streaming": true,
  "function_calling": true,
  "input_modalities": ["text", "image", "document"]
}
```

Values come from a closed set: `text`, `image`, `audio`, `video`, `document`.
`document` means the document modality (PDF, DOCX, PPTX and similar), not PDF
specifically. If you cross-check against models.dev or LiteLLM, note they spell
the same concept `pdf`.

To list only what accepts a given modality, filter the catalogue:

```
GET /v1/models?input_modality=document
```

**Absent means unknown, not "text only".** The key is omitted entirely when the
provider publishes no modality statement. An omitted `input_modalities` is not a
claim that the model rejects your PDF; it means nobody told us either way, and
the reliable test is still to send the request. We omit the field rather than
infer it from model names and be wrong about hundreds of models. For the
same reason, `?input_modality=` **excludes** models with no published statement:
it answers "which routes are documented as accepting this", not "which routes
might".

**A stated list can also be narrower than what a route accepts.** The list
reflects what the provider documents per model, and some providers offer a
capability at the platform level without stating it on the model itself. OpenAI
is the clearest example: its models list `text` and `image`, yet the API accepts
PDFs and bills them as extracted text plus a rendered image of each page. So a
missing `document` means "not stated for this model", never "will be rejected".
Filtering on `?input_modality=document` therefore returns a conservative subset:
every route in it is documented for documents, but routes outside it may still
work.

**It is per provider, not per model.** The same weights can accept different
inputs depending on who serves them. Of the models served by more than one
provider, 39% differ, most often about PDFs. `moonshotai/kimi-k3` is served by
four providers and only one of them takes a PDF. So read the modalities on the
`regions` entry you actually intend to call, not just the model-level value,
which is the union across every route.

When you send a document to an `auto/` id, we prefer a route whose provider
publishes support for it. That is a preference and not a guarantee: routes whose
providers publish nothing stay eligible, because refusing to route over a field
a provider never documented would fail requests that would have worked.

## What we don't claim

- **We do not convert between modalities.** We do not transcribe audio, rasterise
  PDFs, or re-encode images to fit a provider that refuses the format you sent.
  What you send is what the provider receives.
- **We do not guarantee remote URL fetches.** They happen on the provider's
  network under the provider's policy; we can report the failure but cannot fix
  it. The one fetch LowRouter makes itself, for images on Claude via Vertex or
  Bedrock and on Bedrock's other models, has the limits given above, and a
  failure there comes back as a `400` naming the URL.
- **Coverage changes.** Providers add and remove modality support on their own
  schedule. The `400` you get from a live request is always more current than
  this page.
