Input modalities

Text-generation models on LowRouter accept more than text. Images, documents and audio can be sent as content blocks in a message, using the same OpenAI-compatible request shape as everything else.

What varies is what each upstream provider accepts. The answer is not uniform, so this page states exactly where the limits fall.

The short version

Send binaries inline as base64 data URIs. That form works on every provider we route to, for every modality that provider supports.

Remote http(s) URLs are not portable. Some providers fetch them server-side, some refuse them, and some fetch them and fail on their own network policy. If you control the payload, inline it.

Images

JSON
{
  "role": "user",
  "content": [
    { "type": "text", "text": "What is in this image?" },
    {
      "type": "image_url",
      "image_url": { "url": "data:image/png;base64,iVBORw0KGgo..." }
    }
  ]
}

Base64 data URIs work across every provider that serves a vision-capable model.

Remote URLs are per-provider. Where a provider fetches the URL itself, the fetch happens from that provider’s network, not ours, so its success depends on their egress and on whether your host serves them. Where a provider does not support URL sources at all, you get a 400 naming that limitation rather than a silent drop.

Give a remote URL a real file extension (.png, .jpg, .pdf). Some upstream APIs require an explicit media type alongside the URL, and the extension is what we derive it from. What happens to a URL we cannot type depends on the route:

  • Images on Anthropic, and on the providers we reach through their OpenAI-compatible API, are forwarded as you sent them. The provider fetches the image and works out its type itself, so an extensionless CDN or signed URL works (Anthropic and OpenAI both accept one).
  • Documents on Anthropic, and any file on Gemini, are refused with a message saying so: those APIs need the media type with the URL, and we will not guess one.
  • Claude on Vertex AI and on Bedrock, and Bedrock’s other models, take no remote URLs: those APIs accept base64 only. LowRouter fetches the image for you and sends it inline, within the limits Anthropic sets there: https URLs only, on a public host, up to 5 MB each, a JPEG, PNG, GIF or WebP, and at most 100 images and 32 MB in all per request; the extension is not needed. An image you already sent as a data URI is passed on untouched. Documents and audio on those routes still have to be data URIs.

So an untyped URL can succeed on one route and be refused on another; the extension is what makes it portable.

Documents

Two spellings are accepted. The OpenAI file block:

JSON
{
  "type": "file",
  "file": {
    "file_data": "data:application/pdf;base64,JVBERi0xLjcK...",
    "filename": "report.pdf"
  }
}

and the Anthropic document block, which also works on /v1/messages:

JSON
{
  "type": "document",
  "source": {
    "type": "base64",
    "media_type": "application/pdf",
    "data": "JVBERi0xLjcK..."
  }
}

PDF is the widely-supported case, and the only document type some providers accept at all. Plain text, Markdown and CSV reach fewer; office formats (.docx, .xlsx) reach fewer still. On the providers we reach through their native SDKs, a media type the provider cannot express is refused with a 400 naming both the type and the provider. Documents are otherwise always forwarded, and never refused locally on the strength of the route’s published input_modalities, because a published list can be narrower than what the route accepts (see Discovering what a route accepts).

There is one exception to the rule above. On the providers we reach through their OpenAI-compatible API, the content block is forwarded as you sent it. Most of those providers reject a block they do not understand, which gives you a clear 400. A small number instead return 200 and answer as though no document were attached. If a reply looks like the model never saw your file, ask it to quote the document back before trusting the answer, and prefer a provider that fails loudly for anything that matters.

Audio

JSON
{
  "type": "input_audio",
  "input_audio": { "data": "UklGRi...", "format": "wav" }
}

format is the bare container name (wav, mp3, flac, ogg, …), not a media type, and data is raw base64 rather than a data URI. This block follows the OpenAI spelling exactly.

Audio reaches fewer providers than documents do. Some providers have no audio input block in their API at all; a request for one of those returns a 400 naming the provider, so you can route the request elsewhere.

Tool results

A tool result can carry an image or a document too: image_url / file parts in a tool-role message on this surface, image / document blocks inside a tool_result on /v1/messages. This is how Claude Code delivers every screenshot and PDF its Read tool opens. Anthropic and Bedrock routes take those blocks inside the tool result natively. Every OpenAI-compatible provider and Gemini have a text-only tool result, so on those routes the image or document is moved into a user message placed immediately after the result, where the model still sees it, and the response says so in lowrouter_unapplied (tool_result, unsupported_on_this_path). See request parameters.

When a provider cannot accept a modality

You get a 400 with invalid_request_error, naming the modality and the provider, in two situations.

Before the request leaves. When the route’s provider publishes input_modalities for the model and that list excludes an image, audio or video block you sent, the request is refused locally:

JSON
{
  "error": {
    "type": "invalid_request_error",
    "message": "mistral does not accept image input on mistral/mistralai/ministral-3b-2410/global (published input modalities: text); send it to a route whose model lists image among its input modalities"
  }
}

This check acts only on a statement the provider made. A route with no published list is forwarded and the provider answers for itself, and documents are forwarded regardless (a published list can be narrower than what a route accepts; see below). For an auto/ id, routes whose provider states the modality away rank last and are skipped by failover, so the request is refused only when no capable or undocumented route exists.

At the provider. When the request is forwarded and the native API has no block for the modality at all, the adapter refuses it by name:

JSON
{
  "error": {
    "type": "invalid_request_error",
    "message": "anthropic does not accept audio input; send audio to a provider whose model lists audio among its input modalities"
  }
}

On the OpenAI-compatible providers the block reaches the upstream as sent, and a rejection is the upstream’s own message relayed back, which can read as a complaint about your request’s shape rather than about the modality. The local check above exists to catch the common case before that happens. Documents skip that check, so when a provider rejects a request carrying one and its published list leaves document out, the 400 names that as the likely cause and quotes the provider after it:

JSON
{
  "error": {
    "type": "invalid_request_error",
    "message": "Provider mistral rejected the request: mistral/mistralai/mistral-large-2512/global may not accept document input (published input modalities: image, text); the provider said: …"
  }
}

Both refusals are deliberate. A content block that was quietly dropped would produce a confident answer about a document the model never saw, and you would have no way to tell that from a real answer, so the request fails instead.

Billing

Multimodal input is billed from the token counts the provider itself reports, not from an estimate of ours. Providers count images very differently, and the spread is far larger than rounding: the same 640×200 PNG measured at 208 to 550 prompt tokens on Mistral, Scaleway and Bedrock and at 14,188 on openai/openai/gpt-4o-mini (roughly 50×), because that model’s image tokeniser scales small images up before counting them. Your cost for an identical request therefore varies by route by more than the per-token price alone suggests; check usage.prompt_tokens on the first response before committing a batch of images to a route. The per-request cost is on the response (usage.cost) and on /v1/generation/{id} (cost.total_cost).

Discovering what a route accepts

/v1/models reports input modalities in the capabilities block, so you can check before you send:

JSON
"capabilities": {
  "streaming": true,
  "function_calling": true,
  "input_modalities": ["text", "image", "document"]
}

Values come from a closed set: text, image, audio, video, document. document means the document modality (PDF, DOCX, PPTX and similar), not PDF specifically. If you cross-check against models.dev or LiteLLM, note they spell the same concept pdf.

To list only what accepts a given modality, filter the catalogue:

Text
GET /v1/models?input_modality=document

Absent means unknown, not “text only”. The key is omitted entirely when the provider publishes no modality statement. An omitted input_modalities is not a claim that the model rejects your PDF; it means nobody told us either way, and the reliable test is still to send the request. We omit the field rather than infer it from model names and be wrong about hundreds of models. For the same reason, ?input_modality= excludes models with no published statement: it answers “which routes are documented as accepting this”, not “which routes might”.

A stated list can also be narrower than what a route accepts. The list reflects what the provider documents per model, and some providers offer a capability at the platform level without stating it on the model itself. OpenAI is the clearest example: its models list text and image, yet the API accepts PDFs and bills them as extracted text plus a rendered image of each page. So a missing document means “not stated for this model”, never “will be rejected”. Filtering on ?input_modality=document therefore returns a conservative subset: every route in it is documented for documents, but routes outside it may still work.

It is per provider, not per model. The same weights can accept different inputs depending on who serves them. Of the models served by more than one provider, 39% differ, most often about PDFs. moonshotai/kimi-k3 is served by four providers and only one of them takes a PDF. So read the modalities on the regions entry you actually intend to call, not just the model-level value, which is the union across every route.

When you send a document to an auto/ id, we prefer a route whose provider publishes support for it. That is a preference and not a guarantee: routes whose providers publish nothing stay eligible, because refusing to route over a field a provider never documented would fail requests that would have worked.

What we don’t claim

  • We do not convert between modalities. We do not transcribe audio, rasterise PDFs, or re-encode images to fit a provider that refuses the format you sent. What you send is what the provider receives.
  • We do not guarantee remote URL fetches. They happen on the provider’s network under the provider’s policy; we can report the failure but cannot fix it. The one fetch LowRouter makes itself, for images on Claude via Vertex or Bedrock and on Bedrock’s other models, has the limits given above, and a failure there comes back as a 400 naming the URL.
  • Coverage changes. Providers add and remove modality support on their own schedule. The 400 you get from a live request is always more current than this page.