Input modalities
Text-generation models on LowRouter accept more than text. Images, documents and audio can be sent as content blocks in a message, using the same OpenAI-compatible request shape as everything else.
What varies is what each upstream provider accepts. The answer is not uniform, so this page states exactly where the limits fall.
The short version
Send binaries inline as base64 data URIs. That form works on every provider we route to, for every modality that provider supports.
Remote http(s) URLs are not portable. Some providers fetch them
server-side, some refuse them, and some fetch them and fail on their own
network policy. If you control the payload, inline it.
Images
{
"role": "user",
"content": [
{ "type": "text", "text": "What is in this image?" },
{
"type": "image_url",
"image_url": { "url": "data:image/png;base64,iVBORw0KGgo..." }
}
]
}Base64 data URIs work across every provider that serves a vision-capable model.
Remote URLs are per-provider. Where a provider fetches the URL itself, the fetch
happens from that provider’s network, not ours, so its success depends on their
egress and on whether your host serves them. Where a provider does not support
URL sources at all, you get a 400 naming that limitation rather than a silent
drop.
Give a remote URL a real file extension (.png, .jpg, .pdf). Some upstream
APIs require an explicit media type alongside the URL, and the extension is what
we derive it from. What happens to a URL we cannot type depends on the route:
- Images on Anthropic, and on the providers we reach through their OpenAI-compatible API, are forwarded as you sent them. The provider fetches the image and works out its type itself, so an extensionless CDN or signed URL works (Anthropic and OpenAI both accept one).
- Documents on Anthropic, and any file on Gemini, are refused with a message saying so: those APIs need the media type with the URL, and we will not guess one.
- Claude on Vertex AI and on Bedrock, and Bedrock’s other models, take no
remote URLs: those APIs accept base64 only. LowRouter fetches the image for
you and sends it inline, within the limits Anthropic sets there:
httpsURLs only, on a public host, up to 5 MB each, a JPEG, PNG, GIF or WebP, and at most 100 images and 32 MB in all per request; the extension is not needed. An image you already sent as a data URI is passed on untouched. Documents and audio on those routes still have to be data URIs.
So an untyped URL can succeed on one route and be refused on another; the extension is what makes it portable.
Documents
Two spellings are accepted. The OpenAI file block:
{
"type": "file",
"file": {
"file_data": "data:application/pdf;base64,JVBERi0xLjcK...",
"filename": "report.pdf"
}
}and the Anthropic document block, which also works on /v1/messages:
{
"type": "document",
"source": {
"type": "base64",
"media_type": "application/pdf",
"data": "JVBERi0xLjcK..."
}
}PDF is the widely-supported case, and the only document type some providers
accept at all. Plain text, Markdown and CSV reach fewer; office formats
(.docx, .xlsx) reach fewer still. On the providers we reach through their
native SDKs, a media type the provider cannot express is refused with a 400
naming both the type and the provider. Documents are otherwise always
forwarded, and never refused locally on the strength of the route’s published
input_modalities, because a published list can be narrower than what the
route accepts (see Discovering what a route accepts).
There is one exception to the rule above. On the
providers we reach through their OpenAI-compatible API, the content block is
forwarded as you sent it. Most of those providers reject a block they do not
understand, which gives you a clear 400. A small number instead return 200
and answer as though no document were attached. If a reply looks like the model
never saw your file, ask it to quote the document back before trusting the
answer, and prefer a provider that fails loudly for anything that matters.
Audio
{
"type": "input_audio",
"input_audio": { "data": "UklGRi...", "format": "wav" }
}format is the bare container name (wav, mp3, flac, ogg, …), not a
media type, and data is raw base64 rather than a data URI. This block follows
the OpenAI spelling exactly.
Audio reaches fewer providers than documents do. Some providers have no audio
input block in their API at all; a request for one of those returns a 400
naming the provider, so you can route the request elsewhere.
Tool results
A tool result can carry an image or a document too: image_url / file
parts in a tool-role message on this surface, image / document blocks
inside a tool_result on /v1/messages. This is how Claude Code delivers
every screenshot and PDF its Read tool opens. Anthropic and Bedrock routes
take those blocks inside the tool result natively. Every OpenAI-compatible
provider and Gemini have a text-only tool result, so on those routes the
image or document is moved into a user message placed immediately after the
result, where the model still sees it, and the response says so in
lowrouter_unapplied (tool_result, unsupported_on_this_path). See
request parameters.
When a provider cannot accept a modality
You get a 400 with invalid_request_error, naming the modality and the
provider, in two situations.
Before the request leaves. When the route’s provider publishes
input_modalities for the model and that list excludes an image, audio or
video block you sent, the request is refused locally:
{
"error": {
"type": "invalid_request_error",
"message": "mistral does not accept image input on mistral/mistralai/ministral-3b-2410/global (published input modalities: text); send it to a route whose model lists image among its input modalities"
}
}This check acts only on a statement the provider made. A route with no
published list is forwarded and the provider answers for itself, and
documents are forwarded regardless (a published list can be narrower than
what a route accepts; see below). For an auto/ id, routes whose provider
states the modality away rank last and are skipped by failover, so the
request is refused only when no capable or undocumented route exists.
At the provider. When the request is forwarded and the native API has no block for the modality at all, the adapter refuses it by name:
{
"error": {
"type": "invalid_request_error",
"message": "anthropic does not accept audio input; send audio to a provider whose model lists audio among its input modalities"
}
}On the OpenAI-compatible providers the block reaches the upstream as sent,
and a rejection is the upstream’s own message relayed back, which can read
as a complaint about your request’s shape rather than about the modality.
The local check above exists to catch the common case before that happens.
Documents skip that check, so when a provider rejects a request carrying one
and its published list leaves document out, the 400 names that as the
likely cause and quotes the provider after it:
{
"error": {
"type": "invalid_request_error",
"message": "Provider mistral rejected the request: mistral/mistralai/mistral-large-2512/global may not accept document input (published input modalities: image, text); the provider said: …"
}
}Both refusals are deliberate. A content block that was quietly dropped would produce a confident answer about a document the model never saw, and you would have no way to tell that from a real answer, so the request fails instead.
Billing
Multimodal input is billed from the token counts the provider itself reports,
not from an estimate of ours. Providers count images very differently, and the
spread is far larger than rounding: the same 640×200 PNG measured at 208 to 550
prompt tokens on Mistral, Scaleway and Bedrock and at 14,188 on
openai/openai/gpt-4o-mini (roughly 50×), because that model’s image
tokeniser scales small images up before counting them. Your cost for an
identical request therefore varies by route by more than the per-token price
alone suggests; check usage.prompt_tokens on the first response before
committing a batch of images to a route. The per-request cost is on the
response (usage.cost) and on /v1/generation/{id} (cost.total_cost).
Discovering what a route accepts
/v1/models reports input modalities in the capabilities block, so you can
check before you send:
"capabilities": {
"streaming": true,
"function_calling": true,
"input_modalities": ["text", "image", "document"]
}Values come from a closed set: text, image, audio, video, document.
document means the document modality (PDF, DOCX, PPTX and similar), not PDF
specifically. If you cross-check against models.dev or LiteLLM, note they spell
the same concept pdf.
To list only what accepts a given modality, filter the catalogue:
GET /v1/models?input_modality=documentAbsent means unknown, not “text only”. The key is omitted entirely when the
provider publishes no modality statement. An omitted input_modalities is not a
claim that the model rejects your PDF; it means nobody told us either way, and
the reliable test is still to send the request. We omit the field rather than
infer it from model names and be wrong about hundreds of models. For the
same reason, ?input_modality= excludes models with no published statement:
it answers “which routes are documented as accepting this”, not “which routes
might”.
A stated list can also be narrower than what a route accepts. The list
reflects what the provider documents per model, and some providers offer a
capability at the platform level without stating it on the model itself. OpenAI
is the clearest example: its models list text and image, yet the API accepts
PDFs and bills them as extracted text plus a rendered image of each page. So a
missing document means “not stated for this model”, never “will be rejected”.
Filtering on ?input_modality=document therefore returns a conservative subset:
every route in it is documented for documents, but routes outside it may still
work.
It is per provider, not per model. The same weights can accept different
inputs depending on who serves them. Of the models served by more than one
provider, 39% differ, most often about PDFs. moonshotai/kimi-k3 is served by
four providers and only one of them takes a PDF. So read the modalities on the
regions entry you actually intend to call, not just the model-level value,
which is the union across every route.
When you send a document to an auto/ id, we prefer a route whose provider
publishes support for it. That is a preference and not a guarantee: routes whose
providers publish nothing stay eligible, because refusing to route over a field
a provider never documented would fail requests that would have worked.
What we don’t claim
- We do not convert between modalities. We do not transcribe audio, rasterise PDFs, or re-encode images to fit a provider that refuses the format you sent. What you send is what the provider receives.
- We do not guarantee remote URL fetches. They happen on the provider’s
network under the provider’s policy; we can report the failure but cannot fix
it. The one fetch LowRouter makes itself, for images on Claude via Vertex or
Bedrock and on Bedrock’s other models, has the limits given above, and a
failure there comes back as a
400naming the URL. - Coverage changes. Providers add and remove modality support on their own
schedule. The
400you get from a live request is always more current than this page.
