
# Embeddings

`POST /v1/embeddings` turns text into vectors, in the OpenAI response
shape, so any client or RAG framework that speaks the OpenAI embeddings
API can point at LowRouter. It is billed, carbon-accounted and
catalogue-listed like chat.

```bash
curl https://api.lowrouter.ai/v1/embeddings \
  -H "Authorization: Bearer $LOWROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "mistral/mistralai/mistral-embed-2312",
    "input": "Vector databases index embeddings for similarity search."
  }'
```

With the OpenAI SDK, set the base URL and call `embeddings.create` as
usual:

```python
from openai import OpenAI

client = OpenAI(base_url="https://api.lowrouter.ai/v1", api_key="YOUR_KEY")
r = client.embeddings.create(
    model="mistral/mistralai/mistral-embed-2312",
    input=["first passage", "second passage"],
)
print(len(r.data), len(r.data[0].embedding))
```

## Request

| Field | Meaning |
|-------|---------|
| `model` | Required. An embedding model: an explicit `{provider}/{creator}/{model}` ID, an `auto/<creator>/<model>` ID, or an `alias/<name>` on your account. |
| `input` | Required. One string, an array of strings, or an array of token IDs. An empty string or empty array is a `400`. |
| `encoding_format` | `float` (default) for a number array, or `base64`. Forwarded as-is to providers that speak the OpenAI embeddings API; Bedrock and Gemini routes ignore it and always return `float`, so check `data[0].embedding` is a string before decoding base64. |
| `dimensions` | Optional. Ask for a shorter vector; only models that support truncation honour it. |
| `user` | Optional end-user identifier, as on chat. |

**Which models.** Every embedding model in the catalogue is listed by
`GET /v1/models?modality=embedding`; each entry's `modality` is
`embedding`. The model browser's *Modality* filter shows the same set.
An ID whose published modality is anything else is refused before any
upstream call with `400 invalid_request_error`, naming the model and
pointing at that listing, and the reverse holds on `/v1/chat/completions`
for an embedding model. Neither is ever forwarded to fail upstream.

**Routing.** An explicit ID is served by that provider and region,
never substituted. An `auto/` ID picks the route by the same policy as
chat (see [routing](../models/routing)); the response's `model` names
the resolved four-segment route. Embeddings never fail over between
providers: two providers' vectors for the same model are not
interchangeable in an index, so a failed request is returned as the
error it hit rather than silently re-embedded elsewhere.

## Response

```json
{
  "object": "list",
  "data": [
    {"object": "embedding", "index": 0, "embedding": [0.0123, -0.0456, "…"]},
    {"object": "embedding", "index": 1, "embedding": [0.0789, 0.0012, "…"]}
  ],
  "model": "mistral/mistralai/mistral-embed-2312/global",
  "usage": {
    "prompt_tokens": 14,
    "total_tokens": 14,
    "cost": 0.0000014,
    "currency": "EUR",
    "remaining_balance": 9.4187
  },
  "lowrouter_metadata": {
    "provider": "mistral",
    "region": "global",
    "energy_wh": 0.0000053,
    "carbon_gco2e": 0.0000023,
    "carbon_intensity_gco2_per_kwh": 436,
    "estimation_methodology": "calculated_from_size; grid: GLOBAL-AVERAGE",
    "routing_mode": "explicit",
    "eu_sovereign": false,
    "fallback_occurred": false
  }
}
```

**Batch order.** With an array `input`, `data[i].index` is the position
of the input it embeds, and `data` is returned in that order, so
`data[i]` always belongs to `input[i]`.

**Dimensions.** The vector length is the model's own (commonly 768,
1024, 2048 or 4096, per model); ask for less with `dimensions` on a
model that supports it. Vectors from different models, or from the
same model at different dimensions, are not comparable.

**Usage and cost.** `usage.prompt_tokens` is what the provider
counted; there are no completion tokens. `cost` is the amount debited
in `currency` (EUR), from LowRouter's own catalogue price for the
model, and `remaining_balance` is your balance after it. The same
record is on the [Activity page](/app/activity) and in the
[usage export](usage-accounting#scripted-export).

**Per-request metadata.** `lowrouter_metadata` carries the same block
as chat (see [per-request metadata](../models/per-request-metadata)):
the serving provider and region, routing mode, energy and carbon.
One difference in method: EcoLogits models few embedding
architectures, so the energy estimate usually comes from the model's
parameter count and `estimation_methodology` says so
(`calculated_from_size`) instead of `ecologits_formula`. The estimate
covers the input tokens only, since an embedding produces none.

## Pricing

Embedding models are priced per 1M input tokens, listed on each
model's entry in `GET /v1/models` and on its page in the
[model browser](/models). Prices are in EUR, as everywhere else; see
[pricing and FX](pricing-and-fx) for how non-EUR upstream prices are
converted.

## Errors

The [errors reference](errors) applies unchanged: `401` for a bad key,
`402` when the balance is exhausted (checked before the request is
forwarded), `400` for a malformed body or a model that cannot embed,
`404` for an ID that does not exist, and the `5xx` family for an
upstream that fails. A `4xx` from the gateway costs nothing.
