Embeddings

POST /v1/embeddings turns text into vectors, in the OpenAI response shape, so any client or RAG framework that speaks the OpenAI embeddings API can point at LowRouter. It is billed, carbon-accounted and catalogue-listed like chat.

Bash
curl https://api.lowrouter.ai/v1/embeddings \
  -H "Authorization: Bearer $LOWROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "mistral/mistralai/mistral-embed-2312",
    "input": "Vector databases index embeddings for similarity search."
  }'

With the OpenAI SDK, set the base URL and call embeddings.create as usual:

Python
from openai import OpenAI

client = OpenAI(base_url="https://api.lowrouter.ai/v1", api_key="YOUR_KEY")
r = client.embeddings.create(
    model="mistral/mistralai/mistral-embed-2312",
    input=["first passage", "second passage"],
)
print(len(r.data), len(r.data[0].embedding))

Request

model
Required. An embedding model: an explicit {provider}/{creator}/{model} ID, an auto/<creator>/<model> ID, or an alias/<name> on your account.
input
Required. One string, an array of strings, or an array of token IDs. An empty string or empty array is a 400.
encoding_format
float (default) for a number array, or base64. Forwarded as-is to providers that speak the OpenAI embeddings API; Bedrock and Gemini routes ignore it and always return float, so check data[0].embedding is a string before decoding base64.
dimensions
Optional. Ask for a shorter vector; only models that support truncation honour it.
user
Optional end-user identifier, as on chat.

Which models. Every embedding model in the catalogue is listed by GET /v1/models?modality=embedding; each entry’s modality is embedding. The model browser’s Modality filter shows the same set. An ID whose published modality is anything else is refused before any upstream call with 400 invalid_request_error, naming the model and pointing at that listing, and the reverse holds on /v1/chat/completions for an embedding model. Neither is ever forwarded to fail upstream.

Routing. An explicit ID is served by that provider and region, never substituted. An auto/ ID picks the route by the same policy as chat (see routing); the response’s model names the resolved four-segment route. Embeddings never fail over between providers: two providers’ vectors for the same model are not interchangeable in an index, so a failed request is returned as the error it hit rather than silently re-embedded elsewhere.

Response

JSON
{
  "object": "list",
  "data": [
    {"object": "embedding", "index": 0, "embedding": [0.0123, -0.0456, "…"]},
    {"object": "embedding", "index": 1, "embedding": [0.0789, 0.0012, "…"]}
  ],
  "model": "mistral/mistralai/mistral-embed-2312/global",
  "usage": {
    "prompt_tokens": 14,
    "total_tokens": 14,
    "cost": 0.0000014,
    "currency": "EUR",
    "remaining_balance": 9.4187
  },
  "lowrouter_metadata": {
    "provider": "mistral",
    "region": "global",
    "energy_wh": 0.0000053,
    "carbon_gco2e": 0.0000023,
    "carbon_intensity_gco2_per_kwh": 436,
    "estimation_methodology": "calculated_from_size; grid: GLOBAL-AVERAGE",
    "routing_mode": "explicit",
    "eu_sovereign": false,
    "fallback_occurred": false
  }
}

Batch order. With an array input, data[i].index is the position of the input it embeds, and data is returned in that order, so data[i] always belongs to input[i].

Dimensions. The vector length is the model’s own (commonly 768, 1024, 2048 or 4096, per model); ask for less with dimensions on a model that supports it. Vectors from different models, or from the same model at different dimensions, are not comparable.

Usage and cost. usage.prompt_tokens is what the provider counted; there are no completion tokens. cost is the amount debited in currency (EUR), from LowRouter’s own catalogue price for the model, and remaining_balance is your balance after it. The same record is on the Activity page and in the usage export.

Per-request metadata. lowrouter_metadata carries the same block as chat (see per-request metadata): the serving provider and region, routing mode, energy and carbon. One difference in method: EcoLogits models few embedding architectures, so the energy estimate usually comes from the model’s parameter count and estimation_methodology says so (calculated_from_size) instead of ecologits_formula. The estimate covers the input tokens only, since an embedding produces none.

Pricing

Embedding models are priced per 1M input tokens, listed on each model’s entry in GET /v1/models and on its page in the model browser. Prices are in EUR, as everywhere else; see pricing and FX for how non-EUR upstream prices are converted.

Errors

The errors reference applies unchanged: 401 for a bad key, 402 when the balance is exhausted (checked before the request is forwarded), 400 for a malformed body or a model that cannot embed, 404 for an ID that does not exist, and the 5xx family for an upstream that fails. A 4xx from the gateway costs nothing.