Embeddings
POST /v1/embeddings turns text into vectors, in the OpenAI response
shape, so any client or RAG framework that speaks the OpenAI embeddings
API can point at LowRouter. It is billed, carbon-accounted and
catalogue-listed like chat.
curl https://api.lowrouter.ai/v1/embeddings \
-H "Authorization: Bearer $LOWROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "mistral/mistralai/mistral-embed-2312",
"input": "Vector databases index embeddings for similarity search."
}'With the OpenAI SDK, set the base URL and call embeddings.create as
usual:
from openai import OpenAI
client = OpenAI(base_url="https://api.lowrouter.ai/v1", api_key="YOUR_KEY")
r = client.embeddings.create(
model="mistral/mistralai/mistral-embed-2312",
input=["first passage", "second passage"],
)
print(len(r.data), len(r.data[0].embedding))Request
model- Required. An embedding model: an explicit
{provider}/{creator}/{model}ID, anauto/<creator>/<model>ID, or analias/<name>on your account. input- Required. One string, an array of strings, or an array of token IDs. An empty string or empty array is a
400. encoding_formatfloat(default) for a number array, orbase64. Forwarded as-is to providers that speak the OpenAI embeddings API; Bedrock and Gemini routes ignore it and always returnfloat, so checkdata[0].embeddingis a string before decoding base64.dimensions- Optional. Ask for a shorter vector; only models that support truncation honour it.
user- Optional end-user identifier, as on chat.
Which models. Every embedding model in the catalogue is listed by
GET /v1/models?modality=embedding; each entry’s modality is
embedding. The model browser’s Modality filter shows the same set.
An ID whose published modality is anything else is refused before any
upstream call with 400 invalid_request_error, naming the model and
pointing at that listing, and the reverse holds on /v1/chat/completions
for an embedding model. Neither is ever forwarded to fail upstream.
Routing. An explicit ID is served by that provider and region,
never substituted. An auto/ ID picks the route by the same policy as
chat (see routing); the response’s model names
the resolved four-segment route. Embeddings never fail over between
providers: two providers’ vectors for the same model are not
interchangeable in an index, so a failed request is returned as the
error it hit rather than silently re-embedded elsewhere.
Response
{
"object": "list",
"data": [
{"object": "embedding", "index": 0, "embedding": [0.0123, -0.0456, "…"]},
{"object": "embedding", "index": 1, "embedding": [0.0789, 0.0012, "…"]}
],
"model": "mistral/mistralai/mistral-embed-2312/global",
"usage": {
"prompt_tokens": 14,
"total_tokens": 14,
"cost": 0.0000014,
"currency": "EUR",
"remaining_balance": 9.4187
},
"lowrouter_metadata": {
"provider": "mistral",
"region": "global",
"energy_wh": 0.0000053,
"carbon_gco2e": 0.0000023,
"carbon_intensity_gco2_per_kwh": 436,
"estimation_methodology": "calculated_from_size; grid: GLOBAL-AVERAGE",
"routing_mode": "explicit",
"eu_sovereign": false,
"fallback_occurred": false
}
}Batch order. With an array input, data[i].index is the position
of the input it embeds, and data is returned in that order, so
data[i] always belongs to input[i].
Dimensions. The vector length is the model’s own (commonly 768,
1024, 2048 or 4096, per model); ask for less with dimensions on a
model that supports it. Vectors from different models, or from the
same model at different dimensions, are not comparable.
Usage and cost. usage.prompt_tokens is what the provider
counted; there are no completion tokens. cost is the amount debited
in currency (EUR), from LowRouter’s own catalogue price for the
model, and remaining_balance is your balance after it. The same
record is on the Activity page and in the
usage export.
Per-request metadata. lowrouter_metadata carries the same block
as chat (see per-request metadata):
the serving provider and region, routing mode, energy and carbon.
One difference in method: EcoLogits models few embedding
architectures, so the energy estimate usually comes from the model’s
parameter count and estimation_methodology says so
(calculated_from_size) instead of ecologits_formula. The estimate
covers the input tokens only, since an embedding produces none.
Pricing
Embedding models are priced per 1M input tokens, listed on each
model’s entry in GET /v1/models and on its page in the
model browser. Prices are in EUR, as everywhere else; see
pricing and FX for how non-EUR upstream prices are
converted.
Errors
The errors reference applies unchanged: 401 for a bad key,
402 when the balance is exhausted (checked before the request is
forwarded), 400 for a malformed body or a model that cannot embed,
404 for an ID that does not exist, and the 5xx family for an
upstream that fails. A 4xx from the gateway costs nothing.
