Sustainability-first
“Sustainable AI” is a common marketing phrase. This page describes what it means inside LowRouter: what we measure, what we report, and what we have decided not to claim.
What we measure
For every inference request, we estimate two numbers:
- Energy per output token (Wh), derived from the model’s active parameter count using the EcoLogits methodology.
- Grid carbon intensity (gCO₂e/kWh) for the region the provider serves the request from, sourced from Ember Climate annual averages, refined by regional grid-operator data where available.
Their product gives gCO₂e per 1,000 tokens for the request. The exact formula and the confidence we attach to each estimate are in the methodology page.
These numbers are exposed:
- On every API response, in the
lowrouter_metadatablock. - On the dashboard, aggregated by day, model, provider, and region.
- On the public model browser, as a comparable estimate per model.
What we don’t measure (yet)
- Training emissions. We report inference only. Training is a separate, larger, and harder-to-attribute footprint, and folding it into per-request numbers is misleading.
- Hardware embodied carbon. GPU manufacturing has a real footprint; we don’t yet have a defensible per-token number for it.
- Real-time grid mix. We use annual averages by region. Live carbon-aware routing is a future feature, not a current one.
- Embedding workloads. Encoder-only models have a different compute profile and our formula does not yet model them well.
An honest number comes with its limits, so the sustainable-AI page repeats them next to every chart.
Constraints that fall out of this position
A few platform decisions follow from taking energy and carbon seriously:
- The site itself is on a tight transfer budget. Pages like this one ship under 100 KB total. The docs site renders server-side with no client framework. Heavy interactive widgets are not added without a reason.
- The default route is auto-mode rather than “the largest model.” If a smaller model can do the job for a fraction of the energy, we’d rather it be the default.
- Providers and regions with very high grid intensity are deprioritised in routing unless explicitly pinned by the caller. See provider routing.
- We are slower to add features than we’d like. Background jobs, dashboard widgets, and visualisations all carry a per-user energy cost, so we add them only when they are worth it.
What this is not
None of this is a pledge that any single request is green, a claim that the estimate is exact, or an assertion that running an LLM through LowRouter is meaningfully better for the planet than running it directly. It is a measured number, an open formula, and a default that prefers the smaller model and the cleaner grid when other constraints allow.
