Hermes Agent
Hermes Agent is Nous Research’s open-source autonomous agent. It keeps memory and skills across sessions and can run unattended. Any OpenAI-compatible endpoint can serve as its model through a custom provider.
Verified against Hermes Agent v0.21.5 (release v2026.9.24). The
config below is re-run against that release every week, including a
tool call round trip.
Install
Hermes ships through its own installer. It does not build as a regular Python package:
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bashConfigure
Point the model section of ~/.hermes/config.yaml at LowRouter:
model:
provider: custom
base_url: https://api.lowrouter.ai/v1
key_env: LOWROUTER_API_KEY
default: auto/mistralai/mistral-large-2512
context_length: 256000key_env names the environment variable holding the key, so the key
itself stays out of the file. Export it before starting Hermes:
export LOWROUTER_API_KEY=sk-lr-...hermes model offers the same thing interactively under “Custom
endpoint”. Then run a task:
hermes -z "Summarise what this repository does"-z runs one prompt, prints only the final answer and approves
every tool call without asking. hermes alone opens the interactive
session.
Three settings people get wrong:
base_urlends in/v1. Hermes appends only/chat/completions; the bare origin gives a404, which Hermes reports as the provider being “temporarily unavailable” after three retries.- Set
context_lengthto the model’s real context window, which the model browser shows. Without it Hermes guesses, and it plans context compression around that guess. It refuses to start with a model under 64k, because its system prompt and tool schemas need that room. - Model IDs contain slashes. To switch models, change
default:in the config (andcontext_lengthwith it) rather than typing the ID into/modelinside a session.
Picking a model
Hermes does everything through tool calls, so pick a model with the function-calling tag. The model browser filtered to them lists every one, and the model needs at least 64k of context.
Each request carries Hermes’ system prompt and about 25 tool
schemas. In our runs the first request of a task was already around
25,000 input tokens, before any file contents.
Input price matters more than output price here. The example uses
auto/mistralai/mistral-large-2512, served from the EU (256k
context). Models we have run through Hermes’ tool loop:
auto/qwen/qwen3-coder-30b-a3b-instruct, a coding model served from the EU (128k context).auto/anthropic/claude-haiku-4.5, a cheaper Anthropic model (200k context).
Recommended setup
- Use a dedicated key. Hermes runs unattended from its scheduler and messaging gateways, so set the turn cap (below): until per-key spend limits ship, the cap and revoking the key are what end a runaway loop.
- Cap the tool loop.
agent.max_turnsin the config is unlimited by default; set a number for unattended runs. - Narrow the tools with
-t(--toolsets) when a task only needs a few; fewer tool schemas also means fewer input tokens per request.
Troubleshooting
- “custom didn’t answer after 3 attempts” with
HTTP 404:base_urlis missing/v1. 401on the first request:LOWROUTER_API_KEYis not exported in the shell that runs Hermes, orkey_envnames a different variable.- Startup refuses the model for its context size: the model has
under 64k of context, or
context_lengthis set below that.
