Hermes Agent: hardware requirements, local model setup, and context planning
Official-source implementation guide for Hermes Agent. Hermes Agent is a self-hostable AI agent framework with terminal tooling, provider routing, and local-model options including Ollama.
What it is
Hermes Agent describes itself as a self-improving AI agent built by Nous Research, with long-running sessions, tool use, and messaging interfaces.
Its provider system supports hosted APIs and self-hosted endpoints, so you can switch model backends without rewriting agent logic.
Official runtime requirements
- Installer-managed runtime: Current first-party installations run on Python 3.14; installer-managed toolchain includes Python, Node.js, npm, ripgrep, and FFmpeg. (Hermes Installation docs)
- Source install prerequisites: POSIX source install requires Git, curl, tar, and SHA-256 utilities. (Hermes Installation docs)
- Local Ollama guide baseline: Guide table lists 8 GB RAM minimum (3B models), 32+ GB recommended (27B+), 5 GB free storage minimum, and optional NVIDIA GPU with 8+ GB VRAM for faster inference. (Hermes local Ollama setup guide)
Local model connection snippets (official docs)
Ollama
curl -fsSL https://ollama.com/install.sh | sh ollama pull gemma4:31b hermes setup # provider: Custom Endpoint # base URL: http://localhost:11434/v1
Source: Hermes local Ollama setup guide
Ollama
cat > ~/.hermes/cache/scratch/Modelfile << 'EOF' FROM gemma4:31b PARAMETER num_ctx 64000 EOF ollama create gemma4-64k -f ~/.hermes/cache/scratch/Modelfile
From Hermes Ollama guide context-window section.
Source: Hermes local Ollama setup guide
vLLM
model: default: your-model-name provider: custom base_url: http://localhost:8000/v1 api_key: your-key-or-leave-empty-for-local
Hermes docs describe self-hosted endpoints including vLLM through the Custom Endpoint flow.
LM Studio
model: default: your-model-name provider: custom base_url: http://localhost:1234/v1 api_key: your-key-or-leave-empty-for-local
Hermes providers page lists LM Studio and custom endpoint configuration patterns.
llama.cpp
No explicit llama.cpp provider setup snippet appears in the provided Hermes source links.
Kept as a documented gap instead of inventing a command.
Why long context and tool-calling models matter
- Hermes local Ollama guide explicitly says Hermes requires at least 64,000-token context for agentic work with tools.
- Hermes local Ollama guide also states models without tool-call support can chat but cannot perform agent actions such as file edits and command execution.
Prefilled calculator
The fit matrix below is an LLMRAM estimate at Q4_K_M and batch=1 for planning. It is not an official vendor minimum requirement.
VRAM / RAM calculator
Pick a model, quantization, context length, batch size, and target hardware. Results apply to dedicated GPU VRAM, Apple unified memory, and CPU/RAM offload planning.
Custom Hugging Face model id / params (optional)
If repo id is unknown, calculator uses your custom params and labels results as estimated.
| Quant | Weights | KV | Overhead | Total |
|---|---|---|---|---|
| FP16 / BF16 | 50.29 GiB | 16 GiB | 1.78 GiB | 68.07 GiB |
| INT8 / Q8_0 | 25.77 GiB | 16 GiB | 1.78 GiB | 43.55 GiB |
| Q6_K | 19.96 GiB | 16 GiB | 1.78 GiB | 37.74 GiB |
| Q5_K_M | 17.13 GiB | 16 GiB | 1.78 GiB | 34.91 GiB |
| Q4_K_M | 14.46 GiB | 16 GiB | 1.78 GiB | 32.24 GiB |
| Q3_K_M / Q3 | 11 GiB | 16 GiB | 1.78 GiB | 28.78 GiB |
| Q2_K / Q2 | 7.86 GiB | 16 GiB | 1.78 GiB | 25.64 GiB |
| Hardware | Memory | Verdict | Est. tok/s |
|---|---|---|---|
| NVIDIA GeForce RTX 4090 24GB | 24 GiB | Does not fit | n/a (Does not fit) |
| NVIDIA GeForce RTX 5090 32GB | 32 GiB | Does not fit | n/a (Does not fit) |
| NVIDIA GeForce RTX 4080 SUPER 16GB | 16 GiB | Does not fit | n/a (Does not fit) |
| NVIDIA GeForce RTX 4070 Ti SUPER 16GB | 16 GiB | Does not fit | n/a (Does not fit) |
| AMD Radeon RX 7900 XTX 24GB | 24 GiB | Does not fit | n/a (Does not fit) |
| NVIDIA A100 80GB PCIe | 80 GiB | Fits | 84.75 |
| Apple Silicon M3 Max (128GB unified memory) | 128 GiB | Fits | 14.94 |
| Apple Silicon M2 Ultra (192GB unified memory) | 192 GiB | Fits | 29.89 |
| 2× NVIDIA GeForce RTX 4090 (aggregate) | 48 GiB | Fits | 79.45 |
Formula and assumptions
- Weight memory uses resident params × effective_bits / 8 (total params when available, otherwise active/fallback estimate).
- KV cache = 2 × layers × kv_heads × head_dim × kv_bytes × context × batch.
- Runtime overhead = 1.3 GiB base + 3% of KV-cache memory.
- Tokens/sec estimate uses active_params for decode bandwidth and should be treated as a rough directional number.
- Apple Silicon fit checks use an estimated usable unified-memory budget (about 75% by default, about 68% on 16GB systems) aligned with macOS recommendedMaxWorkingSetSize behavior.
- Recommended quantization keeps at least 10% memory headroom relative to usable memory budget.
- CPU/RAM offload (for example llama.cpp partial offload with lower GPU layer count) can reduce VRAM needs at the cost of speed.
- When config values are unavailable, fallback defaults are used and flagged as estimated on the page.
Local model fit table by memory tier (LLMRAM estimate)
The fit matrix below is an LLMRAM estimate at Q4_K_M and batch=1 for planning. It is not an official vendor minimum requirement.
32K context
GPU tiers
| Model profile | Estimated total memory | 8 GB | 12 GB | 16 GB | 24 GB | 32 GB | 48 GB |
|---|---|---|---|---|---|---|---|
| Mistral 7B Instruct v0.3 | 9.17 GiB | ✕ No fit | ✓ Fits | ✓ Fits | ✓ Fits | ✓ Fits | ✓ Fits |
| Qwen 3.8 27B | 24 GiB | ✕ No fit | ✕ No fit | ✕ No fit | ✓ Fits | ✓ Fits | ✓ Fits |
| GLM 5.3 (Flash) | 195.84 GiB | ✕ No fit | ✕ No fit | ✕ No fit | ✕ No fit | ✕ No fit | ✕ No fit |
Mac unified memory tiers
| Model profile | Estimated total memory | 16 GB | 36 GB | 64 GB | 128 GB |
|---|---|---|---|---|---|
| Mistral 7B Instruct v0.3 | 9.17 GiB | ✓ Fits | ✓ Fits | ✓ Fits | ✓ Fits |
| Qwen 3.8 27B | 24 GiB | ✕ No fit | ✓ Fits | ✓ Fits | ✓ Fits |
| GLM 5.3 (Flash) | 195.84 GiB | ✕ No fit | ✕ No fit | ✕ No fit | ✕ No fit |
64K context
GPU tiers
| Model profile | Estimated total memory | 8 GB | 12 GB | 16 GB | 24 GB | 32 GB | 48 GB |
|---|---|---|---|---|---|---|---|
| Mistral 7B Instruct v0.3 | 13.29 GiB | ✕ No fit | ✕ No fit | ✓ Fits | ✓ Fits | ✓ Fits | ✓ Fits |
| Qwen 3.8 27B | 32.24 GiB | ✕ No fit | ✕ No fit | ✕ No fit | ✕ No fit | ✕ No fit | ✓ Fits |
| GLM 5.3 (Flash) | 219.01 GiB | ✕ No fit | ✕ No fit | ✕ No fit | ✕ No fit | ✕ No fit | ✕ No fit |
Mac unified memory tiers
| Model profile | Estimated total memory | 16 GB | 36 GB | 64 GB | 128 GB |
|---|---|---|---|---|---|
| Mistral 7B Instruct v0.3 | 13.29 GiB | ✕ No fit | ✓ Fits | ✓ Fits | ✓ Fits |
| Qwen 3.8 27B | 32.24 GiB | ✕ No fit | ✕ No fit | ✓ Fits | ✓ Fits |
| GLM 5.3 (Flash) | 219.01 GiB | ✕ No fit | ✕ No fit | ✕ No fit | ✕ No fit |
FAQ
Can Hermes run fully local with no API key?
Yes, the Hermes Ollama guide documents a local-only setup using Ollama as the backend and states no API key is required for that local path.
Why do long-context and tool-calling matter for agent frameworks?
Hermes documents that its agentic tool workflows need a large context window and that non-tool-calling models are chat-only for Hermes actions.
Does Hermes force a single provider?
No. Hermes provider docs describe cloud and self-hosted routes, and you can switch through provider/model configuration.
Verification notes
- The runtime requirements section only includes statements present in official framework docs listed on this page.
- Fit tables are LLMRAM planning estimates derived from current model pages and calculator assumptions.
- When a runtime-specific snippet is missing in provided official links (for example Hermes llama.cpp), this page marks the gap explicitly.