Hermes Agent: Hardware-Anforderungen, lokale Modell-Anbindung und Kontextplanung
Offizieller Leitfaden für Hermes Agent. Hermes Agent ist ein selbst hostbares KI-Agent-Framework mit Terminal-Tools, Provider-Routing und lokalen Modellpfaden inklusive Ollama.
Was ist das?
Hermes Agent beschreibt sich als selbstverbessernden KI-Agenten von Nous Research mit Langzeitsitzungen, Tool-Nutzung und Messaging-Schnittstellen.
Das Provider-System unterstützt sowohl gehostete APIs als auch Self-Hosted-Endpunkte, sodass Sie Modell-Backends ohne Umschreiben der Agent-Logik wechseln können.
Offizielle Laufzeit-Anforderungen
- Installer-verwaltete Laufzeit: Aktuelle First-Party-Installationen laufen auf Python 3.14; die verwaltete Toolchain enthält Python, Node.js, npm, ripgrep und FFmpeg. (Hermes Installationsdokumentation)
- Voraussetzungen für Source-Install: Für POSIX-Source-Installationen werden Git, curl, tar und SHA-256-Werkzeuge benötigt. (Hermes Installationsdokumentation)
- Basis laut lokalem Ollama-Guide: Die Guide-Tabelle nennt 8 GB RAM Minimum (3B-Modelle), 32+ GB empfohlen (27B+), mindestens 5 GB freien Speicher und optional NVIDIA-GPU mit 8+ GB VRAM für schnellere Inferenz. (Hermes Leitfaden für lokales Ollama)
Snippets zur lokalen Modell-Anbindung (offizielle Doku)
Ollama
curl -fsSL https://ollama.com/install.sh | sh ollama pull gemma4:31b hermes setup # provider: Custom Endpoint # base URL: http://localhost:11434/v1
Ollama
cat > ~/.hermes/cache/scratch/Modelfile << 'EOF' FROM gemma4:31b PARAMETER num_ctx 64000 EOF ollama create gemma4-64k -f ~/.hermes/cache/scratch/Modelfile
Aus dem Abschnitt zum Kontextfenster im Hermes-Ollama-Guide.
vLLM
model: default: your-model-name provider: custom base_url: http://localhost:8000/v1 api_key: your-key-or-leave-empty-for-local
Hermes dokumentiert Self-Hosted-Endpunkte inklusive vLLM über den Custom-Endpoint-Flow.
LM Studio
model: default: your-model-name provider: custom base_url: http://localhost:1234/v1 api_key: your-key-or-leave-empty-for-local
Die Hermes-Provider-Seite listet Konfigurationsmuster für LM Studio und Custom Endpoints.
llama.cpp
No explicit llama.cpp provider setup snippet appears in the provided Hermes source links.
In den bereitgestellten Hermes-Quellen gibt es kein explizites llama.cpp-Provider-Snippet; daher wird die Lücke klar markiert statt Kommandos zu erfinden.
Warum lange Kontexte und Tool-Calling wichtig sind
- Der lokale Hermes-Ollama-Guide sagt explizit, dass für agentische Tool-Workflows mindestens 64.000 Kontext-Token erforderlich sind.
- Der Guide erklärt außerdem: Modelle ohne Tool-Calling können chatten, aber keine Agent-Aktionen wie Datei-Edits oder Kommandoausführung ausführen.
Vorausgefüllter Rechner
Die Matrix ist eine LLMRAM-Planungsschätzung (Q4_K_M, batch=1), keine offizielle Mindestanforderung.
VRAM / RAM Rechner
Vergleichen Sie Eignung für GPU-VRAM, Apple Unified Memory und CPU/RAM-Offload.
Eigenes Hugging Face Modell / Parameter (optional)
Wenn die Repo-ID unbekannt ist, nutzt der Rechner Ihre Werte als Schätzung.
| Quant | Gewichte | KV | Overhead | Gesamt |
|---|---|---|---|---|
| FP16 / BF16 | 50.29 GiB | 16 GiB | 1.78 GiB | 68.07 GiB |
| INT8 / Q8_0 | 25.77 GiB | 16 GiB | 1.78 GiB | 43.55 GiB |
| Q6_K | 19.96 GiB | 16 GiB | 1.78 GiB | 37.74 GiB |
| Q5_K_M | 17.13 GiB | 16 GiB | 1.78 GiB | 34.91 GiB |
| Q4_K_M | 14.46 GiB | 16 GiB | 1.78 GiB | 32.24 GiB |
| Q3_K_M / Q3 | 11 GiB | 16 GiB | 1.78 GiB | 28.78 GiB |
| Q2_K / Q2 | 7.86 GiB | 16 GiB | 1.78 GiB | 25.64 GiB |
| Hardware | Speicher | Urteil | Tok/s (geschätzt) |
|---|---|---|---|
| NVIDIA GeForce RTX 4090 24GB | 24 GiB | Passt nicht | n/a (Passt nicht) |
| NVIDIA GeForce RTX 5090 32GB | 32 GiB | Passt nicht | n/a (Passt nicht) |
| NVIDIA GeForce RTX 4080 SUPER 16GB | 16 GiB | Passt nicht | n/a (Passt nicht) |
| NVIDIA GeForce RTX 4070 Ti SUPER 16GB | 16 GiB | Passt nicht | n/a (Passt nicht) |
| AMD Radeon RX 7900 XTX 24GB | 24 GiB | Passt nicht | n/a (Passt nicht) |
| NVIDIA A100 80GB PCIe | 80 GiB | Passt | 84.75 |
| Apple Silicon M3 Max (128GB unified memory) | 128 GiB | Passt | 14.94 |
| Apple Silicon M2 Ultra (192GB unified memory) | 192 GiB | Passt | 29.89 |
| 2× NVIDIA GeForce RTX 4090 (aggregate) | 48 GiB | Passt | 79.45 |
Formeln und Annahmen
- Weight memory uses resident params × effective_bits / 8 (total params when available, otherwise active/fallback estimate).
- KV cache = 2 × layers × kv_heads × head_dim × kv_bytes × context × batch.
- Runtime overhead = 1.3 GiB base + 3% of KV-cache memory.
- Tokens/sec estimate uses active_params for decode bandwidth and should be treated as a rough directional number.
- Apple Silicon fit checks use an estimated usable unified-memory budget (about 75% by default, about 68% on 16GB systems) aligned with macOS recommendedMaxWorkingSetSize behavior.
- Recommended quantization keeps at least 10% memory headroom relative to usable memory budget.
- CPU/RAM offload (for example llama.cpp partial offload with lower GPU layer count) can reduce VRAM needs at the cost of speed.
- When config values are unavailable, fallback defaults are used and flagged as estimated on the page.
Fit-Tabelle nach Speicherklasse (LLMRAM-Schätzung)
Die Matrix ist eine LLMRAM-Planungsschätzung (Q4_K_M, batch=1), keine offizielle Mindestanforderung.
32K Kontext
GPU-Klassen
| Modellprofil | Geschätzter Gesamtspeicher | 8 GB | 12 GB | 16 GB | 24 GB | 32 GB | 48 GB |
|---|---|---|---|---|---|---|---|
| Mistral 7B Instruct v0.3 | 9.17 GiB | ✕ Passt nicht | ✓ Passt | ✓ Passt | ✓ Passt | ✓ Passt | ✓ Passt |
| Qwen 3.8 27B | 24 GiB | ✕ Passt nicht | ✕ Passt nicht | ✕ Passt nicht | ✓ Passt | ✓ Passt | ✓ Passt |
| GLM 5.3 (Flash) | 195.84 GiB | ✕ Passt nicht | ✕ Passt nicht | ✕ Passt nicht | ✕ Passt nicht | ✕ Passt nicht | ✕ Passt nicht |
Mac Unified-Memory-Klassen
| Modellprofil | Geschätzter Gesamtspeicher | 16 GB | 36 GB | 64 GB | 128 GB |
|---|---|---|---|---|---|
| Mistral 7B Instruct v0.3 | 9.17 GiB | ✓ Passt | ✓ Passt | ✓ Passt | ✓ Passt |
| Qwen 3.8 27B | 24 GiB | ✕ Passt nicht | ✓ Passt | ✓ Passt | ✓ Passt |
| GLM 5.3 (Flash) | 195.84 GiB | ✕ Passt nicht | ✕ Passt nicht | ✕ Passt nicht | ✕ Passt nicht |
64K Kontext
GPU-Klassen
| Modellprofil | Geschätzter Gesamtspeicher | 8 GB | 12 GB | 16 GB | 24 GB | 32 GB | 48 GB |
|---|---|---|---|---|---|---|---|
| Mistral 7B Instruct v0.3 | 13.29 GiB | ✕ Passt nicht | ✕ Passt nicht | ✓ Passt | ✓ Passt | ✓ Passt | ✓ Passt |
| Qwen 3.8 27B | 32.24 GiB | ✕ Passt nicht | ✕ Passt nicht | ✕ Passt nicht | ✕ Passt nicht | ✕ Passt nicht | ✓ Passt |
| GLM 5.3 (Flash) | 219.01 GiB | ✕ Passt nicht | ✕ Passt nicht | ✕ Passt nicht | ✕ Passt nicht | ✕ Passt nicht | ✕ Passt nicht |
Mac Unified-Memory-Klassen
| Modellprofil | Geschätzter Gesamtspeicher | 16 GB | 36 GB | 64 GB | 128 GB |
|---|---|---|---|---|---|
| Mistral 7B Instruct v0.3 | 13.29 GiB | ✕ Passt nicht | ✓ Passt | ✓ Passt | ✓ Passt |
| Qwen 3.8 27B | 32.24 GiB | ✕ Passt nicht | ✕ Passt nicht | ✓ Passt | ✓ Passt |
| GLM 5.3 (Flash) | 219.01 GiB | ✕ Passt nicht | ✕ Passt nicht | ✕ Passt nicht | ✕ Passt nicht |
FAQ
Kann Hermes komplett lokal ohne API-Key laufen?
Ja. Der Hermes-Ollama-Guide beschreibt einen lokalen-only Pfad mit Ollama-Backend und nennt ausdrücklich, dass dafür kein API-Key nötig ist.
Warum sind lange Kontexte und Tool-Calling wichtig?
Hermes dokumentiert, dass agentische Tool-Workflows ein großes Kontextfenster benötigen und Modelle ohne Tool-Calling in Hermes auf Chat beschränkt sind.
Erzwingt Hermes einen einzelnen Provider?
Nein. Die Provider-Dokumentation beschreibt Cloud- und Self-Hosted-Wege, die per Provider/Model-Konfiguration gewechselt werden können.
Verifizierungsnotizen
- Im Abschnitt zu Laufzeitanforderungen stehen nur Aussagen, die in den auf dieser Seite verlinkten offiziellen Framework-Dokumenten enthalten sind.
- Die Fit-Tabellen sind LLMRAM-Planungsschätzungen auf Basis der aktuellen Modellseiten und Rechnerannahmen.
- Wenn in offiziellen Links ein runtime-spezifisches Snippet fehlt (z. B. Hermes llama.cpp), markieren wir die Lücke ausdrücklich statt Befehle zu erfinden.