OpenClaw: Hardware-Anforderungen, lokale Modell-Anbindung und Kontextplanung
Offizieller Leitfaden für OpenClaw. OpenClaw ist ein Open-Source-Agent-Gateway mit Local-First-Ansatz und unterstützt sowohl Cloud- als auch lokale Modell-Provider.
Was ist das?
OpenClaw positioniert sich als Assistenten-Stack, bei dem Zustand und Zugangsdaten auf Ihren Geräten verbleiben, während Modell- und Runtime-Backends austauschbar sind.
Das Gateway fungiert als lokale Control-Plane für Sessions, Tools, Channels und Modellrouting.
Offizielle Laufzeit-Anforderungen
- Node-Laufzeitanforderung: Die OpenClaw-Installationsdokumentation verlangt Node 24.16+ oder 26.1+ und empfiehlt Node 26. Falls Node fehlt, wird es vom Installer bereitgestellt. (OpenClaw Installationsdokumentation)
- Plattformunterstützung: Die OpenClaw-Installationsdokumentation nennt macOS, Linux und Windows als unterstützte Ziele. (OpenClaw Installationsdokumentation)
- Hinweis zu Local-Model-Speicher: Die OpenClaw-Local-Model-Dokumentation erklärt, dass Speicherbedarf von Modellgewichten, Kontext, Runtime und Hostlast abhängt; verwaltete llama.cpp-Rezepte nutzen 64K-Kontext und das kleinste Rezept hat eine 8-GiB-Hostspeicher-Untergrenze. (OpenClaw Dokumentation für lokale Modelle)
Snippets zur lokalen Modell-Anbindung (offizielle Doku)
Ollama
openclaw onboard # choose Ollama: Cloud + Local, Cloud only, or Local only openclaw models list --provider ollama
Ollama
models: {
providers: {
ollama: {
api: 'ollama',
baseUrl: 'http://host:11434',
},
},
}Die OpenClaw-Ollama-Dokumentation warnt ausdrücklich davor, im nativen Ollama-Provider-Modus /v1 zu verwenden.
LM Studio
models: {
providers: {
lmstudio: {
api: 'openai-responses',
baseUrl: 'http://127.0.0.1:1234/v1',
models: [{ id: 'my-local-model', contextWindow: 196608 }],
},
},
}vLLM
models: {
providers: {
local: {
api: 'openai-completions',
baseUrl: 'http://127.0.0.1:8000/v1',
apiKey: 'sk-local',
},
},
}Die OpenClaw-Local-Model-Dokumentation führt vLLM unter OpenAI-kompatiblen lokalen Proxys auf.
llama.cpp
# Managed local server path from OpenClaw local-model docs openclaw onboard # choose Managed local server (llama.cpp plugin)
Warum lange Kontexte und Tool-Calling wichtig sind
- OpenClaw betont in den Local-Model-Dokumenten, dass ein Modell trotz kurzer Prompt-Tests bei vollständigen Agent-Turns scheitern kann, weil Tool-Schemas, Verlauf und Prompts zusätzlichen Kontext verbrauchen.
- Das OpenClaw-Setup prüft einen echten Tool-Call, bevor das Standard-Local-Model in der verwalteten Einrichtung umgestellt wird.
Vorausgefüllter Rechner
Die Matrix ist eine LLMRAM-Planungsschätzung (Q4_K_M, batch=1), keine offizielle Mindestanforderung.
VRAM / RAM Rechner
Vergleichen Sie Eignung für GPU-VRAM, Apple Unified Memory und CPU/RAM-Offload.
Eigenes Hugging Face Modell / Parameter (optional)
Wenn die Repo-ID unbekannt ist, nutzt der Rechner Ihre Werte als Schätzung.
| Quant | Gewichte | KV | Overhead | Gesamt |
|---|---|---|---|---|
| FP16 / BF16 | 13.04 GiB | 8 GiB | 1.54 GiB | 22.58 GiB |
| INT8 / Q8_0 | 6.68 GiB | 8 GiB | 1.54 GiB | 16.22 GiB |
| Q6_K | 5.17 GiB | 8 GiB | 1.54 GiB | 14.71 GiB |
| Q5_K_M | 4.44 GiB | 8 GiB | 1.54 GiB | 13.98 GiB |
| Q4_K_M | 3.75 GiB | 8 GiB | 1.54 GiB | 13.29 GiB |
| Q3_K_M / Q3 | 2.85 GiB | 8 GiB | 1.54 GiB | 12.39 GiB |
| Q2_K / Q2 | 2.04 GiB | 8 GiB | 1.54 GiB | 11.58 GiB |
| Hardware | Speicher | Urteil | Tok/s (geschätzt) |
|---|---|---|---|
| NVIDIA GeForce RTX 4090 24GB | 24 GiB | Passt | 170.3 |
| NVIDIA GeForce RTX 5090 32GB | 32 GiB | Passt | 302.75 |
| NVIDIA GeForce RTX 4080 SUPER 16GB | 16 GiB | Passt | 124.34 |
| NVIDIA GeForce RTX 4070 Ti SUPER 16GB | 16 GiB | Passt | 113.53 |
| AMD Radeon RX 7900 XTX 24GB | 24 GiB | Passt | 162.19 |
| NVIDIA A100 80GB PCIe | 80 GiB | Passt | 326.91 |
| Apple Silicon M3 Max (128GB unified memory) | 128 GiB | Passt | 57.64 |
| Apple Silicon M2 Ultra (192GB unified memory) | 192 GiB | Passt | 115.28 |
| 2× NVIDIA GeForce RTX 4090 (aggregate) | 48 GiB | Passt | 306.46 |
Formeln und Annahmen
- Weight memory uses resident params × effective_bits / 8 (total params when available, otherwise active/fallback estimate).
- KV cache = 2 × layers × kv_heads × head_dim × kv_bytes × context × batch.
- Runtime overhead = 1.3 GiB base + 3% of KV-cache memory.
- Tokens/sec estimate uses active_params for decode bandwidth and should be treated as a rough directional number.
- Apple Silicon fit checks use an estimated usable unified-memory budget (about 75% by default, about 68% on 16GB systems) aligned with macOS recommendedMaxWorkingSetSize behavior.
- Recommended quantization keeps at least 10% memory headroom relative to usable memory budget.
- CPU/RAM offload (for example llama.cpp partial offload with lower GPU layer count) can reduce VRAM needs at the cost of speed.
- When config values are unavailable, fallback defaults are used and flagged as estimated on the page.
Fit-Tabelle nach Speicherklasse (LLMRAM-Schätzung)
Die Matrix ist eine LLMRAM-Planungsschätzung (Q4_K_M, batch=1), keine offizielle Mindestanforderung.
32K Kontext
GPU-Klassen
| Modellprofil | Geschätzter Gesamtspeicher | 8 GB | 12 GB | 16 GB | 24 GB | 32 GB | 48 GB |
|---|---|---|---|---|---|---|---|
| Mistral 7B Instruct v0.3 | 9.17 GiB | ✕ Passt nicht | ✓ Passt | ✓ Passt | ✓ Passt | ✓ Passt | ✓ Passt |
| Qwen 3.8 27B | 24 GiB | ✕ Passt nicht | ✕ Passt nicht | ✕ Passt nicht | ✓ Passt | ✓ Passt | ✓ Passt |
| GLM 5.3 (Flash) | 195.84 GiB | ✕ Passt nicht | ✕ Passt nicht | ✕ Passt nicht | ✕ Passt nicht | ✕ Passt nicht | ✕ Passt nicht |
Mac Unified-Memory-Klassen
| Modellprofil | Geschätzter Gesamtspeicher | 16 GB | 36 GB | 64 GB | 128 GB |
|---|---|---|---|---|---|
| Mistral 7B Instruct v0.3 | 9.17 GiB | ✓ Passt | ✓ Passt | ✓ Passt | ✓ Passt |
| Qwen 3.8 27B | 24 GiB | ✕ Passt nicht | ✓ Passt | ✓ Passt | ✓ Passt |
| GLM 5.3 (Flash) | 195.84 GiB | ✕ Passt nicht | ✕ Passt nicht | ✕ Passt nicht | ✕ Passt nicht |
64K Kontext
GPU-Klassen
| Modellprofil | Geschätzter Gesamtspeicher | 8 GB | 12 GB | 16 GB | 24 GB | 32 GB | 48 GB |
|---|---|---|---|---|---|---|---|
| Mistral 7B Instruct v0.3 | 13.29 GiB | ✕ Passt nicht | ✕ Passt nicht | ✓ Passt | ✓ Passt | ✓ Passt | ✓ Passt |
| Qwen 3.8 27B | 32.24 GiB | ✕ Passt nicht | ✕ Passt nicht | ✕ Passt nicht | ✕ Passt nicht | ✕ Passt nicht | ✓ Passt |
| GLM 5.3 (Flash) | 219.01 GiB | ✕ Passt nicht | ✕ Passt nicht | ✕ Passt nicht | ✕ Passt nicht | ✕ Passt nicht | ✕ Passt nicht |
Mac Unified-Memory-Klassen
| Modellprofil | Geschätzter Gesamtspeicher | 16 GB | 36 GB | 64 GB | 128 GB |
|---|---|---|---|---|---|
| Mistral 7B Instruct v0.3 | 13.29 GiB | ✕ Passt nicht | ✓ Passt | ✓ Passt | ✓ Passt |
| Qwen 3.8 27B | 32.24 GiB | ✕ Passt nicht | ✕ Passt nicht | ✓ Passt | ✓ Passt |
| GLM 5.3 (Flash) | 219.01 GiB | ✕ Passt nicht | ✕ Passt nicht | ✕ Passt nicht | ✕ Passt nicht |
FAQ
Benötigt OpenClaw zwingend Cloud-APIs?
Nein. Die OpenClaw-Dokumentation beschreibt Local-First-Setups und lokale Modellpfade inklusive Ollama und verwalteter lokaler Server.
Warum dokumentiert OpenClaw sowohl natives Ollama als auch OpenAI-kompatible /v1-Backends?
Für den Ollama-Provider-Modus wird die native Ollama-API genutzt, während vLLM/LM Studio/Proxys über OpenAI-kompatible Provider-Modi konfiguriert werden.
Kann ich aus OpenClaw-Dokumenten ein fixes Speicher-Minimum ableiten?
Nein. OpenClaw sagt explizit, dass der Speicherbedarf von Modell, Kontext, Runtime und Hostbedingungen abhängt; Rezept-Untergrenzen sind keine Fit-Garantie.
Verifizierungsnotizen
- Im Abschnitt zu Laufzeitanforderungen stehen nur Aussagen, die in den auf dieser Seite verlinkten offiziellen Framework-Dokumenten enthalten sind.
- Die Fit-Tabellen sind LLMRAM-Planungsschätzungen auf Basis der aktuellen Modellseiten und Rechnerannahmen.
- Wenn in offiziellen Links ein runtime-spezifisches Snippet fehlt (z. B. Hermes llama.cpp), markieren wir die Lücke ausdrücklich statt Befehle zu erfinden.