OpenClaw: hardware requirements, local model setup, and context planning
Official-source implementation guide for OpenClaw. OpenClaw is an open-source local-first agent gateway that supports hosted and local model providers, with documented flows for Ollama, LM Studio, and other OpenAI-compatible runtimes.
What it is
OpenClaw presents itself as an assistant stack where state and credentials stay on your devices, while model/runtime backends are pluggable.
Its Gateway acts as the local control plane for sessions, tools, channels, and model routing.
Official runtime requirements
- Node runtime requirement: OpenClaw install docs require Node 24.16+ or 26.1+, and recommend Node 26. Installer provisions Node if missing. (OpenClaw install docs)
- Platform support: OpenClaw install docs list macOS, Linux, and Windows as supported installation targets. (OpenClaw install docs)
- Local-model memory note: OpenClaw local-model docs state memory depends on model weights, context, runtime, and host load; managed llama.cpp recipes use 64K context and smallest curated recipe has an 8 GiB host-memory floor. (OpenClaw local models docs)
Local model connection snippets (official docs)
Ollama
openclaw onboard # choose Ollama: Cloud + Local, Cloud only, or Local only openclaw models list --provider ollama
Source: OpenClaw Ollama setup docs
Ollama
models: {
providers: {
ollama: {
api: 'ollama',
baseUrl: 'http://host:11434',
},
},
}OpenClaw Ollama docs explicitly warn not to use /v1 for Ollama native provider mode.
Source: OpenClaw Ollama provider docs
LM Studio
models: {
providers: {
lmstudio: {
api: 'openai-responses',
baseUrl: 'http://127.0.0.1:1234/v1',
models: [{ id: 'my-local-model', contextWindow: 196608 }],
},
},
}Source: OpenClaw local models docs
vLLM
models: {
providers: {
local: {
api: 'openai-completions',
baseUrl: 'http://127.0.0.1:8000/v1',
apiKey: 'sk-local',
},
},
}OpenClaw local-model docs list vLLM under OpenAI-compatible local proxies.
Source: OpenClaw local models docs
llama.cpp
# Managed local server path from OpenClaw local-model docs openclaw onboard # choose Managed local server (llama.cpp plugin)
Source: OpenClaw local models docs
Why long context and tool-calling models matter
- OpenClaw local-model docs emphasize that a model passing short prompts may still fail full agent turns because tool schemas, history, and prompts consume additional context.
- OpenClaw setup verifies a real tool call before switching the default local model in managed setup flow.
Prefilled calculator
The fit matrix below is an LLMRAM estimate at Q4_K_M and batch=1 for planning. It is not an official vendor minimum requirement.
VRAM / RAM calculator
Pick a model, quantization, context length, batch size, and target hardware. Results apply to dedicated GPU VRAM, Apple unified memory, and CPU/RAM offload planning.
Custom Hugging Face model id / params (optional)
If repo id is unknown, calculator uses your custom params and labels results as estimated.
| Quant | Weights | KV | Overhead | Total |
|---|---|---|---|---|
| FP16 / BF16 | 13.04 GiB | 8 GiB | 1.54 GiB | 22.58 GiB |
| INT8 / Q8_0 | 6.68 GiB | 8 GiB | 1.54 GiB | 16.22 GiB |
| Q6_K | 5.17 GiB | 8 GiB | 1.54 GiB | 14.71 GiB |
| Q5_K_M | 4.44 GiB | 8 GiB | 1.54 GiB | 13.98 GiB |
| Q4_K_M | 3.75 GiB | 8 GiB | 1.54 GiB | 13.29 GiB |
| Q3_K_M / Q3 | 2.85 GiB | 8 GiB | 1.54 GiB | 12.39 GiB |
| Q2_K / Q2 | 2.04 GiB | 8 GiB | 1.54 GiB | 11.58 GiB |
| Hardware | Memory | Verdict | Est. tok/s |
|---|---|---|---|
| NVIDIA GeForce RTX 4090 24GB | 24 GiB | Fits | 170.3 |
| NVIDIA GeForce RTX 5090 32GB | 32 GiB | Fits | 302.75 |
| NVIDIA GeForce RTX 4080 SUPER 16GB | 16 GiB | Fits | 124.34 |
| NVIDIA GeForce RTX 4070 Ti SUPER 16GB | 16 GiB | Fits | 113.53 |
| AMD Radeon RX 7900 XTX 24GB | 24 GiB | Fits | 162.19 |
| NVIDIA A100 80GB PCIe | 80 GiB | Fits | 326.91 |
| Apple Silicon M3 Max (128GB unified memory) | 128 GiB | Fits | 57.64 |
| Apple Silicon M2 Ultra (192GB unified memory) | 192 GiB | Fits | 115.28 |
| 2× NVIDIA GeForce RTX 4090 (aggregate) | 48 GiB | Fits | 306.46 |
Formula and assumptions
- Weight memory uses resident params × effective_bits / 8 (total params when available, otherwise active/fallback estimate).
- KV cache = 2 × layers × kv_heads × head_dim × kv_bytes × context × batch.
- Runtime overhead = 1.3 GiB base + 3% of KV-cache memory.
- Tokens/sec estimate uses active_params for decode bandwidth and should be treated as a rough directional number.
- Apple Silicon fit checks use an estimated usable unified-memory budget (about 75% by default, about 68% on 16GB systems) aligned with macOS recommendedMaxWorkingSetSize behavior.
- Recommended quantization keeps at least 10% memory headroom relative to usable memory budget.
- CPU/RAM offload (for example llama.cpp partial offload with lower GPU layer count) can reduce VRAM needs at the cost of speed.
- When config values are unavailable, fallback defaults are used and flagged as estimated on the page.
Local model fit table by memory tier (LLMRAM estimate)
The fit matrix below is an LLMRAM estimate at Q4_K_M and batch=1 for planning. It is not an official vendor minimum requirement.
32K context
GPU tiers
| Model profile | Estimated total memory | 8 GB | 12 GB | 16 GB | 24 GB | 32 GB | 48 GB |
|---|---|---|---|---|---|---|---|
| Mistral 7B Instruct v0.3 | 9.17 GiB | ✕ No fit | ✓ Fits | ✓ Fits | ✓ Fits | ✓ Fits | ✓ Fits |
| Qwen 3.8 27B | 24 GiB | ✕ No fit | ✕ No fit | ✕ No fit | ✓ Fits | ✓ Fits | ✓ Fits |
| GLM 5.3 (Flash) | 195.84 GiB | ✕ No fit | ✕ No fit | ✕ No fit | ✕ No fit | ✕ No fit | ✕ No fit |
Mac unified memory tiers
| Model profile | Estimated total memory | 16 GB | 36 GB | 64 GB | 128 GB |
|---|---|---|---|---|---|
| Mistral 7B Instruct v0.3 | 9.17 GiB | ✓ Fits | ✓ Fits | ✓ Fits | ✓ Fits |
| Qwen 3.8 27B | 24 GiB | ✕ No fit | ✓ Fits | ✓ Fits | ✓ Fits |
| GLM 5.3 (Flash) | 195.84 GiB | ✕ No fit | ✕ No fit | ✕ No fit | ✕ No fit |
64K context
GPU tiers
| Model profile | Estimated total memory | 8 GB | 12 GB | 16 GB | 24 GB | 32 GB | 48 GB |
|---|---|---|---|---|---|---|---|
| Mistral 7B Instruct v0.3 | 13.29 GiB | ✕ No fit | ✕ No fit | ✓ Fits | ✓ Fits | ✓ Fits | ✓ Fits |
| Qwen 3.8 27B | 32.24 GiB | ✕ No fit | ✕ No fit | ✕ No fit | ✕ No fit | ✕ No fit | ✓ Fits |
| GLM 5.3 (Flash) | 219.01 GiB | ✕ No fit | ✕ No fit | ✕ No fit | ✕ No fit | ✕ No fit | ✕ No fit |
Mac unified memory tiers
| Model profile | Estimated total memory | 16 GB | 36 GB | 64 GB | 128 GB |
|---|---|---|---|---|---|
| Mistral 7B Instruct v0.3 | 13.29 GiB | ✕ No fit | ✓ Fits | ✓ Fits | ✓ Fits |
| Qwen 3.8 27B | 32.24 GiB | ✕ No fit | ✕ No fit | ✓ Fits | ✓ Fits |
| GLM 5.3 (Flash) | 219.01 GiB | ✕ No fit | ✕ No fit | ✕ No fit | ✕ No fit |
FAQ
Does OpenClaw require cloud APIs?
No. OpenClaw docs cover local-first setups and local model routes including Ollama and managed local servers.
Why does OpenClaw document both native Ollama and OpenAI-compatible /v1 backends?
OpenClaw uses Ollama native API for Ollama provider mode, while vLLM/LM Studio/proxies are configured through OpenAI-compatible provider modes.
Can I assume one fixed memory minimum from OpenClaw docs?
No. OpenClaw explicitly says memory depends on model, context, runtime, and host conditions; curated recipe floors are not fit guarantees.
Verification notes
- The runtime requirements section only includes statements present in official framework docs listed on this page.
- Fit tables are LLMRAM planning estimates derived from current model pages and calculator assumptions.
- When a runtime-specific snippet is missing in provided official links (for example Hermes llama.cpp), this page marks the gap explicitly.