LLMRAM

OpenClaw: hardware requirements, local model setup, and context planning

Official-source implementation guide for OpenClaw. OpenClaw is an open-source local-first agent gateway that supports hosted and local model providers, with documented flows for Ollama, LM Studio, and other OpenAI-compatible runtimes.

What it is

OpenClaw presents itself as an assistant stack where state and credentials stay on your devices, while model/runtime backends are pluggable.

Its Gateway acts as the local control plane for sessions, tools, channels, and model routing.

Official runtime requirements

  • Node runtime requirement: OpenClaw install docs require Node 24.16+ or 26.1+, and recommend Node 26. Installer provisions Node if missing. (OpenClaw install docs)
  • Platform support: OpenClaw install docs list macOS, Linux, and Windows as supported installation targets. (OpenClaw install docs)
  • Local-model memory note: OpenClaw local-model docs state memory depends on model weights, context, runtime, and host load; managed llama.cpp recipes use 64K context and smallest curated recipe has an 8 GiB host-memory floor. (OpenClaw local models docs)

Local model connection snippets (official docs)

Ollama

openclaw onboard
# choose Ollama: Cloud + Local, Cloud only, or Local only
openclaw models list --provider ollama

Source: OpenClaw Ollama setup docs

Ollama

models: {
  providers: {
    ollama: {
      api: 'ollama',
      baseUrl: 'http://host:11434',
    },
  },
}

OpenClaw Ollama docs explicitly warn not to use /v1 for Ollama native provider mode.

Source: OpenClaw Ollama provider docs

LM Studio

models: {
  providers: {
    lmstudio: {
      api: 'openai-responses',
      baseUrl: 'http://127.0.0.1:1234/v1',
      models: [{ id: 'my-local-model', contextWindow: 196608 }],
    },
  },
}

Source: OpenClaw local models docs

vLLM

models: {
  providers: {
    local: {
      api: 'openai-completions',
      baseUrl: 'http://127.0.0.1:8000/v1',
      apiKey: 'sk-local',
    },
  },
}

OpenClaw local-model docs list vLLM under OpenAI-compatible local proxies.

Source: OpenClaw local models docs

llama.cpp

# Managed local server path from OpenClaw local-model docs
openclaw onboard
# choose Managed local server (llama.cpp plugin)

Source: OpenClaw local models docs

Why long context and tool-calling models matter

  • OpenClaw local-model docs emphasize that a model passing short prompts may still fail full agent turns because tool schemas, history, and prompts consume additional context.
  • OpenClaw setup verifies a real tool call before switching the default local model in managed setup flow.

Prefilled calculator

The fit matrix below is an LLMRAM estimate at Q4_K_M and batch=1 for planning. It is not an official vendor minimum requirement.

VRAM / RAM calculator

Pick a model, quantization, context length, batch size, and target hardware. Results apply to dedicated GPU VRAM, Apple unified memory, and CPU/RAM offload planning.

Custom Hugging Face model id / params (optional)

If repo id is unknown, calculator uses your custom params and labels results as estimated.

Memory by quantization
QuantWeightsKVOverheadTotal
FP16 / BF1613.04 GiB8 GiB1.54 GiB22.58 GiB
INT8 / Q8_06.68 GiB8 GiB1.54 GiB16.22 GiB
Q6_K5.17 GiB8 GiB1.54 GiB14.71 GiB
Q5_K_M4.44 GiB8 GiB1.54 GiB13.98 GiB
Q4_K_M3.75 GiB8 GiB1.54 GiB13.29 GiB
Q3_K_M / Q32.85 GiB8 GiB1.54 GiB12.39 GiB
Q2_K / Q22.04 GiB8 GiB1.54 GiB11.58 GiB
Fits / does not fit by hardware
HardwareMemoryVerdictEst. tok/s
NVIDIA GeForce RTX 4090 24GB24 GiBFits170.3
NVIDIA GeForce RTX 5090 32GB32 GiBFits302.75
NVIDIA GeForce RTX 4080 SUPER 16GB16 GiBFits124.34
NVIDIA GeForce RTX 4070 Ti SUPER 16GB16 GiBFits113.53
AMD Radeon RX 7900 XTX 24GB24 GiBFits162.19
NVIDIA A100 80GB PCIe80 GiBFits326.91
Apple Silicon M3 Max (128GB unified memory)128 GiBFits57.64
Apple Silicon M2 Ultra (192GB unified memory)192 GiBFits115.28
2× NVIDIA GeForce RTX 4090 (aggregate)48 GiBFits306.46

Formula and assumptions

  • Weight memory uses resident params × effective_bits / 8 (total params when available, otherwise active/fallback estimate).
  • KV cache = 2 × layers × kv_heads × head_dim × kv_bytes × context × batch.
  • Runtime overhead = 1.3 GiB base + 3% of KV-cache memory.
  • Tokens/sec estimate uses active_params for decode bandwidth and should be treated as a rough directional number.
  • Apple Silicon fit checks use an estimated usable unified-memory budget (about 75% by default, about 68% on 16GB systems) aligned with macOS recommendedMaxWorkingSetSize behavior.
  • Recommended quantization keeps at least 10% memory headroom relative to usable memory budget.
  • CPU/RAM offload (for example llama.cpp partial offload with lower GPU layer count) can reduce VRAM needs at the cost of speed.
  • When config values are unavailable, fallback defaults are used and flagged as estimated on the page.

Local model fit table by memory tier (LLMRAM estimate)

The fit matrix below is an LLMRAM estimate at Q4_K_M and batch=1 for planning. It is not an official vendor minimum requirement.

32K context

GPU tiers

Model profileEstimated total memory8 GB12 GB16 GB24 GB32 GB48 GB
Mistral 7B Instruct v0.39.17 GiB✕ No fit✓ Fits✓ Fits✓ Fits✓ Fits✓ Fits
Qwen 3.8 27B24 GiB✕ No fit✕ No fit✕ No fit✓ Fits✓ Fits✓ Fits
GLM 5.3 (Flash)195.84 GiB✕ No fit✕ No fit✕ No fit✕ No fit✕ No fit✕ No fit

Mac unified memory tiers

Model profileEstimated total memory16 GB36 GB64 GB128 GB
Mistral 7B Instruct v0.39.17 GiB✓ Fits✓ Fits✓ Fits✓ Fits
Qwen 3.8 27B24 GiB✕ No fit✓ Fits✓ Fits✓ Fits
GLM 5.3 (Flash)195.84 GiB✕ No fit✕ No fit✕ No fit✕ No fit

64K context

GPU tiers

Model profileEstimated total memory8 GB12 GB16 GB24 GB32 GB48 GB
Mistral 7B Instruct v0.313.29 GiB✕ No fit✕ No fit✓ Fits✓ Fits✓ Fits✓ Fits
Qwen 3.8 27B32.24 GiB✕ No fit✕ No fit✕ No fit✕ No fit✕ No fit✓ Fits
GLM 5.3 (Flash)219.01 GiB✕ No fit✕ No fit✕ No fit✕ No fit✕ No fit✕ No fit

Mac unified memory tiers

Model profileEstimated total memory16 GB36 GB64 GB128 GB
Mistral 7B Instruct v0.313.29 GiB✕ No fit✓ Fits✓ Fits✓ Fits
Qwen 3.8 27B32.24 GiB✕ No fit✕ No fit✓ Fits✓ Fits
GLM 5.3 (Flash)219.01 GiB✕ No fit✕ No fit✕ No fit✕ No fit

FAQ

Does OpenClaw require cloud APIs?

No. OpenClaw docs cover local-first setups and local model routes including Ollama and managed local servers.

Why does OpenClaw document both native Ollama and OpenAI-compatible /v1 backends?

OpenClaw uses Ollama native API for Ollama provider mode, while vLLM/LM Studio/proxies are configured through OpenAI-compatible provider modes.

Can I assume one fixed memory minimum from OpenClaw docs?

No. OpenClaw explicitly says memory depends on model, context, runtime, and host conditions; curated recipe floors are not fit guarantees.

Verification notes

  • The runtime requirements section only includes statements present in official framework docs listed on this page.
  • Fit tables are LLMRAM planning estimates derived from current model pages and calculator assumptions.
  • When a runtime-specific snippet is missing in provided official links (for example Hermes llama.cpp), this page marks the gap explicitly.

Official sources used

Internal links