LLMRAM

Hermes Agent: hardware requirements, local model setup, and context planning

Official-source implementation guide for Hermes Agent. Hermes Agent is a self-hostable AI agent framework with terminal tooling, provider routing, and local-model options including Ollama.

What it is

Hermes Agent describes itself as a self-improving AI agent built by Nous Research, with long-running sessions, tool use, and messaging interfaces.

Its provider system supports hosted APIs and self-hosted endpoints, so you can switch model backends without rewriting agent logic.

Official runtime requirements

  • Installer-managed runtime: Current first-party installations run on Python 3.14; installer-managed toolchain includes Python, Node.js, npm, ripgrep, and FFmpeg. (Hermes Installation docs)
  • Source install prerequisites: POSIX source install requires Git, curl, tar, and SHA-256 utilities. (Hermes Installation docs)
  • Local Ollama guide baseline: Guide table lists 8 GB RAM minimum (3B models), 32+ GB recommended (27B+), 5 GB free storage minimum, and optional NVIDIA GPU with 8+ GB VRAM for faster inference. (Hermes local Ollama setup guide)

Local model connection snippets (official docs)

Ollama

curl -fsSL https://ollama.com/install.sh | sh
ollama pull gemma4:31b
hermes setup
# provider: Custom Endpoint
# base URL: http://localhost:11434/v1

Source: Hermes local Ollama setup guide

Ollama

cat > ~/.hermes/cache/scratch/Modelfile << 'EOF'
FROM gemma4:31b
PARAMETER num_ctx 64000
EOF
ollama create gemma4-64k -f ~/.hermes/cache/scratch/Modelfile

From Hermes Ollama guide context-window section.

Source: Hermes local Ollama setup guide

vLLM

model:
  default: your-model-name
  provider: custom
  base_url: http://localhost:8000/v1
  api_key: your-key-or-leave-empty-for-local

Hermes docs describe self-hosted endpoints including vLLM through the Custom Endpoint flow.

Source: Hermes provider integrations docs

LM Studio

model:
  default: your-model-name
  provider: custom
  base_url: http://localhost:1234/v1
  api_key: your-key-or-leave-empty-for-local

Hermes providers page lists LM Studio and custom endpoint configuration patterns.

Source: Hermes provider integrations docs

llama.cpp

No explicit llama.cpp provider setup snippet appears in the provided Hermes source links.

Kept as a documented gap instead of inventing a command.

Source: Hermes provider integrations docs

Why long context and tool-calling models matter

  • Hermes local Ollama guide explicitly says Hermes requires at least 64,000-token context for agentic work with tools.
  • Hermes local Ollama guide also states models without tool-call support can chat but cannot perform agent actions such as file edits and command execution.

Prefilled calculator

The fit matrix below is an LLMRAM estimate at Q4_K_M and batch=1 for planning. It is not an official vendor minimum requirement.

VRAM / RAM calculator

Pick a model, quantization, context length, batch size, and target hardware. Results apply to dedicated GPU VRAM, Apple unified memory, and CPU/RAM offload planning.

Custom Hugging Face model id / params (optional)

If repo id is unknown, calculator uses your custom params and labels results as estimated.

Memory by quantization
QuantWeightsKVOverheadTotal
FP16 / BF1650.29 GiB16 GiB1.78 GiB68.07 GiB
INT8 / Q8_025.77 GiB16 GiB1.78 GiB43.55 GiB
Q6_K19.96 GiB16 GiB1.78 GiB37.74 GiB
Q5_K_M17.13 GiB16 GiB1.78 GiB34.91 GiB
Q4_K_M14.46 GiB16 GiB1.78 GiB32.24 GiB
Q3_K_M / Q311 GiB16 GiB1.78 GiB28.78 GiB
Q2_K / Q27.86 GiB16 GiB1.78 GiB25.64 GiB
Fits / does not fit by hardware
HardwareMemoryVerdictEst. tok/s
NVIDIA GeForce RTX 4090 24GB24 GiBDoes not fitn/a (Does not fit)
NVIDIA GeForce RTX 5090 32GB32 GiBDoes not fitn/a (Does not fit)
NVIDIA GeForce RTX 4080 SUPER 16GB16 GiBDoes not fitn/a (Does not fit)
NVIDIA GeForce RTX 4070 Ti SUPER 16GB16 GiBDoes not fitn/a (Does not fit)
AMD Radeon RX 7900 XTX 24GB24 GiBDoes not fitn/a (Does not fit)
NVIDIA A100 80GB PCIe80 GiBFits84.75
Apple Silicon M3 Max (128GB unified memory)128 GiBFits14.94
Apple Silicon M2 Ultra (192GB unified memory)192 GiBFits29.89
2× NVIDIA GeForce RTX 4090 (aggregate)48 GiBFits79.45

Formula and assumptions

  • Weight memory uses resident params × effective_bits / 8 (total params when available, otherwise active/fallback estimate).
  • KV cache = 2 × layers × kv_heads × head_dim × kv_bytes × context × batch.
  • Runtime overhead = 1.3 GiB base + 3% of KV-cache memory.
  • Tokens/sec estimate uses active_params for decode bandwidth and should be treated as a rough directional number.
  • Apple Silicon fit checks use an estimated usable unified-memory budget (about 75% by default, about 68% on 16GB systems) aligned with macOS recommendedMaxWorkingSetSize behavior.
  • Recommended quantization keeps at least 10% memory headroom relative to usable memory budget.
  • CPU/RAM offload (for example llama.cpp partial offload with lower GPU layer count) can reduce VRAM needs at the cost of speed.
  • When config values are unavailable, fallback defaults are used and flagged as estimated on the page.

Local model fit table by memory tier (LLMRAM estimate)

The fit matrix below is an LLMRAM estimate at Q4_K_M and batch=1 for planning. It is not an official vendor minimum requirement.

32K context

GPU tiers

Model profileEstimated total memory8 GB12 GB16 GB24 GB32 GB48 GB
Mistral 7B Instruct v0.39.17 GiB✕ No fit✓ Fits✓ Fits✓ Fits✓ Fits✓ Fits
Qwen 3.8 27B24 GiB✕ No fit✕ No fit✕ No fit✓ Fits✓ Fits✓ Fits
GLM 5.3 (Flash)195.84 GiB✕ No fit✕ No fit✕ No fit✕ No fit✕ No fit✕ No fit

Mac unified memory tiers

Model profileEstimated total memory16 GB36 GB64 GB128 GB
Mistral 7B Instruct v0.39.17 GiB✓ Fits✓ Fits✓ Fits✓ Fits
Qwen 3.8 27B24 GiB✕ No fit✓ Fits✓ Fits✓ Fits
GLM 5.3 (Flash)195.84 GiB✕ No fit✕ No fit✕ No fit✕ No fit

64K context

GPU tiers

Model profileEstimated total memory8 GB12 GB16 GB24 GB32 GB48 GB
Mistral 7B Instruct v0.313.29 GiB✕ No fit✕ No fit✓ Fits✓ Fits✓ Fits✓ Fits
Qwen 3.8 27B32.24 GiB✕ No fit✕ No fit✕ No fit✕ No fit✕ No fit✓ Fits
GLM 5.3 (Flash)219.01 GiB✕ No fit✕ No fit✕ No fit✕ No fit✕ No fit✕ No fit

Mac unified memory tiers

Model profileEstimated total memory16 GB36 GB64 GB128 GB
Mistral 7B Instruct v0.313.29 GiB✕ No fit✓ Fits✓ Fits✓ Fits
Qwen 3.8 27B32.24 GiB✕ No fit✕ No fit✓ Fits✓ Fits
GLM 5.3 (Flash)219.01 GiB✕ No fit✕ No fit✕ No fit✕ No fit

FAQ

Can Hermes run fully local with no API key?

Yes, the Hermes Ollama guide documents a local-only setup using Ollama as the backend and states no API key is required for that local path.

Why do long-context and tool-calling matter for agent frameworks?

Hermes documents that its agentic tool workflows need a large context window and that non-tool-calling models are chat-only for Hermes actions.

Does Hermes force a single provider?

No. Hermes provider docs describe cloud and self-hosted routes, and you can switch through provider/model configuration.

Verification notes

  • The runtime requirements section only includes statements present in official framework docs listed on this page.
  • Fit tables are LLMRAM planning estimates derived from current model pages and calculator assumptions.
  • When a runtime-specific snippet is missing in provided official links (for example Hermes llama.cpp), this page marks the gap explicitly.

Official sources used

Internal links