LLMRAM

LLMRAM: AI model VRAM & RAM calculator

Use this AI model VRAM and RAM calculator to answer “can I run this model locally?” before you waste time on trial-and-error. Pick a model, set quantization, context length, and batch size, then compare fit results across NVIDIA, AMD, Apple Silicon unified memory, and CPU/RAM offload plans.

VRAM / RAM calculator

Pick a model, quantization, context length, batch size, and target hardware. Results apply to dedicated GPU VRAM, Apple unified memory, and CPU/RAM offload planning.

Custom Hugging Face model id / params (optional)

If repo id is unknown, calculator uses your custom params and labels results as estimated.

Memory by quantization
QuantWeightsKVOverheadTotal
FP16 / BF1650.29 GiB4 GiB3.74 GiB58.03 GiB
INT8 / Q8_025.77 GiB4 GiB2.78 GiB32.56 GiB
Q6_K19.96 GiB4 GiB2.52 GiB26.48 GiB
Q5_K_M17.13 GiB4 GiB2.43 GiB23.56 GiB
Q4_K_M14.46 GiB4 GiB2.46 GiB20.91 GiB
Q3_K_M / Q311 GiB4 GiB2.26 GiB17.26 GiB
Q2_K / Q27.86 GiB4 GiB1.98 GiB13.84 GiB
Fits / does not fit by hardware
HardwareMemoryVerdictEst. tok/s
NVIDIA GeForce RTX 4090 24GB24 GiBFits37.26
NVIDIA GeForce RTX 5090 32GB32 GiBFits66.25
NVIDIA GeForce RTX 4080 SUPER 16GB16 GiBDoes not fit27.21
NVIDIA GeForce RTX 4070 Ti SUPER 16GB16 GiBDoes not fit24.84
AMD Radeon RX 7900 XTX 24GB24 GiBFits35.49
NVIDIA A100 80GB PCIe80 GiBFits71.54
Apple Silicon M3 Max (128GB unified memory)128 GiBFits12.61
Apple Silicon M2 Ultra (192GB unified memory)192 GiBFits25.23
2× NVIDIA GeForce RTX 4090 (aggregate)48 GiBFits67.06

Formula and assumptions

  • Weight memory uses resident params × effective_bits / 8 (total params when available, otherwise active/fallback estimate).
  • KV cache = 2 × layers × kv_heads × head_dim × kv_bytes × context × batch.
  • Runtime overhead includes allocator fragmentation, kernels, and framework buffers.
  • Tokens/sec estimate uses active_params for decode bandwidth and should be treated as a rough directional number.
  • CPU/RAM offload (for example llama.cpp partial offload with lower GPU layer count) can reduce VRAM needs at the cost of speed.
  • When config values are unavailable, fallback defaults are used and flagged as estimated on the page.

Methodology for VRAM, unified memory, and RAM offload

  1. Weights memory = resident params × effective bits / 8, converted to GiB.
  2. For MoE checkpoints, weight residency uses total params; active params are used for rough tokens/sec only.
  3. KV cache = 2 × layers × kv_heads × head_dim × kv_bytes × context × batch.
  4. Overhead includes fragmentation, runtime buffers, and execution graph allocations.
  5. Tokens/sec is estimated from memory bandwidth and decode-time bytes-per-token, clearly labeled as rough.
  6. CPU/RAM partial offload can reduce VRAM pressure with speed trade-offs.
  7. Mac unified memory is treated as a first-class fit target, not a GPU-only fallback.

FAQ

How accurate are the memory estimates?

They are planning-grade estimates, not runtime guarantees. We separate weights, KV cache, and overhead so you can inspect assumptions.

Does this calculator support MoE models?

Yes. We track total and active parameters separately: total parameters drive weight residency memory, while active parameters are used for decode-oriented throughput estimates.

Can I add newly released models quickly?

Yes. The repository includes a script that pulls Hugging Face config.json and generates a typed model entry you can publish within hours.