2× NVIDIA GeForce RTX 4090 (aggregate): what can it run for local AI?
If you searched for “what can dual rtx 4090 run”, this page gives a practical answer with transparent assumptions. We evaluate a curated set of open-source models using a default Q4_K_M profile so you can quickly see what fits and what needs a larger memory budget.
- Memory type: Dedicated VRAM
- Total memory: 48 GiB
- Bandwidth: 1,814 GB/s
- Planning hint: For MoE and multimodal systems, treat these numbers as first-pass planning values and keep extra memory headroom for framework overhead.
Model fit table on 2× NVIDIA GeForce RTX 4090 (aggregate)
Baseline assumption: each model uses its default context and batch with Q4_K_M quantization. Use individual model pages for deeper what-if analysis.
| Model | Estimated total | Verdict | Rough tokens/s | Recommended quant |
|---|---|---|---|---|
| Qwen 3.8 27B | 20.91 GiB | Fits | 79.45 | int8 |
| Qwen 3.8 Flash Next 125B | 78.66 GiB | Does not fit | 357.54 | q2_k |
| MiniMax H3 33B | 22.45 GiB | Fits | 65.01 | int8 |
| Qwen Image 2.1 | 5.38 GiB | Fits | 306.46 | fp16 |
| Kimi K3 | 1,823.65 GiB | Does not fit | 20.63 | No fit |
| GLM 5.3 (Flash) | 215.7 GiB | Does not fit | 119.18 | No fit |
| DeepSeek V4.1 Flash | 334.25 GiB | Does not fit | 134.08 | No fit |
| MiMo V2.6 Pro RL | 625.89 GiB | Does not fit | 51.08 | No fit |
| Bonsai 2 27B | 21.13 GiB | Fits | 78.41 | int8 |
| Mistral 7B Instruct v0.3 | 5.83 GiB | Fits | 306.46 | fp16 |
Quick answer: what can 2× NVIDIA GeForce RTX 4090 (aggregate)?
2× NVIDIA GeForce RTX 4090 (aggregate) can run 5 out of 10 tracked models at the default Q4_K_M profile. If a model does not fit, the table shows a lower quantization suggestion where possible.
For MoE and multimodal systems, treat these numbers as first-pass planning values and keep extra memory headroom for framework overhead.
FAQ
What can 2× NVIDIA GeForce RTX 4090 (aggregate) run in local AI workflows?
Use the baseline Q4_K_M table as a first-pass fit check. For bigger context windows or multimodal pipelines, reserve extra headroom.
Do these numbers include CPU/RAM offload options?
The table assumes in-memory baseline behavior. You can often load bigger models with partial CPU/RAM offload, usually with lower throughput.
How should Mac unified memory be interpreted here?
Unified memory is treated as a first-class target. Fit can improve with large memory pools, but responsiveness still depends on memory bandwidth and runtime kernels.
Source and verification
Hardware spec source: Derived from two RTX 4090 cards (aggregate estimate) (fetched 2026-10-06).