NVIDIA GeForce RTX 5090 32GB: what can it run for local AI?
If you searched for “what can rtx 5090 run”, this page gives a practical answer with transparent assumptions. We evaluate a curated set of open-source models using a default Q4_K_M profile so you can quickly see what fits and what needs a larger memory budget.
- Memory type: Dedicated VRAM
- Total memory: 32 GiB
- Bandwidth: 1,792 GB/s
- Planning hint: For MoE and multimodal systems, treat these numbers as first-pass planning values and keep extra memory headroom for framework overhead.
Model fit table on NVIDIA GeForce RTX 5090 32GB
Baseline assumption: each model uses its default context and batch with Q4_K_M quantization. Use individual model pages for deeper what-if analysis.
| Model | Estimated total | Verdict | Rough tokens/s | Recommended quant |
|---|---|---|---|---|
| Qwen 3.8 27B | 20.91 GiB | Fits | 78.49 | q6_k |
| Qwen 3.8 Flash Next 125B | 78.66 GiB | Does not fit | 353.21 | No fit |
| MiniMax H3 33B | 22.45 GiB | Fits | 64.22 | q6_k |
| Qwen Image 2.1 | 5.38 GiB | Fits | 302.75 | fp16 |
| Kimi K3 | 1,823.65 GiB | Does not fit | 20.38 | No fit |
| GLM 5.3 (Flash) | 215.7 GiB | Does not fit | 117.74 | No fit |
| DeepSeek V4.1 Flash | 334.25 GiB | Does not fit | 132.45 | No fit |
| MiMo V2.6 Pro RL | 625.89 GiB | Does not fit | 50.46 | No fit |
| Bonsai 2 27B | 21.13 GiB | Fits | 77.46 | q6_k |
| Mistral 7B Instruct v0.3 | 5.83 GiB | Fits | 302.75 | fp16 |
Quick answer: what can NVIDIA GeForce RTX 5090 32GB?
NVIDIA GeForce RTX 5090 32GB can run 5 out of 10 tracked models at the default Q4_K_M profile. If a model does not fit, the table shows a lower quantization suggestion where possible.
For MoE and multimodal systems, treat these numbers as first-pass planning values and keep extra memory headroom for framework overhead.
FAQ
What can NVIDIA GeForce RTX 5090 32GB run in local AI workflows?
Use the baseline Q4_K_M table as a first-pass fit check. For bigger context windows or multimodal pipelines, reserve extra headroom.
Do these numbers include CPU/RAM offload options?
The table assumes in-memory baseline behavior. You can often load bigger models with partial CPU/RAM offload, usually with lower throughput.
How should Mac unified memory be interpreted here?
Unified memory is treated as a first-class target. Fit can improve with large memory pools, but responsiveness still depends on memory bandwidth and runtime kernels.
Source and verification
Hardware spec source: NVIDIA product page (fetched 2026-10-06).