Model fit guide / RTX 5090 32GBWeight-based fit estimateReviewed October 7, 2026
32 GB VRAM - Single-card candidate

Local AI model fit for RTX 5090 32GB.

Use this versioned model as a fit candidate, then evaluate it on your workload. Its 35B total language-model parameters imply about 26.3 GB of raw weights at 6 bits. This is an analytical estimate, not a measured allocation or a claim that this model is the latest quality winner.

01 / Recommendation

Run this size class.

Candidate to evaluate

Qwen3.6-35B-A3B (evaluation candidate)

Use Q6 where supported; context and batch must be measured. Estimate the full runtime allocation before deployment. This candidate has not been measured here at a fixed context length or batch size.

Weight estimate

Weight-only estimate: 26.3 GB at 6 bits, before scales, unquantized tensors, vision components, KV cache and runtime buffers.

Evidence

Official model card gives 35B total language-model parameters. Fit labels describe estimated weight headroom only. No tokens-per-second, maximum-context or relative-quality measurement is claimed. Official model card ↗

02 / Alternates

Other deployment options.

A smaller model for batching and longer context

70B at supported quantization after checking its full allocation

Larger MoE only with explicit sharding or CPU offload

03 / More GPUs

Compare another card.

RTX 3060 12GBRTX 4060 Ti 16GBRTX 5080 16GBRTX 3090 24GBRTX 4090 24GB