Hardware · reviewed October 7, 2026

Which AI model fits your GPU?

Start with total weight memory, then budget for the runtime, context and concurrent users. The candidates below are versioned examples for testing; the fit labels describe estimated weight headroom, not measured throughput or a benchmark winner.

Fit matrix
Raw weight memory, before runtime overhead0 GB20 GB40 GB60 GB80 GB100 GB120 GB140 GB8.2B16.48.24.114.8B29.614.87.435B70.035.017.570B140.070.035.0BF16: 2 bytes/parameter · INT8: 1 · 4-bit: 0.5
Arithmetic lower bounds in decimal GB, not measured VRAM. Quantization metadata, unquantized tensors, vision encoders, KV cache and serving buffers add memory. A 1T model needs about 500 GB even at 4-bit, beyond a single 192 GB card.

Single-card fit candidates

GPUVRAMCandidatePrecisionWeight headroomEvidence
RTX 3060 12GB12 GBQwen3-8B (versioned baseline)Q5 where supported; context and batch must be measuredtight estimateModel card ↗

Weight-only estimate: 5.1 GB at 5 bits, before scales, unquantized tensors, vision components, KV cache and runtime buffers.

RTX 4060 Ti 16GB16 GBQwen3-14B (versioned baseline)Q4 where supported; context and batch must be measuredtight estimateModel card ↗

Weight-only estimate: 7.4 GB at 4 bits, before scales, unquantized tensors, vision components, KV cache and runtime buffers.

RTX 5080 16GB16 GBQwen3-14B (versioned baseline)Q4 where supported; context and batch must be measuredtight estimateModel card ↗

Weight-only estimate: 7.4 GB at 4 bits, before scales, unquantized tensors, vision components, KV cache and runtime buffers.

RTX 3090 24GB24 GBQwen3.6-35B-A3B (evaluation candidate)Q4 where supported; context and batch must be measuredtight estimateModel card ↗

Weight-only estimate: 17.5 GB at 4 bits, before scales, unquantized tensors, vision components, KV cache and runtime buffers.

RTX 4090 24GB24 GBQwen3.6-35B-A3B (evaluation candidate)Q4 where supported; context and batch must be measuredtight estimateModel card ↗

Weight-only estimate: 17.5 GB at 4 bits, before scales, unquantized tensors, vision components, KV cache and runtime buffers.

RTX 5090 32GB32 GBQwen3.6-35B-A3B (evaluation candidate)Q6 where supported; context and batch must be measuredtight estimateModel card ↗

Weight-only estimate: 26.3 GB at 6 bits, before scales, unquantized tensors, vision components, KV cache and runtime buffers.

A100 40GB40 GBQwen3.6-35B-A3B (evaluation candidate)Q6 where supported; context and batch must be measuredcomfortable estimateModel card ↗

Weight-only estimate: 26.3 GB at 6 bits, before scales, unquantized tensors, vision components, KV cache and runtime buffers.

A100 80GB80 GBQwen3.6-35B-A3B (evaluation candidate)BF16; context and batch must be measuredcomfortable estimateModel card ↗

Weight-only estimate: 70.0 GB at 16 bits, before scales, unquantized tensors, vision components, KV cache and runtime buffers.

H100 80GB80 GBQwen3.6-35B-A3B (evaluation candidate)BF16; context and batch must be measuredcomfortable estimateModel card ↗

Weight-only estimate: 70.0 GB at 16 bits, before scales, unquantized tensors, vision components, KV cache and runtime buffers.

H200 141GB141 GBQwen3.6-35B-A3B (evaluation candidate)BF16; context and batch must be measuredcomfortable estimateModel card ↗

Weight-only estimate: 70.0 GB at 16 bits, before scales, unquantized tensors, vision components, KV cache and runtime buffers.

B200192 GBQwen3.6-35B-A3B (evaluation candidate)BF16; context and batch must be measuredcomfortable estimateModel card ↗

Weight-only estimate: 70.0 GB at 16 bits, before scales, unquantized tensors, vision components, KV cache and runtime buffers.

MI300X 192GB192 GBQwen3.6-35B-A3B (evaluation candidate)BF16; context and batch must be measuredcomfortable estimateModel card ↗

Weight-only estimate: 70.0 GB at 16 bits, before scales, unquantized tensors, vision components, KV cache and runtime buffers.

Fit is a workload measurement

40 GB does not hold 35B BF16

35 billion parameters at two bytes each need about 70 GB before any overhead. Quantization, sharding or offload is required on a 40 GB A100. Check backend support rather than assuming every GPU accelerates every precision.

MoE memory uses all experts

Kimi K2.6 has 1T total parameters and 32B active. Raw 4-bit weights alone are about 500 GB; FP8 about 1 TB. It cannot reside entirely on one 141 or 192 GB GPU at those precisions.

Kimi model card ↗

Measure peak allocated memory at the intended prompt length, output length and concurrency. Reserve headroom for vision inputs, KV cache, quantization metadata and runtime buffers. Published maximum context is not a promise that it fits your card.

The previous composite-quality and relative-throughput bubbles were withdrawn because they were not measured results.

Keep your own fit notes

Notes and comments in this board are stored in this browser. They are not submitted to a public moderation queue.

GPU cards
Current pick

Qwen3-8B (versioned baseline)

RTX 3060 12GB - Q5 where supported; context and batch must be measured - tight fit

12 GB

Edits are saved in this browser. Use comments below to send corrections for moderation.

Use it for
local inferenceworkload evaluation
Alternates
  • A smaller model for more KV-cache headroom
  • A newer same-size model after workload evaluation
Avoid

Active MoE parameters describe compute, not resident weights. Kimi K2.6 has 1T total parameters: raw 4-bit weights are about 500 GB and FP8 about 1 TB, exceeding a single 141 or 192 GB device.

Comments

What are you actually running on RTX 3060 12GB?

0 local comments
Local-first comments for this prototype.

No local comments yet for this GPU.