| RTX 3060 12GB | 12 GB | Qwen3-8B (versioned baseline) | Q5 where supported; context and batch must be measured | tight estimate | Model card ↗ Weight-only estimate: 5.1 GB at 5 bits, before scales, unquantized tensors, vision components, KV cache and runtime buffers. |
| RTX 4060 Ti 16GB | 16 GB | Qwen3-14B (versioned baseline) | Q4 where supported; context and batch must be measured | tight estimate | Model card ↗ Weight-only estimate: 7.4 GB at 4 bits, before scales, unquantized tensors, vision components, KV cache and runtime buffers. |
| RTX 5080 16GB | 16 GB | Qwen3-14B (versioned baseline) | Q4 where supported; context and batch must be measured | tight estimate | Model card ↗ Weight-only estimate: 7.4 GB at 4 bits, before scales, unquantized tensors, vision components, KV cache and runtime buffers. |
| RTX 3090 24GB | 24 GB | Qwen3.6-35B-A3B (evaluation candidate) | Q4 where supported; context and batch must be measured | tight estimate | Model card ↗ Weight-only estimate: 17.5 GB at 4 bits, before scales, unquantized tensors, vision components, KV cache and runtime buffers. |
| RTX 4090 24GB | 24 GB | Qwen3.6-35B-A3B (evaluation candidate) | Q4 where supported; context and batch must be measured | tight estimate | Model card ↗ Weight-only estimate: 17.5 GB at 4 bits, before scales, unquantized tensors, vision components, KV cache and runtime buffers. |
| RTX 5090 32GB | 32 GB | Qwen3.6-35B-A3B (evaluation candidate) | Q6 where supported; context and batch must be measured | tight estimate | Model card ↗ Weight-only estimate: 26.3 GB at 6 bits, before scales, unquantized tensors, vision components, KV cache and runtime buffers. |
| A100 40GB | 40 GB | Qwen3.6-35B-A3B (evaluation candidate) | Q6 where supported; context and batch must be measured | comfortable estimate | Model card ↗ Weight-only estimate: 26.3 GB at 6 bits, before scales, unquantized tensors, vision components, KV cache and runtime buffers. |
| A100 80GB | 80 GB | Qwen3.6-35B-A3B (evaluation candidate) | BF16; context and batch must be measured | comfortable estimate | Model card ↗ Weight-only estimate: 70.0 GB at 16 bits, before scales, unquantized tensors, vision components, KV cache and runtime buffers. |
| H100 80GB | 80 GB | Qwen3.6-35B-A3B (evaluation candidate) | BF16; context and batch must be measured | comfortable estimate | Model card ↗ Weight-only estimate: 70.0 GB at 16 bits, before scales, unquantized tensors, vision components, KV cache and runtime buffers. |
| H200 141GB | 141 GB | Qwen3.6-35B-A3B (evaluation candidate) | BF16; context and batch must be measured | comfortable estimate | Model card ↗ Weight-only estimate: 70.0 GB at 16 bits, before scales, unquantized tensors, vision components, KV cache and runtime buffers. |
| B200 | 192 GB | Qwen3.6-35B-A3B (evaluation candidate) | BF16; context and batch must be measured | comfortable estimate | Model card ↗ Weight-only estimate: 70.0 GB at 16 bits, before scales, unquantized tensors, vision components, KV cache and runtime buffers. |
| MI300X 192GB | 192 GB | Qwen3.6-35B-A3B (evaluation candidate) | BF16; context and batch must be measured | comfortable estimate | Model card ↗ Weight-only estimate: 70.0 GB at 16 bits, before scales, unquantized tensors, vision components, KV cache and runtime buffers. |