Qwen3-8B (versioned baseline)
Use Q5 where supported; context and batch must be measured. Estimate the full runtime allocation before deployment. This candidate has not been measured here at a fixed context length or batch size.
Weight-only estimate: 5.1 GB at 5 bits, before scales, unquantized tensors, vision components, KV cache and runtime buffers.
Official model card gives 8.2B total language-model parameters. Fit labels describe estimated weight headroom only. No tokens-per-second, maximum-context or relative-quality measurement is claimed. Official model card ↗