Qwen3-14B (versioned baseline)
Use Q4 where supported; context and batch must be measured. Estimate the full runtime allocation before deployment. This candidate has not been measured here at a fixed context length or batch size.
Weight-only estimate: 7.4 GB at 4 bits, before scales, unquantized tensors, vision components, KV cache and runtime buffers.
Official model card gives 14.8B total language-model parameters. Fit labels describe estimated weight headroom only. No tokens-per-second, maximum-context or relative-quality measurement is claimed. Official model card ↗