Model fit guide / H200 141GBWeight-based fit estimateReviewed October 7, 2026
141 GB VRAM - Serving candidate

Local AI model fit for H200 141GB.

Use this versioned model as a fit candidate, then evaluate it on your workload. Its 35B total language-model parameters imply about 70.0 GB of raw weights at 16 bits. This is an analytical estimate, not a measured allocation or a claim that this model is the latest quality winner.

01 / Recommendation

Run this size class.

Candidate to evaluate

Qwen3.6-35B-A3B (evaluation candidate)

Use BF16; context and batch must be measured. Estimate the full runtime allocation before deployment. This candidate has not been measured here at a fixed context length or batch size.

Weight estimate

Weight-only estimate: 70.0 GB at 16 bits, before scales, unquantized tensors, vision components, KV cache and runtime buffers.

Evidence

Official model card gives 35B total language-model parameters. Fit labels describe estimated weight headroom only. No tokens-per-second, maximum-context or relative-quality measurement is claimed. Official model card ↗

02 / Alternates

Other deployment options.

A smaller model for batching and longer context

70B at supported quantization after checking its full allocation

Larger MoE only with explicit sharding or CPU offload

03 / More GPUs

Compare another card.

RTX 3060 12GBRTX 4060 Ti 16GBRTX 5080 16GBRTX 3090 24GBRTX 4090 24GB