Assumptions. A dense decoder-only transformer. Quantized bits include a
scale and metadata allowance (INT8 about 8.5 bits, INT4 about 4.85 bits per weight, in line
with common GGUF and AWQ style formats). GQA is the attention-head to KV-head ratio; classic
multi-head attention is 1, most current models are 4 to 8. 1 GB here means 1024³ bytes,
matching how VRAM is reported. 8-bit KV cache is shown conceptually; support and quality
vary by runtime.
Why real numbers differ. Frameworks, kernels, model architecture,
quantization layout, KV cache paging, multimodal components, speculative decoding drafts,
and OS or display use all shift the true requirement. Treat this as a planning estimate,
leave headroom, and verify with your runtime before you rely on it.