B = BILLION LEARNED WEIGHTS
NOT FACTS · TOKENS · NEURONS · LINES OF CODE
Count is size, not quality. Capability depends on architecture,
data, training, quantization, and the serving stack. Weight file = params × bits ÷ 8
(exact, 1 GB = 10⁹ bytes); running a model needs more (KV cache, activations).
The MoE split (8 experts, top-2) and the training ×3 to ×8 band are ILLUSTRATIVE.