Simplified on purpose. Real serving systems differ in prefill and decode
scheduling, KV cache memory and paging, batch size vs memory limits, priorities, and
admission control. Real API latency also includes network time and provider load.
All numbers here are simulated token steps, not milliseconds.