eqho/LLM Lab
the run
09 / 09
Lesson 09

Inference in production

Why the first word is slow and the rest are fast.

01 / 01
Prefill
The whole prompt runs through the model once. That is time-to-first-token.