eqho
/
LLM Lab
model
lab-tiny
the run
09 / 09
Lesson 09
Inference in production
Why the first word is slow and the rest are fast.
01
How big is a model?
02
Text becomes tokens
03
Tokens become vectors
04
Attention
05
The transformer block
06
Predicting the next token
07
Training
08
From autocomplete to assistant
09
Inference in production
Prefill
Decode
The KV cache
The context window
01 / 01
Prefill
The whole prompt runs through the model once. That is time-to-first-token.
←
→
space