Board contents
Text extracted from this public whiteboard for search and accessibility.
S14 · Open
Session 14 · Training vs generation · why self-attention
Live board · deepen in matheion MA-LLM · Session 14
Today’s story Generating walks forward one token at a time. Training with a causal mask scores every position together. Self-attention wins on parallelism — and pays a quadratic bill.
You should leave able to… • Contrast the two loops • Compare reach / parallelism / cost in plain words • Say when quadratic cost hurts
With matheion Use this board to teach. Open matheion → MA-LLM → Session 14 for full prose, diagrams, and the auto-quiz.
S14 · Two loops
Generation vs training
GENERATION (must walk forward): for each new token: run all layers on the text so far draw the next token (Session 3) append it Tip: store past keys/values so you don’t rebuild them from scratch each step. TRAINING (with causal mask from Session 12): for each layer: … everyone only sees the past … score “predict next token” at EVERY position in one go
S14 · Why self-attn
Why self-attention won (intuition)
Parallelism Self-attention: all pairs in one shot. Recurrence: must finish token 1 before token 2. Training loves parallelism.
Short paths Any two positions talk in one hop. Recurrence needs many steps to carry a signal far — harder to learn.
The bill Cost grows with (number of tokens)² × width. Fine for moderate lengths. Painful for very long documents — hence shorter windows / sparse tricks.
S14 · Trade-off
Honest trade-off
Typical sentences When sequences aren’t huge compared with model width, self-attention is often cheaper and much more parallel than full recurrence.
Very long context Quadratic attention dominates the bill. “How big is the context window?” is mostly asking how much of that bill you can afford.
S14 · Glossary
Glossary · key concepts this session
Keep this frame visible while teaching · say the term, then the plain line
Generation loop Append one sampled token at a time; must walk forward.
Parallel training With a causal mask, score every position’s next token in one forward pass.
KV reuse / cache Store past keys and values when decoding so you don’t rebuild them each step.
Quadratic cost Self-attention cost grows with (number of tokens)² × width.
Context window How many tokens attention can see — limited by that quadratic bill.
Path length (again) Self-attention: one hop between positions; recurrence: many steps.
S14 · Check
Exit ticket
Say these aloud · then quiz in matheion MA-LLM · Session 14
Prompt 1 Draw the generation loop vs the training loop.
Prompt 2 Three reasons self-attention beat recurrence here.
Prompt 3 When does the (tokens)² cost become the wrong tool?