chalkline

Public teaching whiteboard

Train vs generate — how language models run

Training time vs generation time for language models — a clear public teaching whiteboard. · by zlu

Topics: LLM & Transformers

Train vs generate — how language models run
S14 · Open
S14 · Two loops
S14 · Why self-attn
S14 · Trade-off
S14 · Glossary
S14 · Check

Session 14 · Training vs generation · why self-attention

Live board · deepen in matheion MA-LLM · Session 14

Today’s story Generating walks forward one token at a time. Training with a causal mask scores every position together. Self-attention wins on parallelism — and pays a quadratic bill.

You should leave able to… • Contrast the two loops • Compare reach / parallelism / cost in plain words • Say when quadratic cost hurts

With matheion Use this board to teach. Open matheion → MA-LLM → Session 14 for full prose, diagrams, and the auto-quiz.

Generation vs training

text
GENERATION (must walk forward):
  for each new token:
      run all layers on the text so far
      draw the next token (Session 3)
      append it
  Tip: store past keys/values so you don’t rebuild them from scratch each step.

TRAINING (with causal mask from Session 12):
  for each layer:
      … everyone only sees the past …
  score “predict next token” at EVERY position in one go

Why self-attention won (intuition)

Parallelism Self-attention: all pairs in one shot. Recurrence: must finish token 1 before token 2. Training loves parallelism.

Short paths Any two positions talk in one hop. Recurrence needs many steps to carry a signal far — harder to learn.

The bill Cost grows with (number of tokens)² × width. Fine for moderate lengths. Painful for very long documents — hence shorter windows / sparse tricks.

Honest trade-off

Typical sentences When sequences aren’t huge compared with model width, self-attention is often cheaper and much more parallel than full recurrence.

Very long context Quadratic attention dominates the bill. “How big is the context window?” is mostly asking how much of that bill you can afford.

Glossary · key concepts this session

Keep this frame visible while teaching · say the term, then the plain line

Generation loop Append one sampled token at a time; must walk forward.

Parallel training With a causal mask, score every position’s next token in one forward pass.

KV reuse / cache Store past keys and values when decoding so you don’t rebuild them each step.

Quadratic cost Self-attention cost grows with (number of tokens)² × width.

Context window How many tokens attention can see — limited by that quadratic bill.

Path length (again) Self-attention: one hop between positions; recurrence: many steps.

Exit ticket

Say these aloud · then quiz in matheion MA-LLM · Session 14

Prompt 1 Draw the generation loop vs the training loop.

Prompt 2 Three reasons self-attention beat recurrence here.

Prompt 3 When does the (tokens)² cost become the wrong tool?

Drag to pan · scroll to zoom · read-only

Board contents

Text extracted from this public whiteboard for search and accessibility.

S14 · Open

Session 14 · Training vs generation · why self-attention

Live board · deepen in matheion MA-LLM · Session 14

Today’s story Generating walks forward one token at a time. Training with a causal mask scores every position together. Self-attention wins on parallelism — and pays a quadratic bill.

You should leave able to… • Contrast the two loops • Compare reach / parallelism / cost in plain words • Say when quadratic cost hurts

With matheion Use this board to teach. Open matheion → MA-LLM → Session 14 for full prose, diagrams, and the auto-quiz.

S14 · Two loops

Generation vs training

GENERATION (must walk forward): for each new token: run all layers on the text so far draw the next token (Session 3) append it Tip: store past keys/values so you don’t rebuild them from scratch each step. TRAINING (with causal mask from Session 12): for each layer: … everyone only sees the past … score “predict next token” at EVERY position in one go

S14 · Why self-attn

Why self-attention won (intuition)

Parallelism Self-attention: all pairs in one shot. Recurrence: must finish token 1 before token 2. Training loves parallelism.

Short paths Any two positions talk in one hop. Recurrence needs many steps to carry a signal far — harder to learn.

The bill Cost grows with (number of tokens)² × width. Fine for moderate lengths. Painful for very long documents — hence shorter windows / sparse tricks.

S14 · Trade-off

Honest trade-off

Typical sentences When sequences aren’t huge compared with model width, self-attention is often cheaper and much more parallel than full recurrence.

Very long context Quadratic attention dominates the bill. “How big is the context window?” is mostly asking how much of that bill you can afford.

S14 · Glossary

Glossary · key concepts this session

Keep this frame visible while teaching · say the term, then the plain line

Generation loop Append one sampled token at a time; must walk forward.

Parallel training With a causal mask, score every position’s next token in one forward pass.

KV reuse / cache Store past keys and values when decoding so you don’t rebuild them each step.

Quadratic cost Self-attention cost grows with (number of tokens)² × width.

Context window How many tokens attention can see — limited by that quadratic bill.

Path length (again) Self-attention: one hop between positions; recurrence: many steps.

S14 · Check

Exit ticket

Say these aloud · then quiz in matheion MA-LLM · Session 14

Prompt 1 Draw the generation loop vs the training loop.

Prompt 2 Three reasons self-attention beat recurrence here.

Prompt 3 When does the (tokens)² cost become the wrong tool?