chalkline

Public teaching whiteboard

Causal attention in language models

Why GPT-style models use causal (masked) attention so tokens only see the past — public whiteboard. · by zlu

Topics: LLM & Transformers

Causal attention in language models
S10 · Open
S10 · Why mask
S10 · How to mask
S10 · Training gift
S10 · Glossary
S10 · Check

Session 10 · Causal attention · see only the past

Live board · deepen in matheion MA-LLM · Session 10

Today’s story When predicting the next token, you must not peek at future tokens. Hide illegal scores before the softmax — that also unlocks parallel training.

You should leave able to… • Explain why full attention cheats • Show mask-before-softmax on a tiny grid • Say how the mask lets training run all positions together

With matheion Use this board to teach. Open matheion → MA-LLM → Session 10 for full prose, diagrams, and the auto-quiz.

Full attention leaks the future

At generation time When predicting the next piece, only earlier pieces exist. If training lets position 5 read position 8, the model cheats with answers it won’t have later.

Demo sentence “The animal didn’t cross the street because it was too tired.” At the word ‘it’, later words must be invisible.

Hide illegal scores, then softmax

text
For each pair (asker i, candidate j):
  score = query_i · key_j
  if j is in the future (j > i):
      score = a huge negative   # so softmax weight ≈ 0
  then softmax the row

Allowed positions keep weights that still sum to 1.
Future positions get weight ~0.

Pitfall Don’t zero weights after softmax and forget to renormalise. Always hide on the raw scores first.

Why the mask helps training

With the mask The vector at position i only depends on positions ≤ i. So you can score “predict next token” at every position in one forward pass — no need to loop token-by-token while training.

Preview Session 14: generating text still walks forward one token at a time (you can reuse past keys/values to avoid redo work).

Glossary · key concepts this session

Keep this frame visible while teaching · say the term, then the plain line

Causal / autoregressive Position i may only use positions ≤ i — no peeking at the future.

Causal mask Hide illegal (future) scores before softmax so their weights ≈ 0.

Mask before softmax Put a huge negative on forbidden scores, then softmax the row.

Teacher-forcing training With the mask, score every position’s next token in one parallel forward pass.

Generation-time constraint At test time only past tokens exist — training must match that.

Exit ticket

Say these aloud · then quiz in matheion MA-LLM · Session 10

Prompt 1 Why is full self-attention illegal for next-token models?

Prompt 2 Do you hide before or after softmax — and why?

Prompt 3 How does the mask enable parallel training?

Drag to pan · scroll to zoom · read-only

Board contents

Text extracted from this public whiteboard for search and accessibility.

S10 · Open

Session 10 · Causal attention · see only the past

Live board · deepen in matheion MA-LLM · Session 10

Today’s story When predicting the next token, you must not peek at future tokens. Hide illegal scores before the softmax — that also unlocks parallel training.

You should leave able to… • Explain why full attention cheats • Show mask-before-softmax on a tiny grid • Say how the mask lets training run all positions together

With matheion Use this board to teach. Open matheion → MA-LLM → Session 10 for full prose, diagrams, and the auto-quiz.

S10 · Why mask

Full attention leaks the future

At generation time When predicting the next piece, only earlier pieces exist. If training lets position 5 read position 8, the model cheats with answers it won’t have later.

Demo sentence “The animal didn’t cross the street because it was too tired.” At the word ‘it’, later words must be invisible.

S10 · How to mask

Hide illegal scores, then softmax

For each pair (asker i, candidate j): score = query_i · key_j if j is in the future (j > i): score = a huge negative # so softmax weight ≈ 0 then softmax the row Allowed positions keep weights that still sum to 1. Future positions get weight ~0.

Pitfall Don’t zero weights after softmax and forget to renormalise. Always hide on the raw scores first.

S10 · Training gift

Why the mask helps training

With the mask The vector at position i only depends on positions ≤ i. So you can score “predict next token” at every position in one forward pass — no need to loop token-by-token while training.

Preview Session 14: generating text still walks forward one token at a time (you can reuse past keys/values to avoid redo work).

S10 · Glossary

Glossary · key concepts this session

Keep this frame visible while teaching · say the term, then the plain line

Causal / autoregressive Position i may only use positions ≤ i — no peeking at the future.

Causal mask Hide illegal (future) scores before softmax so their weights ≈ 0.

Mask before softmax Put a huge negative on forbidden scores, then softmax the row.

Teacher-forcing training With the mask, score every position’s next token in one parallel forward pass.

Generation-time constraint At test time only past tokens exist — training must match that.

S10 · Check

Exit ticket

Say these aloud · then quiz in matheion MA-LLM · Session 10

Prompt 1 Why is full self-attention illegal for next-token models?

Prompt 2 Do you hide before or after softmax — and why?

Prompt 3 How does the mask enable parallel training?