Board contents
Text extracted from this public whiteboard for search and accessibility.
S10 · Open
Session 10 · Causal attention · see only the past
Live board · deepen in matheion MA-LLM · Session 10
Today’s story When predicting the next token, you must not peek at future tokens. Hide illegal scores before the softmax — that also unlocks parallel training.
You should leave able to… • Explain why full attention cheats • Show mask-before-softmax on a tiny grid • Say how the mask lets training run all positions together
With matheion Use this board to teach. Open matheion → MA-LLM → Session 10 for full prose, diagrams, and the auto-quiz.
S10 · Why mask
Full attention leaks the future
At generation time When predicting the next piece, only earlier pieces exist. If training lets position 5 read position 8, the model cheats with answers it won’t have later.
Demo sentence “The animal didn’t cross the street because it was too tired.” At the word ‘it’, later words must be invisible.
S10 · How to mask
Hide illegal scores, then softmax
For each pair (asker i, candidate j): score = query_i · key_j if j is in the future (j > i): score = a huge negative # so softmax weight ≈ 0 then softmax the row Allowed positions keep weights that still sum to 1. Future positions get weight ~0.
Pitfall Don’t zero weights after softmax and forget to renormalise. Always hide on the raw scores first.
S10 · Training gift
Why the mask helps training
With the mask The vector at position i only depends on positions ≤ i. So you can score “predict next token” at every position in one forward pass — no need to loop token-by-token while training.
Preview Session 14: generating text still walks forward one token at a time (you can reuse past keys/values to avoid redo work).
S10 · Glossary
Glossary · key concepts this session
Keep this frame visible while teaching · say the term, then the plain line
Causal / autoregressive Position i may only use positions ≤ i — no peeking at the future.
Causal mask Hide illegal (future) scores before softmax so their weights ≈ 0.
Mask before softmax Put a huge negative on forbidden scores, then softmax the row.
Teacher-forcing training With the mask, score every position’s next token in one parallel forward pass.
Generation-time constraint At test time only past tokens exist — training must match that.
S10 · Check
Exit ticket
Say these aloud · then quiz in matheion MA-LLM · Session 10
Prompt 1 Why is full self-attention illegal for next-token models?
Prompt 2 Do you hide before or after softmax — and why?
Prompt 3 How does the mask enable parallel training?