chalkline

Public teaching whiteboard

Why attention beats CNNs and RNNs for language

The motivation for attention: what CNNs and RNNs miss on long sequences — free LLM whiteboard. · by zlu

Topics: LLM & Transformers

Why attention beats CNNs and RNNs for language
S7 · Open
S7 · Hard sentence
S7 · Two almost-answers
S7 · Checklist
S7 · Glossary
S7 · Check

Session 7 · Why older sequence models fall short

Live board · deepen in matheion MA-LLM · Session 7

Today’s story Sessions 5–6 built CNNs and RNNs. Now score them on language: long-range, meaning-chosen links. Attention (Session 8+) will hit the checklist.

You should leave able to… • Retell the Alice / She problem • Score the CNN (Session 5) on the checklist • Score the RNN (Session 6) on the checklist

With matheion Use this board to teach. Open matheion → MA-LLM → Session 7 for full prose, diagrams, and the auto-quiz.

The dependency problem

Example “Alice and Bob introduced themselves. She said my name is ___.” Most likely fill: Alice. But that depends on meaning (who ‘She’ refers to), and Alice can sit arbitrarily far back in the text.

We need a mechanism that… 1) reaches any distance 2) chooses which earlier word by meaning 3) does it in few steps 4) can run many positions together when training

Two almost-answers that fail

Sliding window (1D convolution) Look only at nearby neighbours with a fixed window. To reach 20 tokens back you stack many layers. Worse: which past word matters is not a fixed offset — it’s semantic (‘She’ → ‘Alice’, not ‘word − 7’).

Step-by-step recurrence (RNN / LSTM) Carry a summary forward one token at a time. In theory it can remember far back, but: • must process left-to-right (hard to parallelise) • long chains make learning fragile

Checklist attention will hit

Reach Any token can touch any other in one hop.

Choice Weights depend on content (query·key), not a fixed offset.

Parallelism All positions compute together in training.

Short path Far link ≠ deep stack of layers.

Glossary · key concepts this session

Keep this frame visible while teaching · say the term, then the plain line

Long-range dependency A word that needs information from far earlier (or later) in the text.

1D convolution / sliding window Only look at a fixed neighbourhood of nearby tokens.

RNN / LSTM Step-by-step recurrence: update a summary one token at a time, left to right.

Parallelism Compute many positions together instead of waiting token-by-token.

Path length How many steps a signal must travel between two positions in the net.

Attention checklist Reach any distance · choose by meaning · few steps · parallel training.

Exit ticket

Say these aloud · then quiz in matheion MA-LLM · Session 7

Prompt 1 Retell Alice/She and why a fixed window struggles.

Prompt 2 One sentence: sliding-window failure. One sentence: recurrence failure.

Prompt 3 Recite the four checklist items.

Drag to pan · scroll to zoom · read-only

Board contents

Text extracted from this public whiteboard for search and accessibility.

S7 · Open

Session 7 · Why older sequence models fall short

Live board · deepen in matheion MA-LLM · Session 7

Today’s story Sessions 5–6 built CNNs and RNNs. Now score them on language: long-range, meaning-chosen links. Attention (Session 8+) will hit the checklist.

You should leave able to… • Retell the Alice / She problem • Score the CNN (Session 5) on the checklist • Score the RNN (Session 6) on the checklist

With matheion Use this board to teach. Open matheion → MA-LLM → Session 7 for full prose, diagrams, and the auto-quiz.

S7 · Hard sentence

The dependency problem

Example “Alice and Bob introduced themselves. She said my name is ___.” Most likely fill: Alice. But that depends on meaning (who ‘She’ refers to), and Alice can sit arbitrarily far back in the text.

We need a mechanism that… 1) reaches any distance 2) chooses which earlier word by meaning 3) does it in few steps 4) can run many positions together when training

S7 · Two almost-answers

Two almost-answers that fail

Sliding window (1D convolution) Look only at nearby neighbours with a fixed window. To reach 20 tokens back you stack many layers. Worse: which past word matters is not a fixed offset — it’s semantic (‘She’ → ‘Alice’, not ‘word − 7’).

Step-by-step recurrence (RNN / LSTM) Carry a summary forward one token at a time. In theory it can remember far back, but: • must process left-to-right (hard to parallelise) • long chains make learning fragile

S7 · Checklist

Checklist attention will hit

Reach Any token can touch any other in one hop.

Choice Weights depend on content (query·key), not a fixed offset.

Parallelism All positions compute together in training.

Short path Far link ≠ deep stack of layers.

S7 · Glossary

Glossary · key concepts this session

Keep this frame visible while teaching · say the term, then the plain line

Long-range dependency A word that needs information from far earlier (or later) in the text.

1D convolution / sliding window Only look at a fixed neighbourhood of nearby tokens.

RNN / LSTM Step-by-step recurrence: update a summary one token at a time, left to right.

Parallelism Compute many positions together instead of waiting token-by-token.

Path length How many steps a signal must travel between two positions in the net.

Attention checklist Reach any distance · choose by meaning · few steps · parallel training.

S7 · Check

Exit ticket

Say these aloud · then quiz in matheion MA-LLM · Session 7

Prompt 1 Retell Alice/She and why a fixed window struggles.

Prompt 2 One sentence: sliding-window failure. One sentence: recurrence failure.

Prompt 3 Recite the four checklist items.