chalkline

Public teaching whiteboard

RNNs for sequences — sequential NLP models

Recurrent networks for language: how they read sequences and where they struggle — public teaching board. · by zlu

Topics: LLM & Transformers

RNNs for sequences — sequential NLP models
S6 · Open
S6 · Idea
S6 · Hand calc
S6 · Three costs
S6 · LSTM note
S6 · Glossary
S6 · Check

Session 6 · Recurrent networks for sequences

Live board · deepen in matheion MA-LLM · Session 6

Today’s story Keep a memory vector and update it as each token arrives: h_t = f(h_{t−1}, x_t). Unbounded reach in principle — but sequential, bottlenecked, and long-pathed.

You should leave able to… • Unroll three toy hidden-state steps by hand • Count hops from Alice to She • Name the three structural costs

With matheion Use this board to teach. Open matheion → MA-LLM → Session 6 for full prose, diagrams, and the auto-quiz.

A memory that walks left to right

Update rule h_t = f(h_{t−1}, x_t) h_0 = zeros (usually) x_t = embedding at step t h_t = new memory after reading x_t Same f at every step — sharing across time.

Unrolled picture x1 → f → h1 ↓ x2 → f → h2 ↓ x3 → f → h3 Each arrow to the next f is a sequential dependency. You cannot fill h1,h2,h3 of one layer all at once.

Hand calc · scalar memory

Toy: h_t = tanh(0.5·h_{t−1} + x_t) · h_0 = 0 · x = [1, 0, 1]

text
t=1:  input = 0.5·0 + 1 = 1        → h1 = tanh(1) ≈ 0.76
t=2:  input = 0.5·0.76 + 0 ≈ 0.38  → h2 = tanh(0.38) ≈ 0.36
t=3:  input = 0.5·0.36 + 1 ≈ 1.18  → h3 = tanh(1.18) ≈ 0.83

Zero x1 and rerun → h3 changes.
So step 1 still influences step 3 through the chain — “memory”.

Real models: h_t is a vector of width hundreds/thousands; f is a small net.
LSTM/GRU = same skeleton with gates that learn to keep or forget.

Three structural costs

These are why Session 7 will say RNNs struggle at language scale

1 · Sequential h_t needs h_{t−1}. S tokens ⇒ S serial steps. No parallel fill across the sequence.

2 · Bottleneck All of the past must fit in one vector h_t. Many entities compete for the same slots — easy to overwrite Alice.

3 · Long path Alice at pos 1 → She at pos 7: h1 → h2 → … → h7 (six hops). Gradients shrink along the chain. Far links are fragile to learn.

LSTM / GRU in one card

Still recurrent Gates learn to copy or erase pieces of memory. Better long memory in practice. Does NOT remove: sequential steps, one-vector funnel, hop count ≈ distance.

What they get right • Unbounded reach in principle • One f for any length • Natural for streaming text Next session: score CNN vs RNN on Alice/She.

Glossary · key concepts this session

Keep this frame visible while teaching · say the term, then the plain line

RNN Recurrent net: update a hidden state one token at a time with a shared f.

Hidden state h_t The memory vector after reading token t.

Unrolling Drawing the same f once per time step left to right.

Sequential bottleneck Cannot compute all positions of a layer in parallel.

Path length Number of hidden-state hops between two positions (≈ distance).

LSTM / GRU Gated RNNs that learn keep/forget — still sequential recurrence.

Exit ticket

Say these aloud · then quiz in matheion MA-LLM · Session 6

Prompt 1 With h_t=tanh(0.5 h_{t−1}+x_t), h_0=0, x=[0,1] — what is h_2≈?

Prompt 2 Alice at 2, She at 9 — about how many hops?

Prompt 3 Name the three structural costs of RNNs.

Drag to pan · scroll to zoom · read-only

Board contents

Text extracted from this public whiteboard for search and accessibility.

S6 · Open

Session 6 · Recurrent networks for sequences

Live board · deepen in matheion MA-LLM · Session 6

Today’s story Keep a memory vector and update it as each token arrives: h_t = f(h_{t−1}, x_t). Unbounded reach in principle — but sequential, bottlenecked, and long-pathed.

You should leave able to… • Unroll three toy hidden-state steps by hand • Count hops from Alice to She • Name the three structural costs

With matheion Use this board to teach. Open matheion → MA-LLM → Session 6 for full prose, diagrams, and the auto-quiz.

S6 · Idea

A memory that walks left to right

Update rule h_t = f(h_{t−1}, x_t) h_0 = zeros (usually) x_t = embedding at step t h_t = new memory after reading x_t Same f at every step — sharing across time.

Unrolled picture x1 → f → h1 ↓ x2 → f → h2 ↓ x3 → f → h3 Each arrow to the next f is a sequential dependency. You cannot fill h1,h2,h3 of one layer all at once.

S6 · Hand calc

Hand calc · scalar memory

Toy: h_t = tanh(0.5·h_{t−1} + x_t) · h_0 = 0 · x = [1, 0, 1]

t=1: input = 0.5·0 + 1 = 1 → h1 = tanh(1) ≈ 0.76 t=2: input = 0.5·0.76 + 0 ≈ 0.38 → h2 = tanh(0.38) ≈ 0.36 t=3: input = 0.5·0.36 + 1 ≈ 1.18 → h3 = tanh(1.18) ≈ 0.83 Zero x1 and rerun → h3 changes. So step 1 still influences step 3 through the chain — “memory”. Real models: h_t is a vector of width hundreds/thousands; f is a small net. LSTM/GRU = same skeleton with gates that learn to keep or forget.

S6 · Three costs

Three structural costs

These are why Session 7 will say RNNs struggle at language scale

1 · Sequential h_t needs h_{t−1}. S tokens ⇒ S serial steps. No parallel fill across the sequence.

2 · Bottleneck All of the past must fit in one vector h_t. Many entities compete for the same slots — easy to overwrite Alice.

3 · Long path Alice at pos 1 → She at pos 7: h1 → h2 → … → h7 (six hops). Gradients shrink along the chain. Far links are fragile to learn.

S6 · LSTM note

LSTM / GRU in one card

Still recurrent Gates learn to copy or erase pieces of memory. Better long memory in practice. Does NOT remove: sequential steps, one-vector funnel, hop count ≈ distance.

What they get right • Unbounded reach in principle • One f for any length • Natural for streaming text Next session: score CNN vs RNN on Alice/She.

S6 · Glossary

Glossary · key concepts this session

Keep this frame visible while teaching · say the term, then the plain line

RNN Recurrent net: update a hidden state one token at a time with a shared f.

Hidden state h_t The memory vector after reading token t.

Unrolling Drawing the same f once per time step left to right.

Sequential bottleneck Cannot compute all positions of a layer in parallel.

Path length Number of hidden-state hops between two positions (≈ distance).

LSTM / GRU Gated RNNs that learn keep/forget — still sequential recurrence.

S6 · Check

Exit ticket

Say these aloud · then quiz in matheion MA-LLM · Session 6

Prompt 1 With h_t=tanh(0.5 h_{t−1}+x_t), h_0=0, x=[0,1] — what is h_2≈?

Prompt 2 Alice at 2, She at 9 — about how many hops?

Prompt 3 Name the three structural costs of RNNs.