chalkline

Public teaching whiteboard

Transformer architecture stack explained

The full Transformer stack: layers, residuals, and feed-forward blocks — free LLM whiteboard. · by zlu

Topics: LLM & Transformers

Transformer architecture stack explained
S13 · Open
S13 · Blocks
S13 · Three attentions
S13 · Per-token net
S13 · Ends
S13 · Glossary
S13 · Check

Session 13 · The Transformer stack

Live board · deepen in matheion MA-LLM · Session 13

Today’s story Encoder and decoder are repeated blocks: attention, a small per-token network, and residual+normalise. Three attention uses appear in one translation model.

You should leave able to… • Sketch encoder vs decoder blocks in words • Name the three attention uses • Say what the per-token network adds

With matheion Use this board to teach. Open matheion → MA-LLM → Session 13 for full prose, diagrams, and the auto-quiz.

Encoder block vs decoder block

Encoder (repeat N times) 1) Self-attention (full — every source token reads every source token) 2) Add residual + normalise 3) Small per-token network (two linear layers + nonlinearity) 4) Add residual + normalise Builds a rich representation of the source.

Decoder (repeat N times) 1) Self-attention with causal mask (Session 12) 2) Residual + normalise 3) Cross-attention: queries from decoder, keys/values from encoder 4) Residual + normalise 5) Per-token network + residual + normalise

Three uses of multi-head attention

Encoder self Ask/match/deliver all from the source side. Every source token reads every source token.

Decoder self (masked) Target side only, with future hidden. Keeps generation honest.

Cross (encoder–decoder) Queries from the target side; keys/values from the source memory. ‘What to say’ looks up ‘what was read’.

The small per-token network

Often called a feed-forward network (FFN): same weights at every position, different from attention

text
For each token vector x (separately):
  hidden = relu( x × W1 + b1 )     # expand width (paper: 512 → 2048)
  out    = hidden × W2 + b2        # back to model width (2048 → 512)

Attention mixes information ACROSS tokens.
This network mixes features INSIDE one token.

Residual: output_of_step = x + sublayer(x)
then normalise — keeps deep stacks trainable.

Paper base: model width 512, 6 encoder + 6 decoder layers, 8 heads.

In and out of the stack

In Token embeddings + position vectors (Sessions 4 & 12) feed the first layer.

Out Final linear layer → raw scores over the dictionary → softmax lottery (Session 2). Paper often reuses the embedding table weights here.

Glossary · key concepts this session

Keep this frame visible while teaching · say the term, then the plain line

Encoder Stack that reads the source with full self-attention.

Decoder Stack that generates the target with masked self-attention + cross-attention.

Residual connection output = x + sublayer(x) — keeps an identity path for training deep nets.

Layer normalisation Stabilises activations after sublayers (often written with the residual).

Feed-forward / per-token net (FFN) Same small MLP at every position; mixes features inside one token.

Encoder–decoder attention Decoder queries look into encoder keys/values (the source memory).

Weight tying Reuse embedding table weights for the final vocabulary projection.

Exit ticket

Say these aloud · then quiz in matheion MA-LLM · Session 13

Prompt 1 Sketch encoder vs decoder steps in words.

Prompt 2 Name where queries vs keys/values come from in all three attentions.

Prompt 3 What does the per-token network do that attention doesn’t?

Drag to pan · scroll to zoom · read-only

Board contents

Text extracted from this public whiteboard for search and accessibility.

S13 · Open

Session 13 · The Transformer stack

Live board · deepen in matheion MA-LLM · Session 13

Today’s story Encoder and decoder are repeated blocks: attention, a small per-token network, and residual+normalise. Three attention uses appear in one translation model.

You should leave able to… • Sketch encoder vs decoder blocks in words • Name the three attention uses • Say what the per-token network adds

With matheion Use this board to teach. Open matheion → MA-LLM → Session 13 for full prose, diagrams, and the auto-quiz.

S13 · Blocks

Encoder block vs decoder block

Encoder (repeat N times) 1) Self-attention (full — every source token reads every source token) 2) Add residual + normalise 3) Small per-token network (two linear layers + nonlinearity) 4) Add residual + normalise Builds a rich representation of the source.

Decoder (repeat N times) 1) Self-attention with causal mask (Session 12) 2) Residual + normalise 3) Cross-attention: queries from decoder, keys/values from encoder 4) Residual + normalise 5) Per-token network + residual + normalise

S13 · Three attentions

Three uses of multi-head attention

Encoder self Ask/match/deliver all from the source side. Every source token reads every source token.

Decoder self (masked) Target side only, with future hidden. Keeps generation honest.

Cross (encoder–decoder) Queries from the target side; keys/values from the source memory. ‘What to say’ looks up ‘what was read’.

S13 · Per-token net

The small per-token network

Often called a feed-forward network (FFN): same weights at every position, different from attention

For each token vector x (separately): hidden = relu( x × W1 + b1 ) # expand width (paper: 512 → 2048) out = hidden × W2 + b2 # back to model width (2048 → 512) Attention mixes information ACROSS tokens. This network mixes features INSIDE one token. Residual: output_of_step = x + sublayer(x) then normalise — keeps deep stacks trainable. Paper base: model width 512, 6 encoder + 6 decoder layers, 8 heads.

S13 · Ends

In and out of the stack

In Token embeddings + position vectors (Sessions 4 & 12) feed the first layer.

Out Final linear layer → raw scores over the dictionary → softmax lottery (Session 2). Paper often reuses the embedding table weights here.

S13 · Glossary

Glossary · key concepts this session

Keep this frame visible while teaching · say the term, then the plain line

Encoder Stack that reads the source with full self-attention.

Decoder Stack that generates the target with masked self-attention + cross-attention.

Residual connection output = x + sublayer(x) — keeps an identity path for training deep nets.

Layer normalisation Stabilises activations after sublayers (often written with the residual).

Feed-forward / per-token net (FFN) Same small MLP at every position; mixes features inside one token.

Encoder–decoder attention Decoder queries look into encoder keys/values (the source memory).

Weight tying Reuse embedding table weights for the final vocabulary projection.

S13 · Check

Exit ticket

Say these aloud · then quiz in matheion MA-LLM · Session 13

Prompt 1 Sketch encoder vs decoder steps in words.

Prompt 2 Name where queries vs keys/values come from in all three attentions.

Prompt 3 What does the per-token network do that attention doesn’t?