chalkline

Public teaching whiteboard

Scaled dot-product & multi-head attention

Scaled dot-product attention and multi-head attention — the math behind Transformers on a free board. · by zlu

Topics: LLM & Transformers

Scaled dot-product & multi-head attention
S11 · Open
S11 · Why scale
S11 · Multi-head
S11 · Four steps
S11 · Glossary
S11 · Check

Session 11 · Scaling and multi-head attention

Live board · deepen in matheion MA-LLM · Session 11

Today’s story Two upgrades on Sessions 8–9: shrink huge dot products so softmax stays soft, and run several attention ‘views’ in parallel.

You should leave able to… • Explain why we divide by √(key width) • Motivate several heads with the dog/food sentence • List the four steps of one attention block

With matheion Use this board to teach. Open matheion → MA-LLM → Session 11 for full prose, diagrams, and the auto-quiz.

Why divide the scores?

Key width = how many numbers are in each key/query vector (paper often uses 64)

text
Attention(Q, K, V) =
  softmax(  (Q Kᵀ) / √(key_width)  )  × V

Problem: as key_width grows, raw dots get huge
  → softmax becomes nearly one-hot
  → gradients die.

Fix: divide by √(key_width) so the lottery stays soft enough to learn.

Several heads, then combine

Why several? Sentence: “A dog ate the food because it was hungry.” One head might bind ‘it’ → ‘dog’. Another might track ‘ate’ → ‘food’. A single head averages those jobs into mush.

Recipe 1) Split into several smaller attention views (heads). 2) Each head has its own ask/match/deliver maps. 3) Concatenate the head outputs. 4) One final linear mix. Paper base model: 8 heads, each key width 64, model width 512 (= 8×64).

One attention block · mental model

1 · Project Build queries, keys, values (per head).

2 · Score Dots, then divide by √(key width).

3 · Mask Hide future scores if next-token LM.

4 · Mix Softmax → weighted values; concat heads; final mix.

Glossary · key concepts this session

Keep this frame visible while teaching · say the term, then the plain line

Scaled dot-product Softmax( (Q Kᵀ) / √key_width ) × V — scaling keeps softmax soft.

Key width Number of dimensions in each key/query vector (paper often uses 64).

Multi-head attention Several attention views in parallel, then concatenate and mix.

Head One attention view with its own Q/K/V maps.

Model width Main embedding size of the stack (paper base: 512 = 8 heads × 64).

Saturation Softmax nearly one-hot when scores are huge — gradients die.

Exit ticket

Say these aloud · then quiz in matheion MA-LLM · Session 11

Prompt 1 What goes wrong without dividing by √(key width)?

Prompt 2 Why several heads on the dog/food sentence?

Prompt 3 8 heads, model width 512 → each key width is?

Drag to pan · scroll to zoom · read-only

Board contents

Text extracted from this public whiteboard for search and accessibility.

S11 · Open

Session 11 · Scaling and multi-head attention

Live board · deepen in matheion MA-LLM · Session 11

Today’s story Two upgrades on Sessions 8–9: shrink huge dot products so softmax stays soft, and run several attention ‘views’ in parallel.

You should leave able to… • Explain why we divide by √(key width) • Motivate several heads with the dog/food sentence • List the four steps of one attention block

With matheion Use this board to teach. Open matheion → MA-LLM → Session 11 for full prose, diagrams, and the auto-quiz.

S11 · Why scale

Why divide the scores?

Key width = how many numbers are in each key/query vector (paper often uses 64)

Attention(Q, K, V) = softmax( (Q Kᵀ) / √(key_width) ) × V Problem: as key_width grows, raw dots get huge → softmax becomes nearly one-hot → gradients die. Fix: divide by √(key_width) so the lottery stays soft enough to learn.

S11 · Multi-head

Several heads, then combine

Why several? Sentence: “A dog ate the food because it was hungry.” One head might bind ‘it’ → ‘dog’. Another might track ‘ate’ → ‘food’. A single head averages those jobs into mush.

Recipe 1) Split into several smaller attention views (heads). 2) Each head has its own ask/match/deliver maps. 3) Concatenate the head outputs. 4) One final linear mix. Paper base model: 8 heads, each key width 64, model width 512 (= 8×64).

S11 · Four steps

One attention block · mental model

1 · Project Build queries, keys, values (per head).

2 · Score Dots, then divide by √(key width).

3 · Mask Hide future scores if next-token LM.

4 · Mix Softmax → weighted values; concat heads; final mix.

S11 · Glossary

Glossary · key concepts this session

Keep this frame visible while teaching · say the term, then the plain line

Scaled dot-product Softmax( (Q Kᵀ) / √key_width ) × V — scaling keeps softmax soft.

Key width Number of dimensions in each key/query vector (paper often uses 64).

Multi-head attention Several attention views in parallel, then concatenate and mix.

Head One attention view with its own Q/K/V maps.

Model width Main embedding size of the stack (paper base: 512 = 8 heads × 64).

Saturation Softmax nearly one-hot when scores are huge — gradients die.

S11 · Check

Exit ticket

Say these aloud · then quiz in matheion MA-LLM · Session 11

Prompt 1 What goes wrong without dividing by √(key width)?

Prompt 2 Why several heads on the dog/food sentence?

Prompt 3 8 heads, model width 512 → each key width is?