chalkline

Public teaching whiteboard

Self-attention explained for Transformers

Self-attention: every token looks at every token. Free public Transformer teaching whiteboard. · by zlu

Topics: LLM & Transformers

Self-attention explained for Transformers
S9 · Open
S9 · Three roles
S9 · Matrix A
S9 · Self vs cross
S9 · Glossary
S9 · Check

Session 9 · Self-attention and the attention matrix

Live board · deepen in matheion MA-LLM · Session 9

Today’s story Self-attention: every token builds a query, key, and value from the same sequence, then soft-looks-up every token — including itself.

You should leave able to… • Say ask / match / deliver for Q, K, V • Read one row of an attention table • Contrast self vs cross with a translation example

With matheion Use this board to teach. Open matheion → MA-LLM → Session 9 for full prose, diagrams, and the auto-quiz.

One sequence, three roles

Start from the grid of token vectors from Session 4 (call it X)

text
X = stacked token vectors
    (number_of_tokens rows × embedding_width columns)

Learn three maps:
  queries  Q = X × W_ask     # what each token is looking for
  keys     K = X × W_match   # what each token offers to match
  values   V = X × W_deliver # what each token will contribute

Token i’s new vector = soft_lookup(query_i, all keys, all values)
                       (Session 8)

Why three maps? Same content, different jobs: ask · match · deliver. Training specialises each map.

Read a tiny attention table

Four tokens: The · animal · because · it

text
Rows = who is asking.  Columns = who they listen to.
Each row is a lottery (sums to 1) — Session 2 softmax again.

         The   animal  because   it
The     .70     .20      .05    .05
animal  .10     .80      .05    .05
because .05     .10      .75    .10
it      .05     .60      .10    .25   ← ‘it’ mostly listens to ‘animal’

New vector for ‘it’ ≈
  0.05·value_The + 0.60·value_animal + 0.10·value_because + 0.25·value_it

Self vs cross · concrete

Self · same sentence English: “I gave my dog Charlie some food” Queries, keys, values all from that English string. Pronoun-style links resolved inside one language.

Cross · two sequences Translating into French: queries come from the French side; keys/values come from the English side. That’s how a French word can find ‘dog’.

Glossary · key concepts this session

Keep this frame visible while teaching · say the term, then the plain line

Self-attention Queries, keys, and values all come from the same sequence.

Cross-attention Queries from one sequence; keys/values from another.

Attention matrix Table of weights: row i = how token i distributes attention over others.

Q, K, V projections Three learned maps: ask, match, deliver — from the same token matrix X.

W_ask / W_match / W_deliver The three weight matrices that build Q, K, V from X.

Exit ticket

Say these aloud · then quiz in matheion MA-LLM · Session 9

Prompt 1 In plain words: what do the three maps ask / match / deliver do?

Prompt 2 What does one row of the attention table mean?

Prompt 3 Give one self and one cross example.

Drag to pan · scroll to zoom · read-only

Board contents

Text extracted from this public whiteboard for search and accessibility.

S9 · Open

Session 9 · Self-attention and the attention matrix

Live board · deepen in matheion MA-LLM · Session 9

Today’s story Self-attention: every token builds a query, key, and value from the same sequence, then soft-looks-up every token — including itself.

You should leave able to… • Say ask / match / deliver for Q, K, V • Read one row of an attention table • Contrast self vs cross with a translation example

With matheion Use this board to teach. Open matheion → MA-LLM → Session 9 for full prose, diagrams, and the auto-quiz.

S9 · Three roles

One sequence, three roles

Start from the grid of token vectors from Session 4 (call it X)

X = stacked token vectors (number_of_tokens rows × embedding_width columns) Learn three maps: queries Q = X × W_ask # what each token is looking for keys K = X × W_match # what each token offers to match values V = X × W_deliver # what each token will contribute Token i’s new vector = soft_lookup(query_i, all keys, all values) (Session 8)

Why three maps? Same content, different jobs: ask · match · deliver. Training specialises each map.

S9 · Matrix A

Read a tiny attention table

Four tokens: The · animal · because · it

Rows = who is asking. Columns = who they listen to. Each row is a lottery (sums to 1) — Session 2 softmax again. The animal because it The .70 .20 .05 .05 animal .10 .80 .05 .05 because .05 .10 .75 .10 it .05 .60 .10 .25 ← ‘it’ mostly listens to ‘animal’ New vector for ‘it’ ≈ 0.05·value_The + 0.60·value_animal + 0.10·value_because + 0.25·value_it

S9 · Self vs cross

Self vs cross · concrete

Self · same sentence English: “I gave my dog Charlie some food” Queries, keys, values all from that English string. Pronoun-style links resolved inside one language.

Cross · two sequences Translating into French: queries come from the French side; keys/values come from the English side. That’s how a French word can find ‘dog’.

S9 · Glossary

Glossary · key concepts this session

Keep this frame visible while teaching · say the term, then the plain line

Self-attention Queries, keys, and values all come from the same sequence.

Cross-attention Queries from one sequence; keys/values from another.

Attention matrix Table of weights: row i = how token i distributes attention over others.

Q, K, V projections Three learned maps: ask, match, deliver — from the same token matrix X.

W_ask / W_match / W_deliver The three weight matrices that build Q, K, V from X.

S9 · Check

Exit ticket

Say these aloud · then quiz in matheion MA-LLM · Session 9

Prompt 1 In plain words: what do the three maps ask / match / deliver do?

Prompt 2 What does one row of the attention table mean?

Prompt 3 Give one self and one cross example.