Board contents
Text extracted from this public whiteboard for search and accessibility.
S8 · Open
Session 8 · Attention as soft lookup
Live board · deepen in matheion MA-LLM · Session 8
Today’s story Attention is a soft dictionary lookup: ask with a query, match keys, return a blend of values — smooth enough to train with gradients.
You should leave able to… • Tell identical → nearest → soft • Write weights from query·key dots • Hand-compute a 3-key example
With matheion Use this board to teach. Open matheion → MA-LLM → Session 8 for full prose, diagrams, and the auto-quiz.
S8 · Metaphor
Lookup you already know
Dictionary Keys = words Values = definitions Query = the word you want
Phone book Keys = names Values = numbers Query = who you’re calling
Exact match fails in nets Vectors almost never equal exactly. ‘She’ points toward ‘Alice’ — it isn’t identical.
S8 · Three steps
Identical → nearest → soft
1 · Identical Return the value whose key equals the query. Otherwise fail. Never fires with continuous vectors.
2 · Nearest Return the value of the closest key. Always defined — but jumps when the winner flips → no useful gradient.
3 · Soft (attention) Give every value a weight between 0 and 1 (weights sum to 1). Output = weighted mix of all values. Tiny move in the query → tiny move in the output.
S8 · Formula
Where the weights come from
query q · keys k_j · values v_j — same softmax shape as Session 2, candidates are keys
similarity of query to key j = q · k_j (dot product) weight_j = exp(q · k_j) / sum_t exp(q · k_t) output = sum_j weight_j * v_j Weights look like probabilities — but we do NOT draw a key. We average the values.
Roles Query and key must have the same width so dots make sense. Value can have any width — what’s returned needn’t match what’s searched.
S8 · Hand example
Worked example · one query, three keys
q = (1.0, 0.0) key1=(1.0,0.2) value1=(5, 1) → q·key1 = 1.00 key2=(0.9,0.1) value2=(4.5,1.2) → q·key2 = 0.90 key3=(-0.5,0.5) value3=(-1, 3) → q·key3 = -0.50 exp ≈ 2.72, 2.46, 0.61 sum ≈ 5.79 weights ≈ 0.47, 0.425, 0.105 output ≈ (4.15, 1.30) # mostly value1 + value2
Read it Most mass on key1, strong secondary on key2, little on key3. Soft lookup mixes — it doesn’t snap.
S8 · Glossary
Glossary · key concepts this session
Keep this frame visible while teaching · say the term, then the plain line
Query What you’re looking for (the ‘ask’ vector).
Key What can be matched against the query.
Value What gets returned / mixed into the output.
Dot product Similarity score between query and a key (alignment of two vectors).
Soft lookup / attention Weighted average of values; weights from softmax over query·key scores.
Identical lookup Exact key match only — useless with continuous vectors.
Nearest lookup Hard closest key — jumps, not differentiable.
Differentiable Tiny input change → tiny output change; needed for gradient training.
S8 · Check
Exit ticket
Say these aloud · then quiz in matheion MA-LLM · Session 8
Prompt 1 Why nearest lookup blocks learning.
Prompt 2 Write the weight formula and say: average, don’t sample.
Prompt 3 Recompute the 3-key toy from memory.