Board contents
Text extracted from this public whiteboard for search and accessibility.
S4 · Open
Session 4 · Embeddings · one vector per token
Live board · deepen in matheion MA-LLM · Session 4
Today’s story A network multiplies vectors, not integer IDs. An embedding table stores one learned list of numbers per dictionary entry. Look up a sequence → a grid of vectors. Training moves similar meanings close in that space.
You should leave able to… • Look up a sequence by hand and name the shape • Show one-hot × table = pick exactly one row • Compute cosine for tea↔coffee vs tea↔scissors • Count parameters: dictionary size × vector width
With matheion Use this board to teach. Open matheion → MA-LLM → Session 4 for full prose, diagrams, and the auto-quiz.
S4 · Why not the ID
Why not feed the integer itself?
IDs are addresses. Vectors are features.
Problem 1 · labels ≠ magnitudes Dictionary: tea=1, coffee=2, scissors=5. ‘2’ is not ‘more tea-like’ than ‘1’. Neighbouring IDs need not mean neighbouring ideas. Treating the ID as a number invents fake order.
Problem 2 · nothing to multiply A Transformer does matrix maths. An integer has no width — no feature vector. Like pixels needing RGB: tokens need a list of numbers the net can mix.
Fix · one vector per entry Give every dictionary piece a point in space. Training nudges those points so tea drifts toward coffee and away from scissors. That list of numbers = the embedding.
S4 · Running table
Running toy · 5 pieces, width 3
Same table for every demo below · invent numbers, real behaviour
Dictionary size = 5 · vector width = 3 (later people write C × H — say the words first) id piece vector (3 numbers) 1 tea ( 0.9, 0.2, -0.1) 2 coffee ( 0.8, 0.3, 0.0) 3 pot (-0.2, 0.8, 0.4) 4 kettle (-0.1, 0.7, 0.5) 5 scissors ( 0.1, -0.9, 0.2) One row per piece. Shared everywhere: tea at position 1 and tea at position 200 → SAME row.
Vision analogy Image net: one feature vector per pixel. Language net: one feature vector per token. The token is the atom — like the pixel.
Real scale ~50,000 rows × width 768…4096. Same idea as this 5×3 table — just taller and wider.
S4 · Hand lookup
Hand lookup · stack the rows
Sequence: coffee · tea · pot · kettle → IDs [2, 1, 3, 4]
Step 1 — tokeniser (Session 1) already gave IDs: [2, 1, 3, 4] Step 2 — copy each row from the table, in order: position 1 coffee ← row 2 ( 0.8, 0.3, 0.0) position 2 tea ← row 1 ( 0.9, 0.2, -0.1) position 3 pot ← row 3 (-0.2, 0.8, 0.4) position 4 kettle ← row 4 (-0.1, 0.7, 0.5) Result = a 4 × 3 grid 4 = number of tokens in the sequence 3 = vector width THIS grid is the network’s real input from here on.
Shape check Table itself: (# dictionary entries) × width here 5 × 3 Sequence after lookup: (# tokens) × width here 4 × 3 Shape bugs almost always start at lookup.
S4 · One-hot
Lookup = one-hot × table
Same pick-a-row idea — written as a multiply (preview of soft attention)
Want piece ‘pot’ = id 3 in a 5-piece dictionary. Write a selector with a SINGLE 1 (one-hot): selector = (0, 0, 1, 0, 0) ← 1 sits at slot 3 Multiply selector × table: row1·0 + row2·0 + row3·1 + row4·0 + row5·0 = exactly row 3 = (-0.2, 0.8, 0.4) Hard 1 → hard pick of one row. Session 8: replace the hard 1 with soft weights that sum to 1 → soft average of many rows (attention).
Why bother? Same mental model twice: • embedding lookup = hard select • attention = soft select One idea, two hardness levels.
S4 · Cosine
Hand calc · cosine = ‘do they point the same way?’
cosine(u,v) = (u·v) / (|u| |v|) · 1 = same direction, 0 = orthogonal, −1 = opposite
tea u = (0.9, 0.2, -0.1) coffee v = (0.8, 0.3, 0.0) u·v = 0.9·0.8 + 0.2·0.3 + (−0.1)·0 = 0.72 + 0.06 = 0.78 |u| = √(0.81+0.04+0.01) = √0.86 ≈ 0.927 |v| = √(0.64+0.09+0) = √0.73 ≈ 0.854 cos(tea, coffee) ≈ 0.78 / (0.927·0.854) ≈ 0.985 ← almost twins tea vs scissors w = (0.1, -0.9, 0.2): u·w = 0.9·0.1 + 0.2·(−0.9) + (−0.1)·0.2 = 0.09 − 0.18 − 0.02 = −0.11 |w| ≈ 0.927 cos(tea, scissors) ≈ −0.11 / (0.927·0.927) ≈ −0.128 ← unrelated
What training did Drinks that show up in similar sentences drift together. scissors lives elsewhere. ‘Nearby under cosine’ = operational meaning of similar.
S4 · Training moves rows
What training does to the table
Rows are ordinary parameters Backprop nudges every number in every row — same as any weight. Session 2’s loss (how wrong was the next-token lottery?) is what shapes the geometry.
Similar contexts → close vectors tea and coffee often share neighbours → cosine climbs. tea and transistor rarely do → they drift apart. You can ask: ‘who is closest to tea?’ and get a semantic answer.
Subword sharing (Session 1) If friend is one dictionary piece, then: friend friend + ship look up the SAME row for friend. Geometry learned for the piece transfers into compounds automatically.
S4 · Parameter count
Why dictionary size × width is expensive
The embedding table is often a huge slice of the whole model
Parameters in the input embedding table = (# dictionary entries) × (vector width) Example A — our toy: 5 × 3 = 15 numbers (trivial) Example B — smallish real model: 20,000 × 256 = 5,120,000 ≈ 5.1 million …before ANY Transformer layer exists. Example C — common scale: 50,000 × 512 = 25,600,000 ≈ 25.6 million Double the width → double that count. Session 1’s ‘how big is the dictionary?’ was really deciding how fat this final/first table is.
Design tension Wider vectors → more room to separate meanings. Wider vectors → more memory and multiply cost everywhere downstream. Typical modern widths: hundreds → a few thousand.
S4 · Pipeline
Where this sits in the pipeline
text "coffee tea pot kettle" │ ▼ Session 1 tokeniser IDs [2, 1, 3, 4] │ ▼ THIS SESSION embedding lookup grid 4 × 3 (one vector per token) │ ▼ Sessions 8–13 Transformer blocks (mix and transform those rows) │ ▼ Session 2 final linear → logits → softmax lottery next-token probabilities over the whole dictionary Everything after today is ‘how do we mix the rows of this grid?’
S4 · Glossary
Glossary · key concepts this session
Keep this frame visible while teaching · say the term, then the plain line
Embedding A learned list of numbers (a vector) standing for one dictionary piece.
Embedding table Matrix with one row per dictionary entry; lookup by token ID.
Vector width How many numbers per row (e.g. 3 in the toy, 768 in a real model). Later written H.
One-hot Selector with a single 1 — hard-picks exactly one table row.
Cosine similarity (u·v) / (|u||v|) — how aligned two directions are; proxy for ‘similar meaning’.
Sequence matrix Stacked embedding rows for a prompt — shape (# tokens) × width.
Parameter count Dictionary size × width floats just for the input table.
Shared row Same piece always looks up the same vector, wherever it appears.
S4 · Check
Exit ticket
Say these aloud · then quiz in matheion MA-LLM · Session 4
Prompt 1 Look up [5, 3, 1] on the tea/coffee table — write the 3×3 grid and say its shape.
Prompt 2 Compute cos(tea, coffee) roughly — is it near 1 or near 0?
Prompt 3 A dictionary of 20,000 pieces at width 256 — how many embedding parameters?