chalkline

Public teaching whiteboard

Word & token embeddings explained for LLMs

What embeddings are and why language models need them — free public ML teaching whiteboard. · by zlu

Topics: LLM & Transformers

Word & token embeddings explained for LLMs
S4 · Open
S4 · Why not the ID
S4 · Running table
S4 · Hand lookup
S4 · One-hot
S4 · Cosine
S4 · Training moves rows
S4 · Parameter count
S4 · Pipeline
S4 · Glossary
S4 · Check

Session 4 · Embeddings · one vector per token

Live board · deepen in matheion MA-LLM · Session 4

Today’s story A network multiplies vectors, not integer IDs. An embedding table stores one learned list of numbers per dictionary entry. Look up a sequence → a grid of vectors. Training moves similar meanings close in that space.

You should leave able to… • Look up a sequence by hand and name the shape • Show one-hot × table = pick exactly one row • Compute cosine for tea↔coffee vs tea↔scissors • Count parameters: dictionary size × vector width

With matheion Use this board to teach. Open matheion → MA-LLM → Session 4 for full prose, diagrams, and the auto-quiz.

Why not feed the integer itself?

IDs are addresses. Vectors are features.

Problem 1 · labels ≠ magnitudes Dictionary: tea=1, coffee=2, scissors=5. ‘2’ is not ‘more tea-like’ than ‘1’. Neighbouring IDs need not mean neighbouring ideas. Treating the ID as a number invents fake order.

Problem 2 · nothing to multiply A Transformer does matrix maths. An integer has no width — no feature vector. Like pixels needing RGB: tokens need a list of numbers the net can mix.

Fix · one vector per entry Give every dictionary piece a point in space. Training nudges those points so tea drifts toward coffee and away from scissors. That list of numbers = the embedding.

Running toy · 5 pieces, width 3

Same table for every demo below · invent numbers, real behaviour

text
Dictionary size = 5   ·   vector width = 3
(later people write C × H — say the words first)

 id   piece      vector (3 numbers)
  1   tea        ( 0.9,  0.2, -0.1)
  2   coffee     ( 0.8,  0.3,  0.0)
  3   pot        (-0.2,  0.8,  0.4)
  4   kettle     (-0.1,  0.7,  0.5)
  5   scissors   ( 0.1, -0.9,  0.2)

One row per piece. Shared everywhere:
tea at position 1 and tea at position 200 → SAME row.

Vision analogy Image net: one feature vector per pixel. Language net: one feature vector per token. The token is the atom — like the pixel.

Real scale ~50,000 rows × width 768…4096. Same idea as this 5×3 table — just taller and wider.

Hand lookup · stack the rows

Sequence: coffee · tea · pot · kettle → IDs [2, 1, 3, 4]

text
Step 1 — tokeniser (Session 1) already gave IDs:
  [2, 1, 3, 4]

Step 2 — copy each row from the table, in order:

  position 1  coffee   ← row 2   ( 0.8,  0.3,  0.0)
  position 2  tea      ← row 1   ( 0.9,  0.2, -0.1)
  position 3  pot      ← row 3   (-0.2,  0.8,  0.4)
  position 4  kettle   ← row 4   (-0.1,  0.7,  0.5)

Result = a 4 × 3 grid
  4 = number of tokens in the sequence
  3 = vector width

THIS grid is the network’s real input from here on.

Shape check Table itself: (# dictionary entries) × width here 5 × 3 Sequence after lookup: (# tokens) × width here 4 × 3 Shape bugs almost always start at lookup.

Lookup = one-hot × table

Same pick-a-row idea — written as a multiply (preview of soft attention)

text
Want piece ‘pot’ = id 3 in a 5-piece dictionary.

Write a selector with a SINGLE 1 (one-hot):
  selector = (0, 0, 1, 0, 0)   ← 1 sits at slot 3

Multiply selector × table:
  row1·0 + row2·0 + row3·1 + row4·0 + row5·0
  = exactly row 3 = (-0.2, 0.8, 0.4)

Hard 1 → hard pick of one row.
Session 8: replace the hard 1 with soft weights
that sum to 1 → soft average of many rows (attention).

Why bother? Same mental model twice: • embedding lookup = hard select • attention = soft select One idea, two hardness levels.

Hand calc · cosine = ‘do they point the same way?’

cosine(u,v) = (u·v) / (|u| |v|) · 1 = same direction, 0 = orthogonal, −1 = opposite

text
tea    u = (0.9,  0.2, -0.1)
coffee v = (0.8,  0.3,  0.0)

u·v = 0.9·0.8 + 0.2·0.3 + (−0.1)·0 = 0.72 + 0.06 = 0.78
|u| = √(0.81+0.04+0.01) = √0.86 ≈ 0.927
|v| = √(0.64+0.09+0)   = √0.73 ≈ 0.854
cos(tea, coffee) ≈ 0.78 / (0.927·0.854) ≈ 0.985   ← almost twins

tea vs scissors w = (0.1, -0.9, 0.2):
u·w = 0.9·0.1 + 0.2·(−0.9) + (−0.1)·0.2 = 0.09 − 0.18 − 0.02 = −0.11
|w| ≈ 0.927
cos(tea, scissors) ≈ −0.11 / (0.927·0.927) ≈ −0.128   ← unrelated

What training did Drinks that show up in similar sentences drift together. scissors lives elsewhere. ‘Nearby under cosine’ = operational meaning of similar.

What training does to the table

Rows are ordinary parameters Backprop nudges every number in every row — same as any weight. Session 2’s loss (how wrong was the next-token lottery?) is what shapes the geometry.

Similar contexts → close vectors tea and coffee often share neighbours → cosine climbs. tea and transistor rarely do → they drift apart. You can ask: ‘who is closest to tea?’ and get a semantic answer.

Subword sharing (Session 1) If friend is one dictionary piece, then: friend friend + ship look up the SAME row for friend. Geometry learned for the piece transfers into compounds automatically.

Why dictionary size × width is expensive

The embedding table is often a huge slice of the whole model

text
Parameters in the input embedding table
  = (# dictionary entries) × (vector width)

Example A — our toy:
  5 × 3 = 15 numbers   (trivial)

Example B — smallish real model:
  20,000 × 256 = 5,120,000 ≈ 5.1 million
  …before ANY Transformer layer exists.

Example C — common scale:
  50,000 × 512 = 25,600,000 ≈ 25.6 million

Double the width → double that count.
Session 1’s ‘how big is the dictionary?’ was really
deciding how fat this final/first table is.

Design tension Wider vectors → more room to separate meanings. Wider vectors → more memory and multiply cost everywhere downstream. Typical modern widths: hundreds → a few thousand.

Where this sits in the pipeline

text
text  "coffee tea pot kettle"
  │
  ▼  Session 1  tokeniser
IDs   [2, 1, 3, 4]
  │
  ▼  THIS SESSION  embedding lookup
grid  4 × 3   (one vector per token)
  │
  ▼  Sessions 8–13  Transformer blocks
     (mix and transform those rows)
  │
  ▼  Session 2  final linear → logits → softmax lottery
next-token probabilities over the whole dictionary

Everything after today is ‘how do we mix the rows of this grid?’

Glossary · key concepts this session

Keep this frame visible while teaching · say the term, then the plain line

Embedding A learned list of numbers (a vector) standing for one dictionary piece.

Embedding table Matrix with one row per dictionary entry; lookup by token ID.

Vector width How many numbers per row (e.g. 3 in the toy, 768 in a real model). Later written H.

One-hot Selector with a single 1 — hard-picks exactly one table row.

Cosine similarity (u·v) / (|u||v|) — how aligned two directions are; proxy for ‘similar meaning’.

Sequence matrix Stacked embedding rows for a prompt — shape (# tokens) × width.

Parameter count Dictionary size × width floats just for the input table.

Shared row Same piece always looks up the same vector, wherever it appears.

Exit ticket

Say these aloud · then quiz in matheion MA-LLM · Session 4

Prompt 1 Look up [5, 3, 1] on the tea/coffee table — write the 3×3 grid and say its shape.

Prompt 2 Compute cos(tea, coffee) roughly — is it near 1 or near 0?

Prompt 3 A dictionary of 20,000 pieces at width 256 — how many embedding parameters?

Drag to pan · scroll to zoom · read-only

Board contents

Text extracted from this public whiteboard for search and accessibility.

S4 · Open

Session 4 · Embeddings · one vector per token

Live board · deepen in matheion MA-LLM · Session 4

Today’s story A network multiplies vectors, not integer IDs. An embedding table stores one learned list of numbers per dictionary entry. Look up a sequence → a grid of vectors. Training moves similar meanings close in that space.

You should leave able to… • Look up a sequence by hand and name the shape • Show one-hot × table = pick exactly one row • Compute cosine for tea↔coffee vs tea↔scissors • Count parameters: dictionary size × vector width

With matheion Use this board to teach. Open matheion → MA-LLM → Session 4 for full prose, diagrams, and the auto-quiz.

S4 · Why not the ID

Why not feed the integer itself?

IDs are addresses. Vectors are features.

Problem 1 · labels ≠ magnitudes Dictionary: tea=1, coffee=2, scissors=5. ‘2’ is not ‘more tea-like’ than ‘1’. Neighbouring IDs need not mean neighbouring ideas. Treating the ID as a number invents fake order.

Problem 2 · nothing to multiply A Transformer does matrix maths. An integer has no width — no feature vector. Like pixels needing RGB: tokens need a list of numbers the net can mix.

Fix · one vector per entry Give every dictionary piece a point in space. Training nudges those points so tea drifts toward coffee and away from scissors. That list of numbers = the embedding.

S4 · Running table

Running toy · 5 pieces, width 3

Same table for every demo below · invent numbers, real behaviour

Dictionary size = 5 · vector width = 3 (later people write C × H — say the words first) id piece vector (3 numbers) 1 tea ( 0.9, 0.2, -0.1) 2 coffee ( 0.8, 0.3, 0.0) 3 pot (-0.2, 0.8, 0.4) 4 kettle (-0.1, 0.7, 0.5) 5 scissors ( 0.1, -0.9, 0.2) One row per piece. Shared everywhere: tea at position 1 and tea at position 200 → SAME row.

Vision analogy Image net: one feature vector per pixel. Language net: one feature vector per token. The token is the atom — like the pixel.

Real scale ~50,000 rows × width 768…4096. Same idea as this 5×3 table — just taller and wider.

S4 · Hand lookup

Hand lookup · stack the rows

Sequence: coffee · tea · pot · kettle → IDs [2, 1, 3, 4]

Step 1 — tokeniser (Session 1) already gave IDs: [2, 1, 3, 4] Step 2 — copy each row from the table, in order: position 1 coffee ← row 2 ( 0.8, 0.3, 0.0) position 2 tea ← row 1 ( 0.9, 0.2, -0.1) position 3 pot ← row 3 (-0.2, 0.8, 0.4) position 4 kettle ← row 4 (-0.1, 0.7, 0.5) Result = a 4 × 3 grid 4 = number of tokens in the sequence 3 = vector width THIS grid is the network’s real input from here on.

Shape check Table itself: (# dictionary entries) × width here 5 × 3 Sequence after lookup: (# tokens) × width here 4 × 3 Shape bugs almost always start at lookup.

S4 · One-hot

Lookup = one-hot × table

Same pick-a-row idea — written as a multiply (preview of soft attention)

Want piece ‘pot’ = id 3 in a 5-piece dictionary. Write a selector with a SINGLE 1 (one-hot): selector = (0, 0, 1, 0, 0) ← 1 sits at slot 3 Multiply selector × table: row1·0 + row2·0 + row3·1 + row4·0 + row5·0 = exactly row 3 = (-0.2, 0.8, 0.4) Hard 1 → hard pick of one row. Session 8: replace the hard 1 with soft weights that sum to 1 → soft average of many rows (attention).

Why bother? Same mental model twice: • embedding lookup = hard select • attention = soft select One idea, two hardness levels.

S4 · Cosine

Hand calc · cosine = ‘do they point the same way?’

cosine(u,v) = (u·v) / (|u| |v|) · 1 = same direction, 0 = orthogonal, −1 = opposite

tea u = (0.9, 0.2, -0.1) coffee v = (0.8, 0.3, 0.0) u·v = 0.9·0.8 + 0.2·0.3 + (−0.1)·0 = 0.72 + 0.06 = 0.78 |u| = √(0.81+0.04+0.01) = √0.86 ≈ 0.927 |v| = √(0.64+0.09+0) = √0.73 ≈ 0.854 cos(tea, coffee) ≈ 0.78 / (0.927·0.854) ≈ 0.985 ← almost twins tea vs scissors w = (0.1, -0.9, 0.2): u·w = 0.9·0.1 + 0.2·(−0.9) + (−0.1)·0.2 = 0.09 − 0.18 − 0.02 = −0.11 |w| ≈ 0.927 cos(tea, scissors) ≈ −0.11 / (0.927·0.927) ≈ −0.128 ← unrelated

What training did Drinks that show up in similar sentences drift together. scissors lives elsewhere. ‘Nearby under cosine’ = operational meaning of similar.

S4 · Training moves rows

What training does to the table

Rows are ordinary parameters Backprop nudges every number in every row — same as any weight. Session 2’s loss (how wrong was the next-token lottery?) is what shapes the geometry.

Similar contexts → close vectors tea and coffee often share neighbours → cosine climbs. tea and transistor rarely do → they drift apart. You can ask: ‘who is closest to tea?’ and get a semantic answer.

Subword sharing (Session 1) If friend is one dictionary piece, then: friend friend + ship look up the SAME row for friend. Geometry learned for the piece transfers into compounds automatically.

S4 · Parameter count

Why dictionary size × width is expensive

The embedding table is often a huge slice of the whole model

Parameters in the input embedding table = (# dictionary entries) × (vector width) Example A — our toy: 5 × 3 = 15 numbers (trivial) Example B — smallish real model: 20,000 × 256 = 5,120,000 ≈ 5.1 million …before ANY Transformer layer exists. Example C — common scale: 50,000 × 512 = 25,600,000 ≈ 25.6 million Double the width → double that count. Session 1’s ‘how big is the dictionary?’ was really deciding how fat this final/first table is.

Design tension Wider vectors → more room to separate meanings. Wider vectors → more memory and multiply cost everywhere downstream. Typical modern widths: hundreds → a few thousand.

S4 · Pipeline

Where this sits in the pipeline

text "coffee tea pot kettle" │ ▼ Session 1 tokeniser IDs [2, 1, 3, 4] │ ▼ THIS SESSION embedding lookup grid 4 × 3 (one vector per token) │ ▼ Sessions 8–13 Transformer blocks (mix and transform those rows) │ ▼ Session 2 final linear → logits → softmax lottery next-token probabilities over the whole dictionary Everything after today is ‘how do we mix the rows of this grid?’

S4 · Glossary

Glossary · key concepts this session

Keep this frame visible while teaching · say the term, then the plain line

Embedding A learned list of numbers (a vector) standing for one dictionary piece.

Embedding table Matrix with one row per dictionary entry; lookup by token ID.

Vector width How many numbers per row (e.g. 3 in the toy, 768 in a real model). Later written H.

One-hot Selector with a single 1 — hard-picks exactly one table row.

Cosine similarity (u·v) / (|u||v|) — how aligned two directions are; proxy for ‘similar meaning’.

Sequence matrix Stacked embedding rows for a prompt — shape (# tokens) × width.

Parameter count Dictionary size × width floats just for the input table.

Shared row Same piece always looks up the same vector, wherever it appears.

S4 · Check

Exit ticket

Say these aloud · then quiz in matheion MA-LLM · Session 4

Prompt 1 Look up [5, 3, 1] on the tea/coffee table — write the 3×3 grid and say its shape.

Prompt 2 Compute cos(tea, coffee) roughly — is it near 1 or near 0?

Prompt 3 A dictionary of 20,000 pieces at width 256 — how many embedding parameters?