Board contents
Text extracted from this public whiteboard for search and accessibility.
S2 · Open
Session 2 · Predict the next token
Live board · deepen in matheion MA-LLM · Session 2
Today’s story One concrete job: given the tokens so far — e.g. “The cat” — score every piece in the dictionary, turn scores into a lottery, and choose what comes next. Then repeat.
You should leave able to… • Walk the “The cat → ???” example end to end • Say where the candidate words come from (the dictionary) • Compute softmax by hand from three logits
With matheion Use this board to teach. Open matheion → MA-LLM → Session 2 for full prose, diagrams, and the auto-quiz.
S2 · Running example
Running example · what comes after “The cat”?
CONTEXT (tokens so far): "The cat" after Session 1 tokeniser → IDs, say [41, 892] QUESTION the model answers every step: Of all dictionary pieces, which should come NEXT? NOT “pick among sat/on/the only”. The real dictionary has ~50,000 entries. Every one of them gets a score. We zoom in on a few for the hand calculation.
Then repeat Suppose we choose “sat”. New context: “The cat sat” Ask again: what comes next? (maybe “on”, …) That loop is how text is generated.
S2 · Where options come from
Where do “sat / dog / on” come from?
They are ordinary rows in the Session-1 dictionary — not a special shortlist
Full dictionary (toy sketch of a few rows): id piece 0 the 1 a … 55 on 120 sat 200 dog 201 cat 880 mat … 49999 <rare piece> When we write scores for sat / dog / on below, we are LOOKING AT three of those ~50,000 rows. The other ~49,997 also got scores — we just hide them so the arithmetic fits on one slide.
Why those three? Teaching choice only. “sat” — sensible after “The cat” “dog” — related animal word, less fit here “on” — also plausible (The cat on …) A real model scores ALL rows; these three stand in for the whole list.
S2 · Where logits come from
Where do the raw scores (logits) come from?
Pipeline for ONE next-token step: 1. Context "The cat" → token IDs (Session 1) 2. IDs → vectors via embedding table (Session 4) 3. Transformer stack mixes the context (Sessions 8–13) 4. Final linear layer: one number per dictionary row = that row’s logit (raw score) Today we treat steps 2–4 as a black box that ALREADY produced these logits for our zoom-in: sat → 2.0 dog → 1.0 on → 0.0 Your job: turn those three numbers into a lottery.
Logit = raw score Name: logit Meaning: unnormalised preference. Bigger → model likes that next piece more. Not yet a probability — can be negative, and they don’t sum to 1.
S2 · Softmax by hand
Hand calculation · logits → probabilities
Context: “The cat” · zoom-in on three dictionary rows
Step 0 — logits the black-box model gave us: sat = 2.0 dog = 1.0 on = 0.0 Step 1 — exponentiate (make positive, stretch gaps): exp(2.0) ≈ 7.39 exp(1.0) ≈ 2.72 exp(0.0) = 1.00 sum Z = 7.39 + 2.72 + 1.00 = 11.11 Step 2 — divide by Z (now a lottery — this IS softmax): P(sat | "The cat") = 7.39 / 11.11 ≈ 0.665 (66.5%) P(dog | "The cat") = 2.72 / 11.11 ≈ 0.245 (24.5%) P(on | "The cat") = 1.00 / 11.11 ≈ 0.090 ( 9.0%) Check: 0.665 + 0.245 + 0.090 = 1.000 ✓ Step 3 — choose: always-pick-top → sat draw from lottery → sat most times; dog ~1 in 4; on rarely If we picked sat, next step’s context is "The cat sat".
S2 · Habits on these numbers
Three habits · same “The cat” example
Sums to one 0.665 + 0.245 + 0.090 = 1 Every next-token step ends in a lottery over the dictionary (or our zoom-in).
Only gaps matter Add +5 to every logit: sat=7, dog=6, on=5 Softmax probabilities stay exactly 0.665 / 0.245 / 0.090. Only differences between logits count.
Sharp vs flat If logits were sat=5, dog=0, on=0 → almost all mass on sat. If logits were nearly equal → flatter, more uncertain lottery.
S2 · Why the lottery
Why not always pick “sat”?
Always-pick-top Every run: “The cat sat …” Fine for short factual fills. For stories / brainstorming: same bland path every time.
Draw from the lottery Sometimes “The cat sat …” Sometimes “The cat on …” Rarely something else. Variety comes from sampling. Next session: temperature & top-k/p knobs on this same lottery.
S2 · Glossary
Glossary · key concepts this session
Keep this frame visible while teaching · say the term, then the plain line
Context The tokens so far — the input to one next-token step (e.g. “The cat”).
Dictionary / vocabulary All allowed next pieces (~50k). Every row gets a logit each step.
Logit (raw score) Unnormalised preference for one dictionary row. Not a probability yet.
Softmax exp(logit) / sum(exp(logits)) → probabilities that sum to 1.
Next-token probability P(piece | context) after softmax — chance that piece comes next.
Argmax / always-pick-top Choose the single highest P. Deterministic.
Sampling Draw a piece at random according to the probabilities.
Zoom-in toy We show 3 rows for arithmetic; a real model scores the full dictionary.
S2 · Check
Exit ticket
Say these aloud · then quiz in matheion MA-LLM · Session 2
Prompt 1 Given context “The cat” and logits sat=2, dog=1, on=0 — compute the three probabilities.
Prompt 2 Where do sat/dog/on come from — a special shortlist, or dictionary rows?
Prompt 3 If we pick sat, what is the next step’s context?