chalkline

Public teaching whiteboard

Next-token prediction — how LLMs generate text

Predict the next token: the core of language models, explained on a free teaching whiteboard with worked examples. · by zlu

Topics: LLM & Transformers

Next-token prediction — how LLMs generate text
S2 · Open
S2 · Running example
S2 · Where options come from
S2 · Where logits come from
S2 · Softmax by hand
S2 · Habits on these numbers
S2 · Why the lottery
S2 · Glossary
S2 · Check

Session 2 · Predict the next token

Live board · deepen in matheion MA-LLM · Session 2

Today’s story One concrete job: given the tokens so far — e.g. “The cat” — score every piece in the dictionary, turn scores into a lottery, and choose what comes next. Then repeat.

You should leave able to… • Walk the “The cat → ???” example end to end • Say where the candidate words come from (the dictionary) • Compute softmax by hand from three logits

With matheion Use this board to teach. Open matheion → MA-LLM → Session 2 for full prose, diagrams, and the auto-quiz.

Running example · what comes after “The cat”?

text
CONTEXT (tokens so far):
  "The cat"
  after Session 1 tokeniser → IDs, say [41, 892]

QUESTION the model answers every step:
  Of all dictionary pieces, which should come NEXT?

NOT “pick among sat/on/the only”.
The real dictionary has ~50,000 entries.
Every one of them gets a score. We zoom in on a few
for the hand calculation.

Then repeat Suppose we choose “sat”. New context: “The cat sat” Ask again: what comes next? (maybe “on”, …) That loop is how text is generated.

Where do “sat / dog / on” come from?

They are ordinary rows in the Session-1 dictionary — not a special shortlist

text
Full dictionary (toy sketch of a few rows):

  id     piece
  0      the
  1      a
  …
  55     on
  120    sat
  200    dog
  201    cat
  880    mat
  …
  49999  <rare piece>

When we write scores for sat / dog / on below,
we are LOOKING AT three of those ~50,000 rows.
The other ~49,997 also got scores — we just hide them
so the arithmetic fits on one slide.

Why those three? Teaching choice only. “sat” — sensible after “The cat” “dog” — related animal word, less fit here “on” — also plausible (The cat on …) A real model scores ALL rows; these three stand in for the whole list.

Where do the raw scores (logits) come from?

text
Pipeline for ONE next-token step:

  1. Context "The cat" → token IDs          (Session 1)
  2. IDs → vectors via embedding table       (Session 4)
  3. Transformer stack mixes the context     (Sessions 8–13)
  4. Final linear layer:
       one number per dictionary row
       = that row’s logit (raw score)

Today we treat steps 2–4 as a black box that
ALREADY produced these logits for our zoom-in:

       sat → 2.0
       dog → 1.0
       on  → 0.0

Your job: turn those three numbers into a lottery.

Logit = raw score Name: logit Meaning: unnormalised preference. Bigger → model likes that next piece more. Not yet a probability — can be negative, and they don’t sum to 1.

Hand calculation · logits → probabilities

Context: “The cat” · zoom-in on three dictionary rows

text
Step 0 — logits the black-box model gave us:
  sat = 2.0     dog = 1.0     on = 0.0

Step 1 — exponentiate (make positive, stretch gaps):
  exp(2.0) ≈ 7.39
  exp(1.0) ≈ 2.72
  exp(0.0) = 1.00
  sum Z    = 7.39 + 2.72 + 1.00 = 11.11

Step 2 — divide by Z (now a lottery — this IS softmax):
  P(sat | "The cat") = 7.39 / 11.11 ≈ 0.665   (66.5%)
  P(dog | "The cat") = 2.72 / 11.11 ≈ 0.245   (24.5%)
  P(on  | "The cat") = 1.00 / 11.11 ≈ 0.090   ( 9.0%)
  Check: 0.665 + 0.245 + 0.090 = 1.000 ✓

Step 3 — choose:
  always-pick-top → sat
  draw from lottery → sat most times; dog ~1 in 4; on rarely

If we picked sat, next step’s context is "The cat sat".

Three habits · same “The cat” example

Sums to one 0.665 + 0.245 + 0.090 = 1 Every next-token step ends in a lottery over the dictionary (or our zoom-in).

Only gaps matter Add +5 to every logit: sat=7, dog=6, on=5 Softmax probabilities stay exactly 0.665 / 0.245 / 0.090. Only differences between logits count.

Sharp vs flat If logits were sat=5, dog=0, on=0 → almost all mass on sat. If logits were nearly equal → flatter, more uncertain lottery.

Why not always pick “sat”?

Always-pick-top Every run: “The cat sat …” Fine for short factual fills. For stories / brainstorming: same bland path every time.

Draw from the lottery Sometimes “The cat sat …” Sometimes “The cat on …” Rarely something else. Variety comes from sampling. Next session: temperature & top-k/p knobs on this same lottery.

Glossary · key concepts this session

Keep this frame visible while teaching · say the term, then the plain line

Context The tokens so far — the input to one next-token step (e.g. “The cat”).

Dictionary / vocabulary All allowed next pieces (~50k). Every row gets a logit each step.

Logit (raw score) Unnormalised preference for one dictionary row. Not a probability yet.

Softmax exp(logit) / sum(exp(logits)) → probabilities that sum to 1.

Next-token probability P(piece | context) after softmax — chance that piece comes next.

Argmax / always-pick-top Choose the single highest P. Deterministic.

Sampling Draw a piece at random according to the probabilities.

Zoom-in toy We show 3 rows for arithmetic; a real model scores the full dictionary.

Exit ticket

Say these aloud · then quiz in matheion MA-LLM · Session 2

Prompt 1 Given context “The cat” and logits sat=2, dog=1, on=0 — compute the three probabilities.

Prompt 2 Where do sat/dog/on come from — a special shortlist, or dictionary rows?

Prompt 3 If we pick sat, what is the next step’s context?

Drag to pan · scroll to zoom · read-only

Board contents

Text extracted from this public whiteboard for search and accessibility.

S2 · Open

Session 2 · Predict the next token

Live board · deepen in matheion MA-LLM · Session 2

Today’s story One concrete job: given the tokens so far — e.g. “The cat” — score every piece in the dictionary, turn scores into a lottery, and choose what comes next. Then repeat.

You should leave able to… • Walk the “The cat → ???” example end to end • Say where the candidate words come from (the dictionary) • Compute softmax by hand from three logits

With matheion Use this board to teach. Open matheion → MA-LLM → Session 2 for full prose, diagrams, and the auto-quiz.

S2 · Running example

Running example · what comes after “The cat”?

CONTEXT (tokens so far): "The cat" after Session 1 tokeniser → IDs, say [41, 892] QUESTION the model answers every step: Of all dictionary pieces, which should come NEXT? NOT “pick among sat/on/the only”. The real dictionary has ~50,000 entries. Every one of them gets a score. We zoom in on a few for the hand calculation.

Then repeat Suppose we choose “sat”. New context: “The cat sat” Ask again: what comes next? (maybe “on”, …) That loop is how text is generated.

S2 · Where options come from

Where do “sat / dog / on” come from?

They are ordinary rows in the Session-1 dictionary — not a special shortlist

Full dictionary (toy sketch of a few rows): id piece 0 the 1 a … 55 on 120 sat 200 dog 201 cat 880 mat … 49999 <rare piece> When we write scores for sat / dog / on below, we are LOOKING AT three of those ~50,000 rows. The other ~49,997 also got scores — we just hide them so the arithmetic fits on one slide.

Why those three? Teaching choice only. “sat” — sensible after “The cat” “dog” — related animal word, less fit here “on” — also plausible (The cat on …) A real model scores ALL rows; these three stand in for the whole list.

S2 · Where logits come from

Where do the raw scores (logits) come from?

Pipeline for ONE next-token step: 1. Context "The cat" → token IDs (Session 1) 2. IDs → vectors via embedding table (Session 4) 3. Transformer stack mixes the context (Sessions 8–13) 4. Final linear layer: one number per dictionary row = that row’s logit (raw score) Today we treat steps 2–4 as a black box that ALREADY produced these logits for our zoom-in: sat → 2.0 dog → 1.0 on → 0.0 Your job: turn those three numbers into a lottery.

Logit = raw score Name: logit Meaning: unnormalised preference. Bigger → model likes that next piece more. Not yet a probability — can be negative, and they don’t sum to 1.

S2 · Softmax by hand

Hand calculation · logits → probabilities

Context: “The cat” · zoom-in on three dictionary rows

Step 0 — logits the black-box model gave us: sat = 2.0 dog = 1.0 on = 0.0 Step 1 — exponentiate (make positive, stretch gaps): exp(2.0) ≈ 7.39 exp(1.0) ≈ 2.72 exp(0.0) = 1.00 sum Z = 7.39 + 2.72 + 1.00 = 11.11 Step 2 — divide by Z (now a lottery — this IS softmax): P(sat | "The cat") = 7.39 / 11.11 ≈ 0.665 (66.5%) P(dog | "The cat") = 2.72 / 11.11 ≈ 0.245 (24.5%) P(on | "The cat") = 1.00 / 11.11 ≈ 0.090 ( 9.0%) Check: 0.665 + 0.245 + 0.090 = 1.000 ✓ Step 3 — choose: always-pick-top → sat draw from lottery → sat most times; dog ~1 in 4; on rarely If we picked sat, next step’s context is "The cat sat".

S2 · Habits on these numbers

Three habits · same “The cat” example

Sums to one 0.665 + 0.245 + 0.090 = 1 Every next-token step ends in a lottery over the dictionary (or our zoom-in).

Only gaps matter Add +5 to every logit: sat=7, dog=6, on=5 Softmax probabilities stay exactly 0.665 / 0.245 / 0.090. Only differences between logits count.

Sharp vs flat If logits were sat=5, dog=0, on=0 → almost all mass on sat. If logits were nearly equal → flatter, more uncertain lottery.

S2 · Why the lottery

Why not always pick “sat”?

Always-pick-top Every run: “The cat sat …” Fine for short factual fills. For stories / brainstorming: same bland path every time.

Draw from the lottery Sometimes “The cat sat …” Sometimes “The cat on …” Rarely something else. Variety comes from sampling. Next session: temperature & top-k/p knobs on this same lottery.

S2 · Glossary

Glossary · key concepts this session

Keep this frame visible while teaching · say the term, then the plain line

Context The tokens so far — the input to one next-token step (e.g. “The cat”).

Dictionary / vocabulary All allowed next pieces (~50k). Every row gets a logit each step.

Logit (raw score) Unnormalised preference for one dictionary row. Not a probability yet.

Softmax exp(logit) / sum(exp(logits)) → probabilities that sum to 1.

Next-token probability P(piece | context) after softmax — chance that piece comes next.

Argmax / always-pick-top Choose the single highest P. Deterministic.

Sampling Draw a piece at random according to the probabilities.

Zoom-in toy We show 3 rows for arithmetic; a real model scores the full dictionary.

S2 · Check

Exit ticket

Say these aloud · then quiz in matheion MA-LLM · Session 2

Prompt 1 Given context “The cat” and logits sat=2, dog=1, on=0 — compute the three probabilities.

Prompt 2 Where do sat/dog/on come from — a special shortlist, or dictionary rows?

Prompt 3 If we pick sat, what is the next step’s context?