Board contents
Text extracted from this public whiteboard for search and accessibility.
S1 · Open
Session 1 · From text to tokens
Live board · deepen in matheion MA-LLM · Session 1
Today’s story Nets need numbers. We’ll cut one sentence three ways, then build a tiny ‘merge frequent neighbours’ recipe by hand — that’s how modern tokenisers invent pieces like est.
You should leave able to… • Cut one sentence three ways and compare token counts • Merge frequent letter-pairs by hand until ‘est’ appears • Predict how ‘unhappiness’ splits, and why that beats ‘unknown’
With matheion Use this board to teach. Open matheion → MA-LLM → Session 1 for full prose, diagrams, and the auto-quiz.
S1 · Same sentence ×3
Same sentence, three cuttings
Running example: "lowest prices"
TEXT = "lowest prices" 1) CHARACTERS — cut every letter (and the space) l o w e s t _ p r i c e s → 13 pieces. Model must learn that l-o-w-e-s-t means lowest. 2) WHOLE WORDS — cut on spaces only lowest | prices → 2 pieces. But the dictionary must list every form: low, lower, lowest, prices, price, priced, typos… 3) SUBWORDS — reuse shared pieces low | est | prices → 3 pieces. The bit ‘est’ can be reused in newest, highest…
What to say aloud Characters → many pieces, hard learning. Whole words → few pieces, huge dictionary. Subwords → share pieces across words so rare forms still get signal.
S1 · Subword examples
What subword cuts look like in practice
These are the kinds of pieces a trained tokeniser actually emits
word → pieces why it helps ──────────────────────────────────────────────────────────── the → the very common → keep whole playing → play | ing ‘ing’ reused everywhere newest → new | est ‘est’ shared with highest… unhappiness → un | happi | ness prefix + stem + suffix tokenization → token | ization long rare stem, shared ending ChatGPT → Chat | G | PT rare name → fall back to pieces 12345 → 12 | 345 digits often chunked Notice: common stuff stays short; rare stuff reuses known bits.
S1 · Merge recipe
How the vocabulary gets built
Idea first: repeatedly glue the most common neighbouring pair
1 · Start from letters Every character (or byte) is already a legal piece. Any string can be written.
2 · Count neighbours Scan lots of text. Which two pieces sit next to each other most often?
3 · Glue the winner Make that pair a new single piece. Rewrite the text with the glue applied.
4 · Repeat until big enough Stop at a chosen dictionary size (tens of thousands). This merge-by-frequency recipe is what people call BPE (byte-pair encoding). Closely related systems (e.g. WordPiece) use the same ‘build pieces from parts’ idea with a different scoring rule.
S1 · Merge by hand
Worked example · invent ‘est’
Toy corpus: lowest · newest · highest · testing
Start (letters only): l o w e s t n e w e s t h i g h e s t t e s t i n g Neighbour ‘e s’ and then ‘s t’ show up a lot → glue them. After gluing e+s → es, then es+t → est: l o w est n e w est h i g h est t est i n g Later glues may build ‘low’, ‘new’, ‘high’, ‘test’… At cut time, ‘newest’ becomes new | est (2 pieces), not 6 letters.
Student check If we never glued further for ‘best’, how does it cut? b | est Same ‘est’ piece as in newest — that’s why subwords beat a frozen word list.
S1 · IDs & round-trip
Each piece gets an integer ID
text = "newest prices" Suppose the dictionary rows are: 1042 → ‘new’ 881 → ‘est’ 55 → ‘ prices’ (leading space kept as part of the piece) encode → [1042, 881, 55] decode → ‘newest prices’ (must match exactly) Vocabulary size = number of rows ≈ 30,000–100,000 (GPT-2’s dictionary had 50,257 entries). Token count for this string = 3 English word count was 2 — never confuse the two.
Pitfall · teach this hard ‘How many tokens?’ is NOT ‘how many words?’ ‘I love ChatGPT’ might become 5–7 tokens depending on the tokeniser. Always run encode before claiming a length.
S1 · Why this design wins
Design win · one example
Rare word: ‘unhappiness’ If the dictionary only stores whole words and has never seen this one, the model gets an ‘unknown’ placeholder — almost no meaning. With subwords: un | happi | ness Those bits already appeared in unhappy, happiness, kindness… so the pieces already carry signal.
Common word: ‘the’ Stays ONE piece. No point splitting the most frequent English word into t|h|e. Rule of thumb: frequent → keep whole; rare → reuse known bits.
S1 · Glossary
Glossary · key concepts this session
Keep this frame visible while teaching · say the term, then the plain line
Token A piece of text the model treats as one unit — often a subword, not always a whole word.
Tokeniser The program that cuts a string into tokens and maps them to integer IDs (and back).
Vocabulary / dictionary The fixed list of allowed tokens. Size ≈ tens of thousands of rows.
Subword A reusable fragment (ing, est, …). Rare words are built from known fragments.
BPE Byte-pair encoding: repeatedly glue the most frequent neighbouring pair to grow the vocabulary.
Encode / decode String → list of IDs, and IDs → string. Must round-trip exactly.
Token count vs word count How many tokens ≠ how many English words. Always ask the tokeniser.
Unknown placeholder What a whole-word system does for never-seen words — almost no meaning.
S1 · Check
Exit ticket
Say these aloud · then quiz in matheion MA-LLM · Session 1
Prompt 1 Cut ‘lowest prices’ three ways. Who creates the most pieces? Who needs the biggest dictionary?
Prompt 2 From the toy corpus, show how ‘est’ gets invented by gluing neighbours.
Prompt 3 Predict pieces for ‘unhappiness’ and why that’s better than an unknown placeholder.