chalkline

Public teaching whiteboard

Positional encodings in Transformers

How Transformers know token order without recurrence — positional encodings on a public teaching board. · by zlu

Topics: LLM & Transformers

Positional encodings in Transformers
S12 · Open
S12 · Problem
S12 · Sinusoids
S12 · Alternatives
S12 · Glossary
S12 · Check

Session 12 · Positional encodings

Live board · deepen in matheion MA-LLM · Session 12

Today’s story Bare attention doesn’t know order. Add a position vector to each token embedding — the classic paper uses sines and cosines.

You should leave able to… • Show order blindness with a swap • Write the sin/cos recipe in words + formula • Contrast fixed sinusoids vs learned positions

With matheion Use this board to teach. Open matheion → MA-LLM → Session 12 for full prose, diagrams, and the auto-quiz.

Attention alone doesn’t see order

Shuffle the tokens Without position info, ‘Dog bites man’ and ‘Man bites dog’ can look the same to the mechanism — it only sees a bag of vectors, not who came first.

Add a position vector

Each position gets a pattern of sines/cosines; add it to the token’s embedding (same width)

text
For position pos and dimension index i:
  even dims:  sin( pos / 10000^{2i / model_width} )
  odd dims:   cos( pos / 10000^{2i / model_width} )

Then:  input_vector = token_embedding + position_vector

Why sines? For a fixed jump k, the pattern at pos+k
is a simple transform of the pattern at pos — handy for relative distance.

Learned positions?

Alternative Train one vector per position index (clip at a max length). Paper: nearly the same quality. Sinusoids kept partly to help lengths longer than those seen in training.

Modern note Newer models use other position recipes. Same problem (order), newer parameterisations. Sin/cos is enough for this session.

Glossary · key concepts this session

Keep this frame visible while teaching · say the term, then the plain line

Positional encoding A vector added so the model knows token order.

Permutation blindness Bare attention treats a shuffled bag of tokens the same.

Sinusoidal PE Fixed sin/cos pattern per position and dimension (Attention Is All You Need).

Learned positions Train one vector per position index instead of using formulas.

Relative position Information about distance/offset between tokens, not only absolute index.

Exit ticket

Say these aloud · then quiz in matheion MA-LLM · Session 12

Prompt 1 Why must we inject position at all?

Prompt 2 Where do you add the position vector?

Prompt 3 One pro of sinusoids vs a learned table.

Drag to pan · scroll to zoom · read-only

Board contents

Text extracted from this public whiteboard for search and accessibility.

S12 · Open

Session 12 · Positional encodings

Live board · deepen in matheion MA-LLM · Session 12

Today’s story Bare attention doesn’t know order. Add a position vector to each token embedding — the classic paper uses sines and cosines.

You should leave able to… • Show order blindness with a swap • Write the sin/cos recipe in words + formula • Contrast fixed sinusoids vs learned positions

With matheion Use this board to teach. Open matheion → MA-LLM → Session 12 for full prose, diagrams, and the auto-quiz.

S12 · Problem

Attention alone doesn’t see order

Shuffle the tokens Without position info, ‘Dog bites man’ and ‘Man bites dog’ can look the same to the mechanism — it only sees a bag of vectors, not who came first.

S12 · Sinusoids

Add a position vector

Each position gets a pattern of sines/cosines; add it to the token’s embedding (same width)

For position pos and dimension index i: even dims: sin( pos / 10000^{2i / model_width} ) odd dims: cos( pos / 10000^{2i / model_width} ) Then: input_vector = token_embedding + position_vector Why sines? For a fixed jump k, the pattern at pos+k is a simple transform of the pattern at pos — handy for relative distance.

S12 · Alternatives

Learned positions?

Alternative Train one vector per position index (clip at a max length). Paper: nearly the same quality. Sinusoids kept partly to help lengths longer than those seen in training.

Modern note Newer models use other position recipes. Same problem (order), newer parameterisations. Sin/cos is enough for this session.

S12 · Glossary

Glossary · key concepts this session

Keep this frame visible while teaching · say the term, then the plain line

Positional encoding A vector added so the model knows token order.

Permutation blindness Bare attention treats a shuffled bag of tokens the same.

Sinusoidal PE Fixed sin/cos pattern per position and dimension (Attention Is All You Need).

Learned positions Train one vector per position index instead of using formulas.

Relative position Information about distance/offset between tokens, not only absolute index.

S12 · Check

Exit ticket

Say these aloud · then quiz in matheion MA-LLM · Session 12

Prompt 1 Why must we inject position at all?

Prompt 2 Where do you add the position vector?

Prompt 3 One pro of sinusoids vs a learned table.