rumblr Work in progressWIP

● The AI Primer · Lesson 8 · Part 1: how the model works inside

Tokenization

how text becomes the integers a model actually reads

You'll be able to explain BPE from scratch, byte-level tokens, why models miscount letters

Members · open during launch 25 min9 figures and diagrams8 interactive
Guide is what to use and when. How it works builds it from scratch. Math & code adds the formulas and the Python.

The lesson in one minute

What you'll be able to explain

  1. Text is split into subword pieces by BPE: start from characters or bytes and repeatedly merge the most frequent neighbouring pair.
  2. Byte-level BPE can encode anything with no unknown token.
  3. Cost, latency and context limits are counted in tokens; English is about 4 characters (0.75 words) per token, while code and non-English text need more tokens.

Level 1

The practitioner's guide

In one sentence

A tokenizer cuts text into pieces from a fixed vocabulary and hands the model a number for each piece; every price, rate limit, context window and latency figure you meet is counted in those pieces, and where their boundaries fall explains a whole family of odd model behaviour.

When you need it

On every call, because you are billed per token, but you feel it on particular days: when you estimate a bill, when a prompt overflows the context window, when users who write in Hindi or Japanese cost several times what English users cost, when a model miscounts the letters in a word or mis-adds two numbers, when a fine-tuning dataset is quoted in tokens, or when you switch model generations and the same prompt suddenly counts differently. The tell: a number that should be simple (how long is this text?) that you cannot answer without running the model's own tokenizer. This lesson's toy tokenizer, trained on English, spends 1.73 characters per token on an English sentence and 0.38 on the same sentence in Hindi, 4.5 times as many tokens for the same meaning. You never train a tokenizer unless you pretrain a model; you only ever count with one, and choose models partly by theirs.

Your options

From the quickest estimate to the most committed:

Option What it does What it gives you What it costs Where it lives
A rule of thumb About 4 characters or 0.75 words per token of English prose An instant estimate: 1,000 tokens is about 750 words Wrong by several-fold for code, JSON and non-English text Your head, and this lesson's estimate_tokens
Counting with the model's own tokenizer Runs the real tokenizer offline (tiktoken, Hugging Face tokenizers, SentencePiece) Exact counts for any model whose tokenizer ships with its weights The right tokenizer files per model family; counts do not transfer between families Your code
Counting through the API A vendor endpoint counts a whole request: system prompt, tools, images, PDFs Counts for models whose tokenizer is not published; free on Anthropic's API, with its own rate limit One request per count; a close estimate rather than the exact bill The API
Reading usage after each call Every response reports the input and output tokens it consumed The exact billed numbers, and the cache reads Nothing beyond logging it The API response
Shaping the text Prefers prose to nested JSON, trims whitespace and decimals, shortens identifiers in tool outputs Fewer tokens per call at the same meaning Engineering time, and a risk of hurting clarity Your prompt and your tool outputs
Choosing the model by its tokenizer Picks a family whose vocabulary covers your languages Shorter sequences and lower bills for non-English text Larger vocabularies cost embedding parameters, paid by the vendor The model card: vocab_size
Training or extending a vocabulary Learns merges from your own corpus The shortest sequences for your domain The model must be trained or retrained with it Your training stack

How to choose

Start from who runs the tokenizer.

  • A hosted model: count before you send when the size matters (context fit, routing between models) and read usage after every call for the bill. Recount when you change model generations: Anthropic's documentation says the tokenizer introduced with Claude Opus 4.7 produces about 30% more tokens for the same text than earlier models did.
  • An open model: the tokenizer ships with the weights; load the pair together and never count with another family's tokenizer, because the pieces and the ids are unrelated.
  • A multilingual product: budget per language with the real tokenizer, never with the English rule of thumb.
  • A task about letters or digits (spelling, counting characters, arithmetic): give the model a tool or spell the item out, because the tokens hide the letters.
  • A model with a chat template: apply it. Instruction-tuned models expect their special tokens in the right places, and a prompt built by hand produces ids the model was not trained on.
  • Whatever you pick, measure with the tokenizer of the model you will actually call. The rule of thumb is for napkins.

What it costs

Tokens are the unit of four bills.

  • Money. Two meters: this lesson's example call, 2,000 input tokens and a 500-token answer at $3 and $15 per million, costs $0.0135, and a million such calls cost $13,500 (estimate_cost). Hosted APIs price output tokens several times higher than input tokens (5× across Anthropic's price list), so a verbose answer costs more than a long prompt.
  • Latency. Each output token is one full step of generation, so answer length sets the wait; primer.ml.inference puts the step at about 4.78 ms for an 8-billion-parameter model on one GPU.
  • Context. The window is counted in tokens, so the 4.5× gap between English and Hindi in this lesson's run is also a 4.5× gap in how much text fits.
  • Vocabulary. A bigger kit means shorter sequences but a larger embedding table: GPT-2's 50,257 rows of width 768 are 38.6 million parameters, 32% of its smallest model (primer.ml.transformer). Llama 3 grew to a 128K vocabulary, 100K from the tiktoken kit plus 28K for other languages, and its paper reports better compression as a result.

What breaks

  • Letters the model never saw. " strawberry" is two ids in this lesson's tokenizer, straw and berry, not ten letters. Counting the r's means recalling a spelling, not reading one.
  • Digits that split by length. " 1234" is one token and " 12345" is two (" 1234", "5"), and "1234" at the start of a line is three, so place values never line up. Some newer tokenizers split digits into groups of at most three to soften this.
  • A leading space changes the token. " cat" is id 303 and "cat" is ids 99 and 268; the model must learn that both mean cat. A prompt that ends mid-word, or a prefix that gains a trailing space, changes the ids and breaks a cached prefix.
  • The same meaning at several times the price. Hindi in this lesson costs 4.5× English; Petrov et al. (2023) measured differences of up to 15 times between languages on production tokenizers. Budget per language.
  • The wrong tokenizer. Fine-tuning data tokenized with another family's kit trains the model on unrelated ids. Load tokenizer and weights from the same checkpoint.
  • Estimates carried across generations. A count measured on one model can be 30% short on its successor. Recount with the model id you will run.

In the wild

GPT-2 introduced byte-level BPE with a pre-tokenization regex, and most language-model tokenizers still follow that design. OpenAI's tiktoken ships the cl100k_base and o200k_base vocabularies and claims to be 3 to 6 times faster than a comparable open-source tokenizer; Hugging Face tokenizers and Google's SentencePiece cover the open models, SentencePiece writing spaces as the visible symbol ▁. BERT uses WordPiece and T5 uses Unigram. Llama 3's vocabulary is 128K entries, and Mistral's byte-fallback tokenizer guarantees that no character ever becomes an unknown token. Anthropic's count_tokens endpoint counts a full request (system prompt, tools, images and PDFs) free of charge, subject to a rate limit of 5,000 to 20,000 requests per minute by usage tier, and every Messages API response reports its usage. The papers are linked at the end of the lesson.

Go deeper

Level 2 trains a tokenizer by hand on "low lower lowest" in three merges, replays the merges to encode a word it never saw, drops to bytes so that nothing is ever unknown, meets WordPiece and Unigram, and prints the exact pieces behind each quirk above. If you only needed to count and budget, you are done.

Level 2

How it works, from scratch

Level 2 builds the kit of bricks from nothing, starting with what a token is.

Chapter 1

Tokens: building words from a fixed kit of bricks

Cut a sentence with three kits before reading why the middle one won.

Everyday picture A box of Lego with a fixed set of bricks. Common shapes have a ready-made brick; anything unusual gets assembled from smaller bricks. You can build anything, but some builds take many more bricks than others. A tokenizer is that kit for text: a fixed vocabulary of pieces (typically 32,000 to 200,000 of them), each with a number, its token id.

Tiny worked example With a kit that contains the, un, believ and able, "the" takes 1 brick and "unbelievable" takes 3: un + believ + able. The model never sees the letters; it sees a short list of ids such as [262, 403, 11009, 540].

Approach Kit size "unbelievable" Problem
Whole words huge (every form of every word, every name) 1 piece, if it's in the kit words not in the kit become <unk>; typos break
Single bytes 256 12 pieces sequences 4-5× longer, and attention cost grows with length squared
Subwords 32k to 200k ~3 pieces the middle ground every modern LLM uses

Figure 1 · Diagram

Reading it: text enters on the left and never reaches the model as text. The tokenizer cuts it into pieces from its fixed kit and swaps each piece for its number. Those numbers pick rows from the embedding table (see primer.ml.big_picture), and only those vectors reach the model. Everything the model knows about spelling it had to learn through this narrow window. (The ids shown are illustrative.)

In code: trained_tokenizer builds this lesson's toy kit, and ByteBPE.encode turns any text into its list of ids.

Why it matters Everything downstream is counted in tokens: price, speed, and how much fits in the context window.

Chapter 2

Byte Pair Encoding (BPE): learning the kit from data

Run the stenographer's loop yourself: count, merge, write the rule down.

Everyday picture A court stenographer inventing shorthand. Whatever pair of letters they write most often gets its own squiggle. Then the most common pair including that squiggle gets one too, and so on, until they have as many squiggles as they can remember.

Tiny worked example Training text "low lower lowest". Start from single letters and repeatedly merge the most frequent neighbouring pair (train_char_bpe reproduces this exactly):

Step Most frequent pair New token Text becomes
Start none none l o w · l o w e r · l o w e s t
1 l + o (3 times) lo lo w · lo w e r · lo w e s t
2 lo + w (3 times) low low · low e r · low e s t
3 low + e (2 times) lowe low · lowe r · lowe s t

At step 1, l+o and o+w are tied at 3; ties go to the pair seen first.

Figure 2 · Diagram

Reading it: training is a loop around two boxes. "Count adjacent pairs" looks at every neighbouring pair of symbols in the corpus; "Merge" glues the winner into one new symbol everywhere and appends that rule to an ordered list. The loop exits when the vocabulary reaches its target size. The output is not just a vocabulary but the ordered merge list, and encoding depends on that order.

The math and the code Each round picks the pair with the largest count:

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
two symbols that sit side by side l and o
one distinct word of the training text "lower"
"add up the following over every distinct word" low, lower, lowest
how many times word occurs in the text 1 each
how many times is immediately followed by inside 1 in each word
the pair's total, the number BPE ranks by 3

In words: "for each distinct word, count how often the pair appears inside it, multiply by how often the word occurs, and add everything up."

With the numbers: count(l, o) = 1·1 (low) + 1·1 (lower) + 1·1 (lowest) = 3; count(e, r) = 1·1 (lower) = 1. train_char_bpe computes the same totals with a Counter.

Level 3: in Python
# f_w: how often each distinct word occurs
f = {"low": 1, "lower": 1, "lowest": 1}
# n_w(a, b): a immediately followed by b inside w
def n(w, a, b):
    return sum(1 for i in range(len(w) - 1) if w[i] == a and w[i + 1] == b)
def count(a, b):
    # Σ_w f_w · n_w(a, b)
    return sum(f_w * n(w, a, b) for w, f_w in f.items())
count("l", "o"), count("e", "r")  # → (3, 1)

In code: train_char_bpe runs the count-and-merge loop and returns one MergeStep per round, holding the winning pair, its count, the new token and the text after the merge (one row of the table above).

Why it matters Frequent strings end up as single tokens and rare ones stay in pieces, which is exactly why token counts differ from word counts, and why a tokenizer trained mostly on English is cheap for English and expensive for everything else (section 6).

Chapter 3

Encoding new text: replay the merges in order

Everyday picture Following a recipe card: do step 1 everywhere it applies, then step 2, and so on. Skipping ahead would give a different dish.

Tiny worked example Encode "lowest" with the three merges learned above: l o w e s t → (merge 1) lo w e s t → (merge 2) low e s t → (merge 3) lowe s t. Result: 3 tokens. A word never seen in training still works: "slow" → s low; "newer" matches no merge and stays as letters.

Figure 3 · Diagram

Reading it: each box applies one merge rule, in the order the rules were learned, to every place it fits. The word shrinks from 6 symbols to 3. It stops when none of the remaining neighbouring pairs is in the rule list.

The code encode_char_bpe repeatedly finds, among the pairs present, the one learned earliest, merges it, and loops. Training and encoding agree because they apply rules in the same order.

Why it matters Encoding is deterministic: the same text always gives the same ids, which is what makes prompt caching and cost estimates possible.

Chapter 4

Byte-level BPE: nothing is ever unknown

The bytes under the text, and the tokens on top of them.

Everyday picture Every file on a computer, whatever it holds, is a sequence of bytes: numbers from 0 to 255. If the smallest bricks in the kit are those 256 byte values, then no text can ever be unbuildable: emoji, Chinese, source code or garbage all come apart into bytes.

Tiny worked example In UTF-8 (the standard way to store text as bytes) "a" is one byte, 97; "é" is two bytes, 195 169; "🙂" is four bytes, 240 159 153 130. An untrained byte-level tokenizer turns "é🙂" into 6 tokens. After training on English, " password" is one token, id 291, because it was frequent.

Figure 5 · Diagram

Reading it: encoding runs left to right. A regular expression (a text-matching pattern) first cuts the text into chunks, each keeping its leading space; merges never cross a chunk boundary. Each chunk becomes raw bytes (ids 0-255). The tokenizer then applies whichever learned merge came earliest among the pairs present, until none applies. The ids shown are what this module's toy tokenizer produces: "Reset" is two tokens (Re set), " password" is one. Decoding is the reverse lookup: every id maps to a fixed byte string, and concatenating them restores the exact original bytes. That is why a byte-level tokenizer round-trips any text.

The code ByteBPE is train_char_bpe with bytes instead of letters, ids 0-255 for the bytes and 256, 257, ... for each merge. GPT-2, GPT-4's cl100k_base, Llama 3 and most current LLMs work this way.

Figure 4 · Drawn from the lesson's code

250 300 350 400 450 500 vocabulary size (256 bytes + learned merges) 1.0 1.5 2.0 2.5 3.0 characters per token (held-out English) More merges compress text, with diminishing returns

Compression climbs steeply from 1.0 to 2.0 characters per token in the first 44 merges, then flattens near 3.1 by 500 entries

Reading it: the x-axis is vocabulary size: 256 raw bytes plus the number of merges learned. The y-axis is compression, how many characters of held-out English each token covers on average. At 256 (no merges) every byte is a token, so the value is 1. The first few dozen merges capture the most frequent pairs (" t", "he", " the") and the curve climbs steeply; later merges capture rarer strings and the curve flattens. (Our toy corpus runs out of repeated strings near 500; real corpora keep going.) Production tokenizers sit far to the right, at 50k to 200k entries, and reach about 4 characters per token on English. Bigger kits mean shorter sequences but a larger embedding table and output layer.

In code: ByteBPE.train learns the byte merges, ByteBPE.encode and ByteBPE.decode make the round trip, and compression_curve trains tokenizers of growing size and measures each one for the figure above.

Figure 6 · Interactive · computed from the lesson's code

Byte-level BPE tokenizer

Try it: the tokenizer below is this lesson's own, with its 144 learned merges. Drag the merges down to 0 and every byte is its own token; drag them back up and watch frequent pieces such as " the" and " password" fuse, one merge at a time, while characters per token climbs like the curve above. Then type a word the corpus never saw, or an accent or an emoji, and watch it stay in small pieces.

Why it matters No "unknown token" failures, ever, for any input. The price is that unfamiliar scripts fall back to near-byte level and cost many more tokens.

Chapter 5

The cousins: WordPiece and Unigram

Everyday picture BPE glues the most common pair. WordPiece glues the most inseparable pair: two pieces that almost never appear apart, like "Q" and "u" in English. Unigram works the other way round: start with a huge kit and keep throwing out the bricks you'd miss least.

Tiny worked example On "low lower lowest", BPE's first merge is l+o. WordPiece's first merge is s+t: "s" and "t" each appear exactly once, and always together.

Figure 7 · Diagram

Reading it: the two top rows grow a vocabulary bottom-up and differ only in how they rank candidate pairs. The bottom row prunes top-down. All three end with a fixed kit and a rule for cutting new text into pieces from it.

The math WordPiece ranks pairs by

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
how often is immediately followed by count(st) = 1
how often appears at all count(s) = 1
how often appears at all count(t) = 1
fraction bar divide: pairs are rewarded when their parts rarely appear apart

In words: "how often the two appear together, divided by how often each appears at all."

With the numbers: score(s, t) = 1 / (1 · 1) = 1.0; score(l, o) = 3 / (3 · 3) = 0.33. WordPiece merges s+t first (wordpiece_scores).

Level 3: in Python
f = {"low": 1, "lower": 1, "lowest": 1}
# how often a symbol, or a pair written together, appears
def count(piece):
    return sum(f_w * w.count(piece) for w, f_w in f.items())
def score(a, b):
    # count(ab) / (count(a) · count(b))
    return count(a + b) / (count(a) * count(b))
score("s", "t"), round(score("l", "o"), 2)  # → (1.0, 0.33)

BERT marks continuation pieces with ## ("un", "##believ", "##able"). SentencePiece is a library that runs BPE or Unigram directly on raw text, writing spaces as the visible symbol ▁.

Why it matters Different model families segment the same text differently, so their token counts and costs differ. Always count with the tokenizer of the model you'll actually call.

The two rankings from the chapter, side by side on the same pairs.

Chapter 6

Tokens are the unit of cost, speed and memory

Turn the meters before reading the formula behind them.

Everyday picture A taxi meter that ticks per token, not per mile, with a pricier meter for the return trip: output tokens usually cost several times more than input tokens.

Tiny worked example A prompt of 2,000 tokens with a 500-token answer, at example prices of $3 per million input tokens and $15 per million output tokens: 0.006 + 0.0075 = $0.0135. A million such calls cost $13,500.

Figure 9 · Diagram

Reading it: two meters run on every call. Input tokens are counted once on the way in; they also have to fit in the model's context window (the dotted arrow). Output tokens are counted on the way out, and each one also takes a full step of generation, so they drive latency as well as cost.

The math and the code

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
input tokens in the call 2,000
output tokens generated 500
one million, because prices are quoted per million tokens 1,000,000
price per million input tokens, in dollars 3
price per million output tokens, in dollars 15

In words: "input tokens in millions times the input price, plus output tokens in millions times the output price."

With the numbers: 2,000/1,000,000 × 3 + 500/1,000,000 × 15 = 0.006 + 0.0075 = $0.0135 (estimate_cost).

Level 3: in Python
T_in, T_out = 2_000, 500
# dollars per million tokens
p_in, p_out = 3, 15
round(T_in / 10**6 * p_in + T_out / 10**6 * p_out, 4)  # → 0.0135

Rule of thumb: English prose averages about 4 characters per token, or 0.75 words per token, so 1,000 tokens ≈ 750 words (estimate_tokens, tokens_to_words). Use it for quick estimates; count with the real tokenizer for anything that matters.

Figure 8 · Drawn from the lesson's code

English prose German Python code JSON Hindi Emoji 0.0 0.5 1.0 1.5 2.0 2.5 characters per token (higher = cheaper) Same tokenizer, very different cost by content

English prose gets 2.6 characters per token; German, code and JSON 1.2 to 1.4; Hindi and emoji under 0.4

Reading it: each bar is one kind of text encoded by the same English-trained tokenizer from this module; taller bars mean each token carries more text, so the content is cheaper. English prose scores highest because the merges were learned from English. Code and JSON score lower because symbols and unusual identifiers were rarely merged. German shares the alphabet but not the words. Hindi and emoji fall below one character per token: each character is 3-4 UTF-8 bytes and few of those byte pairs were ever merged, the worst case of byte-level fallback.

In code: chars_per_token computes each bar: the length of the text divided by the number of ids ByteBPE.encode returns for it.

Why it matters The same request can differ several-fold in cost, latency and context usage depending on language and content: dense JSON, code, tables and non-English text all need more tokens than English prose.

Chapter 7

Quirks the tokenizer explains

Everyday picture Reading through frosted glass that only shows whole bricks. You can tell which bricks are there, but not the letters printed on them.

Tiny worked example (this module's toy tokenizer):

Text Tokens What it explains
" cat" vs "cat" [303] vs [99, 268] (c at) a leading space makes a different token; the model must learn they mean the same
" 1234" vs " 12345" [' 1234'] vs [' 1234', '5'] digits split by length and context, so place values don't line up: one reason arithmetic is hard
" strawberry" [' straw', 'berry'] 2 ids, not 10 letters: "how many r's?" is hard

Figure 10 · Diagram

Reading it: the question is about letters, but the letters were discarded before the model saw anything. It has to have memorized how straw and berry are spelled from seeing them in text, which is why spelling and letter-counting questions trip up models that handle much harder reasoning.

The code ByteBPE.tokens(text) shows the pieces for any string; the demo prints these cases.

Why it matters These quirks explain surprising behaviour: miscounted letters, shaky arithmetic, odd handling of rare names. Some newer tokenizers split digits into fixed groups of up to 3 to reduce the arithmetic problem.

Test yourself

7 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1Why subwords instead of whole words or characters?Think it through, then reveal

Words give a huge vocabulary and unknown words; characters make sequences several times longer, and attention cost grows with the square of length. Subwords keep common strings short and still encode anything.

Question 2How does BPE training proceed, step by step, on "low lower lowest"?Think it through, then reveal

Split into letters, count neighbouring pairs, and merge the most frequent: l+o (3), then lo+w (3), then low+e (2). Record each merge; to encode, replay the merges in the order learned.

Question 3Why can a byte-level BPE tokenizer never produce an unknown token?Think it through, then reveal

Its base vocabulary is all 256 byte values, and every string is a sequence of UTF-8 bytes, so in the worst case text falls back to single bytes.

Question 4What does WordPiece do differently from BPE?Think it through, then reveal

It merges the pair with the highest count(ab) / (count(a)·count(b)), which favours pairs whose parts rarely appear apart, instead of the most frequent pair.

Question 5Your users write in Hindi. What changes in your estimates?Think it through, then reveal

Expect noticeably more tokens for the same content: higher cost and latency, and less text per context window. Measure with the actual tokenizer.

Question 6Why are LLMs bad at counting letters and at long arithmetic?Think it through, then reveal

They see token ids, not characters, and numbers split into chunks that don't line up with place value.

Question 7What does a call with 2,000 input and 500 output tokens cost at $3/$15 per million?Think it through, then reveal

0.006 + 0.0075 = $0.0135. And 1,000 tokens is about 750 English words.

Primary sources

The papers behind this lesson

Sennrich, Haddow & Birch (2015), Neural Machine Translation of Rare Words with Subword Units.

Brought byte pair encoding, a 1990s compression trick, to language models as a way to build open-vocabulary subword units.

Read the annotated companion →The paper ↗
Radford et al. (2019), Language Models are Unsupervised Multitask Learners (GPT-2).

Introduced byte-level BPE with a pre-tokenization regex, the design most LLM tokenizers still follow.

The paper ↗
Kudo & Richardson (2018), SentencePiece.

A language-independent tokenizer library that works on raw text and encodes spaces as the symbol ▁.

The paper ↗
Kudo (2018), Subword Regularization.

Introduced the Unigram language-model tokenizer that prunes a large vocabulary down instead of merging up.

The paper ↗

Researcher's shelf

Further reading

  • Sennrich et al., Neural Machine Translation of Rare Words with Subword Units (the BPE paper, 2015): https://arxiv.org/abs/1508.07909
  • Andrej Karpathy, Let's build the GPT Tokenizer (video): https://www.youtube.com/watch?v=zduSFxRajkE
  • Karpathy's minbpe (minimal, clean byte-level BPE): https://github.com/karpathy/minbpe
  • OpenAI tiktoken (fast BPE used by GPT models): https://github.com/openai/tiktoken
  • Hugging Face LLM course, BPE chapter: https://huggingface.co/learn/llm-course/chapter6/5
  • Hugging Face LLM course, WordPiece chapter: https://huggingface.co/learn/llm-course/chapter6/6
  • Hugging Face LLM course, Unigram chapter: https://huggingface.co/learn/llm-course/chapter6/7
  • Kudo & Richardson, SentencePiece (2018): https://arxiv.org/abs/1808.06226
  • Kudo, Subword Regularization (the Unigram LM, 2018): https://arxiv.org/abs/1804.10959
  • Petrov et al., Language Model Tokenizers Introduce Unfairness Between Languages (2023): https://arxiv.org/abs/2305.15425

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.