Tokenisation In Depth
BPE: How Text Becomes Integers
Models never see characters or words — they see integer IDs from a fixed vocabulary, and byte-pair encoding (BPE) is how that vocabulary gets built. Start from raw bytes, count which adjacent pairs co-occur most often in a training corpus, merge the most frequent pair into a new token, and repeat tens of thousands of times. Common words end up as single tokens, rare words split into recognisable pieces, and anything at all — any language, emoji, binary junk — remains representable because the 256 base bytes are always in the vocabulary. Tokenisation is a compression scheme with consequences: it fixes how much text fits in a context window, how much a request costs, and where the model's blind spots are, before a single parameter is involved.
- Byte-level BPE guarantees any input is encodable — no out-of-vocabulary failures
- Merges are learned from corpus statistics: frequent strings become cheap, rare ones expensive
- The tokeniser is trained separately, before the model, and frozen forever after
Modern Tokenisers: Bigger Vocabularies, Fairer Coverage
Tokeniser design moved on. The ~100K-vocabulary tokenisers of the GPT-4 era (cl100k) gave way to vocabularies around 200K and beyond — o200k-class tokenisers and comparable designs from other labs — trained on more multilingual and code-heavy corpora. Larger vocabularies compress text into fewer tokens, which directly means more content per context window and lower cost per document, with the biggest gains for non-English languages that older, English-centric tokenisers fragmented badly. A Hindi or Thai sentence that once cost several times its English equivalent now tokenises far more fairly, though inequality hasn't vanished. Code also tokenises efficiently — whitespace runs and common idioms get dedicated tokens. When comparing model costs, remember the unit itself differs: a "token" is not the same amount of text across vendors.
- Modern ~200K vocabularies compress better than the older ~100K generation — cl100k is legacy, not current
- Vocabulary size trades embedding-table memory against sequence length — 200K-ish is the current sweet spot
- Cross-vendor cost comparisons need tokens-per-document tests on your actual data, not list prices
Tokenisation Failure Modes
A surprising share of "the model is being weird" traces to the tokeniser. Character-level questions — count the letter r, reverse this string — are hard because the model literally does not see letters, only multi-character chunks (the famous strawberry problem). Numbers split inconsistently, complicating arithmetic. Glitch tokens — vocabulary entries whose training data was filtered out after the tokeniser was built — can trigger bizarre completions because their embeddings were never properly trained. Trailing whitespace or unusual Unicode can flip which merge path a string takes and measurably change output quality. The diagnostic habit worth building: when a failure looks inexplicable, run the input through a tokeniser visualiser before blaming the model — you'll often find the boundary sitting exactly where the behaviour breaks.
- Character-level tasks fail structurally: the model sees chunks, not letters
- Glitch tokens: undertrained vocabulary entries with unpredictable behaviour
- Tokeniser visualisers are a first-line debugging tool, not a curiosity
Try It Yourself
Everyone quotes the "one token is about four characters" rule of thumb and then budgets as if a token were a word. Measure your own ratio instead — it takes five minutes and it changes your cost model.
Take a real 300–500 word sample of the text your product actually sends — your own writing, a support ticket, a chunk of your code, a document in the language your users write in. Run it through the tokeniser playground on this lesson and record two numbers: W (words, from any word count) and T (tokens). Compute the ratio R = T ÷ W. Then price your real workload with it, using current input pricing from your provider's pricing page — prices change often, so look them up rather than reusing a number you remember.
Sample: W = [words] , T = [tokens] Ratio R = T ÷ W = [____] Workload: [N] requests per month × [P] words of input per request Input tokens/month = N × P × R = [____] Cost/month = (input tokens ÷ 1,000,000) × [price per 1M input tokens] = [____] Now repeat both lines for output tokens — output is priced separately and usually higher.
- You have a number for R, not an assumption — ordinary English prose lands near 1.3 tokens per word, but code, rare names and non-English text run higher
- Redo the cost line with R = 1 (one token per word). The gap between that and your real figure is the size of the mistake most budgets contain
- Your total includes output tokens priced at the output rate, not the input rate
- Tokenise the same content in a second vendor's tokeniser — if T differs, per-million prices are not directly comparable
Paste your own text and watch it split into tokens.
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.