Tokens and Context: The Units It Thinks In
It Does Not See Letters or Words
Before a model reads anything, your text is chopped into tokens — chunks that are usually a short word or a piece of a longer one. "Understanding" might become "under" plus "standing". Spaces, punctuation and emoji all become tokens too. The model never sees the letters inside a token as separate things, which is exactly why models have historically been bad at questions like "how many r's are in this word" or at counting characters. It is not stupid. It is looking at a unit that has the letters welded shut. Once you know this, a whole category of weird failures stops being mysterious.
- A token is roughly a word-fragment; common words are usually one token, rare ones split into several
- Letter-level tasks, precise counting and some rhyming puzzles fight against tokenisation
- Numbers get split oddly too, which is part of why raw arithmetic can go wrong
- Give it a calculator or code to run and the same model gets the sum right — the tool covers the weakness
The Context Window Is Its Whole World
Everything the model can see at once — your message, the conversation so far, any file you pasted, and its own reply as it writes it — has to fit inside a limit called the context window, measured in tokens. Inside that window it is remarkably capable. Outside it, nothing exists. This is why a long chat starts to feel like it has forgotten the beginning: the beginning may literally have fallen out of view, or is still there but competing with everything else for attention. It is also why pasting your actual essay draft in beats describing it, and why starting a fresh chat for a new topic usually gets you sharper answers.
- Context window = the maximum amount of text the model can hold in view for one response
- Modern models hold a lot, but quality often sags for details buried in the middle of very long inputs
- Practical move: paste the source material rather than summarising it from memory
- Practical move: start a new chat when you switch topic, so old text stops crowding the window
Why "Say Nothing About X" Sometimes Backfires
Because generation is prediction over the text in front of it, whatever is in the window influences what comes next — including things you told it to avoid. Mention a wrong answer while asking it not to use that answer, and you have just made those words highly available. The same mechanism works in your favour: put a good example in the window and the output starts to resemble it. This is the real reason "show, don't tell" is the strongest prompting technique there is. You are not persuading the model. You are changing the statistical neighbourhood it is generating inside.
- Positive instructions beat negative ones: describe what you want, not a list of what you do not
- Anything you paste becomes influence, including sloppy notes and half-finished ideas
- One good example in the prompt often outperforms three paragraphs of description
- If a chat has gone badly off track, editing the window (or restarting) beats arguing with it
Paste your own text and watch it split into tokens.
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.