The AI Learning Hub Journal

Tokens and Context: The Units It Thinks In

Tokens, and the window they have to fit ina model does not read words — it reads piecesUnbelievably,tokenisationsplitsordinarywords.one word, three tokenscommon words, one token eachspaces, commas and full stops are tokens too — rare words break into more pieces than common onesEARLY IN THE CHAT — ROOM TO SPARESystem instructionsYour first questionThe answer you got backSPACE STILL FREELATER — THE WINDOW IS FULLYour first questionDROPPEDSeveral early turns, squashed into one summarySUMMARISEDA later answerYour newest questionThe reply being written right nowOnly what is still inside the windowexists as far as the model is concernedWhen it seems to forget, the detailwas dropped or summarised awayRe-paste what matters, or start afresh chat — that is the whole fixA long chat does not make the window bigger — it just decides what gets pushed out of it first
Memory in a chat is not storage — it is a fixed-size window, and anything pushed out of it is gone unless you put it back

It Does Not See Letters or Words

Before a model reads anything, your text is chopped into tokens — chunks that are usually a short word or a piece of a longer one. "Understanding" might become "under" plus "standing". Spaces, punctuation and emoji all become tokens too. The model never sees the letters inside a token as separate things, which is exactly why models have historically been bad at questions like "how many r's are in this word" or at counting characters. It is not stupid. It is looking at a unit that has the letters welded shut. Once you know this, a whole category of weird failures stops being mysterious.

  • A token is roughly a word-fragment; common words are usually one token, rare ones split into several
  • Letter-level tasks, precise counting and some rhyming puzzles fight against tokenisation
  • Numbers get split oddly too, which is part of why raw arithmetic can go wrong
  • Give it a calculator or code to run and the same model gets the sum right — the tool covers the weakness

The Context Window Is Its Whole World

Everything the model can see at once — your message, the conversation so far, any file you pasted, and its own reply as it writes it — has to fit inside a limit called the context window, measured in tokens. Inside that window it is remarkably capable. Outside it, nothing exists. This is why a long chat starts to feel like it has forgotten the beginning: the beginning may literally have fallen out of view, or is still there but competing with everything else for attention. It is also why pasting your actual essay draft in beats describing it, and why starting a fresh chat for a new topic usually gets you sharper answers.

  • Context window = the maximum amount of text the model can hold in view for one response
  • Modern models hold a lot, but quality often sags for details buried in the middle of very long inputs
  • Practical move: paste the source material rather than summarising it from memory
  • Practical move: start a new chat when you switch topic, so old text stops crowding the window

Why "Say Nothing About X" Sometimes Backfires

Because generation is prediction over the text in front of it, whatever is in the window influences what comes next — including things you told it to avoid. Mention a wrong answer while asking it not to use that answer, and you have just made those words highly available. The same mechanism works in your favour: put a good example in the window and the output starts to resemble it. This is the real reason "show, don't tell" is the strongest prompting technique there is. You are not persuading the model. You are changing the statistical neighbourhood it is generating inside.

  • Positive instructions beat negative ones: describe what you want, not a list of what you do not
  • Anything you paste becomes influence, including sloppy notes and half-finished ideas
  • One good example in the prompt often outperforms three paragraphs of description
  • If a chat has gone badly off track, editing the window (or restarting) beats arguing with it
◆ See it for yourself
Open the Tokeniser in the library →

Paste your own text and watch it split into tokens.

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.