Trained Once, Used Millions of Times
Two Completely Different Phases
There is a moment when a model is built and a moment when it is used, and they have almost nothing in common. Training happens once, takes months of work, runs on enormous clusters of specialised chips, and costs an amount of money that only large organisations can spend. It produces a fixed set of numbers — the model's weights. After that, using the model is a comparatively small computation that happens every time anyone sends a message. When you chat, you are not training anything. You are running a finished artefact, the same frozen set of numbers everyone else is running.
- Training: one-time, enormous, produces the weights that define the model
- Inference: what happens on every single message, fast and comparatively cheap
- Your conversation does not update the weights — the model does not "learn from you" mid-chat
- Newer versions come from new training runs, not from the model quietly improving on its own
What "Learning From Data" Actually Means
It does not mean the model stored the internet somewhere and looks things up. Training adjusts billions of numerical weights so that the model gets better at one task: predicting the next token in real text. Statistical regularities in language get compressed into those weights — grammar, facts that appear consistently, the shape of a good argument, how code is structured, and also every bias and error that recurs in the source material. What survives compression is what was common and consistent. Rare details get blurred or lost, which is precisely where confident invention creeps in later.
- The output of training is weights — numbers — not a stored copy of the training text
- Frequently repeated, consistent information survives compression well; one-off details often do not
- Patterns in the data become patterns in the output, including the unfair ones
- A later stage tunes the model with human feedback so it answers helpfully rather than just continuing text
Knowledge Cutoffs and Why Memory Is a Product Feature
Because the weights were fixed at the end of training, a model's built-in knowledge stops at a certain point. Anything after that is invisible unless the product goes and fetches it — which is what happens when a tool searches the web, reads a file you upload, or pulls from a company's documents. Similarly, when an assistant "remembers" your name across sessions, that is the app storing text and quietly re-inserting it into the context window. Both memory and up-to-date knowledge are things built around the model, not properties of the model itself. Knowing the difference tells you what to trust.
- Knowledge cutoff: the model has no built-in awareness of events after training ended
- Search, file upload and retrieval add fresh information by putting it into the context window
- Persistent "memory" is stored text replayed to the model, not the model recalling you
- When accuracy on recent events matters, use a tool that cites sources and check them
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.