The AI Learning Hub Journal

LLM Mechanisms

Input text:"Cybersecurity is hard"Tokenized:Cybersecurity is hard~750 words ≈ 1,000 tokens. A 1M token context = roughly 750K words.
One word, multiple tokens — pricing and limits are measured in tokens, not words

Tokens: The Unit of Everything

LLMs read and write in tokens — typically 3–4 characters of English text. Pricing, throughput, and context limits are all measured in tokens, not words. Rule of thumb: 750 words ≈ 1,000 tokens. Tokens extend beyond text: images are tokenized as patches (a 1024×1024 image can cost hundreds to thousands of tokens), video as frames over time (30 seconds of video can exceed tens of thousands of tokens), and audio as discrete sound units. Every modality maps to tokens — every token has a cost.

  • A typical phishing email plus headers ≈ 500–1,500 tokens; a SIEM alert with enrichment ≈ 200–800 tokens
  • A 10-page incident report ≈ 4,000–6,000 tokens; a 500-page IR engagement document ≈ 200,000+ tokens
  • A 30-second deepfake video sample for analysis can exceed 50,000 tokens — very different cost profile than text analysis

Context Windows and the Lost-in-the-Middle Problem

The context window is the maximum tokens a model can see at once — input plus output combined. Modern frontier models offer 200K to 2M+ tokens. But longer context does not equal better answers: models often degrade in the middle of very long contexts (the "lost in the middle" problem). Understanding this prevents the common mistake of treating context window size as a substitute for retrieval architecture.

  • Large context window does not mean large context quality — attention degrades non-linearly at long ranges
  • The right answer for large documents is usually retrieving the relevant slice at query time (RAG), not stuffing everything in the prompt
  • Test long-context performance at the p90 document size of your customer's use case — degradation is non-linear and will surprise you

Temperature: Controlling Output Randomness

Temperature determines how deterministic the model's output is. Low temperature: the model reliably picks the most probable next token — useful for structured extraction, classification, and code. High temperature: the model samples more broadly, producing varied responses. Most production security deployments use low to moderate temperature. This explains why analysts sometimes get different answers to the same query — and why a vendor's temperature configuration is a meaningful security question.

  • Temperature 0: maximally deterministic — the model always picks the most probable token; right for security classification tasks
  • Temperature > 0: introduces variance — the same prompt can yield different outputs on every call; risky for high-stakes decisions
  • Nucleus sampling (top-p): an alternative to temperature that limits token candidates to a cumulative probability threshold

Embeddings: Meaning as Math

Embeddings are not just a RAG tool — they are the mechanic at the heart of how LLMs process language. Before generating any token, the model converts every input into a high-dimensional numerical vector. Semantically similar items end up near each other in this vector space. This is what lets models "understand" meaning rather than match strings, find related concepts across paraphrases, and power the retrieval layer in RAG systems. The same mechanism that powers LLM understanding also powers semantic search and threat similarity detection.

  • Every token is converted to an embedding before the model processes it — embeddings are the model's internal language
  • Semantic similarity = geometric closeness: related concepts cluster together in vector space even with no shared words
  • A "phishing email" and "credential harvesting message" cluster together — enabling detection of novel variants with no keyword overlap

Vector Search: Semantic Detection at Scale

Once content is embedded as vectors, you can ask: what items are closest to this query in vector space? That is vector search — also called semantic search. Unlike keyword search (which matches literal words), it matches meaning. A query for credential theft will surface documents about phishing, password dumps, and token harvesting even when those exact words do not appear. This capability is foundational for finding similar past incidents, detecting novel attack variants, and surfacing related threat intelligence.

  • Semantic search powers "find similar incidents" — the most requested SOC analyst workflow after natural-language querying
  • Vector search detects novel phishing variants with no shared keywords — this is a quantifiable detection coverage gain
  • Ask vendors: what embedding model does the product use, and how often is it retrained on new threat data? Stale embeddings mean degraded detection
phishing emailcredential harvestingBEC fraudmalware sampleransomware payloadtrojan dropperfirewall lognetwork telemetryflow dataIdentity attack clusterMalware clusterNetwork telemetry cluster2D projection of high-dimensional embedding space
Semantically similar concepts cluster together — embeddings power similarity search
◆ See it for yourself
Open the Tokeniser in the library →

Paste your own text and watch it split into tokens.

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.