The AI Learning Hub Journal

Identity, Audit Trails, and Defence in Depth

Every layer has a gap — the arrangement is the defenceattribution, a full trace, and detection are what turn six defeatable controls into a positionIDENTITY — THIS AGENT, ON BEHALF OF THIS USER, THIS RUNa distinct identity per agentowner · purpose · review date · decommission pathon behalf of this user, expressedauthorisation applies the user’s scopea run id on every downstream callattribution: which run, which user, which versionTHE TRACE — WHAT DID IT DO, WHY, AND WHAT INFLUENCED ITthe prompt as assembled, post-templatingretrieved items with source identifierstool calls: arguments, identity, resultpolicy + guardrail decisions — allows tooapproval events as displayedmemory writes with their contextcorrelated by run id — forensic when read afterwards, a control when it feeds detectionthe signals are sequence and destination: privileged call after ingestion, first-seen egress, canariesEACH CONTROL'S GAP, AND THE CONTROL EXPECTED TO COVER ITCONTROLWHAT IT DOES NOT COVEREXPECTED TO CATCH THATleast privilege [S]misuse within the granted scope remainsgates + deterministic policy checkssandbox [S]legitimate tool calls sit outside itleast privilege · approval gatesdefault-deny egress [S]delete, spend, send need no egressgates on irreversible actionsapproval gates [H]splitting, fatigue, misleading rationalepolicy checks · gate metrics watchedconstrained decoding [S]shapes structure, not intent or valuesvalidation + authz at every consumerguardrail classifiers [P]error rates; an evadable input surfacethe structural layers beneath it[S] STRUCTURAL — holds by construction[P] PROBABILISTIC — error rates; evadable[H] HUMAN — depends on attentionNO SINGLE FAILURE IS SUFFICIENT — WRITE DOWN EACH GAP AND NAME WHAT COVERS ITnever weight a probabilistic control like a structural one — a stack described without its gaps is marketing
Give each agent an identity and a full trace, then layer individually defeatable controls so no single failure is sufficient — and know which kind of control each one is before you lean on it.

Agents Need Identities, Not Borrowed Ones

An agent acting under a shared service identity is unattributable by construction: you cannot tell which run, which user, or which version produced an action, which makes both incident scoping and routine access review impossible. Give each agent a distinct identity, and make each run distinguishable — a run identifier that appears on every downstream call, in every log entry, and in every record the agent creates. Where a call acts for a user, the identity must express both facts: this agent, on behalf of this user, so that authorisation applies the user's scope while attribution records the agent. Treat these identities with the same lifecycle discipline as human accounts: an owner, a purpose, a review date, and a decommissioning path. Agent identities that outlive the feature they were created for are the AI-era version of the orphaned service account.

  • Distinct identity per agent, distinct correlatable identifier per run
  • Express both parties: this agent acting on behalf of this user
  • Authorisation uses the user's scope; attribution records the agent
  • Owner, purpose, review date, decommissioning path — same as any account

What a Usable Trace Contains

After an incident you need to answer three questions: what did the system do, why did it decide to, and what influenced that decision. That determines what to capture. Every model call with the prompt as actually assembled, after templating, retrieval and any context management, because a reconstruction from templates will not match what the model saw. Every retrieved item with its source and identifier, so a poisoned document can be located. Every tool call with full arguments, the acting identity, and the result. Every policy or guardrail decision, including allows, not only blocks. Approval events with what was displayed to the approver. Memory writes with their context. All correlated by run identifier and timestamped consistently. Then handle the trace as sensitive: it contains everything the model saw, which usually means user data, so apply access control, retention limits, and redaction of secrets at write time.

  • Capture the assembled prompt, not the template — the gap is where incidents hide
  • Log retrieved items with source identifiers so poisoned content can be located
  • Record allows as well as blocks; a policy decision is evidence either way
  • Traces contain everything the model saw — protect and retain them accordingly

Detection Built on the Trace

A trace that is only read after an incident is a forensic asset; a trace that feeds detection is a control. The signals worth building are mostly about sequence and destination rather than content. An unusual tool sequence for a given feature, particularly a privileged call following ingestion of external content in the same run. Outbound requests to destinations not previously seen for that agent. Retrieval of documents unrelated to the user's query. Sharp changes in refusal, error, or gate-approval rates, which often move before anything else is noticed. Canary records or documents that should never legitimately be retrieved or transmitted, which are cheap to plant and unambiguous when they fire. Volume anomalies in run length or token consumption. Each of these needs a defined owner and response, or it becomes another unread dashboard.

  • Privileged tool call after untrusted ingestion in the same run is the highest-value signal
  • Alert on new outbound destinations per agent, not just on blocked ones
  • Canary records and documents fire unambiguously and cost almost nothing
  • Rate changes in refusals, errors and approvals often move first

Defence in Depth, Stated Honestly

No control in this module holds on its own, and saying so plainly is what makes a security position credible. Filtering is evadable. Guardrails have error rates. Approval gates can be split, misled, or worn down. Constrained decoding shapes structure but not intent. Sandboxes contain code but not legitimate tool calls. Least privilege reduces reach but does not prevent misuse within the granted scope. The design goal is therefore not a control that cannot fail but an arrangement in which no single failure is sufficient: capability limited so a persuaded model can do little, egress denied so what it can reach cannot leave, gates on the actions that matter, structure enforced where output is consumed, and a trace good enough to detect and scope what still gets through. Write down, for each control, what it does not cover and which other control is supposed to catch that.

  • Every control here is individually defeatable — the arrangement is the defence
  • Aim for no single failure being sufficient rather than for a control that cannot fail
  • Document each control's gap and name the control that covers it
  • A stack described without its gaps is marketing, not a security position

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.