Identity, Audit Trails, and Defence in Depth
Agents Need Identities, Not Borrowed Ones
An agent acting under a shared service identity is unattributable by construction: you cannot tell which run, which user, or which version produced an action, which makes both incident scoping and routine access review impossible. Give each agent a distinct identity, and make each run distinguishable — a run identifier that appears on every downstream call, in every log entry, and in every record the agent creates. Where a call acts for a user, the identity must express both facts: this agent, on behalf of this user, so that authorisation applies the user's scope while attribution records the agent. Treat these identities with the same lifecycle discipline as human accounts: an owner, a purpose, a review date, and a decommissioning path. Agent identities that outlive the feature they were created for are the AI-era version of the orphaned service account.
- Distinct identity per agent, distinct correlatable identifier per run
- Express both parties: this agent acting on behalf of this user
- Authorisation uses the user's scope; attribution records the agent
- Owner, purpose, review date, decommissioning path — same as any account
What a Usable Trace Contains
After an incident you need to answer three questions: what did the system do, why did it decide to, and what influenced that decision. That determines what to capture. Every model call with the prompt as actually assembled, after templating, retrieval and any context management, because a reconstruction from templates will not match what the model saw. Every retrieved item with its source and identifier, so a poisoned document can be located. Every tool call with full arguments, the acting identity, and the result. Every policy or guardrail decision, including allows, not only blocks. Approval events with what was displayed to the approver. Memory writes with their context. All correlated by run identifier and timestamped consistently. Then handle the trace as sensitive: it contains everything the model saw, which usually means user data, so apply access control, retention limits, and redaction of secrets at write time.
- Capture the assembled prompt, not the template — the gap is where incidents hide
- Log retrieved items with source identifiers so poisoned content can be located
- Record allows as well as blocks; a policy decision is evidence either way
- Traces contain everything the model saw — protect and retain them accordingly
Detection Built on the Trace
A trace that is only read after an incident is a forensic asset; a trace that feeds detection is a control. The signals worth building are mostly about sequence and destination rather than content. An unusual tool sequence for a given feature, particularly a privileged call following ingestion of external content in the same run. Outbound requests to destinations not previously seen for that agent. Retrieval of documents unrelated to the user's query. Sharp changes in refusal, error, or gate-approval rates, which often move before anything else is noticed. Canary records or documents that should never legitimately be retrieved or transmitted, which are cheap to plant and unambiguous when they fire. Volume anomalies in run length or token consumption. Each of these needs a defined owner and response, or it becomes another unread dashboard.
- Privileged tool call after untrusted ingestion in the same run is the highest-value signal
- Alert on new outbound destinations per agent, not just on blocked ones
- Canary records and documents fire unambiguously and cost almost nothing
- Rate changes in refusals, errors and approvals often move first
Defence in Depth, Stated Honestly
No control in this module holds on its own, and saying so plainly is what makes a security position credible. Filtering is evadable. Guardrails have error rates. Approval gates can be split, misled, or worn down. Constrained decoding shapes structure but not intent. Sandboxes contain code but not legitimate tool calls. Least privilege reduces reach but does not prevent misuse within the granted scope. The design goal is therefore not a control that cannot fail but an arrangement in which no single failure is sufficient: capability limited so a persuaded model can do little, egress denied so what it can reach cannot leave, gates on the actions that matter, structure enforced where output is consumed, and a trace good enough to detect and scope what still gets through. Write down, for each control, what it does not cover and which other control is supposed to catch that.
- Every control here is individually defeatable — the arrangement is the defence
- Aim for no single failure being sufficient rather than for a control that cannot fail
- Document each control's gap and name the control that covers it
- A stack described without its gaps is marketing, not a security position
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.