The AI Learning Hub Journal

Excessive Agency and Exfiltration Paths

More capability than the task requires — in three distinct wayseach dimension has its own fix; an approval prompt does not shrink the service account behind itEXCESSIVE FUNCTIONALITYthe toolset includes actions thetask never needs — a serverexposed a family; all registeredFIX: TRIM THE TOOL LISTEXCESSIVE PERMISSIONcredentials carry broader scopethan the tool needs — an inheritedservice account, reusedFIX: SCOPE THE CREDENTIALEXCESSIVE AUTONOMYconsequential actions execute withno check beyond the model’sown decisionFIX: GATE THE ACTIONAUDIT WHAT IS GRANTED, NOT WHAT WAS INTENDEDpull the live tool list · resolve effective permissions, inherited and group-derived grants includedwhose authority applies at each call? service-account retrieval is the classic confused deputyTHE EXFILTRATION INVENTORY — EVERY PATH BYTES CAN LEAVE BYagent holdingprivate contextdirect network calls from tools or coderendered output fetching remote resourceswrites to docs, repos, tickets, channelsmessages — email, chat, notificationsquery strings to third-party serviceslogs and telemetry leaving the boundaryerror strings returned to callersthe visible responsecan look normalPROVE CLOSURE WITH A MARKER, NOT AN ARGUMENTplant a benign marker in private context and attempt each path — closure claims decay as features are added
Functionality, permission and autonomy are fixed by trim, scope and gate respectively — then enumerate every outbound path and prove each one closed.

Three Dimensions of Excessive Agency

Excessive agency is the structural condition in which an agent holds more capability than its task requires, and it decomposes into three dimensions that are fixed in different ways. Excessive functionality: the toolset includes actions the task never needs, often because a server exposes a family of operations and all of them were registered. Excessive permission: the credentials behind the tools carry broader scope than the tool needs, typically because a service account accumulated grants over time and is now reused. Excessive autonomy: the runtime executes consequential actions without an independent check, so the model's decision is the final control. Each dimension has a distinct remedy — trim the tool list, scope the credential, add a gate — and conflating them produces the common outcome where a team adds an approval prompt and leaves a broadly privileged service account untouched behind it.

  • Functionality: tools registered that the task never needs
  • Permission: credentials broader than the tool requires, usually an inherited service account
  • Autonomy: consequential actions executed with no check beyond the model's decision
  • Each has a different fix; a gate does not shrink a credential
  • On the standard maps: OWASP's LLM Top 10 lists this as excessive agency; the leak side is its sensitive information disclosure entry

Audit Granted, Not Intended

The reliable way to find excessive agency is to stop reading design documents and enumerate what is actually granted. Pull the live tool list the agent receives, including anything added by a registered server, and compare it against the tasks the feature performs. Pull the effective permissions of every credential in use — not the role name, the resolved permission set including inherited and group-derived grants — and check what data each could reach if the agent asked. Determine whose authority applies at each tool call: the requesting user's, or the agent's own. That last check finds confused-deputy conditions, where an agent answers using access the asking user does not have, and it is one of the most common serious findings in real systems because the shortcut of using a service account for retrieval is convenient and rarely revisited.

  • Enumerate the live tool list, including tools added by registered servers
  • Resolve effective permissions, including inherited and group-derived grants
  • Determine whose authority applies per call — the user's or the agent's
  • Service-account retrieval is the usual root of cross-user disclosure findings

The Exfiltration Inventory

Outbound paths are more numerous than teams expect, and enumerating them is the prerequisite for containing them. Direct network calls from tools or executed code. Content rendered in a client that fetches remote resources — images, link previews, stylesheets, embedded frames — where the destination can be assembled from data the model produced. Writes to shared destinations: documents, repositories, tickets, channels, calendar entries, anything another party can read. Messages: email, chat, notifications, and their metadata. Tool side-effects such as creating a record whose fields carry data outward, or querying a third-party service where the query string itself is the payload. Logs and telemetry that leave the trust boundary or are readable by a broader audience. And error messages returned to a caller. Any of these can carry data even when the response the user sees looks entirely normal.

  • Rendered output is an exfiltration channel whenever a client fetches a remote resource
  • Writes to shared stores and messages count as outbound, not internal
  • Query strings to third-party services carry data even when nothing is returned
  • Logs, telemetry and error strings leave the boundary more often than teams realise

Closing Paths and Proving It

Containment works better here than detection, because each path has a concrete infrastructural control. Default-deny egress with a narrow allowlist removes direct network paths and is enforced outside the application. Disabling automatic fetching of remote resources in output rendering, and constructing any outbound URL from allowlisted templates rather than model-produced strings, closes the rendering path. Scoping write access removes shared-destination paths. Restricting which fields of a tool call may contain free text limits side-effect channels. Then prove closure empirically: place a distinctive benign marker in private context and attempt to route it out through each enumerated path in a controlled test, verifying the marker never appears at the destination. Re-run that test after changes. Arguing that a path is closed is not the same as demonstrating that a marker cannot traverse it.

  • Default-deny egress is the highest-leverage control because it lives outside the app
  • Never build outbound URLs from model-produced strings; use allowlisted templates
  • Prove closure with a benign marker routed against each enumerated path
  • Re-test after every change — closure claims decay as features are added

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.