The AI Learning Hub Journal

Data-Flow Diagramming an LLM Feature

One worked feature: the internal support assistantordinary DFD notation plus three annotations — trust level inbound, acting identity on tool edges, outbound reachstaff userRETRIEVAL INDEXwiki — staff-editedticket text — user-writtenvendor doc sync — unreviewedper-user memory (store)context window+ modelquestion · authenticatedchunks · mixed trustmemory · user-scopedmodel-requested writecustomer record lookupacting id: service acct — whose scope?create ticketwrite — reversible, loggedpost to shared channeloutbound — visible beyond the teamanswer → rich text + linksrenders in browser — leaves your controlauthenticated / revieweduntrusted / unreviewedper-user storecoloured annotations = the three additionsTHE COMBINATION TO FINDan untrusted inbound flow, a private datastore, and an outbound flow that leavesyour control — this process has all three,and that combination is the design testFOUR MORE READS OF THE SAME DIAGRAMtool edges under the wrong identity ·stores written by one user, read byanother · untrusted sources upstream ofirreversible actions · unlogged crossingsTHE DIAGRAM IS A QUESTION-GENERATING DEVICE — ITS QUESTIONS ARE THE DELIVERABLEevery finding becomes a named change with an owner and a date — or an entry in the accepted-risk register
Annotate every flow with trust level, acting identity and outbound reach, then hunt the process that combines untrusted input, private data and an exit.

A Worked Subject

Take a concrete feature so the method stays honest: an internal support assistant. It answers staff questions using a retrieval index built from a wiki, a ticket system, and a vendor documentation sync; it can look up a customer record, create a ticket, and post a message to a shared channel; it keeps per-user memory of preferences and prior issues; and its answers render as rich text with links in a web client. That description is already enough to draw the diagram, and drawing it will surface questions the description hid — who can edit the wiki, whether the vendor sync is reviewed, whether the customer lookup applies the asking user's permissions or the service account's, and whether the shared channel is visible outside the team. A diagram is worth building precisely because it forces those questions into the open at design time rather than after an incident.

  • Pick a real feature with real tools; abstract examples produce abstract findings
  • Write the one-paragraph description first, then draw what it implies
  • The questions the diagram raises are the deliverable, as much as the diagram itself
  • Ambiguity about whose permissions a tool uses is the most common early discovery

Notation That Earns Its Keep

Use ordinary data-flow notation — external entities, processes, data stores, flows, and trust boundaries — with three AI-specific additions. Annotate each flow into the context window with its trust level, so the diagram shows at a glance how many untrusted sources reach a single prompt. Annotate each tool edge with the identity and scope used, so a service account with broad rights is visible rather than implied. And annotate each outbound flow with whether it leaves your control — to a user's browser, a third-party API, a shared store, a log another team reads. Keep the notation minimal; the aim is a diagram an engineer will update, not a formal specification. If it cannot be redrawn on a whiteboard in ten minutes, it will be stale within a month and misleading within two.

  • Standard DFD elements plus three annotations: trust level, identity and scope, outbound reach
  • Show every flow that reaches the context window, including metadata and tool results
  • Make the acting identity explicit on every tool edge
  • Optimise for redrawability — a diagram nobody maintains is worse than none

Reading the Diagram for Findings

The diagram is a question-generating device, and a few reads produce most of the value. First, find every process that has an untrusted inbound flow, a private data store, and an outbound flow leaving your control — that combination is the design test the next module builds on. Second, look for identity mismatches: a tool edge using a service account where the user's own scope should apply. Third, look for stores written by one user and read by another, which is where cross-tenant contamination lives. Fourth, trace each irreversible action back to the flows that can influence it and count how many untrusted sources are upstream. Fifth, check whether every edge that crosses a boundary is logged well enough to reconstruct a decision afterwards. Each read produces findings phrased as design changes, which is exactly what you want to hand a team.

  • Find processes combining untrusted input, private data, and outbound reach
  • Look for tool edges acting under the wrong identity
  • Stores written by one user and read by another are cross-tenant exposure
  • Trace irreversible actions upstream and count the untrusted sources that can steer them

From Diagram to Backlog

A diagram that does not become work is decoration. Convert each finding into a specific, assignable change with an owner and a date: narrow this tool to read-only, split this agent so the one reading external content has no customer data access, apply the requesting user's scope to this lookup, add provenance tagging on this ingestion path, gate this action, deny egress here by default. Record the ones you are not going to do and why, in the same list — that becomes the accepted-risk register covered later in this module. Then pick the two or three findings with the widest reach and confirm them empirically before building anything, because a design review can be wrong about what the running system actually does. The order that works is model, verify, then remediate — not model, remediate, and discover later that the diagram was aspirational.

  • Every finding becomes a named change with an owner and a date
  • Findings you decline go into the accepted-risk register, not into silence
  • Verify the top findings against the running system before funding the fix
  • Diagrams describe intent; traces describe behaviour — reconcile them early

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.