The AI Learning Hub Journal

Security Regression Tests in CI

Closed means a test now fails if the path reopenswrite the test against the condition, not the attempt — the attempt gets patched, the condition recursfinding from theexerciseregression test againstthe conditionruns on everychange, in CIstays closed acrossmodel + prompt changesplant benign marker content in the test corpus → assert no privileged tool call appears in the traceINVARIANTS — ZERO ACROSS ALL RUNS· no privileged call after untrusted ingestion· no canary record in any output· no request to a non-allowlisted destinationa single violation in any run is a real failurePROBABILISTIC — THRESHOLD PLUS TREND· run each scenario several times, assert the aggregate· drifting from rarely to often succeeding is aregression even while the suite still passessmall and fast on every change beats comprehensive and weeklySTRUCTURAL TESTS — NO MODEL CALL, NEVER FLAKYtool list matches theapproved manifestschema text hashesmatch approved valuescredential scopes stayinside the declared setegress allowlist growthneeds a change recordprivileged interfaceaccepts typed fields onlymilliseconds, never flaky — they catch the configuration drift behind most real exposure; build these firstWHAT BLOCKS THE RELEASEinvariant or structural failureblocks — a known-closed path reopenedthreshold movementreview, not auto-block — models movea suite that cries wolfgets bypassed — you lose the controlASSERT ON TRACES — TOOLS, IDENTITY, DESTINATIONS — NOT ON OUTPUT TEXT THAT DRIFTSkeep markers controlled and rotated, and run against the deployed configuration, guardrails included
Every accepted finding becomes a CI test against the condition — zero tolerance for invariants, thresholds with trend for probabilistic properties, and cheap structural checks built first.

Every Finding Becomes a Test

The mechanism that converts a point-in-time exercise into a durable control is simple: no finding is closed until a test exists that would fail if it returned. Write the test against the condition rather than the specific attempt, because the attempt will be patched and the condition will recur. If the finding was that content retrieved from an external source could cause a privileged tool call, the test plants benign marker content in a test corpus and asserts that no privileged tool appears in the resulting trace. If it was that an agent answered using data the requesting user could not reach, the test issues a request as a restricted user and asserts the response contains no canary record. Assertions on the trace — which tools were called, with what identity, which destinations were contacted — are more durable than assertions on output text, which drifts with every model and prompt change.

  • A finding is closed when a failing-if-it-returns test exists, not when the ticket is updated
  • Test the condition, not the specific attempt that revealed it
  • Assert on traces: tools called, acting identity, destinations contacted
  • Output-text assertions drift; structural assertions survive model changes

Handling Non-Determinism Without Flakiness

Security tests against a stochastic system need statistical treatment or they become noise that teams disable. Run each scenario several times and assert on the aggregate, but choose the aggregate to match the property. For an invariant that must always hold — no privileged tool call after untrusted ingestion, no canary in output, no request to a non-allowlisted destination — assert zero occurrences across all runs, because a single violation is a real failure. For probabilistic properties such as resistance to a category of persuasion, assert a threshold and track the trend, since a scenario drifting from rarely succeeding to often succeeding is a regression even while it still passes. Keep the run count high enough to be meaningful and the scenario set small enough to run on every change; a comprehensive suite that only runs weekly catches regressions after they ship.

  • Invariants get zero-tolerance assertions across all runs
  • Probabilistic properties get thresholds plus trend tracking
  • A rate moving in the wrong direction is a regression even while it passes
  • Fast and small on every change beats comprehensive and weekly

Structural Tests That Need No Model at All

A large share of the most valuable security tests are deterministic and cheap, because they check configuration rather than behaviour. Assert the tool list matches an approved manifest, so a newly registered server fails the build instead of appearing silently. Assert the hash of every tool schema text matches the approved value. Assert no credential in the agent's environment carries a scope outside the declared set. Assert that the egress allowlist has not grown without a corresponding change record. Assert that the privileged component's interface accepts only typed fields, so a change that starts passing free text through fails. These run in milliseconds, never flake, and catch the configuration drift that produces most real exposure — which makes them a better first investment than an elaborate behavioural suite.

  • Approved tool manifest and schema hashes checked in CI catch silent additions
  • Assert credential scopes and egress allowlists against declared values
  • Assert the typed interface between unprivileged and privileged components
  • Deterministic, fast, and never flaky — build these before behavioural tests

Gating and Test Data Hygiene

Decide deliberately what blocks a release. Invariant violations and structural assertions should block, because they represent a known-closed path reopening. Threshold regressions on probabilistic properties usually warrant review rather than an automatic block, since a legitimate model or prompt change can move them, and a suite that blocks too often gets bypassed — which costs you the whole control. Keep the test corpus and its markers in a controlled repository with the same access restrictions as any sensitive material, and keep it clear of anything derived from real customer data. Rotate marker values so a system tuned to recognise them does not pass artificially. And run the suite against the deployed configuration including guardrails, not against a stripped-down harness, or you will be measuring a system nobody ships.

  • Block on invariants and structural checks; review on threshold movement
  • A suite that blocks too often gets bypassed, which forfeits the control entirely
  • Keep test corpora and markers controlled and free of real customer data
  • Rotate markers and test the deployed configuration, guardrails included

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.