The AI Learning Hub Journal

CI for Prompts and Agents

Continuous integration for prompts, tools and agentsa change lands, the suite runs hermetically, results are diffed against the baseline, and a gate decidesPROMPT TEXTA wording change is abehaviour change.TOOL DEFINITIONNew tool, new schema, ora new description.MODEL VERSIONA pinned version moved,or the provider shipped.RETRIEVAL CONFIGChunking, index contents,top-k, or the reranker.Any one of these changes behaviour, so any one of them triggers the same runHERMETIC RUN — NO OUTSIDE WORLDSame inputs, same fixtures, same versions — so a difference inthe output means a difference in the system, not in the weather.External calls mockedTool calls hit recorded fixtures, not live services.Environment pinnedModel version, decoding settings and seeds all fixed.Reruns must agreeAn unchanged system that scores differently is broken.DIFF AGAINST THE STORED BASELINEbar = this run · tick = baseline · every suite compared, not the meanGrounding suitesameRefusal boundarybetterTool selectionREGRESSIONMulti-turn recoverysamePrompts are code — version-controlled, reviewed, diffed, never edited straight into productionTHE GATE — WHAT ACTUALLY HAPPENS TO THE MERGENO REGRESSION → MERGEEvery suite at or above thebaseline, so it merges withthe run attached to the PR.REGRESSION → BLOCKEDA suite dropped. The mergestops until it is fixed orexplicitly signed off.INTENDED CHANGE → REBASELINEThe new behaviour is thewanted one, so the baselinemoves — as a reviewed commit.A gate that never blocks anything is not a gate — it is a report nobody readsWithout a hermetic suite and a stored baseline, “it feels better” is your only evidenceMocked externals are what make two runs comparable — a live dependency turns the suite into noise
Ordinary software engineering discipline — the only new part is that the unit under test is a prompt

Prompts Are Code

The behaviour of an AI system is determined as much by its prompts, tool descriptions, and rubrics as by its code, yet these artefacts routinely live outside every engineering control — edited in a console, stored in a database row, changed without review. Bring them in. Prompts belong in version control with meaningful diffs, code review by someone other than the author, and a deployment path that is auditable and revertible. Tool descriptions deserve particular attention because they are model-facing instructions that shape behaviour strongly and are usually written once and never reviewed. The test is simple and worth applying honestly: can you answer, for any output your system produced last Tuesday, exactly which prompt, tool schema, model, and rubric version were in play? If not, your evals cannot be reproduced and your incidents cannot be investigated.

  • Version prompts, tool schemas, rubrics, and grader code together with application code
  • Review prompt diffs like code diffs; small edits have non-local behavioural effects
  • Console-edited prompts break reproducibility and post-incident analysis

Two Tiers and a Budget

Eval runs cost money and time, which makes tiering a design requirement rather than an optimisation. The pre-merge tier is a small, fast, mostly deterministic subset — a few minutes, single runs per case, focused on catching gross breakage and the specific regressions this change might plausibly cause. The nightly or pre-release tier is the full suite: every case, multiple runs each, judges included, results tracked as a trend. Put an explicit cost ceiling on CI evaluation and monitor it, because judge calls multiplied by cases multiplied by runs multiplied by pull requests grows faster than anyone expects. Cache aggressively where inputs and versions are unchanged, and make the cost of a full run visible so the trade-off between coverage and spend is a decision rather than an accident.

  • Fast deterministic subset pre-merge; full statistical suite nightly
  • Cases times runs times judges times PRs — model the cost before it surprises you
  • Cache results keyed on input, prompt, and model version
  • Make eval spend a visible line item so coverage decisions are deliberate

Hermetic Environments for Agent Evals

Agent evals are integration tests against a world, so the world has to be controlled or the results mean nothing. Every run needs a known starting state and full reset afterwards, because state leaking between runs produces failures that depend on execution order — the hardest class of eval bug to diagnose. Decide deliberately between mocked and live dependencies: mocks are fast, deterministic, and free, but they encode your assumptions about how the external system behaves, and agents fail in production precisely where those assumptions are wrong. Recorded interactions replayed back are a strong middle path. Keep a small live-integration tier that runs less often against real systems in a sandbox account, because the divergence between your mocks and reality is itself a failure mode worth measuring.

  • Known initial state and guaranteed teardown, or your results depend on run order
  • Mocks encode assumptions; record-and-replay keeps realism without flakiness
  • Keep a smaller live tier against sandboxed real systems to catch mock drift
  • Never point agent evals at production systems or real customer accounts

Gating Policy

Deciding what blocks a merge is a policy question that deserves an explicit answer rather than a default. Hard gates belong on things that are unambiguous and consequential: safety criteria, schema validity, a floor on the critical slice, and any case protecting a past incident. Soft signals — a small aggregate movement, a judge score wobble within its noise band — should be reported in the pull request without blocking, because gates that fire on noise are disabled within a month and then nothing is gated at all. Post the diff where the reviewer already is: which cases flipped in each direction, with links to the traces, so a human can judge whether a regression is acceptable. And write down who can override a gate and what they must record when they do.

  • Hard gate: safety, schema validity, critical-slice floors, incident regression cases
  • Report-only: aggregate wobble inside the noise band — gating on noise kills the gate
  • Post flipped cases with trace links in the PR; make the diff readable
  • Define the override path explicitly, including what gets recorded

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.