The AI Learning Hub Journal
◆ Free course · 5 modules · 29 lessons

Agent Engineering: Building the Harness

The scaffolding that turns a model into a working system — loops, context, tools, and the failures that only appear in production

Read it in the library →
Module 1 · 5 lessons

The Loop

The control flow you own rather than the one the model implies: what a step actually is, how a run ends on purpose, the four budgets that keep it bounded, detecting a stuck or oscillating agent, and the frequent case where a deterministic script is simply the better system.

  1. The Loop Is Yours, Not the Model's
  2. Ending a Run on Purpose
  3. Budgets and Cost Ceilings
  4. Detecting a Stuck or Looping Agent
  5. When a Script Beats an Agent
Module 2 · 6 lessons

Context Engineering

The window as a budget you spend on purpose: assembling context in code rather than by accretion, deciding what earns its place, ordering for attention, compacting without losing the thread, memory tiers and their lifetimes, and cache-aware layout as a first-order cost lever.

  1. Context Is Assembled, Not Accumulated
  2. What Earns Its Place
  3. Ordering, Recency and Attention
  4. Compaction Without Losing the Thread
  5. Memory by Lifetime
  6. Cache-Aware Layout and What Actually Moves Quality
Module 3 · 6 lessons

The Harness

Everything around the model that determines whether it can work: tools designed as APIs for a reader who cannot ask questions, descriptions and schemas that steer selection, return values that enable recovery, retry and idempotency for actions with side effects, verification built into the loop, and delegation that adds capability rather than failure modes.

  1. Tool Design Is API Design
  2. Descriptions and Schemas the Model Actually Reads
  3. Return Values That Enable Recovery
  4. Errors, Retries and Side Effects
  5. Verification Inside the Loop
  6. Sub-Agents and the Cost of Delegation
Module 4 · 6 lessons

Making It Hold Up

Operating an agent that runs unattended: designing around non-determinism, tracing well enough to reconstruct a failure, the failure taxonomy you will actually meet and what each one signals, cost and latency as requirements rather than reports, degrading gracefully instead of stopping dead, and what belongs in a log.

  1. Non-Determinism as an Engineering Constraint
  2. Tracing a Reconstructable Run
  3. The Failure Taxonomy You Will Actually Meet
  4. Cost and Latency as Requirements
  5. Degrading Gracefully
  6. What to Log and What Never To
Module 5 · 6 lessons

Keeping Humans In It

Oversight as a designed part of the system: gates people engage with rather than clear, escalation as a first-class outcome, presenting a run so a human can actually judge it, widening autonomy along an evidence-backed trust curve, the honest limits of review at volume, and writing down what the agent is not allowed to do.

  1. Designing a Gate Someone Actually Uses
  2. Escalation as a First-Class Outcome
  3. Presenting a Run So a Human Can Judge It
  4. The Trust Curve
  5. The Honest Limits of Oversight at Volume
  6. Writing Down What the Agent May Not Do

Every lesson is free, with no sign-up. Reading happens in the library, where your progress is saved on your device.