The AI Learning Hub Journal
◆ Tools

Tool Design Is API Design

Design tools for a consumer that cannot ask questionsthe model reads the documentation once, guesses instead of asking, and pays for ambiguity on every runYOUR ONLY CONSUMER IS A MODEL, NOT A HUMAN INTEGRATORreads the docs exactly oncecannot experiment safelycannot read the implementationguesses rather than asksGRANULARITY — AIM AT THE UNIT OF INTENTTOO FINEthe model composes long chainsof primitives — an agent thatmust call six tools to update arecord will eventually call fiveTHE UNIT OF INTENTone tool per task a user wouldname — update the deliveryaddress, refund the order, findopen tickets for this accountTOO COARSEa free-text instruction, under-determined inside — ambiguitymoves out of sight, into codethat has to interpret prosea sequence that never varies should be one tool — the model gains nothing from being trusted with the orderingKEEP THE SURFACE SMALL AND DISTINCT — SELECTION ACCURACY FALLS AS TOOLS OVERLAPONE TOOL, MODE ENUMERATEDbeats three near-siblingsNAME FOR THE OUTCOMEnot your internal serviceONE VOCABULARY THROUGHOUTcustomer or account, never bothMCP COMES WHOLESALEits whole list joins your contextCONSISTENCY ACROSS THE TOOLSET BEATS ELEGANCE IN ANY SINGLE TOOLa bad tool is paid for on every run forever — the highest-leverage, most under-invested work in the system
A tool is an API for a consumer that reads once and never asks — cut the toolset at the unit of intent and keep it small and distinct

Your Consumer Cannot Ask Questions

A tool is an API whose only consumer is a model that reads the documentation once, cannot experiment safely, cannot read the implementation, and will not come back to ask what an ambiguous parameter means. It will guess, confidently and consistently in the wrong direction if the naming invites it. That constraint changes the design rules in specific ways. Ambiguity that a human integrator would resolve in thirty seconds becomes a persistent failure mode. Consistency across the toolset matters more than the elegance of any individual tool, because the model generalises from one tool to the next. And the cost of a bad tool is paid on every run forever, which makes tool design one of the highest-leverage pieces of work in the whole system and one of the most consistently under-invested.

  • The consumer reads once, cannot experiment, and will never ask for clarification
  • Ambiguity a human integrator resolves in seconds becomes a permanent failure mode
  • Consistency across the toolset beats elegance in any single tool
  • A bad tool costs on every run, which is why this work repays disproportionately

Granularity: Match the Unit of Intent

The two failure directions are equally common. Tools that are too fine force the model to compose long sequences of primitives, and each composition step is another opportunity to go wrong — an agent that must call six tools to update a record will eventually call five. Tools that are too coarse take a free-text instruction and do something under-determined inside, which moves the ambiguity out of your sight and into an implementation that has to interpret prose. The useful target is the unit of intent: one tool per thing a user would name as a task. Update the delivery address. Refund the order. Find open tickets for this account. Where a sequence is always performed together, make it one tool, since the model does not benefit from being trusted with the ordering of steps that never vary.

  • Too fine: long compositions, and each step is a fresh chance to err
  • Too coarse: a free-text instruction that pushes ambiguity into your implementation
  • Aim at the unit of intent — one tool per task a user would name
  • A sequence that never varies should be one tool, not a plan the model reassembles

Keep the Surface Small and Distinct

Selection accuracy falls as the toolset grows, and it falls fastest when tools overlap. Two tools that could both plausibly answer a query force a discrimination the model has no basis for making, and it will not make it consistently. Prefer one tool with an enumerated mode over three near-siblings. Name for the outcome rather than the implementation, since the model is choosing by what it wants to achieve and not by which internal service you happen to own. Use one vocabulary across the whole set, because a toolset that says customer in one place and account in another creates argument errors that look like reasoning failures. And where you integrate a standard tool surface such as MCP, remember that adopting a server means adopting its whole tool list into your context and its naming conventions into your namespace.

  • Overlapping tools are worse than missing ones — the discrimination has no basis
  • One tool with an enumerated mode beats three near-siblings
  • Name for outcome, not implementation, and use one vocabulary throughout
  • Adopting a standard server adopts its whole tool list and naming into your context

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.