AI & LLMs
Tool Contracts for Reliable AI Agents
Design narrow, typed, permission-aware tools that make agent actions easier to validate, observe, retry, and audit.
An agent becomes useful when it can act, and dangerous when those actions are vague. A good tool contract limits what can happen and gives the runtime enough structure to validate every request.
This guide focuses on the decisions that survive contact with production: clear boundaries, observable behavior, and a feedback loop that reveals when an assumption is wrong.
The problem worth solving
Tools that accept arbitrary commands or large untyped objects move policy into the prompt. Prompts are guidance; authorization and validation belong in code at the tool boundary.
The useful move is to make the hidden constraint explicit. Write down what must stay correct, what can be delayed, and how the system should behave when a dependency fails. That turns a vague idea into something a team can test.
A practical implementation
Give each tool one verb, validate inputs, return structured errors, and separate read operations from writes. Require an idempotency key for repeatable writes and attach the authenticated actor on the server.
const CreateIssue = z.object({
title: z.string().min(5).max(120),
projectId: z.string().uuid(),
idempotencyKey: z.string().uuid(),
})
async function createIssue(input: unknown, actor: Actor) {
return issues.create(CreateIssue.parse(input), actor)
}
The example is intentionally small. In a real project, add structured logs, metrics around the failure path, and tests for retries or partial results. Keep the interface narrow so the implementation can change without forcing every caller to change too.
What to measure
Measure the outcome rather than activity. For software, that may be latency, error rate, queue depth, or recovery time. For product work, it may be activation, retention, or the number of useful conversations. Review the signal on a regular cadence and record what changed.
Takeaway
Agent reliability begins at the interface. Narrow tools produce clearer plans, safer permissions, and failures a human can understand.
Start with the smallest version that can teach you something, make its behavior visible, and improve it from evidence. That rhythm is more dependable than trying to design the final answer in one pass.
Further reading
Explore more AI & LLMs articles from this journal.
Tushar Sharma