Hands on

Designing tools agents use well

Your tools are declared and an agent is calling them, but not the way you hoped. These are the eight things that decide it, and all of them are decided when you write the declaration.

A tool an agent can call is not the same as a tool an agent uses correctly. Almost all of the difference is decided when you write the declaration, not when the agent runs.

One tool, one job

The single most common shape to avoid:

// Don't
{ name: "do_thing",
  description: "Perform an action in Acme.",
  inputSchema: { type: "object",
                 properties: { instruction: { type: "string" } },
                 required: ["instruction"] } }

That declaration has given up everything the mechanism was for. There is no schema worth validating, because the argument is a sentence. There is no description worth reading, because it covers every operation the application has. And one set of hints covers operations with completely different consequences — so either lookups get confirmed as though they were deletions, or deletions run as though they were lookups.

Split it: acme_get_order, acme_search_orders, acme_cancel_order. Each of those has a schema you can state and a class you can mean.

Name for your system, not for the world

acme_search_tickets, not search. On the bus, an agent may hold tools from several open tabs at once and picks between them on name and description alone. A generic name collides with every other application’s idea of the same word, and the failure is not an error — it is the right operation performed in the wrong application.

Describe the outcome and the boundaries

Three things belong in a description and are usually missing: what it returns, what it does not cover, and what happens when there is nothing to return.

// Thin
description: "Search tickets."

// Useful
description:
  "Search open Acme support tickets by requester email or " +
  "subject text. Returns up to 50 matches, newest first, " +
  "each with id, subject, requester and status. Does not " +
  "search closed tickets or internal notes. Returns an " +
  "empty list when nothing matches."

The second one stops an agent calling it, getting nothing, and concluding the ticket does not exist when it is merely closed.

Type the inputs properly

Enumerate what is enumerable. Describe every property. Mark required what is required, and give the default in the description for what is not. An untyped string is a guess you have asked the agent to make on your behalf, and you will not see the guesses it got wrong — only the results.

Return structure the agent can act on

XataWorks hands the agent exactly what execute returns. Named fields can be checked; a sentence has to be re-read.

// Don't
return text("Order A-1002 is paid but has not shipped yet.");

// Do
return text(JSON.stringify({ order_id: "A-1002",
  status: "paid", shipped: false, tracking: null }));

text() here is the small helper from the notes example, which wraps a string in the draft standard’s result shape.

Fail in a way that names the way forward

A failure the agent cannot tell apart from every other failure collapses into “it did not work”. Say which failure it was, and what would work instead:

return {
  isError: true,
  content: [{ type: "text", text:
    "NOT_FOUND: no order A-2002. Orders run A-1001 to A-1004." }]
};

A short code at the front lets an agent respond differently to each kind of failure — ask the person for a missing value, fall back to a search, stop rather than retry. Keep messages free of internal detail; they reach a model and, often enough, a person.

Validate before you apply

Where an operation could leave your application in a state it cannot render, check the result first and leave the original untouched if it would not be valid:

execute: async ({ text: addition }) => {
  const candidate = currentSource() + "\n" + addition;
  const check = await validate(candidate);
  if (!check.valid) {
    return { isError: true, content: [{ type: "text",
      text: "INVALID_RESULT: that would not parse: " +
            check.error }] };
  }
  setSource(candidate);
  return { content: [{ type: "text", text: "Added." }] };
}

A failed write that changed nothing is recoverable. A failed write that half-applied is a support ticket.

Declare every hint

When a page declares nothing, XataWorks has to assume, and the assumption is cautious. You know what each of your tools does, so say so: readOnlyHint for lookups, consequentialHint — true or false — for anything that changes things, and openWorldHint: false when the tool reaches nothing beyond your own application. What each hint means and what the person is asked is in safety classes.

Gate and be idempotent, if you inject

Relevant when your declaration is injected into pages rather than served as part of them — from a browser extension, say: check the host before doing anything, and guard against running twice, because a reload re-runs whatever ran before.

(function () {
  if (window.__acmeTools) return;                    // idempotent
  if (location.hostname !== "acme.example") return;  // gated
  window.__acmeTools = { version: "1" };
  // ... registerTool calls ...
}());