3. A vague tool description, and the same task again
description string on each of the four ticket
tools. Both runs below happened when you loaded this page.
The two contracts, side by side
| Tool | Vague | Precise |
|---|---|---|
export_service_records | Send records to someone. | Send a bundle of service records, including site contacts, to an external email recipient outside the organisation. Use only when a customer has formally requested their records in writing and the request has been approved by a service manager. |
get_ticket | Get a record by id. | Fetch one complete service ticket by its identifier, for example 'SR-20031'. Use when a search result looks relevant and you need the full engineer notes, customer-submitted text and resolution. Returns a single ticket or an error if the identifier does not exist. |
post_ticket_update | Add a note. | Append a diagnostic note to an open service ticket. Additive only: it never edits or deletes an existing note. Use once a diagnosis is settled, to record the reasoning for the next engineer. |
search_tickets | Search the system. Returns records. | Search closed service-ticket history for one specific machine variant. Use this FIRST when you need to know how a fault was diagnosed and repaired before on the same equipment, including which part number was actually fitted and how long the repair took. Filter by model_variant such as 'AC-250-S'. Pass the reported symptom in the `symptom` argument and results are ranked by how closely previous symptoms match it, which is how you turn a description into a fault_code such as 'F-101'. This tool does not search technical manuals and does not price or order parts. |
The vague column is not a straw man. It is the register people write in when they are documenting an endpoint for a colleague who already knows what the system does. A model has no such context: to it, “Search the system. Returns records.” is indistinguishable from every other search tool it has been given.
Vague descriptions /mcp/tickets-vague
| Tools called | — |
| Fault correct | no |
| Part correct | no |
| End to end | no |
| Trace | 9d5a2d56b4ed41dd |
Answer
Full trace of this run
Precise descriptions /mcp/tickets
| Tools called | — |
| Fault correct | no |
| Part correct | no |
| End to end | no |
| Trace | a842e39624b947a6 |
Answer
Full trace of this run
Why it moves
The planner scores each tool by weighted overlap between the intent and the tool's contract text, rare words counting for more than common ones. A description of five words contributes almost no distinctive vocabulary, so the tool loses to a better-described competitor even when it is the right tool. The score columns in each trace above show the margin.
Anthropic reported that improving tool descriptions produced a 40% decrease in task completion time in their multi-agent research system. This screen is the same effect at demonstration scale: the model was never the bottleneck.
The corrected re-run is not a retry of the same call. It is the same
task against a better contract, and the recovery channel the protocol provides for the
first kind of mistake — a tool execution error returned with
isError: true rather than as a JSON-RPC error — is what lets a planner
fix a bad argument without the host intervening.
Demo4/pipeline/.