How to test tool-calling accuracy in AI agents
OpenRouter outlines separate checks for tool selection, argument structure and values, and multi-step call trajectories. It recommends keeping cases, graders, settings, and routing consistent when comparing models.
SessionWatcher editorial · Published · Updated · Source announcement: 2026-09-30
Separate tool choice from argument correctness
Test whether an agent chose the appropriate tool separately from whether its arguments are correct. Include cases where no tool should be called; a response containing a tool call is not automatically a pass.
For arguments, use schema validation to detect malformed JSON, missing fields, wrong types, invalid enum values, and undeclared parameters. Then check values against known expected values: a schema-valid order ID can still be the wrong one.
Choose graders for the question
Use deterministic checks when the expected tool or argument value is known. A reference-free LLM judge can assess context-dependent choices when several tools or values may be reasonable, but its instructions should be clear and its decisions checked against reviewed examples.
For multi-step workflows, compare trajectories only as strictly as the requirement demands. If call order can vary, grade required calls or the resulting state instead of enforcing one exact sequence.
Keep model comparisons consistent
Run the same cases, tools, grading rules, model settings, and routing configuration for each candidate. Check that candidates support the parameters you set, and repeat cases rather than relying on a single response.
Test realistic edge cases, including missing information, similar tool descriptions, multiple calls, and requests where no tool is needed. For multi-step agents, evaluate the full trace or resulting state, not just one tool-calling turn.
AI assisted reporting, checked against the linked official sources. Source pages checked 2026-09-30. Editorial process and corrections.