Programmatic tool calling

Agentic AIProtocols and integrationPublished By Simon Budziak

Programmatic tool calling lets a model write code that calls tools inside a sandbox instead of requesting them one at a time through the conversation. Intermediate results stay out of the context window, so an agent can chain many tool calls at lower latency and token cost.

What problem does programmatic tool calling solve?

Round trips. Classic tool calling sends every result back through the model, so a task that touches thirty records pays thirty context updates, and the raw data crowds the context window whether or not the model needs it. Letting the model orchestrate tools in code turns that loop into one execution: fetch, filter, aggregate, return the answer. Anthropic’s public-beta implementation runs the code in its execution container, and Anthropic reports double-digit accuracy gains with materially fewer input tokens on search-heavy benchmarks.

How does a programmatic tool call actually run?

The model emits a script rather than a tool request. The harness hands that script to a sandbox with the permitted tools exposed as functions; in Anthropic’s public beta the sandbox is the provider’s own execution container. The script calls tools like ordinary functions, loops over results, branches on what comes back, and prints only what the model should see. Only that final print re-enters the conversation. To the model, the excursion looks like one tool call that got a lot done. To the operator it is code: the orchestration logic is inspectable and replayable, which a chain of opaque model decisions never is, and with deterministic replay a failed run can be re-executed exactly.

What does it require from your tools?

Tools that are safe to call without a person watching each call. The code runs inside sandboxed code execution, but a script can fan out to many calls before anyone reviews the batch, so side-effecting tools still need a tool approval gate ahead of anything that moves money or messages. Mark which tools code may call and keep the rest conversational; the agent harness is where that split is enforced.

How do you debug an agent that calls tools from code?

Treat the script as the unit of failure, not the tool. A bad run leaves three artifacts to pull: the code the model wrote, its stdout and stderr, and the tool calls the sandbox dispatched, which is where LLM tracing earns its keep. The failure modes differ from conversational tool use. A script can catch a tool error, continue, and return a confident wrong answer, so make tools fail loudly rather than return empty results a loop will happily iterate past. And since a whole batch executes before anyone reviews it, keep writes idempotent: the natural retry after a failed script is running the script again.

Frequently asked questions

How is programmatic tool calling different from normal tool calling?

In normal tool calling every tool result returns to the model, costing a round trip and context tokens each time. In programmatic tool calling the model writes a script that calls the tools, filters and combines their outputs in code, and hands back only the final result.

When is programmatic tool calling the wrong choice?

When a human or a policy must see each action before it runs, since calls made from code skip the per-call review a conversational loop gives you, and when a task needs only one or two tool calls, where the code execution setup buys nothing.

Summarize this page with

See this working in a system we built