4.7×
Consistency across runs
Structured, low-level tools improve consistency across repeated attempts by up to 4.7×.
TOOL INTERFACES AND AGENT BEHAVIOR
We study how tool interfaces affect coding-agent
consistency, exploration, and efficiency.
EMPIRICAL RESULTS
With similar tool capabilities,
agent behavior varies across interfaces.
4.7×
Structured, low-level tools improve consistency across repeated attempts by up to 4.7×.
>11%
Natural-language search broadens exploration and increases access to relevant files by more than 11%.
56.3%
Python interfaces use 56.3% fewer tokens and 41.6% fewer steps, with similar task performance.
Results reported in the paper relative to BashOnly; effects vary by actor and task. The full study evaluates six architectures, three actor models, and 11,700 trajectories.
EXPERIMENTAL DESIGN
We compare four sets of tools with the same capabilities
and different interfaces.
01 / CONSISTENCY · ATOMIC
Two runs use the same model on the same task.
Bash repeatedly repairs invalid edits;
Atomic resolves the task.
Keep the message’s name when converting it to a dictionary.
Loading the recorded executions…
Editing and recovery. Bash inserts code by line number, removes a method, and repeatedly encounters invalid Python. Atomic anchors its change to the existing code and preserves the name on its first source edit.
02 / EXPLORATION · FILE ACCESS
Each run starts with an empty visited set.
The replay records returned filenames,
opened files, and edits.
Loading file visits…
The setting is used in two code paths.
NL search returns a relevant file
that the Bash run does not inspect.
Make max_cpu_count=0 use all available CPUs.
Loading the recorded exploration…
Access to relevant code. The search subagent returns CMake’s use of the same setting. The agent then inspects that file, updates its logic, and tests it. Bash updates the direct MSBuild helper and leaves CMake untouched.
A selected example of effective exploration. The missed branch plausibly contributes to the outcome difference. This pair does not measure diversity across repeats or prove causation.
03 / EFFICIENCY · PYTHON
Both runs resolve the same task.
Shared milestones align their progress;
the bars show cumulative token usage.
Loading recorded steps and token usage…
STUDY DETAILS
We evaluate six tool architectures and three actor models
across 11,700 coding-agent trajectories.
The devil is in the interface: Tool interface shapes agent behavior.
Purdue University / Microsoft Research / The University of Chicago