4.7×
More consistent.
Structured, low-level tools improve consistency across repeated attempts by up to 4.7×.
SAME CAPABILITIES. DIFFERENT INTERFACES.
Consistency, exploration, and efficiency change
when we change how tools are organized.
MEET THE SETUPS
Search a repository. Inspect some code. Make a change.
Watch how the interface changes the way an agent gets there.
01 / THE MECHANISM · ATOMIC
Same model. Same bug.
One run gets caught repairing its edits.
The other makes the change and moves on.
Keep the message’s name when converting it to a dictionary.
Loading the recorded executions…
A small edit. A long detour. Bash inserts code by line number, removes a method, and repeatedly encounters invalid Python. Atomic anchors its change to the existing code and preserves the name on its first source edit.
A selected pair of real runs. Atomic also encounters environment and test-helper errors; those remain visible. Final outcomes come from the benchmark evaluation.
02 / THE MECHANISM · NL SEARCH
One setting. Two code paths.
Natural-language search surfaces
the branch the Bash run misses.
Make max_cpu_count=0 use all available CPUs.
Loading the recorded exploration…
Find it. Follow it. Fix it. The search subagent returns CMake’s use of the same setting. The agent then inspects that file, updates its logic, and tests it. Bash updates the direct MSBuild helper and leaves CMake untouched.
A selected example of effective exploration. The missed branch plausibly contributes to the outcome difference. This pair does not measure diversity across repeats or prove causation.
02 / EXPLORATION · THE FILE SETS
Two visited sets, starting empty.
Watch filenames become reads,
and reads become edits.
Loading file visits…
03 / THE MECHANISM · PYTHON
Both runs resolve the task.
Follow their shared milestones,
and the tokens spent getting there.
Loading recorded steps and token usage…
THE INTERFACE MAKES A DIFFERENCE
The underlying capabilities stay similar.
The way agents use them does not.
4.7×
Structured, low-level tools improve consistency across repeated attempts by up to 4.7×.
>11%
Natural-language search broadens exploration and increases access to relevant files by more than 11%.
56.3%
Python interfaces use 56.3% fewer tokens and 41.6% fewer steps, with similar task performance.
Highlights reported in the paper relative to BashOnly; effects vary by actor and task. The full study evaluates six architectures, three actor models, and 11,700 trajectories.
THE DEVIL IS IN THE INTERFACE
Six architectures. Three actor models.
11,700 trajectories of coding-agent behavior.
The Devil Is in the Interface: Evaluating How Tool Architecture Shapes Coding Agent Behavior
Purdue University / Microsoft Research / The University of Chicago