More autonomy for agents is not always better
A lot of agent design assumes the goal is to keep the agent running for as long as possible. This article makes the opposite case. Once a task gets long enough, small errors start to stack. OpenAI’s own numbers show boundary flags going from 8.6% at five tasks to 19.7% at ten. So the better system may be one that stops more often. Clear scope, short runs, then a check before it keeps going. The smartest agent might be the one that knows when to hand the work back to the human
Read original source ↗