With auto-stop, the robot decides when clarification is complete instead of asking for approval every round. BAL stops once its posterior entropy plateaus. The LLM baselines stop when they believe the intent is clear.
The LLM-based methods lose substantial success under auto-stop. They end with both low success and few rounds, which shows they often stop too early because they are overconfident in their intent estimates. BAL keeps a similar level of performance, using a few more rounds so its posterior uncertainty can converge. External uncertainty estimation gives an explicit signal for when clarification is still needed.



