Bayesian Active Learning for Intent Disambiguation in Interactive Robot Planning

Huao Li1, Carson Sobolewski1, Augustinos Saravanos1, William Tan2, John Karigiannis2, Chuchu Fan1
1Massachusetts Institute of Technology,  2GE Vernova
CoRL 2026 Spotlight
Human-Robot Dialogue Workshop, IROS 2026
BAL framework overview: LLM proposes STL candidates, grammar VAE embeds them, a Gaussian process models utility, and an acquisition function selects clarification questions before a formal planner executes the inferred specification.

BAL treats clarification as Bayesian active learning over grounded STL task specifications. The LLM proposes candidate intents and phrases questions; a Gaussian process over a learned STL latent space decides what to ask and when to stop; a formal planner executes the result.

Abstract

Interactive robot planning requires robots to infer and execute human intentions from natural language instructions that are often ambiguous, incomplete, or underspecified. Although large language models (LLMs) provide a powerful interface for clarification, relying on the generative model to drive a multi-turn conversation can introduce systematic failures. We propose a Bayesian framework that treats clarification as an active learning problem over grounded Signal Temporal Logic (STL) task specifications. Our method uses LLMs to initialize candidate formal specifications and translate informative contrasts into natural-language clarification questions, while Bayesian optimization maintains uncertainty estimates over user intent and selects queries that maximize information gain. Because queries are selected in a learned STL embedding space, the system can also ask about plausible specifications that the LLM did not propose. After convergence, the inferred STL specification is passed to a formal planner to synthesize a verifiable robot trajectory. Across four simulated and real-world task domains, our approach generally achieves higher task satisfaction and requires fewer clarification rounds than LLM baselines, while helping smaller models close the performance gap against larger reasoning models.

Environments

Franka Panda tabletop

Franka Panda (sim)

City map

City (sim)

School scan

School (real)

Factory scan

Factory (real)

Video

Results

BAL asks more informative questions

Choosing each question to maximize expected information gain about the user's intent lets BAL recover that intent more efficiently than letting the LLM run the dialogue. Across both simulation domains and all six base LLMs, BAL consistently outperforms the baselines in success rate and generally requires fewer clarification rounds. The pattern holds for every base model, so the gains come from the formal specification representation, Bayesian intent modeling, and active query selection, not from a particular LLM.

Success rate (%, ↑) and rounds of clarification (↓), mean ± standard error. Bold marks the per-domain best. LM: LLM clarification that outputs the plan directly. LM+P: LLM clarification, then translation to STL for the same planner BAL uses. All methods share the base LLM, simulated user, and context.

Base modelPanda — successCity — success
BALLMLM+PBALLMLM+P
GPT-5.491.1±4.248.9±7.582.2±5.762.5±9.929.2±9.320.8±8.3
GPT-5.4-mini93.2±3.864.4±7.186.7±5.164.7±11.620.8±8.312.5±6.8
Claude-Opus-4-691.1±4.264.4±7.182.2±5.750.0±10.229.2±9.320.8±8.3
Claude-Sonnet-4-691.1±4.273.3±6.686.7±5.141.7±10.129.2±9.38.3±5.6
Gemini-3.1-Pro91.1±4.282.2±5.791.1±4.254.2±10.233.3±9.625.0±8.8
Gemini-3.1-Flash-Lite93.3±3.764.4±7.191.1±4.266.7±9.629.2±9.312.5±6.8
Base modelPanda — roundsCity — rounds
BALLMLM+PBALLMLM+P
GPT-5.42.33±0.416.16±0.623.51±0.546.54±0.579.14±0.319.17±0.30
GPT-5.4-mini2.09±0.365.40±0.583.24±0.516.46±0.709.42±0.309.53±0.28
Claude-Opus-4-62.93±0.474.69±0.623.38±0.527.25±0.538.75±0.398.94±0.41
Claude-Sonnet-4-63.09±0.484.49±0.603.02±0.508.11±0.498.67±0.399.39±0.33
Gemini-3.1-Pro2.36±0.423.47±0.493.16±0.457.78±0.568.83±0.399.06±0.34
Gemini-3.1-Flash-Lite2.44±0.355.42±0.623.13±0.476.14±0.598.83±0.429.53±0.23

BAL makes LLMs less overconfident

With auto-stop, the robot decides when clarification is complete instead of asking for approval every round. BAL stops once its posterior entropy plateaus. The LLM baselines stop when they believe the intent is clear.

The LLM-based methods lose substantial success under auto-stop. They end with both low success and few rounds, which shows they often stop too early because they are overconfident in their intent estimates. BAL keeps a similar level of performance, using a few more rounds so its posterior uncertainty can converge. External uncertainty estimation gives an explicit signal for when clarification is still needed.

Success rate and clarification rounds under human-approval vs auto-stop on Panda

Human-approval vs. auto-stop on Panda. BAL maintains high task success under auto-stop, while LLM-based baselines terminate earlier but suffer from substantially lower success rates.

BAL helps smaller models

Latent-space inference and Bayesian optimization add computation, but more efficient clarification often offsets it, while larger reasoning models depend on long reasoning traces, tool calls, and repeated dialogue. As a result, BAL with a smaller model can match or outperform direct clarification with a larger model at lower wall-clock cost.

BAL moves much of the reasoning burden to its neuro-symbolic components. The LLM handles semantic interpretation and dialogue. The probabilistic model and formal planner handle uncertainty estimation, query selection, and planning.

Pareto plot of success vs time cost across methods and base models

Trade-off between time cost and performance. BAL works best with smaller models near the top-left corner.

BibTeX

@inproceedings{li2026bal,
  title     = {Bayesian Active Learning for Intent Disambiguation in Interactive Robot Planning},
  author    = {Li, Huao and Sobolewski, Carson and Saravanos, Augustinos and Tan, William and Karigiannis, John and Fan, Chuchu},
  booktitle = {Conference on Robot Learning (CoRL)},
  year      = {2026}
}