Epistemic Uncertainty as an Early Warning of Unsafe Robot Navigation Around People
Abstract
Indoor service robots carry medication, lab samples, and supplies through hospital corridors and move goods through warehouse aisles. They share this space with people who are not trained operators and who are often not paying attention to the robot. Navigation policies for such robots are increasingly learned through reinforcement learning or imitation. When a learned policy meets a situation outside its training data, it still produces an action, and nothing in that action tells the robot or its operator that the policy may be wrong.
Supervision
- Prakash Aryan
- Sebastiano Panichella
Motivation
Bayesian deep learning distinguishes two kinds of predictive uncertainty. Aleatoric uncertainty reflects noise inherent in the observations, such as the next step of a person whose movement is partly random. Epistemic uncertainty reflects what the model does not know, and it shrinks as more training data becomes available. If epistemic uncertainty rises before a policy does something unsafe, the robot could use it as a warning and slow down or stop in time. This thesis measures whether the signal rises early enough to be useful, and at what false-alarm rate, when the unexpected situation is caused by a person who stops suddenly, turns around, or walks while distracted.
Goal
The student will train a navigation policy for a TurtleBot3 in Unreal Robotics Lab (URLab) with separate estimates of both kinds of uncertainty, and test whether these estimates give a useful early warning around unpredictable people. The thesis is expected to deliver:
- a navigation policy with an epistemic estimate from deep ensembles (with MC-Dropout as a cheaper alternative) and a learned-variance head for aleatoric uncertainty, both published at runtime on a ROS 2 topic
- simulated hospital corridor and warehouse aisle scenes with parameterised human behaviours such as sudden stops, distracted walking, and sharp changes of direction
- measurements of warning time and false-alarm rate for four classes of unsafe event: collision, personal-space intrusion, startling a person, and the robot freezing in place
- a comparison between a two-stage safety filter that reacts differently to each kind of uncertainty and a single-threshold fallback
- a small Meta Quest 3 user study testing whether the calibration obtained with simulated pedestrians carries over to real people in the same virtual scene
Requirements
The student should have a solid background in machine learning with PyTorch and experience with reinforcement learning or imitation learning. Familiarity with ROS 2 is expected, and experience with Unreal Engine or another robot simulator is helpful. Knowledge of basic statistics is needed to evaluate calibration and false-alarm rates.
Pointers
- A. Kendall and Y. Gal, “What Uncertainties Do We Need in Bayesian Deep Learning for Computer Vision?,” in Advances in Neural Information Processing Systems (NIPS), 2017.
- B. Lakshminarayanan, A. Pritzel, and C. Blundell, “Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles,” in Advances in Neural Information Processing Systems (NIPS), 2017.
- D. Flögel, M. Gómez Villafañe, J. Ransiek, and S. Hohmann, “Disentangling Uncertainty for Safe Social Navigation using Deep Reinforcement Learning,” in Proc. IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2025.
- J. Embley-Riches, J. Liu, S. Julier, and D. Kanoulas, “Unreal Robotics Lab: A High-Fidelity Robotics Simulator with Advanced Physics and Rendering,” arXiv:2504.14135, 2025. ,