Computer Science

Dharmendra Kumar

2026.2.17Evolutionary Intelligence

DOI: 10.1007/s12065-026-01153-y

tlooto Summary

ETPNav is presented as a robust vision-language navigation (VLN) system that uses a combination of an online topological map with a transformer model for predicting waypoints based on natural language instructions in order to create a reliable method for following instructions in complex and changing indoor environments.

Abstract

In this study we present ETPNav as a robust vision-language navigation (VLN) system that uses a combination of an online topological map with a transformer model for predicting waypoints based on natural language instructions in order to create a reliable method for following instructions in complex and changing indoor environments. In contrast to many other VLN models that are trained using static representations of scenes, ETPNav learns a graph of how spaces are related by constructing it incrementally using RGB-D data during exploration. This enables ETPNav to reason about the relationships between known spaces and unknown spaces in its environment. The high-level planner is a transformer that predicts waypoints that correspond to where the agent should move in order to reach a goal, and the low-level controller generates smooth motion commands to guide the agent through the environment based on odometry and sensor readings. We evaluate ETPNav using two benchmarks that test both the ability of the model to follow instructions and the robustness of the model to changes in the environment. We find that ETPNav performs significantly better than prior

Citation format

KUMAR, Dharmendra. Etpnav: Robust vision-language navigation via online topological mapping. Evolutionary Intelligence, 2026, 19(2).