Systematic Multi-Agent Vision-and-Language Navigation: Formulation, Benchmark, and Method
Author1, Author2, Author3, Author4, Author5
Abstract
This paper introduces a systematic formalization of multi-agent vision-and-language navigation as a constrained coordination problem, addressing the need for teams of robots to perform complex tasks.
Reality Card
The paper presents the first systematic formalization of multi-agent VLN as a constrained coordination problem and introduces a comprehensive baseline system, TRISS, for coordinating multiple agents.
The MAVLN benchmark comprises 11,724 episodes across 145 scenes with teams of up to four agents.
The challenges of coordinating under MAVLN task constraints may hinder reproducibility and generalization of results.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.