Vision-dialog navigation by exploring Cross-modal Memory

Yi Zhu, Fengda Zhu, Zhaohuan Zhan, Bingqian Lin, Jianbin Jiao, Xiaojun Chang, Xiaodan Liang

Research output: Chapter in Book/Report/Conference proceedingConference PaperResearchpeer-review

1 Citation (Scopus)

Abstract

Vision-dialog navigation posed as a new holy-grail task in vision-language disciplinary targets at learning an agent endowed with the capability of constant conversation for help with natural language and navigating according to human responses. Besides the common challenges faced in visual language navigation, vision-dialog navigation also requires to handle well with the language intentions of a series of questions about the temporal context from dialogue history and co-reasoning both dialogs and visual scenes. In this paper, we propose the Cross-modal Memory Network (CMN) for remembering and understanding the rich information relevant to historical navigation actions. Our CMN consists of two memory modules, the language memory module (L-mem) and the visual memory module (V-mem). Specifically, L-mem learns latent relationships between the current language interaction and a dialog history by employing a multi-head attention mechanism. V-mem learns to associate the current visual views and the cross-modal memory about the previous navigation actions. The cross-modal memory is generated via a vision-to-language attention and a language-to-vision attention. Benefiting from the collaborative learning of the L-mem and the V-mem, our CMN is able to explore the memory about the decision making of historical navigation actions which is for the current step. Experiments on the CVDN dataset show that our CMN outperforms the previous state-of-the-art model by a significant margin on both seen and unseen environments. 1

Original languageEnglish
Title of host publicationProceedings - 33th IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2020
EditorsCe Liu, Greg Mori, Kate Saenko, Silvio Savarese
Place of PublicationPiscataway NJ USA
PublisherIEEE, Institute of Electrical and Electronics Engineers
Pages10727-10736
Number of pages10
ISBN (Electronic)9781728171685
ISBN (Print)9781728171692
DOIs
Publication statusPublished - 2020
EventIEEE Conference on Computer Vision and Pattern Recognition 2020 - Virtual, China
Duration: 14 Jun 202019 Jun 2020
http://cvpr2020.thecvf.com (Website )
https://openaccess.thecvf.com/CVPR2020 (Proceedings)
https://ieeexplore.ieee.org/xpl/conhome/9142308/proceeding (Proceedings)

Publication series

NameProceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition
PublisherIEEE, Institute of Electrical and Electronics Engineers
ISSN (Print)1063-6919
ISSN (Electronic)2575-7075

Conference

ConferenceIEEE Conference on Computer Vision and Pattern Recognition 2020
Abbreviated titleCVPR 2020
CountryChina
CityVirtual
Period14/06/2019/06/20
Internet address

Cite this