Skip to main navigation Skip to search Skip to main content

Integrating Gaze and Speech for Enabling Implicit Interactions

Research output: Chapter in Book/Report/Conference proceedingConference PaperResearchpeer-review

Abstract

Gaze and speech are rich contextual sources of information that, when combined, can result in effective and rich multimodal interactions. This paper proposes a machine learning-based pipeline that leverages and combines users' natural gaze activity, the semantic knowledge from their vocal utterances and the synchronicity between gaze and speech data to facilitate users' interaction. We evaluated our proposed approach on an existing dataset, which involved 32 participants recording voice notes while reading an academic paper. Using a Logistic Regression classifier, we demonstrate that our proposed multimodal approach maps voice notes with accurate text passages with an average F1-Score of 0.90. Our proposed pipeline motivates the design of multimodal interfaces that combines natural gaze and speech patterns to enable robust interactions.

Original languageEnglish
Title of host publicationProceedings of the 2022 CHI Conference on Human Factors in Computing Systems
EditorsSteven Drucker, Julie Williamson, Koji Yatani
Place of PublicationNew York NY USA
PublisherAssociation for Computing Machinery (ACM)
Number of pages14
ISBN (Electronic)9781450391573
DOIs
Publication statusPublished - 2022
Externally publishedYes
EventInternational Conference on Human Factors in Computing Systems 2022 - New Orleans, United States of America
Duration: 29 Apr 20225 May 2022
Conference number: 40th
https://dl.acm.org/doi/proceedings/10.1145/3491102 (ACM Digital Library - proceedings)
https://chi2022.acm.org/ (Website)
https://dl.acm.org/doi/proceedings/10.1145/3491101 (Extended Abstracts)

Conference

ConferenceInternational Conference on Human Factors in Computing Systems 2022
Abbreviated titleCHI 2022
Country/TerritoryUnited States of America
CityNew Orleans
Period29/04/225/05/22
Internet address

Keywords

  • implicit annotation
  • natural gaze
  • natural language processing
  • semantic similarity
  • voice interfaces

Cite this