Microsoft ends support for Internet Explorer on June 16, 2022.
We recommend using one of the browsers listed below.

  • Microsoft Edge(Latest version) 
  • Mozilla Firefox(Latest version) 
  • Google Chrome(Latest version) 
  • Apple Safari(Latest version) 

Please contact your browser provider for download and installation instructions.

Open search panel Close search panel Open menu Close menu

October 7, 2026

Information

3 papers from NTT accepted to ICMI2026, a leading international conference on multimodal AI and human-society interaction

From October 5 to 9, 2026, the 28th ACM International Conference on Multimodal Interaction (ICMI 2026), a top international conference in the field of multimodal interaction, will be held in Naples, Italy. Three papers from NTT Laboratories have been accepted to the main conference. ICMI is a leading international conference on multimodal interaction, which focuses on the integrated processing and understanding of multiple types of information (modalities), such as speech, language, facial expressions, gaze, and body movements, involved in human-human and human-AI/robot communication. ICMI 2026 will present cutting-edge research on topics including multimodal dialogue understanding and technologies for enabling natural interactions between humans and AI systems or robots.

Abbreviated names of the laboratories:
HI:Human Informatics Labs., NTT, Inc.

■Self-Distillation as a Structure-Dependent Mechanism in Multimodal Dialogue via Multi-Context Modeling

Ryo Ishii (HI), Chihiro Takayama (HI), Jiro Nagao (HI), Toshiki Onishi (HI), Yukiko I. Nakano (Seikei University), Junichi Sawase (HI)

For AI to understand multiparty conversations, it is important to capture not only what is being said, but also information such as vocal and facial cues, interactions among participants, and each participant’s communication behavior. In this study, we proposed a multi-context model that separately captures interaction dynamics across the entire conversation and the behavior of individual participants. We further demonstrated that the effectiveness of “self-distillation,” which uses the model’s own predictions as learning signals, depends on how dialogue information is represented, and that appropriately combining these approaches achieves high performance in dialogue understanding.

■Causal Temporal Padding for Low-Latency Real-Time Gesture Generation

Ryo Ishii (HI), Shinichiro Eitoku (HI), Jiro Nagao (HI), Junichi Sawase (HI)

For AI agents and robots that interact with humans, it is important to generate body movements synchronized with speech without delay. In this study, we identified that conventional gesture generation models implicitly require future information during internal processing, which contributes to latency. To address this issue, we proposed a method that sequentially generates movements using only past and current information. Our evaluation confirmed that the proposed method substantially reduces processing latency while maintaining gesture quality and improves the responsiveness perceived by users.

One other paper was accepted.

Information is current as of the date of issue of the individual topics.
Please be advised that information may be outdated after that point.