A research group including Professor Hideki Koike of the Department of Computer Science, the School of Computing at the Institute of Science Tokyo, and Professor Kris Kitani of the Robotics Institute at Carnegie Mellon University in the United States, has developed "DyaDiT," a deep learning model that generates natural bodily movements (gestures) from conversational audio. The model can automatically generate gestures that take into account the content of conversation, the relationship between speakers, and their personalities. This technology is expected to be applied to the communication systems of conversational AI agents. The results were presented at the international conference "The IEEE/CVF Conference on Computer Vision and Pattern Recognition 2026," which began on June 3.
Because most conventional gesture generation technologies generate gestures by inputting only a single speaker's audio, it has been difficult to generate natural conversational gestures. By utilizing social information such as the relationship between speakers and their personality traits, in addition to the conversational audio of two people, this model can generate appropriate reaction gestures in real time as a listener while the user is speaking. Furthermore, by taking the user's gestures and reactions as inputs, the model can generate natural movements corresponding to the situational context of the conversation.
The developed technology is expected to serve as a foundational technology for digital humans and conversational AI agents to communicate more naturally with humans. Moving forward, the team aims to generate richer conversational behaviors, including not only upper-body gestures but also full-body movements and facial expressions. If interfaces that enable natural communication between humans and AI are realized, applications are expected in a wide range of fields, such as education, customer service, and remote communication.
Koike stated, "With the development of AI, virtual conversational agents that interact naturally with users are being actively developed. Most conventional technologies returned uniform, identical gestures in response to a user's speech. However, the DyaDiT model we developed can generate different gestures depending on the relationship with the user (that with friend, family, stranger, etc.) and the user's reactions. This will make it possible to realize more natural conversational agents. The research was conducted as a joint study with Carnegie Mellon University in the United States. Although the timeframe from the start of the research to the writing of the paper was very short, researchers from Japan and the U.S. collaborated closely to successfully complete the project."
This article has been translated by JST with permission from The Science News Ltd. (https://sci-news.co.jp/). Unauthorized reproduction of the article and photographs is prohibited.

