Latest News

sciencenews.png

Speech production via physics-informed machine learning

2026.07.31

Physics-Informed Neural Networks (PINNs), a method that embeds partial differential equations representing physical laws into neural network training, have attracted considerable attention in recent years. A research group consisting of Assistant Professor Kazuya Yokota of the Department of Mechanical Engineering, Associate Professor Ryosuke Harakawa of the Department of Electrical, Electronics and Information Engineering, Associate Professor Masaaki Baba of the Department of Mechanical Engineering, and Professor Masahiro Iwahashi of the Department of Electrical, Electronics and Information Engineering at Nagaoka University of Technology has developed a novel speech production method that couples vocal fold vibration with vocal tract resonance using PINNs. The study was published in IEEE Transactions on Audio, Speech, and Language Processing.

By training a neural network on a physical model of human speech production, it becomes possible to synthesize speech and estimate the state of the vocal folds from speech.
Provided by Assistant Professor Kazuya Yokota, Nagaoka University of Technology.

The team introduced several architectures into the model, including a framework to learn the complex mechanics of vocal fold closure, a method to estimate pitch from lung pressure via neural network training, and a network topology capable of handling the mutual interaction between the vocal folds and the vocal tract.

To validate the model, the researchers performed both forward and inverse analyses. In the forward analysis, they successfully produced speech waveforms for the vowels "a" and "u", confirming the validity of the approach through a comparison with conventional numerical simulation methods. This marks the first reported case of speech production utilizing PINNs that have learned the physical laws governing the vocal folds and vocal tract. In the inverse analysis, under conditions where the vocal tract shape and vocal fold parameters were known, they demonstrated that vocal fold movement, glottal flow, and subglottal pressure could be accurately estimated from simulation-produced speech waveforms.

This study proposes a new speech production methodology to understand how human speech is generated by linking observational data with physical laws. In the future, by estimating the state of vocal organs such as the vocal folds and vocal tract directly from audio data, this approach is expected to evolve into a new speech production and analysis tool in fields such as elucidating vocalization mechanisms, analyzing voice disorders, and supporting vocal training.

Journal Information
Publication: IEEE Transactions on Audio, Speech and Language Processing
Title: Physics-Informed Neural Networks for Speech Production
DOI: 10.1109/TASLPRO.2026.3700036

This article has been translated by JST with permission from The Science News Ltd. (https://sci-news.co.jp/). Unauthorized reproduction of the article and photographs is prohibited.

Back to Latest News

Latest News

Recent Updates

    Most Viewed