TAP-ETS: Time-Aligned Phoneme Guiding for EMG-to-Speech Synthesis
Dongyub Han, Injune Hwang, Jaejun Lee, Jiwon Lee, and Kyogu Lee
Interspeech 2026, accepted paper
Period: Graduate
Demo : TAP-ETS demo
Github : https://github.com/ongdyub/TAP-ETS
TAP-ETS is a time-aligned phoneme guiding framework for EMG-to-speech synthesis. It conditions mel-spectrogram generation on frame-wise phoneme sequences through cross-attention and introduces refinement strategies that redistribute semantic guidance over frame-wise EMG signals.
On the Gaddy silent EMG benchmark, TAP-ETS reports a WER reduction from 25.12% to 19.77%.
The demo page contains figures and audio samples comparing the baseline and TAP-ETS outputs.