TAP-ETS: Time-Aligned Phoneme Guiding for EMG-to-Speech Synthesis

Publication

TAP-ETS: Time-Aligned Phoneme Guiding for EMG-to-Speech Synthesis
Dongyub Han, Injune Hwang, Jaejun Lee, Jiwon Lee, and Kyogu Lee
Interspeech 2026, accepted paper
Period: Graduate

Demo : TAP-ETS demo

Github : https://github.com/ongdyub/TAP-ETS


Summary

TAP-ETS is a time-aligned phoneme guiding framework for EMG-to-speech synthesis. It conditions mel-spectrogram generation on frame-wise phoneme sequences through cross-attention and introduces refinement strategies that redistribute semantic guidance over frame-wise EMG signals.

On the Gaddy silent EMG benchmark, TAP-ETS reports a WER reduction from 25.12% to 19.77%.


Demo

The demo page contains figures and audio samples comparing the baseline and TAP-ETS outputs.

Open TAP-ETS demo