Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Improved Speech Reconstruction from Silent Video

Aug 29, 2017

Ariel Ephrat, Tavi Halperin, Shmuel Peleg

Figure 1 for Improved Speech Reconstruction from Silent Video

Figure 2 for Improved Speech Reconstruction from Silent Video

Figure 3 for Improved Speech Reconstruction from Silent Video

Figure 4 for Improved Speech Reconstruction from Silent Video

Share this with someone who'll enjoy it:

Abstract:Speechreading is the task of inferring phonetic information from visually observed articulatory facial movements, and is a notoriously difficult task for humans to perform. In this paper we present an end-to-end model based on a convolutional neural network (CNN) for generating an intelligible and natural-sounding acoustic speech signal from silent video frames of a speaking person. We train our model on speakers from the GRID and TCD-TIMIT datasets, and evaluate the quality and intelligibility of reconstructed speech using common objective measurements. We show that speech predictions from the proposed model attain scores which indicate significantly improved quality over existing models. In addition, we show promising results towards reconstructing speech from an unconstrained dictionary.

* Accepted to ICCV 2017 Workshop on Computer Vision for Audio-Visual Media. Supplementary video: https://www.youtube.com/watch?v=Xjbn7h7tpg0. arXiv admin note: text overlap with arXiv:1701.00495

View paper on

Share this with someone who'll enjoy it:

Title:Improved Speech Reconstruction from Silent Video

Paper and Code