Lombard Speech Synthesis Using Transfer Learning in a Tacotron Text-to-Speech System

Currently, there is increasing interest to use sequence-tosequence models in text-to-speech (TTS) synthesis with attention like that in Tacotron models. These models are end-to-end, meaning that they learn both co-articulation and duration properties directly from text and speech. Since these models are entirely data-driven, they need large amounts of data to generate synthetic speech of good quality. However, in challenging speaking styles, such as Lombard speech, it is difficult to record… 
