text to speech neural network

Help | Advanced Search

Electrical Engineering and Systems Science > Audio and Speech Processing

Title: a survey on neural speech synthesis.

Abstract: Text to speech (TTS), or speech synthesis, which aims to synthesize intelligible and natural speech given text, is a hot research topic in speech, language, and machine learning communities and has broad applications in the industry. As the development of deep learning and artificial intelligence, neural network-based TTS has significantly improved the quality of synthesized speech in recent years. In this paper, we conduct a comprehensive survey on neural TTS, aiming to provide a good understanding of current research and future trends. We focus on the key components in neural TTS, including text analysis, acoustic models and vocoders, and several advanced topics, including fast TTS, low-resource TTS, robust TTS, expressive TTS, and adaptive TTS, etc. We further summarize resources related to TTS (e.g., datasets, opensource implementations) and discuss future research directions. This survey can serve both academic researchers and industry practitioners working on TTS.

Submission history

Access paper:.

Other Formats

References & Citations

Google Scholar
Semantic Scholar

BibTeX formatted citation

Bibliographic and Citation Tools

Code, data and media associated with this article, recommenders and search tools.

Institution

arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs .

IMAGES

Speech Recognition using Convolutional Neural Networks
Neural Network Architecture for Text-to-Speech Synthesis
Speech Synthesis Techniques using Deep Neural Networks
Azure AI milestone: New Neural Text-to-Speech models more closely
Convolutional Neural Network Architecture for Speech Recognition
语音到文本引擎的声学模型训练-腾讯云开发者社区-腾讯云

VIDEO

Decoding the neural processing of speech
How to get started with neural text to speech in Azure
Voicer App
RNN for text generation?
Luganda text-to-speech: breaking barriers and promoting accessibility for visualliy impaired
NVIDIA Riva Automatic Speech Recognition for AudioCodes VoiceAI Connect Users

COMMENTS

Text-To-Speech Synthesis
Efficiently Trainable Text-to-Speech System Based on Deep Convolutional Networks with Guided Attention. coqui-ai/TTS • • 24 Oct 2017. This paper describes a novel text-to-speech (TTS) technique based on deep convolutional neural networks (CNN), without use of any recurrent units.
[2106.15561] A Survey on Neural Speech Synthesis
This paper provides a comprehensive overview of neural TTS, a research topic that aims to synthesize natural and intelligible speech from text. It covers the key components, advanced topics, and future directions of neural TTS, as well as the related resources and references.