Shared by automation-1 using Learnlo
Create your own pack βPick a topic to learn or start your exam journey.
0/20 topics mastered
Speech synthesis is the artificial production of human speech. A computer system that performs this task is called a speech synthesizer, and it can be implemented in either software or hardware. In particular, text-to-speech (TTS) systems convert normal language text into spoken output, while other speech-synthesis systems can render symbolic linguistic representations (such as phonetic transcriptions) into sound. Speech synthesizers can generate speech in different ways. One common approach is to concatenate segments of recorded speech stored in a database, where the choice and size of stored units (e.g., phones or diphones versus whole words/sentences) affect the trade-off between output range and clarity. Another approach is to model the vocal tract and other human voice characteristics to produce a more fully synthetic voice. Overall quality is judged by how natural the output sounds and how intelligible it is. A typical TTS system is organized into two parts: a front-end and a back-end. The front-end normalizes text (e.g., expanding symbols like numbers and abbreviations), converts text to phonetic transcriptions (grapheme-to-phoneme or text-to-phoneme conversion), and assigns prosodic structure such as phrases, clauses, and sentences. The back-end (the synthesizer) then converts this symbolic representation into audio, sometimes also computing target prosody such as pitch contour and phoneme durations before producing the final speech waveform.
0/2 modes complete
0/2 modes complete