Shared by automation-1 using Learnlo
Create your own pack →Pick a topic to learn or start your exam journey.
0/20 topics mastered
Speech recognition (automatic speech recognition, ASR; computer speech recognition; speech-to-text, STT) is a computational linguistics sub-field focused on methods and technologies that convert spoken language into text or other interpretable outputs. Its scope includes translating audio speech signals into symbolic representations (e.g., words or phonemes) and supporting downstream uses such as voice user interfaces, transcription, dictation, and searching within audio recordings. It can also support related tasks like analyzing speaker characteristics (e.g., native language cues) and, in a different but related area, voice recognition (speaker identification) for authentication or simplifying recognition when systems are trained for a particular speaker’s voice. Within its scope, speech recognition systems typically combine acoustic modeling (how speech sounds map to units) and language modeling (how word sequences are likely) to produce the most probable transcription. Historically, the field has evolved from early single-speaker and limited-vocabulary systems toward large-vocabulary, continuous, and speaker-independent recognition, with major progress driven by statistical models (notably hidden Markov models and dynamic time warping) and later by deep learning approaches (e.g., LSTMs, CTC-trained models, and attention-based transformers). Performance evaluation is commonly framed in terms of accuracy (such as word error rate) and speed, and recognition accuracy depends on factors like vocabulary size, speaker dependence, speech type (isolated vs. continuous), task constraints, and adverse conditions such as noise and echoes.
0/2 modes complete
0/2 modes complete