Speech processing
Study of speech signals and their digital processing methods.
Speech processing examines how speech signals are captured and analyzed, typically in digital form, making it a subset of digital signal processing focused on spoken language. It covers how these signals are obtained, modified, stored, transmitted, and played back. Common tasks include recognizing words, generating synthetic speech, identifying who is speaking, cleaning up noisy audio, and verifying a speaker's identity.
Early work in this field concentrated on recognizing simple vowel sounds. In 1952, researchers at Bell Labs built a system that could identify spoken digits from a single person. During the 1940s, studies had already explored speech recognition by examining sound spectra. The linear predictive coding (LPC) algorithm was introduced in 1966 by researchers in Japan, and later refined at Bell Labs in the 1970s. LPC became the foundation for voice-over-IP technology and for speech synthesis chips, like those used in the Speak & Spell toys released in 1978. The first commercial speech recognition product, Dragon Dictate, came out in 1990. Two years later, Bell Labs technology enabled AT&T to route calls automatically using voice commands. By then, these systems had vocabularies larger than an average person's.
Around the early 2000s, the main approach in speech processing began moving from hidden Markov models toward neural networks and deep learning. In 2012, a team at the University of Toronto showed that deep neural networks could beat traditional HMM-based systems on large-vocabulary continuous speech recognition. This led to widespread industry adoption of deep learning. By the mid-2010s, major tech companies had integrated advanced speech recognition into virtual assistants like Google Assistant, Cortana, Alexa, and Siri, using deep learning for more natural interactions. Transformer-based models such as BERT and GPT later improved context-aware understanding of speech. More recently, end-to-end speech recognition models have become popular, converting audio directly into text without separate feature extraction or acoustic modeling steps.
Dynamic time warping measures similarity between two sequences that may differ in speed, finding the optimal alignment between them with minimal cost, calculated as the sum of absolute differences between matched indices.
A hidden Markov model estimates an unobserved variable over time based on observed data. It relies on the Markov property: the current hidden state depends only on the previous one, and each observation depends only on the current hidden state.
An artificial neural network consists of connected units called artificial neurons, loosely modeled after biological neurons. Connections transmit signals as real numbers, and each neuron's output is a nonlinear function of its summed inputs.
Phase in speech signals is often treated as random but carries useful information. Phase wrapping occurs due to periodic jumps of 2π, and phase unwrapping expresses the phase as the sum of a linear component and a residual component.
- field
- Digital signal processing, speech recognition, speech synthesis
- known_for
- Enabling voice-operated systems, virtual assistants, and call center automation through techniques like linear predictive coding, hidden Markov models, and deep neural networks
Lore & Background
Early attempts at speech processing and recognition were primarily focused on understanding a handful of simple phonetic elements such as vowels. Biddulph, and K. H. Davis—developed a system that could recognize digits spoken by a single speaker. Pioneering works in the field of speech recognition using analysis of its spectrum were reported in the 1940s. Further developments in LPC technology were made by Bishnu S. Atal and Manfred R. Schroeder at Bell Labs during the 1970s.
Reader's Guide
Speech processing has evolved from early vowel recognition to sophisticated systems that power modern virtual assistants. The development of linear predictive coding in the 1960s and 1970s laid the groundwork for voice-over-IP and speech synthesis. By the early 2000s, the dominant strategy shifted from hidden Markov models to neural networks and deep learning. In 2012, Geoffrey Hinton and his team at the University of Toronto demonstrated that deep neural networks could significantly outperform traditional HMM-based systems on large vocabulary continuous speech recognition tasks, leading to widespread industry adoption. By the mid-2010s, companies like Google, Microsoft, Amazon, and Apple integrated advanced speech recognition into virtual assistants such as Google Assistant, Cortana, Alexa, and Siri. Transformer-based models like BERT and GPT further pushed boundaries, enabling more context-aware understanding. End-to-end speech recognition models have recently gained popularity by directly converting audio input into text output, streamlining development and improving performance.
More in Communication & Media 1-19
Spotted an error? Know more?
This is a living reference — every entry is fact-audited, and reader corrections feed straight into our audit queue. Suggest an edit · See this site's audit record
