100% found this document useful (1 vote)

251 views39 pages

Understanding Speech Recognition Technology

This document discusses speech recognition technologies and the challenges of automatic speech recognition. It describes the main tasks of speech recognition as digitizing the speech signal, separating speech from noise, identifying individual phonemes despite variability, and interpreting prosodic features. The main approaches are described as template matching, rule-based systems using phonetic and linguistic rules, and statistical approaches using machine learning on large datasets. Key difficulties discussed include variability between individuals, distinguishing similar phonemes, interpreting prosodic cues, and handling disfluencies. Evaluation of speech recognition systems considers factors like vocabulary size, speaking style, and recognition method. Remaining challenges include improving robustness, adaptability, and handling out-of-vocabulary words.

Uploaded by

chayan_m_shah

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PPT, PDF, TXT or read online on Scribd

100% found this document useful (1 vote)

251 views39 pages

Understanding Speech Recognition Technology

Uploaded by

chayan_m_shah

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PPT, PDF, TXT or read online on Scribd

Speech recognition ?

Speech Recognition are technologies of particular interest, for their support of direct communication between humans and computers, through a communications mode, humans commonly use among themselves and at which they are highly skilled.

Rudnicky, Hauptman, and Lee

Types of speech recognition

Isolated words
Connected words Continuous speech Spontaneous speech (automatic speech

recognition) Voice verification and identification

Speech recognition uses and applications

Dictation
Command and control Telephony Medical/disabilities

Challenges of speech recognition

Ease of use Robust performance Automatic learning of new words and sounds Grammar for spoken language Control of synthesized voice quality Integrated learning for speech recognition and synthesis

Automatic speech recognition

What is the task?
What are the main difficulties? How is it approached? How good is it? How much better could it be?

What is the task?

Getting a computer to understand spoken

language By understand we might mean

React appropriately

Convert the input speech into another medium,

e.g. text

Several variables impinge on this (see later)

How do humans do it?

Articulation produces sound waves which the ear conveys to the

brain for processing

How might computers do it?

Acoustic waveform Acoustic signal

Digitization Acoustic analysis of the

speech signal Linguistic interpretation

Speech recognition

Whats hard about that?

Digitization
Converting analogue signal into digital representation

Signal processing

Separating speech from background noise

Phonetics
Variability in human speech

Phonology
Recognizing individual sound distinctions (similar phonemes)

Lexicology and syntax

Disambiguating homophones Features of continuous speech

Syntax and pragmatics

Interpreting prosodic features

Pragmatics
Filtering of performance errors (disfluencies)

Digitization
Analogue to digital conversion Sampling and quantizing Use filters to measure energy levels for various points on the frequency spectrum Knowing the relative importance of different frequency bands (for speech) makes this process more efficient E.g. high frequency sounds are less informative, so can be sampled using a broader bandwidth (log scale)

Separating speech from background noise Noise cancelling microphones

Two mics, one facing speaker, the other facing away
Ambient noise is roughly same for both mics

Knowing which bits of the signal relate to

speech

Spectrograph analysis

Variability in individuals speech

Variation among speakers due to
Vocal range (f0, and pitch range see later) Voice quality (growl, whisper, physiological

elements such as nasality, adenoidality, etc) ACCENT !!! (especially vowel systems, but also consonants, allophones, etc.)

Variation within speakers due to

Health, emotional state Ambient conditions

Speech style: formal read vs spontaneous

Speaker-(in)dependent systems Speaker-dependent systems

Require training to teach the system your individual

idiosyncracies

The more the merrier, but typically nowadays 5 or 10 minutes is enough User asked to pronounce some key words which allow computer to infer details of the users accent and voice Fortunately, languages are generally systematic

More robust But less convenient And obviously less portable

Speaker-independent systems
Language coverage is reduced to compensate need to be

flexible in phoneme identification Clever compromise is to learn on the fly

Identifying phonemes
Differences between some phonemes are

sometimes very small

May be reflected in speech signal (eg vowels

have more or less distinctive f1 and f2) Often show up in coarticulation effects (transition to next sound)

e.g. aspiration of voiceless stops in English

Allophonic variation

Disambiguating homophones
Mostly differences are recognised by humans by context and need to make sense
Its hard to wreck a nice beach What dimes a necks drain to stop port?

Systems can only recognize words that are in their lexicon, so limiting the lexicon is an obvious ploy Some ASR systems include a grammar which can help disambiguation

(Dis)continuous speech
Discontinuous speech much easier to

recognize
clearly

Single words tend to be pronounced more

Continuous speech involves contextual

coarticulation effects
Weak forms Assimilation

Contractions

Interpreting prosodic features

Pitch, length and loudness are used to

indicate stress All of these are relative

On a speaker-by-speaker basis And in relation to context

Pitch and length are phonemic in some

languages

Pitch
Pitch contour can be extracted from speech

signal

But pitch differences are relative One mans high is another (wo)mans low Pitch range is variable

Pitch contributes to intonation But has other functions in tone languages Intonation can convey meaning

Lengtheasy to measure but difficult to Length is

interpret Again, length is relative It is phonemic in many languages Speech rate is not constant slows down at the end of a sentence

Loudness
Loudness is easy to measure but difficult to

interpret Again, loudness is relative

Performance errors
Performance errors include Non-speech sounds Hesitations False starts, repetitions Filtering implies handling at syntactic level or

above Some disfluencies are deliberate and have pragmatic effect this is not something we can handle in the near future

Approaches to ASR
Template matching
Knowledge-based (or rule-based) approach Statistical approach: Noisy channel model + machine learning

Template-based approach
Store examples of units (words, phonemes),

then find the example that most closely fits the input Extract features from speech signal, then its just a complex similarity matching problem, using solutions developed for all sorts of applications OK for discrete utterances, and a single user

Template-based approach
Hard to distinguish very similar templates
And quickly degrades when input differs from

templates Therefore needs techniques to mitigate this degradation:

More subtle matching techniques Multiple templates which are aggregated

Taken together, these suggested

Rule-based approach
Use knowledge of phonetics and linguistics to

guide search process Templates are replaced by rules expressing everything (anything) that might help to decode:
Phonetics, phonology, phonotactics Syntax

Pragmatics

Rule-based based on blackboard approach Typical approach is

architecture:
At each decision point, lay out the possibilities Apply rules to determine which sequences are

permitted

Poor performance due to s h i: p i Difficulty to express rules t Difficulty to make rules interact Difficulty to know how to improve the system

t h s

Identify individual phonemes Identify words Identify sentence structure and/or meaning Interpret prosodic features (pitch, loudness, length)

Statistics-based approach
Can be seen as extension of template-based

approach, using more powerful mathematical and statistical tools Sometimes seen as anti-linguistic approach
Fred Jelinek (IBM, 1988): Every time I fire a

linguist my system improves

Statistics-based approach
Collect a large corpus of transcribed speech

recordings Train the computer to learn the correspondences (machine learning) At run time, apply statistical processes to search through the space of all possible solutions, and pick the statistically most likely one

Machine learning
Acoustic and Lexical Models Analyse training data in terms of relevant features Learn from large amount of data different possibilities

different phone sequences for a given word different combinations of elements of the speech signal for a given phone/phoneme

Combine these into a Hidden Markov Model

expressing the probabilities

HMMs for some words

Language model
Models likelihood of word given previous

word(s) n-gram models:

Build the model by calculating bigram or

trigram probabilities from text training corpus Smoothing issues

The Noisy Channel Model

Search through space of all possible sentences Pick the one that is most probable given the waveform

The Noisy Channel Model

Use the acoustic model to give a set of likely

phone sequences Use the lexical and language models to judge which of these are likely to result in probable word sequences The trick is having sophisticated algorithms to juggle the statistics A bit like the rule-based approach except that it is all learned automatically from data

Evaluation
Funders have been very keen on competitive

quantitative evaluation Subjective evaluations are informative, but not cost-effective For transcription tasks, word-error rate is popular (though can be misleading: all words are not equally important) For task-based dialogues, other measures of understanding are needed

Comparing ASR systems

Factors include
Speaking mode: isolated words vs continuous

speech Speaking style: read vs spontaneous Enrollment: speaker (in)dependent Vocabulary size (small <20 large > 20,000) Equipment: good quality noise-cancelling mic telephone Size of training set (if appropriate) or rule set Recognition method

Remaining problems
Robustness graceful degradation, not catastrophic failure Portability independence of computing platform Adaptability to changing conditions (different mic,

background noise, new speaker, new task domain, new language even) Language Modelling is there a role for linguistics in improving the language models? Confidence Measures better methods to evaluate the absolute correctness of hypotheses. Out-of-Vocabulary (OOV) Words Systems must have some method of detecting OOV words, and dealing with them in a sensible way. Spontaneous Speech disfluencies (filled pauses, false starts, hesitations, ungrammatical constructions etc) remain a problem. Prosody Stress, intonation, and rhythm convey important information for word recognition and the user's intentions (e.g., sarcasm, anger) Accent, dialect and mixed language non-native speech is a huge problem, especially where code-switching is commonplace

THANKS

39/34

Challenges in Automatic Speech Recognition
No ratings yet
Challenges in Automatic Speech Recognition
34 pages
Speech Recognition UTHM
No ratings yet
Speech Recognition UTHM
30 pages
Understanding Automatic Speech Recognition
No ratings yet
Understanding Automatic Speech Recognition
3 pages
Speechrecognitionfinalpresentation 141124072610 Conversion Gate01
No ratings yet
Speechrecognitionfinalpresentation 141124072610 Conversion Gate01
30 pages
Feature Extraction Using PCA
No ratings yet
Feature Extraction Using PCA
36 pages
Speech Recognition Technology Overview
No ratings yet
Speech Recognition Technology Overview
19 pages
Understanding Speech Recognition Technology
100% (1)
Understanding Speech Recognition Technology
20 pages
Understanding Speech Recognition Technology
100% (1)
Understanding Speech Recognition Technology
17 pages
Overview of Automatic Speech Recognition
No ratings yet
Overview of Automatic Speech Recognition
35 pages
Understanding Speech Recognition Systems
100% (2)
Understanding Speech Recognition Systems
26 pages
Understanding Speech Recognition Basics
0% (1)
Understanding Speech Recognition Basics
27 pages
Understanding Speech Recognition Technology
No ratings yet
Understanding Speech Recognition Technology
26 pages
Speech Recognition PPT F
100% (3)
Speech Recognition PPT F
16 pages
Speech Recognition
No ratings yet
Speech Recognition
4 pages
Indian Accent Speech Recognition Report
No ratings yet
Indian Accent Speech Recognition Report
14 pages
Speech Recognition Course Guide
No ratings yet
Speech Recognition Course Guide
74 pages
ASR
No ratings yet
ASR
13 pages
Understanding Automatic Speech Recognition
No ratings yet
Understanding Automatic Speech Recognition
9 pages
A Review On Different Approaches For Speech - Recognition System
No ratings yet
A Review On Different Approaches For Speech - Recognition System
6 pages
(IJCST-V4I2P62) :Dr.V.Ajantha Devi, Ms.V.Suganya
No ratings yet
(IJCST-V4I2P62) :Dr.V.Ajantha Devi, Ms.V.Suganya
6 pages
Overview of Speech Recognition Systems
No ratings yet
Overview of Speech Recognition Systems
14 pages
Understanding Automatic Speech Recognition
No ratings yet
Understanding Automatic Speech Recognition
30 pages
Speech Technology
No ratings yet
Speech Technology
5 pages
Speech Recognition for Developers
No ratings yet
Speech Recognition for Developers
38 pages
Machine Learning in Speech Recognition
No ratings yet
Machine Learning in Speech Recognition
70 pages
Speech Recognition
No ratings yet
Speech Recognition
17 pages
Voice-Enabled Phone Directory System
No ratings yet
Voice-Enabled Phone Directory System
13 pages
Understanding Automatic Speech Recognition
No ratings yet
Understanding Automatic Speech Recognition
49 pages
Speech Processing and Recognition Overview
No ratings yet
Speech Processing and Recognition Overview
26 pages
Overview of Speech Recognition Systems
No ratings yet
Overview of Speech Recognition Systems
16 pages
Automatic Speech Recognition (ASR) : Omar Khalil Gómez - Università Di Pisa
100% (1)
Automatic Speech Recognition (ASR) : Omar Khalil Gómez - Università Di Pisa
65 pages
Speech Recognition and ASR Overview
No ratings yet
Speech Recognition and ASR Overview
11 pages
Speech Recognition Technology Overview
No ratings yet
Speech Recognition Technology Overview
11 pages
Speech Recognition System Overview
No ratings yet
Speech Recognition System Overview
7 pages
Multilingual Automatic Speech Recognition
No ratings yet
Multilingual Automatic Speech Recognition
17 pages
Speech Recognition Fundamentals Explained
No ratings yet
Speech Recognition Fundamentals Explained
39 pages
Speech Recognition1
No ratings yet
Speech Recognition1
24 pages
Speech Recognition with Neural Networks
No ratings yet
Speech Recognition with Neural Networks
23 pages
Automatic Speech Recognition Study
100% (1)
Automatic Speech Recognition Study
2 pages
Final Report
No ratings yet
Final Report
35 pages
Ai Project Sona-1 (1) - 250630 - 194118
No ratings yet
Ai Project Sona-1 (1) - 250630 - 194118
10 pages
Speech Recognition Seminar Overview
No ratings yet
Speech Recognition Seminar Overview
26 pages
Next-Gen Speech Recognition Systems
No ratings yet
Next-Gen Speech Recognition Systems
7 pages
Four Types of Speech Recognition
No ratings yet
Four Types of Speech Recognition
17 pages
Automatic Speech Recognition Documentation
No ratings yet
Automatic Speech Recognition Documentation
24 pages
Robust Speech Recognition Using Articulatory Information: Der Technischen Fakult at Der Universit at Bielefeld
100% (1)
Robust Speech Recognition Using Articulatory Information: Der Technischen Fakult at Der Universit at Bielefeld
148 pages
Speech Synthesis and Recognition Course
No ratings yet
Speech Synthesis and Recognition Course
6 pages
Lecture 04
No ratings yet
Lecture 04
543 pages
Speech Signal Processing Course Overview
No ratings yet
Speech Signal Processing Course Overview
48 pages
NLP 1.3.1 - Speed Recogmnition
No ratings yet
NLP 1.3.1 - Speed Recogmnition
20 pages
Overview of Automatic Speech Recognition
No ratings yet
Overview of Automatic Speech Recognition
77 pages
Unit 5 UA
No ratings yet
Unit 5 UA
19 pages
CCS369 - TSS-Unit 5
No ratings yet
CCS369 - TSS-Unit 5
23 pages
AI Applications in Speech Recognition
No ratings yet
AI Applications in Speech Recognition
10 pages
Thesis-Speech Recognition Markov
No ratings yet
Thesis-Speech Recognition Markov
65 pages
Speech Recognition On Mobile Devices
No ratings yet
Speech Recognition On Mobile Devices
27 pages
Overview of Speech Recognition Systems
No ratings yet
Overview of Speech Recognition Systems
12 pages
Prasar Bharti 2013 Recruitment Details
No ratings yet
Prasar Bharti 2013 Recruitment Details
1 page
Comprehensive Guide to OLED Technology
No ratings yet
Comprehensive Guide to OLED Technology
4 pages
2012 Norton Cybercrime Report
No ratings yet
2012 Norton Cybercrime Report
2 pages
Blu-ray Disc: Revolution in Storage Technology
No ratings yet
Blu-ray Disc: Revolution in Storage Technology
12 pages
Pheonix Led Television
No ratings yet
Pheonix Led Television
24 pages
Brain-Machine Interfaces: Past, Present and Future: Mikhail A. Lebedev and Miguel A.L. Nicolelis
No ratings yet
Brain-Machine Interfaces: Past, Present and Future: Mikhail A. Lebedev and Miguel A.L. Nicolelis
11 pages
Speech Processing
No ratings yet
Speech Processing
16 pages
Pheonix Led Television
No ratings yet
Pheonix Led Television
24 pages
Stealth Technology
No ratings yet
Stealth Technology
2 pages
Understanding Programmable Logic Controllers
No ratings yet
Understanding Programmable Logic Controllers
33 pages
Laser Fundamentals in Photonics Module
No ratings yet
Laser Fundamentals in Photonics Module
45 pages
Instrument Landing System
100% (1)
Instrument Landing System
7 pages
Understanding Rocket Science Basics
No ratings yet
Understanding Rocket Science Basics
10 pages
Overview of Control Valves in Systems
100% (1)
Overview of Control Valves in Systems
26 pages
Understanding Television Systems and Signals
No ratings yet
Understanding Television Systems and Signals
158 pages
Understanding Television Systems and Signals
No ratings yet
Understanding Television Systems and Signals
158 pages
These Free Access Uploder Are Not Working Properly
No ratings yet
These Free Access Uploder Are Not Working Properly
1 page
Bilingual Infants' Speech Production Insights
No ratings yet
Bilingual Infants' Speech Production Insights
2 pages
ECPE Speaking Sample Commentary
100% (1)
ECPE Speaking Sample Commentary
4 pages
Phonetics Assignment: Intonation & Prosody
No ratings yet
Phonetics Assignment: Intonation & Prosody
9 pages
PURPOSIVE COMMUNICATION Final
No ratings yet
PURPOSIVE COMMUNICATION Final
18 pages
English Curriculum Map for Grade 7
100% (1)
English Curriculum Map for Grade 7
16 pages
Discourse Analysis in English Linguistics
No ratings yet
Discourse Analysis in English Linguistics
64 pages
Supra Segment A Ls
No ratings yet
Supra Segment A Ls
47 pages
Infixation
No ratings yet
Infixation
5 pages
Computer Assisted Pronunciation Training A Systematic Review
No ratings yet
Computer Assisted Pronunciation Training A Systematic Review
21 pages
Lpgrade7 10
No ratings yet
Lpgrade7 10
17 pages
Prosody's Role in Foreign Accent Perception
No ratings yet
Prosody's Role in Foreign Accent Perception
4 pages
The Role of Segmental and Suprasegmental Features in Spoken English Language
100% (1)
The Role of Segmental and Suprasegmental Features in Spoken English Language
10 pages
English 10 Module 2 For Students
No ratings yet
English 10 Module 2 For Students
5 pages
Phonology & Phonetics Bibliography
No ratings yet
Phonology & Phonetics Bibliography
22 pages
Purcom Semi Finals Coverage
No ratings yet
Purcom Semi Finals Coverage
19 pages
Pali Prosody by Charles Duroiselle
No ratings yet
Pali Prosody by Charles Duroiselle
10 pages
Suprasegmental Phonology Book 2
No ratings yet
Suprasegmental Phonology Book 2
57 pages
Frankel Becker Rowe Pearson JofE 2016 196 WhatisLiteracy
No ratings yet
Frankel Becker Rowe Pearson JofE 2016 196 WhatisLiteracy
13 pages
Formulating VB MAPP Interventions
No ratings yet
Formulating VB MAPP Interventions
200 pages
English8 Q2 Week1 1
No ratings yet
English8 Q2 Week1 1
20 pages
Understanding Intonation in Linguistics
No ratings yet
Understanding Intonation in Linguistics
2 pages
English Daily Lesson Log for Grade VII
100% (8)
English Daily Lesson Log for Grade VII
28 pages
UC Berkeley: The CATESOL Journal
No ratings yet
UC Berkeley: The CATESOL Journal
21 pages
English
No ratings yet
English
2 pages
Intonation in Please-Requests Study
No ratings yet
Intonation in Please-Requests Study
29 pages
Prosodic Features PDF
No ratings yet
Prosodic Features PDF
3 pages
Lessons 3-4 Suprasegmental Features
No ratings yet
Lessons 3-4 Suprasegmental Features
9 pages
Summary and Analysis Assignment Guide
No ratings yet
Summary and Analysis Assignment Guide
5 pages
English Stress: Definitions and Functions
No ratings yet
English Stress: Definitions and Functions
13 pages
Klɛrɪhju / Edmund Clerihew Bentley
No ratings yet
Klɛrɪhju / Edmund Clerihew Bentley
5 pages

Understanding Speech Recognition Technology

Uploaded by

Understanding Speech Recognition Technology

Uploaded by

Speech recognition ?

Rudnicky, Hauptman, and Lee

Types of speech recognition

recognition) Voice verification and identification

Speech recognition uses and applications

Challenges of speech recognition

Automatic speech recognition

What is the task?

language By understand we might mean

Convert the input speech into another medium,

Several variables impinge on this (see later)

How do humans do it?

Articulation produces sound waves which the ear conveys to the

brain for processing

How might computers do it?

Digitization Acoustic analysis of the

speech signal Linguistic interpretation

Whats hard about that?

Separating speech from background noise

Lexicology and syntax

Syntax and pragmatics

Separating speech from background noise Noise cancelling microphones

Knowing which bits of the signal relate to

Variability in individuals speech

Variation within speakers due to

Speech style: formal read vs spontaneous

Speaker-(in)dependent systems Speaker-dependent systems

More robust But less convenient And obviously less portable

flexible in phoneme identification Clever compromise is to learn on the fly

sometimes very small

May be reflected in speech signal (eg vowels

e.g. aspiration of voiceless stops in English

Single words tend to be pronounced more

Continuous speech involves contextual

Interpreting prosodic features

indicate stress All of these are relative

On a speaker-by-speaker basis And in relation to context

Pitch and length are phonemic in some

Lengtheasy to measure but difficult to Length is

interpret Again, loudness is relative

templates Therefore needs techniques to mitigate this degradation:

Taken together, these suggested

Rule-based based on blackboard approach Typical approach is

linguist my system improves

Combine these into a Hidden Markov Model

expressing the probabilities

HMMs for some words

word(s) n-gram models:

Build the model by calculating bigram or

trigram probabilities from text training corpus Smoothing issues

The Noisy Channel Model

The Noisy Channel Model

Comparing ASR systems

You might also like