语音信号处理讲义demo-sen_第1页
语音信号处理讲义demo-sen_第2页
语音信号处理讲义demo-sen_第3页
语音信号处理讲义demo-sen_第4页
语音信号处理讲义demo-sen_第5页
已阅读5页,还剩20页未读 继续免费阅读

付费下载

下载本文档

版权说明:本文档由用户提供并上传,收益归属内容提供方,若内容存在侵权,请进行举报或认领

文档简介

1、Report:2002-2003Dr. Zhang Sen Chinese Academy of SciencesBeijing, China2021/8/151952 Bell Labs DigitsFirst word (digit) recognizerApproximates energy in formants (vocal tract resonances) over wordAlready has some robust ideas(insensitive to amplitude, timing variation)Worked very wellMain weakness w

2、as technological (resistorsand capacitors)ASR Progress Overview50S ISOLATED DIGIT RECOGNITION (BELL LAB) 60S : HARDWARE SPEECH SEGMENTATOR (JAPAN) DYNAMIC PROGRAMMING (U.S.S.R) 70S : CLUSTERING ALGORITHM (SPEAKER INDEPENDECY) DTW 80S: HMM, DARPA, SPHINX 90S : ADAPTION, ROBUSTNESSThe 60sBetter digit

3、recognitionBreakthroughs: Spectrum Estimation (FFT, cepstra, LPC), Dynamic Time Warp (DTW), and Hidden Markov Model (HMM) theoryHARDWARE SPEECH SEGMENTATOR (JAPAN)1971-76 ARPA ProjectFocus on Speech UnderstandingMain work at 3 sites: System DevelopmentCorporation, CMU and BBNOther work at Lincoln, S

4、RI, BerkeleyGoal was 1000-word ASR, a few speakers,connected speech, constrained grammar,less than 10% semantic errorResultsOnly CMU Harpy fulfilled goals - used LPC, segments, lots of high levelknowledge, learned from Dragon *(Baker)* The CMU system done in the early 70s; as opposed to the company

5、formed in the 80sAchieved by 1976Spectral and cepstral features, LPCSome work with phonetic featuresIncorporating syntax and semanticsInitial Neural Network approachesDTW-based systems (many)HMM-based systems (Dragon, IBM)Dynamic Time WarpOptimal time normalization with dynamic programmingProposed b

6、y Sakoe and Chiba, circa 1970Similar time, proposal by ItakuraProbably Vintsyuk was first (1968)HMMs for SpeechMath from Baum and others, 1966-1972Applied to speech by Baker in theoriginal CMU Dragon System (1974)Developed by IBM (Baker, Jelinek, Bahl,Mercer,.) (1970-1993) Extended by others in the

7、mid-1980sThe 1980sCollection of large standard corporaFront ends: auditory models, dynamicsEngineering: scaling to large vocabulary continuous speech Second major (D)ARPA ASR projectHMMs e ready for prime timeStandard Corpora CollectionBefore 1984, chaosTIMITRM (later WSJ) ATISNIST, ARPA, LDCFront E

8、nds in the 1980sMel cepstrum (Bridle, Mermelstein)PLP (Hermansky)Delta cepstrum (Furui) Auditory models (Seneff, Ghitza, others)Dynamic Speech Features temporal dynamics useful for ASR local time derivatives of cepstra “delta features estimated over multiple frames (typically 5) usually augments sta

9、tic features can be viewed as a temporal filterHMMs for Continuous SpeechUsing dynamic programming for cts speech(Vintsyuk, Bridle, Sakoe, Ney.)Application of Baker-Jelinek ideas to continuous speech (IBM, BBN, Philips, .)Multiple groups developing major HMMsystems (CMU, SRI, Lincoln, BBN, ATT) Engi

10、neering development - coping with data, fast computers2nd (D)ARPA ProjectCommon taskFrequent evaluationsConvergence to good, but similar, systems Lots of engineering development - now up to 60,000 word recognition, in real time, on aworkstation, with less than 10% word errorCompetition inspired othe

11、rs not in project -Cambridge did HTK, now widely distributedSome 1990s IssuesIndependence to long-term spectrumAdaptationEffects of spontaneous speechInformation retrieval/extraction withbroadcast material Query-style systems (e.g., ATIS)Applying ASR technology to relatedareas (language ID, speaker

12、verification)We still needWe still need scienceNeed language, intelligenceAcoustic robustness still poorPerceptual research, modelsFundamentals of statistical patternrecognition for sequencesRobustness to accent, stress,rate of speech, .Progress in past YearsFrom digits to 60,000 wordsFrom single sp

13、eakers to manyFrom isolated words to continuousspeech From no products to many products,some systems actually saving LOTSof moneyReal UsesTelephone: phone company services(collect versus credit card)Telephone: call centers for queryinformation (e.g., stock quotes, parcel tracking)Dictation products:

14、 continuous recognition, speaker dependent/adaptiveButStill 97% accurate on “yes” for telephoneUnexpected rate of speech causes doublingor tripling of error rateUnexpected accent hurts badly Accuracy on unrestricted speech at 60%Dont know when we knowFew advances in basic understandingWhy is ASR Har

15、d?Natural speech is continuousNatural speech has disfluenciesNatural speech is variable over:global rate, local rate, pronunciationwithin speaker, pronunciation acrossspeakers, phonemes in differentcontexts Why is ASR Hard?(continued)Large vocabularies are confusableOut of vocabulary words inevitabl

16、eRecorded speech is variable over:room acoustics, channel characteristics,background noiseLarge training times are not practicalUser expectations are for equal to orgreater than “human performance”Main Causes of Speech VariabilityEnvironmentSpeakerInputEquipment Speech - correlated noisereverberatio

17、n, reflectionUncorrelated noiseadditive noise(stationary, nonstationary) Attributes of speakersdialect, gender, age Manner of speakingbreath & lip noisestressLombard effectratelevelpitchcooperativenessMicrophone (Transmitter)Distance from microphoneFilterTransmission systemdistortion, noise, echoRecording equipmentASR DimensionsSpeaker dependent, independentIsolated, continuous, keywor

温馨提示

  • 1. 本站所有资源如无特殊说明,都需要本地电脑安装OFFICE2007和PDF阅读器。图纸软件为CAD,CAXA,PROE,UG,SolidWorks等.压缩文件请下载最新的WinRAR软件解压。
  • 2. 本站的文档不包含任何第三方提供的附件图纸等,如果需要附件,请联系上传者。文件的所有权益归上传用户所有。
  • 3. 本站RAR压缩包中若带图纸,网页内容里面会有图纸预览,若没有图纸预览就没有图纸。
  • 4. 未经权益所有人同意不得将文件中的内容挪作商业或盈利用途。
  • 5. 人人文库网仅提供信息存储空间,仅对用户上传内容的表现方式做保护处理,对用户上传分享的文档内容本身不做任何修改或编辑,并不能对任何下载内容负责。
  • 6. 下载文件中如有侵权或不适当内容,请与我们联系,我们立即纠正。
  • 7. 本站不保证下载资源的准确性、安全性和完整性, 同时也不承担用户因使用这些下载资源对自己和他人造成任何形式的伤害或损失。

评论

0/150

提交评论