Global ETD Search

271	A Design of Mandarin Speech Recognition System for Addresses in Taiwan Cheng, Chi-Feng 31 August 2005 (has links) A Mandarin speech recognition system for addresses in Taiwan, based on end-point detection, MFCC and HMM, is proposed and implemented in this thesis. It includes both phrase and monosyllable recognition tasks. For the phrase recognition part, we select the initial candidates before the final recognition stage to tremendously reduce the computational time. On the other side, for the monosyllable recognition part, we further refine the recognition details to improve the correct rate under easily confused circumstances. The final system can achieve 85% correct identification rate, and the address recognition can be completed within 2 seconds in the laboratory environment for speaker-dependent case. Hidden Markov model(HMM) Mel-frequency cepstrum(MFCC) End-point detection Mel-frequency cepstrum
272	A System Design of Chinese Resume by Speech Construction Chen, Yue-sheng 28 August 2006 (has links) A system of Chinese resume by speech construction is developed by the use of a novel segmentation mechanism and the classical Hidden Markov Model. The recognition system is based on both mono-syllable HMM's and speech-text alignment schemes. Experimental results indicate that the amount of training materials used for feature extraction can be greatly reduced, and the text content of the recorded speech training data can be different from those of the recognition tasks as well. Each phrase in the resume can be identified within one second, that is approximately the same as the graduate did last year. Furthermore, the user interface of the resume system has been redesigned and polished by the GTK toolkit in order to enable event-driven X-window operations. Speech-text alignment Hidden Markov model(HMM)
273	A Design of Speech Recognition System for Chinese Names of Historical Figures Around the World Lin, Wei-Ci 07 September 2006 (has links) A design of speech recognition system for Chinese names of historical figures around the world is proposed in this thesis. A speech database of approximately forty-six thousand Chinese names is collected and recorded twice for system evaluation. This system applies Mel-frequency cepstrum coefficients, monosyllable HMM¡¦s and speech-text alignment scheme to accomplish initial candidate selection. A Mandarin pitch identification mechanism is then followed to increase the correct rate and obtain the final answer. The experimental results indicate that a 90% correct identification rate can be achieved, under the condition that the first session recording material is used for training and the second one for testing. For the speaker dependent case, the correct name can be recognized within 1.5 seconds, using a PC with an Intel Celeron 2.4 GHz CPU and RedHat Linux 9.0 Operation System. Hidden Markov model(HMM) Endpoint detection
274	The Continuous Speech Recognition System Base on Hidden Markov Models with One-Stage Dynamic Programming Algorithm. Hsieh, Fang-Yi 03 July 2003 (has links) Based on Hidden Markov Models (HMM) with One-Stage Dynamic Programming Algorithm, a continuous-speech and speaker-independent Mandarin digit speech recognition system was designed in this work. In order to implement this architecture to fit the performance of hardware, various parameters of speech characteristics were defined to optimize the process. Finally, the ¡§State Duration¡¨ and the ¡§Tone Transition Property Parameter¡¨ were extracted from speech temporal information to improve the recognition rate. Via using the test database, experimental results show that this new ideal of one-stage dynamic programming algorithm , with ¡§state duration¡¨ and ¡§ tone transition property parameter¡¨ , will have 18% recognition rate increase when compare to the conventional one. For speaker-independent and connect-word recognition, this system will achieve recognition rate to 74%. For speaker-independent but isolate-word recognition, it will have recognition rate higher than 96%. Recognition rate of 92% is obtained as this system is applied to the connect-word speaker-dependent recognition. Hidden Markov Models Continuous Speech Recognition One-Stage Dynamic Programming Algorithm
275	A Design and Applications of Mandarin Keyword Spotting System Hou, Cheng-Kuan 11 August 2003 (has links) A Mandarin keyword spotting system based on MFCC, discrete-time HMM and Viterbi algorithm with DTW is proposed in this thesis. Joining with a dialogue system, this keyword spotting platform is further refined to a prototype of natural speech patient registration system of Kaohsiung Veterans General Hospital. After the ID number is asked by the computer-dialogue attendant in the registration process, the user can finish all relevant works in one sentence. Functions of searching clinical doctors, making and canceling registration are all built in this system. In a laboratory environment, the correct rate of this speaker-independent patient registration system can reach 97% and all registration process can be completed within 75 seconds. Mel-frequency cepstrum coefficients phrase recognition Dynamic Time Warping Keyword spotting Hidden Markov model
276	A Design of Speech Recognition System under Noisy Environment Cheng, Po-Wen 11 August 2003 (has links) The objective of this thesis is to build a phrase recognition system under noisy environment that can be used in real-life. In this system, the noisy speech is first filtered by the enhanced spectral subtraction method to reduce the noise level. Then the MFCC with cepstral mean subtraction is applied to extract the speech features. Finally, hidden Markov model (HMM) is used in the last stage to build the probabilistic model for each phrase. A Mandarin microphone database of 514 company names that are in Taiwan¡¦s stock market is collected. A speaker independent noisy phrase recognition system is then implemented. This system has been tested under various noise environments and different noise strengths. cepstral mean subtraction spectral subtraction speaker-independent phrase recognition hidden Markov model
277	A Design of Japanese Speech Recognition System Chen, Meng-yang 24 August 2009 (has links) This thesis investigates the design and implementation strategies for a Japanese speech recognition system. It utilizes the speech features of the 188 common Japanese mono-syllables as the major training and recognition methodology. A training database of 10 utterances per mono-syllable is established by applying Japanese pronunciation rules. These 10 utterances are collected through reading 5 rounds of 188 mono-syllables, where every mono-syllable is consecutively read twice in each round. Mel-frequency cepstrum coefficients, linear predicted cepstrum coefficients, and hidden Markov model are used as the two feature models and the recognition model respectively. Under the Pentium 2.4 GHz personal computer and Ubuntu 8.04 operating system environment, a correct phrase recognition rate of 87% can be reached for a 34,000 Japanese phrase database. The average computation time for each phrase is about 1.5 seconds. Linear predicted cepstrum coefficients Hidden Markov model Speech recognition Mel-frequency cepstrum coefficients
278	A Design of Recognition Rate Improving Strategy for Mandarin Speech Recognition System - A Case Study on Address Inputting System and Phrase Recognition System Hsieh, Wen-kuang 24 August 2009 (has links) This thesis investigates the recognition rate improvement strategies for a Mandarin speech recognition system. Both automatic tone recognition and consonant correction schemes are studied and applied to the Mandarin address inputting system and the Mandarin 2, 3, 4-word phrase recognition systems. For automatic tone recognition scheme, the acoustic properties of the four tones in the Mandarin training database are estimated statistically by 4 sets of parameters within 6 minutes. These automatically generated parameters can greatly increase the tone recognition accuracy, and at the same time reduce the amount of time spent in the manual tone parameter adjustment, that is about 8 hours in general. For consonant correction scheme, the sub-syllable models are developed to enhance the consonant recognition accuracy, and hence further improve the overall correct rate for the whole Mandarin phrases. Experimental results indicate that over 90% correct rate can be achieved for the Mandarin address inputting system with 180 thousand place names by applying the above two schemes. Furthermore, the recognition rates for the Mandarin 2, 3, 4-word phrase recognition systems with 116 thousand phrases in total can be improved from 77%, 94% and 97.5%, to 85%, 96% and 98% respectively. speech recognition hidden Markov model sub-syllable model automatic tone recognition LPCC MFCC
279	A Design of Taiwanese Speech Recognition System Jhu, Hao-fu 24 August 2009 (has links) This thesis investigates the design and implementation strategies for a Taiwanese speech recognition system. It adopts a 4 plus 1¡]five times¡^recording strategy, where the 1st four recordings are used for speech feature training and the last recording for speech recognition simulation. Mel-frequency cepstrum coefficients and hidden Markov model are used as the feature model and the recognition model respectively. Under the Intel Celeron 2.4 GHz personal computer and Red Hat Linux 9.0 operating system environment, a correct phrase recognition rate of 90% can be reached for a 4200 Taiwanese phrase database. Speech recognition Mel-frequency cepstrum coefficients Gaussian distribution Hidden Markov model
280	A Design of English Speech Recognition System Chen, Yung-ming 24 August 2009 (has links) This thesis investigates the design and implementation strategies for a English speech recognition system. Two speech inputting methods, the spelling inputting and the reading inputting, are implemented for English word recognition and query. Mel-frequency cepstrum coefficients, linear predicted cepstrum coefficients, and hidden Markov model are used as the two feature models and the recognition model respectively. Under the Pentium 1.6 GHz personal computer and Ubuntu 8.04 operating system environment, a 95% correct recognition rate can be obtained for a 110 thousand English word database by the spelling inputting method; and a 93% correct recognition rate can be achieved for a 1,500 English word database by the reading inputting method. The average computation time for each word using either inputting method is about 1.5 seconds. Linear predicted cepstrum coefficients Hidden Markov model Mel frequency cepstrum coefficients Speech recognition

Search results