Global ETD Search

Return to search

Speech Recognition Using a Synthesized Codebook

Speech sounds generated by a simple waveform synthesizer were used to create a vector quantization codebook for use in speech recognition. Recognition was tested over the TI-20 isolated word data base using a conventional DTW matching algorithm. Input speech was band limited to 300 - 3300 Hz, then passed through the Scott Instruments Corp. Coretechs process, implemented on a VET3 speech terminal, to create the speech representation for matching. Synthesized sounds were processed in software by a VET3 signal processing emulation program. Emulation and recognition were performed on a DEC VAX 11/750.
The experiments were organized in 2 series. A preliminary experiment, using no vector quantization, provided a baseline for comparison.
The original codebook contained 109 vectors, all derived from 2 formant synthesized sounds. This codebook was decimated through the course of the first series of experiments, based on the number of times each vector was used in quantizing the training data for the previous experiment, in order to determine the smallest subset of vectors suitable for coding the speech data base. The second series of experiments altered several test conditions in order to evaluate the applicability of the minimal synthesized codebook to conventional codebook training.
The baseline recognition rate was 97%. The recognition rate for synthesized codebooks was approximately 92% for sizes ranging from 109 to 16 vectors. Accuracy for smaller codebooks was slightly less than 90%. Error analysis showed that the primary loss in dropping below 16 vectors was in coding of voiced sounds with high frequency second formants. The 16 vector synthesized codebook was chosen as the seed for the second series of experiments.
After one training iteration, and using a normalized distortion score, trained codebooks performed with an accuracy of 95.1%. When codebooks were trained and tested on different sets of speakers, accuracy was 94.9%, indicating that very little speaker dependence was introduced by the training.

Automatic speech recognition.

Speech processing systems.

Identifer	oai:union.ndltd.org:unt.edu/info:ark/67531/metadc332203
Date	08 1900
Creators	Smith, Lloyd A. (Lloyd Allen)
Contributors	Swigger, Kathleen M., Brazile, Robert Pershing, 1941-, Jacob, Roy Thomas, Mackey, H. J., Conrady, Denis A.
Publisher	University of North Texas
Source Sets	University of North Texas
Language	English
Detected Language	English
Type	Thesis or Dissertation
Format	x, 164 leaves: ill., Text
Rights	Public, Smith, Lloyd A. (Lloyd Allen), Copyright, Copyright is held by the author, unless otherwise noted. All rights reserved.

Page generated in 0.0028 seconds

Speech Recognition Using a Synthesized Codebook

Description

Links & Downloads

Tags

Additional Fields