Search PubMed⌕ Search

Biomedical subjects

D G Childers

Publications and source records attributed to D G Childers.

At least 19 recordsLinked to original sources

Labelling and discrimination of a synthetic fricative continuum in noise: a study of absolute duration and relative onset time cues.

Categorical perception was evaluated for a nine-token voice onset time (VOT) continuum with endpoint tokens /feil/-/veil/. The synthetic speech continuum was presented in a random-level noise masker at different signal-to-noise ratios (SNR = 0, + 6, +12 dB) and overall presentation levels (50 and 70 dB HL). Overall labelling performance deteriorated as the SNR was reduced. Labelling results for the +12-dB-SNR condition reflected a category boundary at 87 ms for listeners with normal hearing sensitivity. The companion two-step discrimination function revealed better-than-chance performance between pairs of tokens labelled fail, chance performance between pairs of tokens labelled vail, and a slight performance peak at the labelling boundary between fail and vail. Listeners with high-frequency audiometric deficits produced labelling results for the +12-dB-SNR condition that were similar to normal functions measured for the 0-dB-SNR condition. These listeners were unable to discriminate two-step differences in voicing duration, but they produced a normal temporal labelling boundary. To try to understand the noncategorical discrimination data, a psychoacoustic analog for the speech continuum was evaluated. Relative onset time (ROT) difference limens (DLs) were measured as a function of the temporal onset delay of a low-frequency sawtooth waveform relative to the onset of a high-frequency noise burst. The ROT cue was used only when absolute stimulus duration could not be relied upon as a consistent cue, under conditions where a large range of random overall duration was presented to the listener. The ROT DLs were relatively invariant over a range of standard delays from 50 to 110 ms. The average DL was about 30 ms, which is consistent with the small performance peak in the synthetic speech discrimination function.

Humans↗

Modeling the glottal volume-velocity waveform for three voice types.

The purpose of this study was to model features of the glottal volume-velocity waveform for three voice types: modal voice, vocal fry, and breathy voice. The study analyzed data measured from two sustained vowels and one sentence uttered by nine adult, male subjects who represented examples of the three voice types. The primary analysis procedure was glottal inverse filtering, which estimated the glottal volume-velocity waveform. The estimated glottal volume-velocity waveform was then fit to an LF model waveform. Four parameters of the LF model were adjusted to minimize the mean-squared error between the estimated glottal waveform and the LF model waveform. Statistical averages and standard deviations of the four parameters of the LF glottal waveform model were calculated using the data for each voice type. The four LF model parameters characterize important low-frequency features of the glottal waveform, namely, the glottal pulse width, pulse skewness, abruptness of closure of the glottal pulse, and the spectral tilt of the glottal pulse. Statistical analysis included ANOVA and multiple linear regression analysis. The ANOVA results demonstrated that there was a difference in three of the four LF model parameters for the three voice types. The linear regression analysis between the four LF model parameters and a formal rating by a listening test of the quality of the three voice types was used to determine the most significant LF model parameters for each voice type. A simple rule was devised for synthesizing the three voice types with a formant synthesizer using the LF glottal waveform model. Listener evaluations of the synthesized speech tended to confirm the results determined by the analysis procedures.

Algorithms↗

Measuring and modeling vocal source-tract interaction.

The quality of synthetic speech is affected by two factors: intelligibility and naturalness. At present, synthesized speech may be highly intelligible, but often sounds unnatural. Speech intelligibility depends on the synthesizer's ability to reproduce the formants, the formant bandwidths, and formant transitions, whereas speech naturalness is thought to depend on the excitation waveform characteristics for voiced and unvoiced sounds. Voiced sounds may be generated by a quasiperiodic train of glottal pulses of specified shape exciting the vocal tract filter. It is generally assumed that the glottal source and the vocal tract filter are linearly separable and do not interact. However, this assumption is often not valid, since it has been observed that appreciable source-tract interaction can occur in natural speech. Previous experiments in speech synthesis have demonstrated that the naturalness of synthetic speech does improve when source-tract interaction is simulated in the synthesis process. The purpose of this paper is two-fold: 1) to present an algorithm for automatically measuring source-tract interaction for voiced speech, and 2) to present a simple speech production model that incorporates source-tract interaction into the glottal source model. This glottal source model controls: 1) the skewness of the glottal pulse, and 2) the amount of the first formant ripple superimposed on the glottal pulse. A major application of the results of this paper is the modeling of vocal disorders.

Algorithms↗

Speech synthesis by glottal excited linear prediction.

This paper describes a linear predictive (LP) speech synthesis procedure that resynthesizes speech using a 6th-order polynomial waveform to model the glottal excitation. The coefficients of the polynomial model form a vector that represents the glottal excitation waveform for one pitch period. A glottal excitation code book with 32 entries for voiced excitation is designed and trained using two sentences spoken by different speakers. The purpose for using this approach is to demonstrate that quantization of the glottal excitation waveform does not significantly degrade the quality of speech synthesized with a glottal excitation linear predictive (GELP) synthesizer. This implementation of the LP synthesizer is patterned after both a pitch-excited LP speech synthesizer and a code excited linear predictive (CELP) speech coder. In addition to the glottal excitation codebook, we use a stochastic codebook with 256 entries for unvoiced noise excitation. Analysis techniques are described for constructing both codebooks. The GELP synthesizer, which resynthesizes speech with high quality, provides the speech scientist a simple speech synthesis procedure that uses established analysis techniques, that is able to reproduce all speed sounds, and yet also has an excitation model waveform that is related to the derivative of the glottal flow and the integral of the residue. It is conjectured that the glottal excitation codebook approach could provide a mechanism for quantitatively comparing the differences in glottal excitation codebooks for male and female speakers and for speakers with vocal disorders and for speakers with different voice types such as breathy and vocal fry voices. Conceivably, one could also convert the voice of a speaker with one voice type, e.g., breathy, to the voice of a speaker with another voice type, e.g., vocal fry, by synthesizing speech using the vocal tract LP parameters for the speaker with the breathy voice excited by the glottal excitation codebook trained for vocal fry.

Communication Devices for People with Disabilities↗

Detection of laryngeal function using speech and electroglottographic data.

The purpose of this research was to develop quantitative measures for the assessment of laryngeal function using speech and electroglottographic (EGG) data. We developed two procedures for the detection of laryngeal pathology: 1) a spectral distortion measure using pitch synchronous and asynchronous methods with linear predictive coding (LPC) vectors and vector quantization (VQ) and 2) analysis of the EGG signal using time interval and amplitude difference measures. The VQ procedure was conjectured to offer the possibility of circumventing the need to estimate the glottal volume velocity wave-form by inverse filtering techniques. The EGG procedure was to evaluate data that was "nearly" a direct measure of vocal fold vibratory motion and thus was conjectured to offer the potential for providing an excellent assessment of laryngeal function. A threshold based procedure gave 75.9 and 69.0% probability of pathological detection using procedures 1) and 2), respectively, for 29 patients with pathological voices and 52 normal subjects. The false alarm probability was 9.6% for the normal subjects.

Adult↗

Gender recognition from speech. Part I: Coarse analysis.

The purpose of this research was to investigate the potential effectiveness of digital speech processing and pattern recognition techniques for the automatic recognition of gender from speech segments. In this paper "coarse" acoustic coefficients (autocorrelation, linear prediction, cepstrum, and reflection) were used to form test and reference templates for vowels, voiced fricatives, and unvoiced fricatives. The effects of different distance measures, filter orders, recognition schemes, and vowels and fricatives were comparatively assessed to determine their effectiveness for the task of gender recognition from speech segments. The results showed that most of the acoustic parameters worked well for gender recognition. A within-gender and within-subject averaging technique was important for generating appropriate test and reference templates. The Euclidean distance measure appeared to be the most robust as well as the simplest of the distance measures. The results from this study implied that the gender information is time invariant, phoneme independent, and speaker independent for a given gender. One recognition scheme achieved 100% correct speaker gender classification for a database of 52 talkers (27 male and 25 female). In part II of this paper [D.G. Childers and K. Wu, J. Acoust. Soc. Am. 90, 1841-1856 (1991); hereafter referred to as paper II] the detailed features of ten vowels that appeared responsible for distinguishing a speaker's gender were examined statistically. Included in paper II is a replication of part of the classical study of Peterson and Barney [J. Acoust. Soc. Am. 24, 175-184 (1952)] of vowel characteristics.

Adult↗

Gender recognition from speech. Part II: Fine analysis.

The purpose of this research was to investigate the potential effectiveness of digital speech processing and pattern recognition techniques for the automatic recognition of gender from speech. In part I Coarse Analysis [K. Wu and D. G. Childers, J. Acoust. Soc. Am. 90, 1828-1840 (1991)] various feature vectors and distance measures were examined to determine their appropriateness for recognizing a speaker's gender from vowels, unvoiced fricatives, and voiced fricatives. One recognition scheme based on feature vectors extracted from vowels achieved 100% correct recognition of the speaker's gender using a database of 52 speakers (27 male and 25 female). In this paper a detailed, fine analysis of the characteristics of vowels is performed, including formant frequencies, bandwidths, and amplitudes, as well as speaker fundamental frequency of voicing. The fine analysis used a pitch synchronous closed-phase analysis technique. Detailed formant features, including frequencies, bandwidths, and amplitudes, were extracted by a closed-phase weighted recursive least-squares method that employed a variable forgetting factor, i.e., WRLS-VFF. The electroglottograph signal was used to locate the closed-phase portion of the speech signal. A two-way statistical analysis of variance (ANOVA) was performed to test the differences between gender features. The relative importance of grouped vowel features was evaluated by a pattern recognition approach. Numerous interesting results were obtained, including the fact that the second formant frequency was a slightly better recognizer of gender than fundamental frequency, giving 98.1% versus 96.2% correct recognition, respectively. The statistical tests indicated that the spectra for female speakers had a steeper slope (or tilt) than that for males. The results suggest that redundant gender information was imbedded in the fundamental frequency and vocal tract resonance characteristics. The feature vectors for female voices were observed to have higher within-group variations than those for male voices. The data in this study were also used to replicate portions of the Peterson and Barney [J. Acoust. Soc. Am. 24, 175-184 (1952)] study of vowels for male and female speakers.

Computer Graphics↗

Vocal quality factors: analysis, synthesis, and perception.

The purpose of this study was to examine several factors of vocal quality that might be affected by changes in vocal fold vibratory patterns. Four voice types were examined: modal, vocal fry, falsetto, and breathy. Three categories of analysis techniques were developed to extract source-related features from speech and electroglottographic (EGG) signals. Four factors were found to be important for characterizing the glottal excitations for the four voice types: the glottal pulse width, the glottal pulse skewness, the abruptness of glottal closure, and the turbulent noise component. The significance of these factors for voice synthesis was studied and a new voice source model that accounted for certain physiological aspects of vocal fold motion was developed and tested using speech synthesis. Perceptual listening tests were conducted to evaluate the auditory effects of the source model parameters upon synthesized speech. The effects of the spectral slope of the source excitation, the shape of the glottal excitation pulse, and the characteristics of the turbulent noise source were considered. Applications for these research results include synthesis of natural sounding speech, synthesis and modeling of vocal disorders, and the development of speaker independent (or adaptive) speech recognition systems.

Adult↗

Electroglottography and vocal fold physiology.

The electroglottogram (EGG) is known to be related to vocal fold motion. A major hypothesis undergoing examination in several research centers is that the EGG is related to the area of contact of the vocal folds. This hypothesis is difficult to substantiate with direct measurements using human subjects. However, other supporting evidence can be offered. For this study we made measurements from synchronized ultra high-speed laryngeal films and from EGG waveforms collected from subjects with normal larynges and patients with vocal disorders. We compare certain features of the EGG waveform to (a) the instant of the opening of the glottis, (b) the instant of the closing of the glottis, and (c) the instant of the maximum opening of the glottis. In addition, we compare both the open quotient and the relative average perturbation measured from the glottal area to that estimated from the EGG. All of these comparisons indicate that vocal fold vibratory characteristics are reflected by features of the EGG waveform. This makes the EGG useful for speech analysis and synthesis as well as for modeling laryngeal behavior. The limitations of the EGG are discussed.

Electrodiagnosis↗

Acoustic correlates of vocal quality.

We have investigated the relationship between various voice qualities and several acoustic measures made from the vowel /i/ phonated by subjects with normal voices and patients with vocal disorders. Among the patients (pathological voices), five qualities were investigated: overall severity, hoarseness, breathiness, roughness, and vocal fry. Six acoustic measures were examined. With one exception, all measures were extracted from the residue signal obtained by inverse filtering the speech signal using the linear predictive coding (LPC) technique. A formal listening test was implemented to rate each pathological voice for each vocal quality. A formal listening test also rated overall excellence of the normal voices. A scale of 1-7 was used. Multiple linear regression analysis between the results of the listening test and the various acoustic measures was used with the prediction sums of squares (PRESS) as the selection criteria. Useful prediction equations of order two or less were obtained relating certain acoustic measures and the ratings of pathological voices for each of the five qualities. The two most useful parameters for predicting vocal quality were the Pitch Amplitude (PA) and the Harmonics-to-Noise Ratio (HNR). No acoustic measure could rank the normal voices.

Adult↗

Cochannel speech separation.

The multisignal minimum-cross-entropy spectral analysis (multisignal MCESA) is applied to the problem of separating the speech signals of two talkers speaking simultaneously on a single channel, e.g., when two talkers use a single microphone. A new two-stage approach to the problem is proposed in which a spectral separator is followed by a spectral tailoring procedure. The spectral separator produces an initial estimate of the speech spectrum for each talker. Then the spectral tailoring procedure employs the multisignal MCESA technique to adjust the initial spectral estimates to account for the characteristics of the known cochannel composite speech signal. The research emphasis is placed on the implementation and evaluation of the spectral tailoring procedure, i.e., the use of the multisignal MCESA in the proposed scheme. Its usefulness is evaluated and validated by listening tests and by comparing the spectral distortions of the estimated voices before and after the multisignal MCESA processing.

Algorithms↗

Brain potentials related to seeing one's own name.

Subjects were assigned an assumed name and then shown a series of statements of the form, "My name / is / X", where X was the assumed name, their own first name, or one of a set of other false names. Their task was to respond positively to the "assumed" name and reject as false all other names, including their own. An N380 feature of the averaged task-related brain potentials, considered to be inversely related to the degree of contextual priming, was greatly enhanced for the false names compared to the assumed name. The N380 to one's own name was more similar to that of the false than the assumed name, indicating that the sentence context's priming of various names was under the subjects' attentional control, and that the late negativity could be modulated by this attention. In contrast, a large P510 feature distinguished one's own name from the false name, and this difference was unaffected by practice. Even in cases, then, where the context allows anticipation of one verbal event (here, the assumed name), a highly overlearned and salient stimulus such as one's own name continues to produce a distinctive neural response.

Adolescent↗

Event-related potentials: a critical review of methods for single-trial detection.

The analysis of ERP data has followed several lines over the last 20 years. The most prevalent method is simply to average ERPs for a given class of stimuli. The ERPs are compared for differences across classes of stimuli. Little other special data processing is used. The ERP comparisons are usually performed using visual examination of the wave-shapes. Sometimes statistics are calculated such as means, variances, and confidence limits. Linear filtering is used to reduce interference. Another approach is to model or analyze the ERP as a sequence of vectors or frames of data samples. These samples may be of the ERP time waveform or they may be of the frequency transform of the ERP waveform. The frames of data vary in length from the entire ERP waveform (500 to 1000 msec) to frames as short as ten sample points (100 msec). Recognition of an event in the ERP is achieved by computing a distance measure between parameter vectors for one class of stimuli and corresponding parameter vectors for another class of stimuli. Recognition is achieved by selecting the ERP with the lowest distance score. This approach is "pattern matching" and relies on two assumptions: adjacent frames of data are uncorrelated, and the variability of the data can be accounted for by the distance measured for all stimuli in the classes presented. Subject variability is generally not accounted for, other than to assume it is the same for all classes of stimuli. The data are clustered into a variety of reference patterns that represent particular manifestations of a particular stimulus. Another approach is "feature-based" recognition. The idea is to identify and automatically extract features of the data that can provide a characterization of stimuli. The features selected may be abstract. They are calculated from the data or transforms of the data.

Biometry↗

A model for vocal fold vibratory motion, contact area, and the electroglottogram.

The electroglottogram (EGG) has been conjectured to be related to the area of contact between the vocal folds. This hypothesis has been substantiated only partially via direct and indirect observations. In this paper, a simple model of vocal fold vibratory motion is used to estimate the vocal fold contact area as a function of time. This model employs a limited number of vocal fold vibratory features extracted from ultra high-speed laryngeal films. These characteristics include the opening and closing vocal fold angles and the lag (phase difference) between the upper and lower vocal fold margins. The electroglottogram is simulated using the contact area, and the EGG waveforms are compared to measured EGGs for normal male voices producing both modal and pulse register tones. The model also predicts EGG waveforms for vocal fold vibration associated with a nodule or polyp.

Glottis↗

Brain potentials during sentence verification: automatic aspects of comprehension.

College students learned a set of facts relating fictitious people and their occupations (e.g. 'Matthew is a lawyer'). Event-related brain potentials (ERPs) were recorded while they subsequently viewed a series of such statements presented in segments (e.g. 'Matthew/is a/dentist'). ERPs to occupations completing statements falsely were significantly more negative than those to true statements in an interval 200-420 msec poststimulus (peak N320), whether subjects were required to make a decision about each statement or passively view the presented segments (Experiments 1 and 2). A later ERP positivity was observed during 'response' trials that was of longer latency for false than true completions; but this positive component was greatly attenuated during 'no-response' trials. The enhanced N320 for false completions was not affected by requiring subjects on some trials to respond incorrectly (Experiment 3). It is concluded that attending to a presented word results in an automatic analysis of its meaning in the context of a preceding verbal input, and that ERPs can indicate the nature of the output of that analysis.

Brain↗

A critical review of electroglottography.

The technique of electroglottography is reviewed from the perspective of a laboratory instrument for assessing laryngeal function, a device to assist speech and speaker recognition, and as a potential diagnostic aid in the clinic. A description of the electronic functioning of the electroglottograph (EGG) is provided. Considerable emphasis is given to contemporary research which has focused on laryngeal assessment using the EGG. Methods for validating and aiding the interpretation or reading of the EGG are discussed, including photoglottography, stroboscopy, ultrahigh-speed laryngeal cinematography, and others. The relationship of the EGG to glottal area and glottal volume velocity estimated by inverse filtering is presented. An elementary model of the EGG is described and used to predict characteristic features of the EGG waveform. Clinical data as well as data obtained from subjects with a normal functioning larynx are analyzed. Applications of the EGG to speech processing are outlined, including real-time detection of voicing, voiced and unvoiced speech segments, and silence intervals. The EGG device has potential for assisting speech and speaker recognition systems in certain applications.

Aged↗

Brain potentials during sentence verification: late negativity and long-term memory strength.

Subjects decided whether self-referential statements were true or false. Event-related potentials (ERPs) associated with final words creating false statements displayed a late negativity (N340) relative to ERPs for true completions. The size of this difference between true and false statements was greater for highly familiar statements (e.g. "My name is Ira") than for less familiar ones (e.g. "I go to bed late") even after all the statements had been practised a number of times. The late negativity appears to be associated with a discrepancy between presented and remembered information, and its magnitude reflects the long-term familiarity or strength of the remembered information.

Adult↗