Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sound Spectrography”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 919 records · Page 51Linked to original sources

Production of intonation and contrastive stress in electrolaryngeal speech.

Acoustical investigations of intonation and contrastive stress patterns in speech produced with electronic artificial larynges were completed. High-quality tape recordings of sentences spoken by 4 normal speakers, 3 users of the Western Electric 5A electrolarynx, and 2 users of the Servox electrolarynx were subjected to acoustic analysis. All electrolarynx users distinguished stressed from unstressed syllables by varying the duration of syllables and contiguous pauses. One Western Electric 5A speaker also controlled fundamental frequency. This speaker distinguished statements from questions by varying the rate and extent of the initial rising portion of fundamental frequency contours. Findings are interpreted in relation to their implications for clinical intervention and in terms of suggestions for improved design of artificial larynges.

Humans↗

Speaking clearly for the hard of hearing I: Intelligibility differences between clear and conversational speech.

This paper is concerned with variations in the intelligibility of speech produced for hearing-impaired listeners under two conditions. Estimates were made of the magnitude of the intelligibility differences between attempts to speak clearly and attempts to speak conversationally. Five listeners with sensorineural hearing losses were tested on groups of nonsense sentences spoken clearly and conversationally by three male talkers as a function of level and frequency-gain characteristic. The average intelligibility difference between clear and conversational speech averaged across talker was found to be 17 percentage points. To a first approximation, this difference was independent of the listener, level, and frequency-gain characteristic. Analysis of segmental-level errors was only possible for two listeners and indicated that improvements in intelligibility occurred across all phoneme classes.

Adult↗

Quantitative spectral evaluation of shimmer and jitter.

A vowel [a]-like, synthesized speech wave was perturbated by defined and comparable jitter and shimmer levels. The signal-to-noise ratio was calculated from the speech wave spectra. Noise emerges in those spectral regions in which the harmonics have high amplitudes, that is, at low frequencies and in the formant regions. Jitter created noise levels significantly higher than shimmer. To verify the theoretical findings, the voices of 32 women with functional voice disorders were analyzed for shimmer and jitter. It was found that only jitter is relevant for differentiating between hypo- and hyperfunctional voice disorders. Jitter was reduced in hyperfunctional voice disorder. This is presumed to be an effect of the high vocal fold tension found in the disorder.

Adult↗

Correspondence between an accelerometric nasal/voice amplitude ratio and listeners' direct magnitude estimations of hypernasality.

Miniature accelerometers were used to transduce nasal and anterior-neck tissue vibrations of 12 hypernasal and 3 normal children. The accelerometric voltages provided an analog implementation of Horii's (1980) nasal/voice ratio. Simultaneous audio recordings were later evaluated for hypernasality by listeners. Listeners' direct magnitude estimations (DME) of hypernasality were highly correlated with the accelerometric nasal/voice ratio when the stimulus sentences contained obstruents, nonnasal semivowels, and vowels. No correlation existed between DME and accelerometric values when the stimulus sentences contained primarily nasal semivowels and vowels.

Child↗

A comparison of temporal measures of speech using spectrograms and digital oscillograms.

To determine whether any systematic differences occur as a result of using spectrograms versus digital oscillograms to make durational measurements, a number of temporal features (e.g., voice onset time, vowel duration, and consonant closure duration) for 3 speakers were independently measured by 2 different investigators. Both experimenters measured the same intervals with conventional spectrograms and with digital oscillograms, separated by at least a 2-week interval. Oscillograms tended to reveal slightly longer vowel durations and more voicing during consonant closure, while spectrograms evidenced slightly longer consonant closure durations. In general, variations between the two types of instrumentation were no more than 8 to 10 ms and are, therefore, of primary consequence only for studies in which quite small temporal differences are critical.

Adult↗

Speaking clearly for the hard of hearing. II: Acoustic characteristics of clear and conversational speech.

The first paper of this series (Picheny, Durlach, & Braida, 1985) presented evidence that there are substantial intelligibility differences for hearing-impaired listeners between nonsense sentences spoken in a conversational manner and spoken with the effort to produce clear speech. In this paper, we report the results of acoustic analyses performed on the conversational and clear speech. Among these results are the following. First, speaking rate decreases substantially in clear speech. This decrease is achieved both by inserting pauses between words and by lengthening the durations of individual speech sounds. Second, there are differences between the two speaking modes in the numbers and types of phonological phenomena observed. In conversational speech, vowels are modified or reduced, and word-final stop bursts are often not released. In clear speech, vowels are modified to a lesser extent, and stop bursts, as well as essentially all word-final consonants, are released. Third, the RMS intensities for obstruent sounds, particularly stop consonants, is greater in clear speech than in conversational speech. Finally, changes in the long-term spectrum are small. Thus, speaking clearly cannot be regarded as equivalent to the application of high-frequency emphasis.

Hearing Loss, Sensorineural↗

Dynamic aspects of phonatory control in spasmodic dysphonia.

To understand the voluntary laryngeal movement disorder in spasmodic dysphonia (SD), SD patients were compared with normal controls on speech tasks with different laryngeal motor-control demands. Nine patients with idiopathic chronic SD and no other speech, otolaryngologic, neurologic, or psychiatric disorders were compared with 15 control subjects who were free of such disorders. Speech production tasks required different degrees of dynamic and precise control of vocal fold movement. Phonatory off times were increased in the SD patients, while maximum phonation time, phonatory on time, frequency and intensity control, and reaction times for CV syllables were not affected. On a reaction-time task, the onset of laryngeal movement was not delayed in the SD patients, however, the time between the onset of laryngeal movement and phonatory onset was significantly increased in the SD patients in comparison with the controls. Therefore, SD patients had no difficulty with the onset of laryngeal movement but were slow to achieve phonation, indicating a movement-control disorder affecting vocal fold adduction for phonation onset.

Adult↗

Prediction of vocal severity within and across voice types.

Fifty-one subjects representing diverse laryngeal etiologies recorded /a/ and /i/ to provide a study sample of 102 vowel sounds. Listeners categorized each vowel on the basis of four voice types (normal, breathy, hoarse, unclassified) and evaluated the degree of vocal abnormality on a 7-point scale. In addition to spectrographic noise (SN) classification, several acoustic measures based on period variability were entered into a multiple regression analysis for the prediction of vocal severity across and within voice types. In general, spectrographic noise and curvilinear derivatives of the period standard deviation (PSD) provided the best predictions of disorder severity. Different variables were the major predictors for different voice types. Several variables used in previous studies were inefficient as predictors of severity.

Adolescent↗

Training effects on vowel production by two profoundly hearing-impaired speakers.

Two profoundly hearing-impaired adolescents received systematic speech training to improve their production of the vowels /i/ and /ae/. Acoustic measures of F1, F2, and duration, and listener judgements of vowel acceptability, were used to quantify vowel production before and after training. Both subjects demonstrated significant changes in their production of the two vowels at the acoustic and perceptual levels following treatment. The changes were highly individualized. For some features, significant improvement occurred posttreatment with differences between the hearing-impaired subject and a control group of subjects with normal hearing no longer present. There was a significant improvement in the acceptability of the two vowels in each subject's speech after training. Vowel duration remained unchanged in the speech of one subject whereas it increased in the speech of the other subject following training. There was a trend toward reduced token-to-token variation in the posttreatment samples. Acoustic and perceptual measures also were obtained on two vowels not directly trained in the program. Significant changes occurred in the production of these segments but some of the changes resulted in greater deviation in the post- than in the pretreatment samples.

Adolescent↗

A methodological study of perturbation and additive noise in synthetically generated voice signals.

There is a relatively large body of research that is aimed at finding a set of acoustic measures of voice signals that can be used to: (a) aid in the detection, diagnosis, and evaluation of voice-quality disorders; (b) identify individual speakers by their voice characteristics; or (c) improve methods of voice synthesis. Three acoustic parameters that have received a relatively large share of attention, especially in the voice-disorders literature, are pitch perturbation, amplitude perturbation, and additive noise. The present study consisted of a series of simulations using a general-purpose formant synthesizer that were designed primarily to determine whether these three parameters could be measured independent of one another. Results suggested that changes in any single dimension can affect measured values of all three parameters. For example, adding noise to a voice signal resulted not only in a change in measured signal-to-noise ratio, but also in measured values of pitch and amplitude perturbation. These interactions were quite large in some cases, especially in view of the fact that the perturbation phenomena that are being measured are generally quite small. For the most part, the interactions appear to be readily explainable when the measurement techniques are viewed in relation to what is known about the acoustics of voice production.

Communication Devices for People with Disabilities↗

Temporal characteristics of the speech of normal elderly adults.

A number of physical and psychological changes occur as a result of the normal aging process. These changes often result in an increase in the time subjects require to perform various motor (and sensory) tasks. Although the effects of aging upon a variety of behaviors have been quite well documented, considerably less information is available concerning how normal aging may affect speech production. The present study examined temporal characteristics of the speech of 10 normal, elderly adults and 10 young adults who produced a variety of words and sentences at both normal and fast speaking rates. Acoustic analyses indicated that the elderly adults' segment, syllable, and sentence durations were 20 to 25% longer than those of the young adults at both the normal and the fast rates of speech. In addition to comparisons that were made between these two groups of subjects, comparisons were also made with durations of the speech of young children studied in previous research. It was observed that the elderly subjects tended to produce durations comparable to those of 6- and 7-year-old children.

Adult↗

Least mean square measures of voice perturbation.

A signal processing technique is described for measuring the jitter, shimmer, and signal-to-noise ratio of sustained vowels. The measures are derived from the least mean square fit of a waveform model to the digitized speech waveform. The speech waveform is digitized at an 8.3 kHz sampling rate, and an interpolation technique is used to improve the temporal resolution of the model fit. The ability of these procedures to measure low levels of perturbation is evaluated both on synthetic speech waveforms and on the speech recorded from subjects with normal voice characteristics.

Communication Devices for People with Disabilities↗

Changes in voice fundamental frequency following discharge of single motor units in cricothyroid and thyroarytenoid muscles.

This investigation was designed to measure voice Fo changes related to single motor unit (SMU) contractions in the cricothyroid and thyroarytenoid muscles. Four subjects (3 men and 1 woman) were recorded producing a prolonged vowel at modal pitch and loudness levels while simultaneous recordings of electromyograms (EMG) from the muscles were obtained. Voice Fo changes unrelated to SMU firings in the muscles were eliminated using an averaging method previously described by Baer (1979). Results indicate that the time between discharge of the SMU and the peak in Fo change ("Fo Latency") was variable and ranged from 5 to 20 ms for the thyroarytenoid and 6 to 75 ms for the cricothyroid muscle. Distinct oscillations in Fo were always present in recordings from the woman subject and from the men when they phonated at higher-than-modal pitch levels. The findings are discussed in relation to SMU contraction times, biomechanics of the vocal folds, and the presence of jitter in normal voices.

Adult↗

Composite speech spectrum for hearing and gain prescriptions.

Average long-term RMS 1/3-octave band speech spectra were generated for 30 male and 30 female talkers. The two spectra were significantly different in both low and high frequency bands but were similar in the mid-frequency region. It was concluded that a single spectrum could validly be used to represent both male and female speech in the frequency region important for hearing aid gain prescriptions: 250 Hz through 6300 Hz. In addition, the male and female spectra were compared with analogous spectra reported by Byrne (1977) and Pearsons, Bennett, and Fidell (1977). For each sex, significant differences were found among the three spectra in a few frequency bands. The best estimate of the average speech spectrum for each sex was obtained from a weighted average of the three sets of data, excluding the significantly different data points. The long-term RMS 1/3-octave band speech spectrum for male and female talkers combined was derived for use in hearing aid gain prescriptions.

Adult↗

Similarities between tactual and auditory speech perception.

Perception of synthetic speech continua through the sense of touch and audition was compared utilizing a 32-channel spectrally oriented electrocutaneous display and standard auditory psychophysical procedures. Results indicated a close correspondence between tactual and auditory discrimination and identification for a vowel (/a/-/e/) and a consonant (/sta/-/sa/) continuum. These results suggest that at least some aspects of speech perception are amodal.

Humans↗

Acoustic and perceptual analysis of word-initial stop consonants in phonologically disordered children.

Spectrographic measures of voice onset time (VOT) were made for phonologically disordered children in whom a voicing contrast was just beginning to emerge. These temporal measures were related to adult listeners' perception of voicing of the initial stop consonant to determine how well VOT could predict perceived voicing. In general, the predictive utility of VOT was not very high. The relation between VOT as produced by the phonologically disordered children and perceived voicing ranged from 0.31 to 0.43. A finer-grained analysis was conducted to determine what other acoustic cues might have influenced the listeners' judgments of voicing. Although no one acoustic cue could be found to explain all listeners' responses, spectral cues such as fundamental and F1 frequencies at the onset of voicing, as well as the burst and aspiration amplitude relative to the vowel onset amplitude accounted for the perceived voicing of about half of the tokens that were not differentiated by VOT. Rather than relying solely on the temporal characteristics of the VOT interval, a matrix of acoustic cues may influence how a listener perceives word-initial voicing as produced by phonologically disordered children.

Adult↗

Automatic phonetogram recording supplemented with acoustical voice-quality parameters.

A new method for automatic voice-quality registration is presented. The method is based on a technique called phonetography, which is the registration of the dynamic range of a voice as a function of fundamental frequency. In the new phonetogram-recording method fundamental frequency (Fo) and sound-pressure level (SPL) are automatically measured and represented in an XY-diagram. Three additional acoustical voice-quality parameters are measured simultaneously with Fo and SPL: (a) jitter in the Fo as a measure for roughness, (b) the SPL difference between the 0-1.5 kHz and the 1.5-5 kHz bands as a measure for sharpness, and (c) the vocal-noise level above 5 kHz as a measure for breathiness. With this method, the voice-quality parameter values, which may change substantially as a function of Fo and SPL, are pinned to a reference position in the patient's total vocal range. Seen as a reference tool, the phonetogram opens the possibility for a more meaningful comparison of voice-quality data. Some examples, demonstrating the dependence of the chosen quality parameters on Fo and SPL are given.

Humans↗

Technical considerations in computation of spectral harmonics-to-noise ratios for sustained vowels.

This paper explores technical issues affecting computed measures of the relative level of noise in the frequency spectrum of a vowel. This type of measure has been proposed for quantification of hoarseness in pathological speakers. An analysis of synthesized vowels was used to test the influence of vowel type, fundamental frequency, perturbation type, perturbation level and quantization. The algorithms were shown to be highly sensitive to errors in pitch-period demarcation, and a dependency on jitter perturbations, fundamental frequency, and vowel type was demonstrated. Relationships between algorithm performance and methods of spectrum estimation were discussed, and approaches for reducing the dependencies were proposed. Finally, a method for achieving a significant reduction in computation time was described.

Algorithms↗