Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sound Spectrography”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 973 records · Page 54Linked to original sources

Spectral distribution of prosodic information.

Prosodic speech cues for rhythm, stress, and intonation are related primarily to variations in intensity, duration, and fundamental frequency. Because these cues make use of temporal properties of the speech waveform they are likely to be represented broadly across the speech spectrum. In order to determine the relative importance of different frequency regions for the recognition of prosodic cues, identification of four prosodic features, syllable number, syllabic stress, sentence intonation, and phrase boundary location, was evaluated under six filter conditions spanning the range from 200-6100 Hz. Each filter condition had equal articulation index (AI) weights, AI = 0.01; p(C)isolated words approximately equal to 0.40. Results obtained with normally hearing subjects showed that there was an interaction between filter condition and the identification of specific prosodic features. For example, information from high-frequency regions of speech was particularly useful in the identification of syllable number and stress, whereas information from low-frequency regions was helpful in identifying intonation patterns. In spite of these spectral differences, overall listeners performed remarkably well in identifying prosodic patterns, although individual differences were apparent. For some subjects, equivalent levels of performance across the six filter conditions were achieved. These results are discussed in relation to auditory and auditory-visual speech recognition.

Adolescent↗

Discriminability and perceptual weighting of some acoustic cues to speech perception by 3-year-olds.

Studies of children's speech perception have shown that young children process speech signals differently than adults. Specifically, the relative contributions made by various acoustic parameters to some linguistic decisions seem to differ for children and adults. Such findings have led to the hypothesis that there is a developmental shift in the perceptual weighting of acoustic parameters that results from experience with a native language (i.e., the Developmental Weighting Shift). This developmental shift eventually leads the child to adopt the optimal perceptual weighting strategy for the native language being learned (i.e., one that allows the listener to make accurate decisions about the phonemic structure or his or her native language). Although this proposal has intuitive appeal, there is at least one serious challenge that can be leveled against it: Perhaps age-related differences in speech perception can appropriately be explained by age-related differences in basic auditory-processing abilities. That is, perhaps children are not as sensitive as adults to subtle differences in acoustic structure and so make linguistic decisions based on the acoustic information that is most perceptually salient. The present study tested this hypothesis for the acoustic cues relevant to fricative identity in fricative-vowel syllables. Results indicated that 3-year-olds were not as sensitive to changes in these acoustic cues as adults are, but that these age-related differences in auditory sensitivity could not entirely account for age-related differences in perceptual weighting strategies.

Adult↗

Activity of intrinsic laryngeal muscles in fluent and disfluent speech.

The goal of the present experiment was to determine if stuttering is associated with unusually high levels of activity in laryngeal muscles. Qualitative and quantitative analyses of thyroarytenoid and cricothyroid recordings from 4 stuttering and 3 nonstuttering adults revealed the following: Compared to periods of fluent speech, intervals of disfluent speech are not typically characterized by higher levels of activity in these muscles; and when EMG levels during conversational speech are compared to maximal activation levels for these muscles (e.g., those observed during singing and the Valsalva maneuver), normally fluent adults show robust and sometimes near maximal recruitment during conversational speech. The adults who stutter had a lower operating range for these muscles during conversational speech, and their disfluencies did not produce relatively high activation levels. In summary, the present data require us to reject the claim that adults with a history of chronic stuttering routinely produce excessive levels of intrinsic laryngeal muscle activity. These results suggest that the use of botulinum toxin injections into the vocal folds to treat stuttering should be questioned.

Adult↗

The specific relation between perception and production errors for place of articulation in developmental apraxia of speech.

Developmental apraxia of speech is a disorder of phonological and articulatory output processes. However, it has been suggested that perceptual deficits may contribute to the disorder. Identification and discrimination tasks offer a fine-grained assessment of central auditory and phonetic functions. Seventeen children with developmental apraxia (mean age 8:9, years:months) and 16 control children (mean age 8:0) were administered tests of identification and discrimination of resynthesized and synthesized monosyllabic words differing in place-of-articulation of the initial voiced stop consonants. The resynthetic and synthetic words differed in the intensity of the third formant, a variable potentially enlarging their clinical value. The results of the identification task showed equal slopes for both subject groups, which indicates no phonetic processing deficit in developmental apraxia of speech. The hypothesized effect of the manipulation of the intensity of the third formant of the stimuli was not substantiated. However, the children with apraxia demonstrated poorer discrimination than the control children, which suggests affected auditory processing. Furthermore, analyses of discrimination performance and articulation data per apraxic subject demonstrated a specific relation between the degree to which auditory processing is affected and the frequency of place-of-articulation substitutions in production. This indicates the interdependence of perception and production. The results also suggest that the use of perceptual tasks has significant clinical value.

Apraxias↗

Speaking clearly for the hard of hearing IV: Further studies of the role of speaking rate.

The contribution of reduced speaking rate to the intelligibility of "clear" speech (Picheny, Durlach, & Braida, 1985) was evaluated by adjusting the durations of speech segments (a) via nonuniform signal time-scaling, (b) by deleting and inserting pauses, and (c) by eliciting materials from a professional speaker at a wide range of speaking rates. Key words in clearly spoken nonsense sentences were substantially more intelligible than those spoken conversationally (15 points) when presented in quiet for listeners with sensorineural impairments and when presented in a noise background to listeners with normal hearing. Repeated presentation of conversational materials also improved scores (6 points). However, degradations introduced by segment-by-segment time-scaling rendered this time-scaling technique problematic as a means of converting speaking styles. Scores for key words excised from these materials and presented in isolation generally exhibited the same trends as in sentence contexts. Manipulation of pause structure reduced scores both when additional pauses were introduced into conversational sentences and when pauses were deleted from clear sentences. Key-word scores for materials produced by a professional talker were inversely correlated with speaking rate, but conversational rate scores did not approach those of clear speech for other talkers. In all experiments, listeners with normal hearing exposed to flat-spectrum background noise performed similarly to listeners with hearing loss.

Adult↗

Speech timing in apraxia of speech versus conduction aphasia.

This study examined temporal parameters of speech in subjects with apraxia of speech, conduction aphasia, and normal speech. They were asked to repeat target words in a carrier phrase 10 times. Acoustic analyses involved measurement of stop gap duration, voice onset time, vowel nucleus duration, and consonant-vowel (CV) duration. Speakers with apraxia of speech had longer and more variable stop gap, vowel, and CV durations than did subjects with aphasia or normal speech. Speakers with conduction aphasia had longer vowel durations and CV durations than subjects with normal speech. Also, subjects with apraxia of speech showed greater token-to-token variability than the other subject groups. The variability shown by subjects with apraxia of speech was significantly correlated with perceptual judgments of their speech. The significance of these results is discussed in the context of motoric and phonological explanations for apraxia of speech and conduction aphasia.

Adult↗

Effects of single-band syllabic amplitude compression on temporal speech information in nonsense syllables and in sentences.

The effects of single-band amplitude compression on the use by subjects with normal hearing of temporal speech information were assessed using speech stimuli that had been processed to remove most spectral information before being compressed. The resulting signal-related-noise (SRN) stimuli isolated the effects of compression on the temporal information in speech by making it impossible for subjects to identify stimulus items on the basis of spectral speech information. Subjects with normal hearing listened to /aCa/ SRN disyllables that had been subjected to single-band compression at various combinations of compression ratio (CR) and time constants (TC). Performance was reduced only in the most severe compression condition (CR = 8; TC = 50), and then only slightly. Additional testing showed that subjects could use both periodicity and compression-overshoot artifactual information--in addition to envelope information--to identify the compressed /aCa/ stimuli. When a list of 10 context-controlled sentences was converted to SRN and compressed at CR = 8 and TC = 50, the ability of subjects with normal hearing to identify the sentences was significantly affected. Results established that (a) subjects with normal hearing differ widely in their abilities to use temporal information for speech identification, even after training; (b) subjects can learn to use both temporal envelope and periodicity information for identification if disyllables, even though; (c) subjects with normal hearing need envelope but not periodicity information to identify SRN sentences in a closed set. These results suggest that single-band compression at CR = 8 and TC = 50 would be undesirable for persons with limited ability to resolve speech spectral information. It is currently not known how less severe compression conditions would affect envelope information in sentences.

Adult↗

Consonant confusions in amplitude-expanded speech.

The perceptual consequences of expanding the amplitude variations in speech were studied under conditions in which spectral information was obscured by signal correlated noise that had an envelope correlated with the speech envelope, but had a flat amplitude spectrum. The noise samples, created individually from 22 vowel-consonant-vowel nonsense words, were used as maskers of those words, with signal-to-noise ratios ranging from -15 to 0 dB. Amplitude expansion was by a factor of 3.0 in terms of decibels. In the first experiment, presentation level for speech peaks was 80 dB SPL. Consonant recognition performance for expanded speech by 50 listeners with normal hearing was as much as 30 percentage points poorer than for unexpanded speech and the types of errors were dramatically different, especially in the midrange of S-N ratios. In a second experiment presentation level was varied to determine whether reductions in consonant levels produced by expansion were responsible for the differences between conditions. Recognition performance for unexpanded speech at 40 dB SPL was nearly equivalent to that for expanded speech at 80 dB SPL. The error patterns obtained in these two conditions were different, suggesting that the differences between conditions in Experiment 1 were due largely to expanded amplitude envelopes rather than differences in audibility.

Humans↗

Velocity profiles of lip protrusion across changes in speaking rate.

The effects of speaking rate manipulation were examined in the velocity profiles of anticipatory lip protrusion gestures. Systematic changes in the shape, symmetry, and smoothness of the velocity profiles were observed as speaking rate was modulated across a wide range of self-selected rates, from fast to slow. Velocity profiles of movements produced at slower than normal speaking rates demonstrated greater asymmetry, irregularity, and differences in geometric form, compared to a normal and faster-than-normal rates. Subjects evidenced both inter- and intrasubject variability in the accomplishment of lip protrusion and rate manipulations. These results indicate that the velocity profiles of lip protrusion gestures do not necessarily remain invariant across changes in speaking rate. Rather, the data suggest that distinct movement patterns may be generated for slow speaking rates, with select characteristics of the movement pattern being maintained across normal and fast speaking rates.

Adult↗

Lip and jaw kinematics in bilabial stop consonant production.

This paper reports two experiments, each designed to clarify different aspects of bilabial stop consonant production. The first one examined events during the labial closure using kinematic recordings in combination with records of oral air pressure and force of labial contact. The results of this experiment suggested that the lips were moving at a high velocity when the oral closure occurred. They also indicated mechanical interactions between the lips during the closure, including tissue compression and the lower lip moving the upper lip upward. The second experiment studied patterns of upper and lower lip interactions, movement variability within and across speakers, and the effects on lip and jaw kinematics of stop consonant voicing and vowel context. Again, the results showed that the lips were moving at a high velocity at the onset of the oral closure. No consistent influences of stop consonant voicing were observed on lip and jaw kinematics in five subjects, nor on a derived measure of lip aperture. The overall results are compatible with the hypothesis that one target for the lips in bilabial stop production is a region of negative lip aperture. A negative lip aperture implies that to reach their virtual target, the lips would have to move beyond each other. Such a control strategy would ensure that the lips will form an air light seal irrespective of any contextual variability in the onset positions of their closing movements.

Female↗

Effect of acoustic cues on labeling fricatives and affricates.

Previous studies have shown that manipulation of frication amplitude relative to vowel amplitude in the third formant frequency region affects labeling of place of articulation for the fricative contrast /s/-/integral of/ [Hedrick & Ohde, 1993; Stevens, 1985]. The current study examined the influence of this relative amplitude manipulation in conjunction with presentation level, frication duration, and formant transition cues for labeling fricative place of articulation by listeners with normal hearing and listeners with sensorineural hearing loss. Synthetic consonant-vowel (CV) stimuli were used in which the amplitude of the frication relative to vowel onset amplitude in the third formant frequency region was manipulated across a 20 dB range. The listeners with hearing loss appeared to have more difficulty using the formant transition component than the relative amplitude component for the labeling task than most listeners with normal hearing. A second experiment was performed with the same stimuli in which the listeners were given one additional labeling response alternative, the affricate /t integral of/. Results from this experiment showed that listeners with normal hearing gave more /t integral of/ labels as relative amplitude and presentation level increased and frication duration decreased. There was a significant difference between the two groups in the number of affricate responses, as listeners with hearing loss gave fewer /t integral of/ labels.

Adult↗

Long-term phonatory instability in individuals with multiple sclerosis.

This paper uses a new approach to describe and quantify the long-term phonatory instability of speakers with MS. Sustained vowel phonations of 20 individuals with a definite diagnosis of multiple sclerosis (MS) and 20 age- and gender-matched individuals with normal speech were recorded. The phonations were f0 and intensity analyzed and subjected to spectral analysis using the Fast Fourier Transform. Three methods for analyzing the instabilities are presented, compared, and related to perceptual judgments: (a) coefficients of variation, (b) magnitude-based analysis of spectral energy, and (c) frequency-based analysis of spectral components. All measures reliably distinguished between individuals with MS and persons with normal speech. A single factor based on a linear discriminant analysis of the frequency-based measures was especially useful in distinguishing these groups. Critical frequency bands of instability, corresponding to wow (1-2 Hz), tremor (around 8 Hz), and flutter (17-18 Hz), distinguished the MS group from those of the control group.

Adult↗

Development of a two-stage procedure for the automatic recognition of dysfluencies in the speech of children who stutter: II. ANN recognition of repetitions and prolongations with supplied word segment markers.

This program of work is intended to develop automatic recognition procedures to locate and assess stuttered dysfluencies. This and the preceding article focus on developing and testing recognizers for repetitions and prolongations in stuttered speech. The automatic recognizers classify the speech in two stages: In the first the speech is segmented and in the second the segments are categorized. The units segmented are words. The current article describes results for an automatic recognizer intended to classify words as fluent or containing a repetition or prolongation in a text read by children who stutter that contained the three types of words alone. Word segmentations are supplied and the classifier is an artificial neural network (ANN). Classification performance was assessed on material that was not used for training. Correct performance occurred when the ANN placed a word into the same category as the human judge whose material was used to train the ANNs. The best ANN correctly classified 95% of fluent, and 78% of dysfluent words in the test material.

Adolescent↗

Acoustic measures of temporal intervals across speaking rates: variability of syllable- and phrase-level relative timing.

It has been suggested previously that at least some levels of the temporal organization for speech production are characterized by proportional timing. The proportional timing model maintains that the duration of temporal intervals within a sequence would remain proportionally invariant across changes in overall duration of the sequence. In order to test this hypothesis for the acoustic level of speech production, 18 women produced three trials of the utterance "Buy Bobby a poppy" at each of three speaking rates (i.e., slow, normal, fast). Acoustically derived temporal intervals were paired to form ratios reflecting either syllable-level or phrase-level relative timing. Findings indicated that ratios of temporal intervals at both the syllable-level and phrase-level did not remain invariant across speaking rates. Rather, statistically significant changes in the relative duration of both types of intervals were observed as a function of overall rate of production. For most of the obtained ratios, the direction of these changes was highly consistent across individual subjects.

Adult↗

Speech recognition as a function of the number of electrodes used in the SPEAK cochlear implant speech processor.

Speech recognition was measured in listeners with the Nucleus-22 SPEAK speech processing strategy as a function of the number of electrodes. Speech stimuli were analyzed into 20 frequency bands and processed according to the usual SPEAK processing strategy. In the normal clinical processor each electrode is assigned to represent the output of one filter. To create reduced-electrode processors the output of several adjacent filters were directed to a single electrode, resulting in processors with 1, 2, 4, 7, 10, and 20 electrodes. The overall spectral bandwidth was preserved, but the number of active electrodes was progressively reduced. After a 2-day period of adjustment to each processor, speech recognition performance was measured on medial consonants, vowels, monosyllabic words, and sentences. Performance with a single electrode processor was poor in all listeners, and average performance increased dramatically on all test materials as the number of electrodes was increased from 1 to 4. No differences in average performance were observed on any test in the 7-, 10-, and 20-electrode conditions. On sentence and consonant tests there was no difference between average performance with the 4-electrode and 20-electrode processors. This pattern of results suggests that cochlear implant listeners are not able to make full use of the spectral information on all 20 electrodes. Further research is necessary to understand the reasons for this limitation and to understand how to increase the amount of spectral information in speech received by implanted listeners.

Adolescent↗

Effect of relative amplitude and formant transitions on perception of place of articulation by adult listeners with cochlear implants.

Previous studies have shown that manipulation of a particular frequency region of the consonantal portion of a syllable relative to the amplitude of the same frequency region in an adjacent vowel influences the perception of place of articulation. This manipulation has been called the relative amplitude cue. Earlier studies have examined the effect of relative amplitude and formant transition manipulations upon labeling place of articulation for fricatives and stop consonants in listeners with normal hearing. The current study sought to determine if (a) the relative amplitude cue is used by adult listeners wearing a cochlear implant to label place of articulation, and (b) adult listeners wearing a cochlear implant integrated the relative amplitude and formant transition information differently than listeners with normal hearing. Sixteen listeners participated in the study, 12 with normal hearing and 4 postlingually deafened adults wearing the Nucleus 22 electrode Mini Speech Processor implant with the multipeak processing strategy. The stimuli used were synthetic consonant-vowel (CV) syllables in which relative amplitude and formant transitions were manipulated. The two speech contrasts examined were the voiceless fricative contrast /s/-"sh" and the voiceless stop consonant contrast /p/-/t/. For each contrast, listeners were asked to label the consonant sound in the syllable from the two response alternatives. Results showed that (a) listeners wearing this implant could use relative amplitude to consistently label place of articulation, and (b) listeners with normal hearing integrated the relative amplitude and formant transition information more than listeners wearing a cochlear implant, who weighted the relative amplitude information as much as 13 times that of the transition information.

Adult↗

Enhancement of electrolaryngeal speech by adaptive filtering.

Artificial larynges provide a means of verbal communication for people who have either lost or are otherwise unable to use their larynges. Although they enable adequate communication, the resulting speech has an unnatural quality and is significantly less intelligible than normal speech. One of the major problems with the widely used Transcutaneous Artificial Larynx (TAL) is the presence of a steady background noise caused by the leakage of acoustic energy from the TAL, its interface with the neck, and the surrounding neck tissue. The severity of the problem varies from speaker to speaker, partly depending upon the characteristics of the individual's neck tissue. The present study tests the hypothesis that TAL speech is enhanced in quality (as assessed through listener preference judgments) and intelligibility by removal of the inherent, directly radiated background signal. In particular, the focus is on the improvement of speech over the telephone or through some other electronic communication medium. A novel adaptive filtering architecture was designed and implemented to remove the background noise. Perceptual tests were conducted to assess speech, from two individuals with a laryngectomy and two normal speakers using the Servox TAL, before and after processing by the adaptive filter. A spectral analysis of the adaptively filtered TAL speech revealed a significant reduction in the amount of background source radiation yet preserved the acoustic characteristics of the vocal output. Results from the perceptual tests indicate a clear preference for the processed speech. In general, there was no significant improvement or degradation in intelligibility. However, the processing did improve the intelligibility of word-initial non-nasal consonants.

Adult↗

Characterizing knowledge deficits in phonological disorders.

To aid the development of finer-grained measures of phonological competence within a representation-based approach to phonology, two aspects of nonsymbolic phonological knowledge (knowledge of the acoustic/perceptual space and of the articulatory/production space) were examined in 6 preschool-age children with phonological disorders and 6 typically developing age peers. To evaluate perceptual knowledge, gating and noise-center tasks were used. Children with phonological disorders recognized significantly fewer words than age peers on both tasks. To evaluate production knowledge, spectral and temporal measures were obtained for CV sequences involving both lingual and labial stop consonants. Group differences on this task (such as larger transition slope values from lingual consonants to vowels for children with phonological disorders) were also observed. These differerences were interpreted as indicating that the children with phonological disorders were less able to maneuver jaw and tongue body separately or that they used "ballistic" (i.e., less controlled) gestures from lingual consonants to vowels than their age peers. These results suggest that phonological knowledge is multifaceted, and that seemingly categorical deficits at one level can be linked to less robust representations at other levels.

Articulation Disorders↗