Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sound Spectrography”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

The effect of anchors and training on the reliability of perceptual voice evaluation.

Perceptual voice evaluation is a common clinical tool for rating the severity of vocal quality impairment. However, the evaluation process involves subjective judgment, and reliability is therefore a major issue that needs to be considered. When listeners are asked to judge the quality of a voice signal, they use their own internal standards as the references. These internal standards can be variable, as different individuals may have acquired different standards in prior situations. In order to improve the reliability of the perceptual voice evaluation process, external anchors and training are provided to counteract the effect of these internal standards. This study investigated to what extent the provision of anchors and a training program would improve the reliability of perceptual voice evaluation by naive listeners. The results show, in general, that anchors and training helped to improve the reliability of perceptual voice evaluation, especially in the rating of male voices. Furthermore, it was found that anchors made up of synthesized signals combined with training were more effective in improving reliability in judging perceptual roughness and breathiness than natural voice anchors.

Adult↗

Defining and measuring speech movement events.

A long-held view in speech research is that utterances are built up from a series of discrete units joined together. However, it is difficult to reconcile this view with the observation that speech movement waveforms are smooth and continuous. Developing methods for reliable identification of speech movement units is necessary for describing speech motor behavior and for addressing theoretically relevant questions about its organization. We describe a simple method of parsing movement signals into a series of individual movement "strokes," where a stroke is defined as the period between two successive local minima in the speed history of an articulator point, and use that method to segment speech-related movement of marker points placed on the tongue blade, tongue dorsum, lower lip, and jaw in a group of healthy young speakers. Articulator fleshpoints could be distinguished on the basis of kinematic features (i.e., peak and boundary speed, duration and distance) of the strokes they produce. Further, tongue blade and jaw fleshpoint strokes identified to temporally overlap with acoustic events identified as alveolar fricatives could be distinguished from speech strokes in general on the basis of a number of kinematic measures. Finally, the acoustic timing of alveolar fricatives did not appear to be related to the kinematic features of strokes presumed to be related to their production in any direct way. The advantages and disadvantages of this simple approach to defining movement units are discussed.

Adult↗

Final consonant discrimination in children: effects of phonological disorder, vocabulary size, and articulatory accuracy.

Preschool-age children with phonological disorders were compared to their typically developing age peers on their ability to discriminate CVC words that differed only in the identity of the final consonant in whole-word and gated conditions. The performance of three age groups of typically developing children and adults was also assessed on the same task. Children with phonological disorders performed more poorly than age-matched peers, and younger typically developing children performed more poorly than older children and adults, even when the entire CVC word was presented. Performance in the whole-word condition was correlated with receptive vocabulary size and a measure of articulatory accuracy across all children. These results suggest that there is a complex relationship among word learning skills, the ability to attend to fine phonetic detail, and the acquisition of articulatory-acoustic and acoustic-auditory representations.

Adult↗

Acoustic patterns of infant vocalizations expressing emotions and communicative functions.

The present study aimed at identifying the acoustic pattern of vocalizations, produced by 7- to 11-month-old infants, that were interpreted by their mothers as expressing emotions or communicative functions. Participants were 6 healthy, first-born English infants, 3 boys and 3 girls, and their mothers. The acoustic analysis of the vocalizations was performed using a pattern recognition (PR) software system. A PR system not only calculates signal features, it also automatically detects patterns in the arrangement of such features. The following results were obtained: (a) the PR system distinguished vocalizations interpreted as emotions from vocalizations interpreted as communicative functions with an overall accuracy of 87.34%; (b) the classification accuracy of the PR system for vocalizations that convey emotions was 85.4% and for vocalizations that convey communicative functions was 89.5%; and (c) compared to vocalizations that express emotions, vocalizations that express communicative functions were shorter, displayed lower fundamental frequency values, and had greater overall intensity. These findings suggest that in the second half of the first year, infants possess a vocal repertoire that contributes to regulating cooperative interaction with their mothers, which is considered one of the major prerequisites for language acquisition.

Adult↗

Fundamental frequency onset and offset behavior: a comparative study of children and adults.

Short-term changes in vowel fundamental frequency (F0) immediately preceding (F0 offset) and following (F0 onset) production of voiceless obstruents were examined in groups of 4-year-olds, 8-year-olds, and 21-year-olds. Definitive patterns of laryngeal behavior were observed for each measure F0 was found to significantly lower at vowel offset across age groups, with no significant differences noted between groups, suggesting that F0 offset is simply an acoustic consequence of producing a voiceless obstruent preceded by a vowel. The F0 at vowel onset was high and significantly decreased thereafter. Age-related differences were identified for F0 onset with 4-year-olds in that their F0 rose to a lesser degree than that of adults. However, adult females demonstrated a greater change in both F0 onset and F offset behavior than adult males and children, suggesting that age-related differences in F0 behavior are likely to be influenced by sex. The results are discussed with regard to the physiologic constraints of F0 surrounding voiceless obstruent production in children and adults.

Adult↗

Speech synthesis using damped sinusoids.

A speech synthesizer was developed that operates by summing exponentially damped sinusoids at frequencies and amplitudes corresponding to peaks derived from the spectrum envelope of the speech signal. The spectrum analysis begins with the calculation of a smoothed Fourier spectrum. A masking threshold is then computed for each frame as the running average of spectral amplitudes over an 800-Hz window. In a rough simulation of lateral suppression, the running average is then subtracted from the smoothed spectrum (with negative spectral values set to zero), producing a masked spectrum. The signal is resynthesized by summing exponentially damped sinusoids at frequencies corresponding to peaks in the masked spectra. If a periodicity measure indicates that a given analysis frame is voiced, the damped sinusoids are pulsed at a rate corresponding to the measured fundamental period. For unvoiced speech, the damped sinusoids are pulsed on and off at random intervals. A perceptual evaluation of speech produced by the damped sinewave synthesizer showed excellent sentence intelligibility, excellent intelligibility for vowels in /hVd/ syllables, and fair intelligibility for consonants in CV nonsense syllables.

Communication Devices for People with Disabilities↗

Acoustic variations in reading produced by speakers with spasmodic dysphonia pre-botox injection and within early stages of post-botox injection.

Acoustic analysis of a reading passage was used to identify the abnormal phonatory events associated with adductor spasmodic dysphonia (ADSD) pre- and postinjection of Botulinum Toxin A (Botox). Thirty-one patients (age 22 to 74 years) diagnosed with ADSD were included for study. All patients were new recipients of Botox, and the examination of their voice occurred before and after their initial injection of Botox. Acoustic events were identified from reading samples of the Rainbow Passage produced by each of the patients. These events were examined from sentences containing primarily voiced sound segments. Dependent variables included the number of phonatory breaks, frequency shifts, and aperiodic segments--all variables previously defined by the investigators. Additionally, calculated variables were made of the percentage of time these events occurred relative to the duration of the cumulative voiced segments. A sex- and age-matched control group (+/-2 years) was included for statistical comparison. Results indicated that those with ADSD produced more aberrant acoustic events than the controls. Aperiodicity was the predominant acoustic event produced during the reading, followed by frequency shifts and phonatory breaks. Within the ADSD group, the number of atypical acoustic events decreased following Botox injection. It is important that the occurrence of specific abnormal acoustic events was sufficient to differentiate the disordered speakers from the controls following as well as preceding initial Botox injection, as indicated by discriminant function analysis. This paper complements our previous work using this acoustic analysis method for defining the abnormal events present in the voice of those with ADSD and further suggests that these measures can be used in conjunction with perceptual impressions to differentiate speakers on the basis of initial severity.

Adult↗

Prosodic control in severe dysarthria: preserved ability to mark the question-statement contrast.

Speakers with severe dysarthria are known to have reduced range in prosody. Consistent control within that range, however, has largely been ignored. In earlier investigations speakers with severe dysarthria were able to control pitch and duration for sustained vowel production despite reduced flexibility of control (Patel, 1998). The present experiment examined whether 8 speakers with severe dysarthria due to cerebral palsy used prosodic parameters of pitch contour and syllable duration for phrase-level productions. Speakers with dysarthria (N = 8) produced 3-syllable phrases as questions and statements. Naïve listeners (N = 48) classifed dysarthric productions as either questions or statements. Listeners were able to distinguish questions from statements with accuracy levels ranging from 81% to 98%. We were also interested in studying how dysorthric speakers marked the question-statement contrast. Prosodic features of pitch contour and syllable duration were systematically removed from the original recorded vocalizations to examine the salience of these features on listener classification. Removal of pitch contour cues dramatically reduced listener accuracy scores to almost chance performance. Listeners found pitch contour cues to be information-bearing cues in dysarthric vocalizations even though the range of frequency control in these speakers may be reduced. That speakers with dysarthria were able to exert sufficient control to signal the question-statement contrast has implications for diagnostic and intervention practices aimed to optimally exploit prosodic control for enhancing communication efficiency.

Adult↗

Intelligibility of modified speech for young listeners with normal and impaired hearing.

Exposure to modified speech has been shown to benefit children with language-learning impairments with respect to their language skills (M. M. Merzenich et al., 1998; P. Tallal et al., 1996). In the study by Tallal and colleagues, the speech modification consisted of both slowing down and amplifying fast, transitional elements of speech. In this study, we examined whether the benefits of modified speech could be extended to provide intelligibility improvements for children with severe-to-profound hearing impairment who wear sensory aids. In addition, the separate effects on intelligibility of slowing down and amplifying speech were evaluated. Two groups of listeners were employed: 8 severe-to-profoundly hearing-impaired children and 5 children with normal hearing. Four speech-processing conditions were tested: (1) natural, unprocessed speech; (2) envelope-amplified speech; (3) slowed speech; and (4) both slowed and envelope-amplified speech. For each condition, three types of speech materials were used: words in sentences, isolated words, and syllable contrasts. To degrade the performance of the normal-hearing children, all testing was completed with a noise background. Results from the hearing-impaired children showed that all varieties of modified speech yielded either equivalent or poorer intelligibility than unprocessed speech. For words in sentences and isolated words, the slowing-down of speech had no effect on intelligibility scores whereas envelope amplification, both alone and combined with slowing-down, yielded significantly lower scores. Intelligibility results from normal-hearing children listening in noise were somewhat similar to those from hearing-impaired children. For isolated words, the slowing-down of speech had no effect on intelligibility whereas envelope amplification degraded intelligibility. For both subject groups, speech processing had no statistically significant effect on syllable discrimination. In summary, without extensive exposure to the speech processing conditions, children with impaired hearing and children with normal hearing listening in noise received no intelligibility advantage from either slowed speech or envelope-amplified speech.

Audiometry, Pure-Tone↗

Developmental change in variability of lip muscle activity during speech.

Compared to adults, children's speech production measures sometimes show higher trial-to-trial variability in both kinematic and acoustic analyses. A reasonable hypothesis is that this variability reflects variations in neural drive to muscles as the developing system explores different solutions to achieving vocal tract goals. We investigated that hypothesis in the present study by analyzing EMG waveforms produced across repetitions of a phrase spoken by 7-year-olds, 12-year-olds, and young adults. The EMG waveforms recorded via surface electrodes at upper lip sites were clearly modulated in a consistent manner corresponding to lip closure for the bilabial consonants in the utterance. Thus we were able to analyze the amplitude envelope of the rectified EMG with a phrase-level variability index previously used with kinematic data. Both the 7- and 12-year-old children were significantly more variable on repeated productions than the young adults. These results support the idea that children are using varying combinations of muscle activity to achieve phonetic goals. Even at age 12 years, these children were not adult-like in their performance. These and earlier kinematic studies of the oral motor system suggest that children retain their flexibility, employing more degrees of freedom than adults, to dynamically control lip aperture during speech. This strategy is adaptive given the many neurophysiological and biomechanical changes that occur during the transition from adolescence to adulthood.

Adult↗

"Pitch" accent in alaryngeal speech.

Highly proficient alaryngeal speakers are known to convey prosody successfully. The present study investigated whether alaryngeal speakers not selected on grounds of proficiency were able to convey pitch accent (a pitch accent is realized on the word that is in focus, cf. Bolinger, 1958). The participating speakers (10 tracheoesophageal, 9 esophageal, and 10 laryngeal [control] speakers) produced sentences in which accent was cued by the preceding context. For each utterance, a group of listeners identified which word conveyed accent. All speakers were able to convey accent. Acoustic analyses showed that some alaryngeal speakers had little or no control over fundamental frequency. Contrary to expectation, these speakers did not compensate by using nonmelodic cues, whereas speakers using F0 did use nonmelodic cues. Thus, temporal and intensity cues are concomitant with the use of F0; if F0 is affected, these nonmelodic cues will be as well. A pitch perception experiment confirmed that alaryngeal speakers who had no control over F0 and who did not use nonmelodic cues were nevertheless able to produce pitch movements. Speakers with no control over F0 apparently relied on an alternative pitch system to convey accents and other pitch movements.

Adult↗

Influence of hearing loss on the perceptual strategies of children and adults.

To accommodate growing vocabularies, young children are thought to modify their perceptual weights as they gain experience with speech and language. The purpose of the present study was to determine whether the perceptual weights of children and adults with hearing loss differ from those of their normal-hearing counterparts. Adults and children with normal hearing and with hearing loss served as participants. Fricative and vowel segments within consonant-vowel-consonant stimuli were presented at randomly selected levels under two conditions: unaltered and with the formant transition removed. Overall performance for each group was calculated as a function of segment level. Perceptual weights were also calculated for each group using point-biserial correlation coefficients that relate the level of each segment to performance. Results revealed child-adult differences in overall performance and also revealed an effect of hearing loss. Despite these performance differences, the pattern of perceptual weights was similar across all four groups for most conditions.

Adult↗

Spectral characteristics of speech at the ear: implications for amplification in children.

This study examined the long- and short-term spectral characteristics of speech simultaneously recorded at the ear and at a reference microphone position (30 cm at 0 degrees azimuth). Twenty adults and 26 children (2-4 years of age) with normal hearing were asked to produce 9 short sentences in a quiet environment. Long-term average speech spectra (LTASS) were calculated for the concatenated sentences, and short-term spectra were calculated for selected phonemes within the sentences (/m/, /n/, /s/, [see text], /f/, /a/, /u/, and /i/). Relative to the reference microphone position, the LTASS at the ear showed higher amplitudes for frequencies below 1 kHz and lower amplitudes for frequencies above 2 kHz for both groups. At both microphone positions, the short-term spectra of the children's phonemes revealed reduced amplitudes for /s/ and [see text] and for vowel energy above 2 kHz relative to the adults' phonemes. The results of this study suggest that, for listeners with hearing loss (a) the talker's own voice through a hearing instrument would contain lower overall energy at frequencies above 2 kHz relative to speech originating in front of the talker, (b) a child's own speech would contain even lower energy above 2 kHz because of adult-child differences in overall amplitude, and (c) frequency regions important to normal speech development (e.g., high-frequency energy in the phonemes /s/ and [see text]) may not be amplified sufficiently by many hearing instruments.

Adult↗

Changes in the human vocal tract due to aging and the acoustic correlates of speech production: a pilot study.

This investigation used a derivation of acoustic reflection (AR) technology to make cross-sectional measurements of changes due to aging in the oral and pharyngeal lumina of male and female speakers. The purpose of the study was to establish preliminary normative data for such changes and to obtain acoustic measurements of changes due to aging in the formant frequencies of selected spoken vowels and their long-term average spectra (LTAS) analysis. Thirty-eight young men and women and 38 elderly men and women were involved in the study. The oral and pharyngeal lumina of the participants were measured with AR technology, and their formant frequencies were analyzed using the Kay Elemetrics Computerized Speech Lab. The findings have delineated specific and similar patterns of aging changes in human vocal tract configurations in speakers of both genders. Namely, the oral cavity length and volume of elderly speakers increased significantly compared to their young cohorts. The total vocal tract volume of elderly speakers also showed a significant increment, whereas the total vocal tract length of elderly speakers did not differ significantly from their young cohorts. Elderly speakers of both genders also showed similar patterns of acoustic changes of speech production, that is, consistent lowering of formant frequencies (especially F1) across selected vowel productions. Although new research models are still needed to succinctly account for the speech acoustic changes of the elderly, especially for their specific patterns of human vocal tract dimensional changes, this study has innovatively applied the noninvasive and cost-effective AR technology to monitor age-related human oral and pharyngeal lumina changes that have direct consequences for speech production.

Adolescent↗

Developmental effects in the masking-level difference.

Adults and children (aged 5 years 1 month to 10 years 8 months) were tested in a masking-level difference (MLD) paradigm in which detection of brief signals was contrasted for signal placement in masker envelope maxima versus masker envelope minima. Maskers were 50-Hz-wide noise bands centered on 500 Hz, and the signals were So or Sp 30-ms, 500-Hz tones. In agreement with previous studies, it was found that MLDs were greater for masker envelope minima placement than for masker envelope maxima placement. Across the age range of the children tested here, the binaural advantage associated with the masker envelope minima increased with the age of the child. One interpretation of the present results is that there is a developmental improvement in binaural temporal resolution over the age range tested here.

Acoustic Stimulation↗

Variability in /s/ production in children and adults: evidence from dynamic measures of spectral mean.

Previous research has found developmental decreases in temporal variability in speech. Relatively less work has examined spectral variability, and, in particular, variability in consonant spectra. This article examined variability in productions of the consonant /s/ by adults and by 3 groups of children, with mean ages of 3;11 (years;months), 5;04, and 8;04. Specifically, it measured the influence of age, phonetic context, and syllabic context on variability. Spectral variability was estimated by measuring dynamic spectral characteristics of multiple productions of /s/ in sV, spV, and swV sequences, where the vowel was either /a/ or /u/. Mean duration, variability in duration, and coarticulation were also measured. Children were found to produce /s/ with greater temporal and spectral variability than adults. Duration and coarticulation were comparable across the 4 age groups. Spectral variability was greater in swV contexts than in sV or spV sequences. The lack of consistent effects of phonetic context on spectral variability suggests that the developmental differences were related to subtle variability in place of articulation for /s/ in the children's productions.

Adult↗

How well can children recognize speech features in spectrograms? Comparisons by age and hearing status.

Real-time spectrographic displays (SDs) have been used in speech training for more than 30 years with adults and children who have severe and profound hearing impairments. Despite positive outcomes from treatment studies, concerns remain that the complex and abstract nature of spectrograms may make these speech training aids unsuitable for use with children. This investigation examined how well children with normal hearing sensitivity and children with impaired hearing can recognize spectrographic cues for vowels and consonants, and the ages at which these visual cues are distinguished. Sixty children (30 with normal hearing sensitivity, 30 with hearing impairments) in 3 age groups (6-7, 8-9, and 10-11 years) were familiarized with the spectrographic characteristics of selected vowels and consonants. The children were then tested on their ability to select a match for a model spectrogram from among 3 choices. Overall scores indicated that spectrographic cues were recognized with greater-than-chance accuracy by all age groups. Formant contrasts were recognized with greater accuracy than consonant manner contrasts. Children with normal hearing sensitivity and those with hearing impairment performed equally well.

Age Factors↗

Effect of F2 intensity on identity of /u/ in degraded listening conditions.

The current study investigated the influence of the second formant (F2) intensity on vowel labeling along a /u/-/i/ continuum. Twenty-two listeners with normal-hearing (NH) sensitivity and 14 listeners with sensorineural hearing impairment (HI) were initially presented 2 stimuli for which the F2 intensity differed by 20 dB. The listeners were asked to label the 2 stimuli categorically as /u/ or /i/. After passing this criterion test, listeners were presented 9 stimuli whose F2 intensity varied within the 20-dB range. The 9 stimuli were evaluated in 3 listening conditions: in quiet, in the presence of a continuous speech spectrum noise (0-dB signal-to-noise ratio), and in the presence of reverberation (T = 1.0 s). The intensity manipulation altered the vowel labeling of NH listeners and yielded a differential effect in noise versus reverberation. Only 5 of the HI listeners were able to pass the criterion test, and of these 5, only 2 were able to label the 9 stimuli categorically. Results from HI listeners suggest problems in categorizing spectral shape.

Adult↗