Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sound Spectrography”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 991 records · Page 55Linked to original sources

The role of consonant-vowel amplitude ratio in the recognition of voiceless stop consonants by listeners with hearing impairment.

Several authors have evaluated consonant-to-vowel ratio (CVR) enhancement as a means to improve speech recognition in listeners with hearing impairment, with the intention of incorporating this approach into emerging amplification technology. Unfortunately, most previous studies have enhanced CVRs by increasing consonant energy, thus possibly confounding CVR effects with consonant audibility. In this study, we held consonant audibility constant by reducing vowel transition and steady-state energy rather than increasing consonant energy. Performance-by-intensity (PI) functions were obtained for recognition of voiceless stop consonants (/p/, /t/, /k/) presented in isolation (burst and aspiration digitally separated from the vowel) and for consonant-vowel syllables, with readdition of the vowel /a/. There were three CVR conditions: normal CVR, vowel reduction by 6 dB, and vowel reduction by 12 dB. Testing was conducted in broadband noise fixed at 70 dB SPL and at 85 dB SPL. Six adults with sensorineural hearing impairment and 2 adults with normal hearing served as listeners. Results indicated that CVR enhancement did not improve identification performance when consonant audibility was held constant, except at the higher noise level for one listener with hearing impairment. The re-addition of the vowel energy to the isolated consonant did, however, produce large and significant improvements in phoneme identification.

Adult↗

Acoustic examination of preterm and full-term infant cries: the long-time average spectrum.

The acoustic characteristics of crying behavior displayed in 2 groups of newborn infants are reported. The crying episodes of 10 full-term and 10 preterm infants were audio recorded and analyzed with regard to the long-time average spectrum (LTAS) characteristics. An LTAS display was created for each infant's non-partitioned crying episode, as well as for 3 equidurational partitions of the crying episode. Measures of first spectral peak, mean spectral energy, and spectral tilt were revealing of differences between full-term and preterm infants' non-partitioned crying episodes. In addition, the full-term infants demonstrated significant changes in their crying behavior across partitions, whereas the preterm infants changed little across the crying episode. Discussion focuses on possible differences between full-term and preterm infants in their neurophysiological maturity, and the subsequent impact on their speech development. The importance of examining entire crying episodes when evaluating the crying behavior of infants is also discussed.

Anthropometry↗

The effects of a flattened fundamental frequency on intelligibility at the sentence level.

The purpose of this preliminary experiment was to evaluate the effect of a flattened fundamental frequency (F0) contour on sentence intelligibility. The perceptual dimension monotone pitch is frequently used to describe the speech of persons with dysarthria, and relatively flat F0 contours have been noted in several acoustic studies of dysarthria. To determine the independent effect of a flattened F0 contour on sentence intelligibility a resynthesis technique was used that held timing and spectral characteristics of utterances constant while allowing parametric control over successive pitch periods. Two male speakers produced low-probability utterances selected from the SPIN test, which were then resynthesized with a flattened F0 contour. Speech intelligibility was assessed using two measures: one involving word transcription and the other interval scaling. These measures were collected from 10 listeners. The results showed that both measures were significantly lower when the F0 contour was flattened, as compared with naturally varying contours. Several different explanations are proposed for this effect, which can and should be explored in greater detail using the resynthesis technique given the prominence of this characteristic in dysarthria.

Adult↗

Speech and oral motor learning in individuals with cerebellar atrophy.

The purpose of this study was to determine whether cerebellar pathology interferes with motor learning for either speech or novel tasks. Practice effects were contrasted between persons with cerebellar cortical atrophy (CCA) and control participants on previously learned real speech, nonsense speech, and novel nonspeech oral-movement tasks. Studies of limb motor learning suggested that control participants would evidence reduced variability, increased speed of movement, and reduced movement amplitude with practice as compared with the CCA group. No significant differences were found between the real- and nonsense-speech tasks. For both speech tasks, although neither group reduced their movement variability with practice, both groups significantly reduced jaw closing displacement and velocity with practice. For the novel nonspeech oral-movement task, no change with practice was observed in either group in terms of variability, amplitude, or peak velocity. No effects of cerebellar pathology were seen in either the speech- or oral-movement tasks. These results demonstrated that with practice of speech tasks, a previously learned motor skill, movement speed and displacement decreased in both groups. Therefore, the effects of practice differed between previously learned speech tasks and the novel oral-movement task regardless of cerebellar pathology.

Adult↗

Identification of pathological voices using glottal noise measures.

We investigated the abilities of four fundamental frequency (F0)-dependent and two F0-independent measures to quantify vocal noise. Two of the F0-dependent measures were computed in the time domain, and two were computed using spectral information from the vowel. The F0-independent measures were based on the linear prediction (LP) modeling of vowel samples. Tests using a database of sustained vowel samples, collected from 53 normal and 175 pathological talkers, showed that measures based on the LP model were much superior to the other measures. A classification rate of 96.5% was achieved by a parameter that quantifies the spectral flatness of the unmodeled component of the vowel sample.

Adult↗

Older listeners' use of temporal cues altered by compression amplification.

This study compared the ability of younger and older listeners to use temporal information in speech when that information was altered by compression amplification. Recognition of vowel-consonant-vowel syllables was measured for four groups of adult listeners (younger normal hearing, older normal hearing, younger hearing impaired, older hearing impaired). There were four conditions. Syllables were processed with wide-dynamic range compression (WDRC) amplification and with linear amplification. In each of those conditions, recognition was measured for syllables containing only temporal information and for syllables containing spectral and temporal information. Recognition of WDRC-amplified speech provided an estimate of the ability to use altered amplitude envelope cues. Syllables were presented with a high-frequency masker to minimize confounding differences in high-frequency sensitivity between the younger and older groups. Scores were lower for WDRC-amplified speech than for linearly amplified speech, and older listeners performed more poorly than younger listeners. When spectral information was unrestricted, the age-related decrement was similar for both amplification types. When spectral information was restricted for listeners with normal hearing, the age-related decrement was greater for WDRC-amplified speech than for linearly amplified speech. When spectral information was restricted for listeners with hearing loss, the age-related decrement was similar for both amplification types. Clinically, these results imply that when spectral cues are available (i.e., when the listener has adequate spectral resolution) older listeners can use WDRC hearing aids to the same extent as younger listeners. For older listeners without hearing loss, poorer scores for compression-amplified speech suggest an age-related deficit in temporal resolution.

Adult↗

Respiratory control in stuttering speakers: evidence from respiratory high-frequency oscillations.

This study tested the hypothesis that, in stuttering speakers, relations between the neural control systems for speech and life support, or metabolic breathing, may differ from relations previously observed in normally fluent subjects. Bilaterally coherent high-frequency oscillations in inspiratory-related EMGs, measured as maximum coherence in the frequency band of 60-110 Hz (MC-HFO), were used as indicators of participation by the brainstem controller for metabolic breathing in 10 normally fluent and 10 stuttering speakers. In all controls and most stuttering subjects, MC-HFO for speech was higher than or comparable to MC-HFO for deep breathing. For 4 stuttering subjects, higher MC-HFO was observed for speech than for deep breathing. Comparison of deep breathing to a speechlike breathing task yielded similar results. No relationship between MC-HFO during speech and severity of disfluency was observed. We conclude that in some stuttering speakers, the relations between respiratory controllers are atypical, but that high participation by the HFO-producing circuitry in the brainstem during speech is not sufficient to disrupt fluency.

Adult↗

Laryngeal factors in voiceless consonant production in men, women, and 5-year-olds.

Voicing control in stop consonants has often been measured by means of voice onset time (VOT) and discussed in terms of interarticulator timing. However, control of voicing also involves details of laryngeal setting and management of sub- and supraglottal pressure levels, and many of these factors are known to undergo developmental change. Mechanical and aerodynamic conditions at the glottis may therefore vary considerably in normal populations as functions of age and/or sex. The current study collected oral airflow, intraoral pressure, and acoustic signals from normal English-speaking adults and children producing stop consonants and /h/ embedded in a short carrier utterance. Measures were made of stop VOTs, /h/ voicing and flow characteristics, and subglottal pressure during /p/ closures. Clear age and gender effects were observed for /h/: Fully voiced /h/ was most common in men, and /h/ voicing and flow data showed the highest variability among the 5-year-olds. For individual participants, distributional measures of VOT in /p t/ were correlated with distributional measures of voicing in /h/. The data indicate that one cannot assume comparable laryngeal conditions across speaker groups. This, in turn, implies that VOT acquisition in children cannot be interpreted purely in terms of developing interarticulator timing control, but must also reflect growing mastery over voicing itself. Further, differences in laryngeal structure and aerodynamic quantities may require men and women to adopt somewhat different strategies for achieving distinctive consonantal voicing contrasts.

Adult↗

Ataxic dysarthria.

Although ataxic dysarthria has been studied with various methods in several languages, questions remain concerning which features of the disorder are most consistent, which speaking tasks are most sensitive to the disorder, and whether the different speech production subsystems are uniformly affected. Perceptual and acoustic data were obtained from 14 individuals (seven men, seven women) with ataxic dysarthria for several speaking tasks, including sustained vowel phonation, syllable repetition, sentence recitation, and conversation. Multidimensional acoustic analyses of sustained vowel phonation showed that the largest and most frequent abnormality for both men and women was a long-term variability of fundamental frequency. Other measures with a high frequency of abnormality were shimmer and peak amplitude variation (for both sexes) and jitter (for women). Syllable alternating motion rate (AMR) was typically slow and irregular in its temporal pattern. In addition, the energy maxima and minima often were highly variable across repeated syllables, and this variability is thought to reflect poorly coordinated respiratory function and inadequate articulatory/voicing control. Syllable rates tended to be slower for sentence recitation and conversation than for AMR, but the three rates were highly similar. Formant-frequency ranges during sentence production were essentially normal, showing that articulatory hypometria is not a pervasive problem. Conversational samples varied considerably across subjects in intelligibility and number of words/ morphemes in a breath group. Qualitative analyses of unintelligible episodes in conversation showed that these samples generally had a fairly well-defined syllable pattern but subjects differed in the degree to which the acoustic contrasts typical of consonant and vowel sequences were maintained. For some individuals, an intelligibility deficit occurred in the face of highly distinctive (and contrastive) acoustic patterns.

Adolescent↗

Multivariate statistical analysis of flat vowel spectra with a view to characterizing dysphonic voices.

The aim of this article is to show how dysphonic voices can be characterized by means of a multivariate statistical analysis of flat vowel spectra. The spectral contour was obtained by means of a wavelet transform of the logarithmic magnitude spectrum, which was subsequently flattened to remove interspeaker variability related to the excitation and vocal tract filter functions. The results of the statistical analysis of flat spectra were the following. Firstly, principal components analysis produced markers that separated noisy from clean spectra. Secondly, the heuristic search for harmonic peaks or interharmonic dips could be omitted. Thirdly, conventional spectral markers of noise appeared as special instances of the markers that were derived statistically. Fourthly, the levels of visually assigned hoarseness and the first two principal components were significantly correlated. The assignment of different levels of (visual) hoarseness to different vowel timbres could be explained by the variability associated with the spectral contour.

Algorithms↗

A comparison of speech training methods with deaf adolescents: spectrographic versus noninstrumental instruction.

The effects of speech training with real-time spectrographic displays (SDs) were examined and compared to the effects of noninstrumental (NI) instruction (i.e., training without computerized displays of speech) for deaf adolescents. A single-case modified alternating-treatment experimental design with replication across subjects and speech targets was used to examine within-subject performance in establishing, maintaining, and generalizing target consonants. Comparisons between the two approaches were accomplished by determining how frequently each method resulted in improvement, maintenance of improvement, and generalization to untrained words. Each of the 4 subjects demonstrated improvement under both forms of instruction in a relatively short time. Maintenance of improvement was observed 6 weeks post-treatment for two NI-trained targets and one SD-trained target. Two subjects showed better generalization for their SD-trained target than their NI target. There was little difference in generalization scores for the remaining subjects. All subjects either regained their highest previous levels of acceptability or maintained high-level acceptability following brief, independent practice with SDs 10 weeks after training was discontinued. The expediency of independent practice with SDs was discussed.

Adolescent↗

A new acoustic method of differentiating palatal from non-palatal snoring.

Palatal snoring produces explosive peaks of sound at very low frequency (approximately 20 Hz). Using a digital sound trace a ratio of peak amplitude to root mean square amplitude can be calculated. This Peak Factor Ratio is significantly higher for palatal snores than non-palatal snores (P < 0.01). This acoustic method will be useful for selecting patients for palatal surgery as it is non-invasive and could be used in a home monitor.

Acoustics↗

Sound frequency analysis and the site of snoring in natural and induced sleep.

The aim of this study was to compare the snoring sounds induced during sleep nasendoscopy, and to compare them with those of natural sleep using sound frequency spectra. The snoring of 16 subjects was digitally recorded during natural and induced sleep, noting the site of vibration during sleep nasendoscopy. Patients with palatal snoring during sleep nasendoscopy had a median peak frequency at 137 Hz (118 snore samples). The peak frequency of tongue-base snoring was 1243 Hz (10 snore samples), and simultaneous palate and tongue was 190 Hz (six snore samples). The median power ratios were 7, 0.2 and 5 respectively. The centre frequencies were 371, 1094 and 404 Hz respectively. Epiglottic snores had a peak frequency of 490 Hz (five snore samples). Comparison of the induced (n = 118) and natural (n = 300) snore samples of the 12 palatal snorers showed a significant difference in both the power ratio and centre frequencies (P = 0.031 and P = 0.049). The peak frequency position was similar (P = 0.34). Our results indicate that induced snores contain a higher frequency component of sound, not evident during natural snoring. This is consistent with an element of tongue-base snoring. Although there is good correlation generally, sleep nasendoscopy may not accurately reflect natural snoring.

Endoscopy↗

The analysis of frequency of occlusal sounds in patients with periodontal diseases and gnathic dysfunction.

A study is presented of occlusosonograms obtained from healthy subjects and patients with periodontal diseases and gnathic problems. Gnathosonic examinations were carried out by means of a set of equipment whose basic elements were two piezoceramic transducers mounted on a specially made holder and an IBM PC with an ADC card. The evaluation of occlusosonograms was by means of mathematical analysis of the frequency pattern. In pathological stages, the occlusosonograms exhibit more high frequency components. In the case of healthy subjects, the recorded signal took the form of a strong impact of low frequencies and its subsequent exponential decay.

Adolescent↗

Validation of a recording protocol for assessing temporomandibular sounds and a method for assessing jaw position.

Sounds are often produced by the temporomandibular joint (TMJ) during movement in both symptomatic and asymptomatic subjects. However, subjective methods of describing these sounds have been shown to have poor inter- and intra-observer reliability. In this study, a low cost system in which TMJ sounds were detected using the loudspeakers of lightweight in-ear headphones as microphones is evaluated. The sounds were recorded on tape and then analysed using a computer. Sounds were elicited by asking subjects to bring their teeth together with sufficient force to produce a tooth contact sound, then open their mouths as far as possible and then close again. Placing the microphones in the ears attenuated ambient sounds by 58%, thus providing a degree of immunity from ambient noise. Sampling was performed on the left microphone only at 3.4 kHz and from the left and right microphones together at 1.7 kHz for 60 TMJ sounds and 60 tooth contact sounds. Spectral analysis of sounds recorded at the two sample rates revealed no significant differences. Therefore, a sample rate of 1.7 kHz is adequate to resolve the frequency components present in the TMJ sounds. Although simply recording TMJ sounds does not give a direct measurement of the position of the mandible, using this protocol allows the length of the open close cycle to be determined. If the envelope of movement is assumed to approximate a sinusoid, then the direction of mandibular movement can be assumed to reverse at the half way point in the cycle. The accuracy of this assumption was calculated by comparing the mid-point of the cycle to the point of maximum gape in 129 cycles from nine subjects. The mean difference expressed as percentage of cycle length was 1.3 +/- 0.9%.

Adolescent↗

Monitoring the state of the occlusion--gnathosonics can be reliable.

Monitoring the state of the occlusion should include records that can be compared. Gnathosonics has not become generally accepted as an everyday method of assessing the quality of the occlusion. It is suggested that this may be due to inconsistencies in the results obtained by various workers in the field. It is further suggested that such inconsistencies could be due to the manner of data gathering in the areas of equipment, transfer of data from the patient and interpretation of the records. A reliable method of data gathering using accelerometers is suggested and, taking account of the complexities of sound transmission through the skull, concludes that the overall envelope of the sounds produced by the occlusion of the teeth is more informative that the actual frequencies generated. Ways in which gnathosonics could be placed on a scientific footing and areas where further investigation might be useful are suggested. Gnathosonics, properly controlled, can provide a simple, quick and reliable method of making permanent records of occlusal sounds for comparison and assessment of stability or change in the state of the occlusion.

Adult↗

Autocorrelation of acoustic signals from the temporomandibular joint.

Although there have been many investigations of TMJ sounds in the time and frequency domains, no previous reports have been found of investigations of the autocorrelation spectra of these sounds. In the present study, TMJ sounds were digitized at 1.7 kHz and 300 ms samples containing either clicks (single short duration sounds), crepitus (long duration continuous sounds) or creaks (a series of two or more clicks) were selected. These samples were compared with sounds of known origin: tooth impact sounds, frictional sounds elicited by scratching the head, and bruxing sounds resulting from stick-slip friction as teeth were slid against one another under high pressure. There were clear qualitative and quantitative differences between the autocorrelation spectra of the three types of TMJ sounds. Clicks were similar to tooth impact contact sounds, creaks were similar to the bruxing sounds, and crepitus was similar to the scratching sounds. The repetition rate of creaks was 16 Hz (s.d. 9 Hz), this being similar to the resonance of the mandible about the condylar axis. It is suggested that the creaks are due to stick-slip friction in the lower joint compartment of the TMJ.

Acoustics↗

Resonant characteristics of the human head in relation to temporomandibular joint sounds.

Frequency analysis of the sounds produced by the temporomandibular joint (TMJ) has been claimed to be of diagnostic value. In this study the resonant behaviour of the human skull has been characterized. Assuming that impact is one major mechanism generating the sound, it is shown that the frequencies seen in TMJ sounds relate to the resonant modes of the skull. Where the rise time of the impact is sufficiently short, higher resonant modes are excited.

Adolescent↗