Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sound Spectrography”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,531 records · Page 85Linked to original sources

Effects of low pass filtering on the intelligibility of speech in noise for people with and without dead regions at high frequencies.

People with high-frequency sensorineural hearing loss differ in the benefit they gain from amplification of high frequencies when listening to speech. Using vowel-consonant-vowel (VCV) stimuli in quiet that were amplified and then low pass filtered with various cutoff frequencies, Vickers et aL [J. Acoust. Soc. Am. 110, 1164-1175 (2001)] found that the benefit from amplification of high-frequency components was related to the presence or absence of a cochlear dead region at high frequencies. For hearing-impaired subjects without dead regions, performance improved with increasing cutoff frequency up to 7.5 kHz (the highest value tested). Subjects with high-frequency dead regions showed no improvement when the cutoff frequency was increased above about 1.7 times the edge frequency of the dead region. The present study was similar to that of Vickers et al. but used VCV stimuli presented in background noise. Ten subjects with high-frequency hearing loss, including eight from the study of Vickers et al., were tested. Five had dead regions starting below 2 kHz, and five had no dead regions. Speech stimuli at a nominal level of 65 dB were mixed with spectrally matched noise, amplified according to the "Cambridge" prescriptive formula for each subject and then low pass filtered. The noise level was chosen separately for each subject to give a moderate reduction in intelligibility relative to listening in quiet. For subjects without dead regions, performance generally improved with increasing cutoff frequency up to 7.5 kHz, on average more so in noise than in quiet. For most subjects with dead regions, performance improved with cutoff frequency up to 1.5-2 times the edge frequency of the dead region, but hardly changed with further increases. Calculations of speech audibility using a modified version of the articulation index showed that application of the Cambridge formula was at least partially successful in making high-frequency components of the speech audible for subjects with dead regions, and that such subjects often failed to benefit from increased audibility of the speech at high frequencies.

Aged↗

The intelligibility of speech with "holes" in the spectrum.

The intelligibility of speech having either a single "hole" in various bands or having two "holes" in disjoint or adjacent bands in the spectrum was assessed with normal-hearing listeners. In experiment 1, the effect of spectral "holes" on vowel and consonant recognition was evaluated using speech processed through six frequency bands, and synthesized as a sum of sine waves. Results showed a modest decrease in vowel and consonant recognition performance when a single hole was introduced in the low- and high-frequency regions of the spectrum, respectively. When two spectral holes were introduced, vowel recognition was sensitive to the location of the holes, while consonant recognition remained constant around 70% correct, even when the middle- and high-frequency speech information was missing. The data from experiment 1 were used in experiment 2 to derive frequency-importance functions based on a least-squares approach. The shapes of the frequency-importance functions were found to be different for consonants and vowels in agreement with the notion that different cues are used by listeners to identify consonants and vowels. For vowels, there was unequal weighting across the various channels, while for consonants the frequency-importance function was relatively flat, suggesting that all bands contributed equally to consonant identification.

Humans↗

Factors underlying the speech-recognition performance of elderly hearing-aid wearers.

This paper reports the aided and unaided speech-recognition scores from a group of 171 elderly hearing-aid wearers. All hearing-aid wearers were fit with identical instruments (linear Class-D amplifiers with output-limiting compression) and evaluated with a standard protocol. In addition to including multiple measures of speech recognition, an extensive set of physiological and perceptual measures of auditory function, as well as general measures of cognitive function, were completed prior to the hearing-aid fitting. Comparison of the results from this study to available norms suggested that this group of participants was fairly typical or representative for their hearing loss and age. Approaches to the prediction of general speech-recognition performance that were examined included methods based on an acoustical index, the Speech Intelligibility Index (SII), and others based on linear-regression statistical analysis. The latter approach proved to be the most successful, accounting for about two-thirds of the variance in speech-recognition performance, with the primary predictive factors being measures of hearing loss and cognitive function.

Aged↗

Perception of synthesized voice quality in connected speech by Cantonese speakers.

Perceptual voice analysis is a subjective process. However, despite reports of varying degrees of intrajudge and interjudge reliability, it is widely used in clinical voice evaluation. One of the ways to improve the reliability of this procedure is to provide judges with signals as external standards so that comparison can be made in relation to these "anchor" signals. The present study used a Klatt speech synthesizer to create a set of speech signals with varying degree of three different voice qualities based on a Cantonese sentence. The primary objective of the study was to determine whether different abnormal voice qualities could be synthesized using the "built-in" synthesis parameters using a perceptual study. The second objective was to determine the relationship between acoustic characteristics of the synthesized signals and perceptual judgment. Twenty Cantonese-speaking speech pathologists with at least three years of clinical experience in perceptual voice evaluation were asked to undertake two tasks. The first was to decide whether the voice quality of the synthesized signals was normal or not. The second was to decide whether the abnormal signals should be described as rough, breathy, or vocal fry. The results showed that signals generated with a small degree of aspiration noise were perceived as breathiness while signals with a small degree of flutter or double pulsing were perceived as roughness. When the flutter or double pulsing increased further, tremor and vocal fry, rather than roughness, were perceived. Furthermore, the amount of aspiration noise, flutter, or double pulsing required for male voice stimuli was different from that required for the female voice stimuli with a similar level of perceptual breathiness and roughness. These findings showed that changes in perceived vocal quality could be achieved by systematic modifications of synthesis parameters. This opens up the possibility of using synthesized voice signals as external standards or "anchors" to improve the reliability of clinical perceptual voice evaluation.

Communication Devices for People with Disabilities↗

Combined feedback-feedforward active noise-reducing headset--the effect of the acoustics on broadband performance.

Active noise-reducing headsets that employ analog feedback control and provide good broadband attenuation are commercially available for a wide range of applications. Recent studies have explored the integration of an adaptive digital feedforward controller with the analog feedback controller to provide additional attenuation of periodic noise components. This paper presents an experimental study of such a combined control system, but with both feedback- and feedforward controllers attenuating broadband noise. Good performance is demonstrated in a reverberant sound field, while under direct sound-field conditions the attenuation performance of the feedforward controller is shown to be dependent on head position. The paper concludes with an analysis of the forward path delay showing how the passive attenuation mechanism improves broadband performance.

Computers↗

Spectral and temporal cues to pitch in noise-excited vocoder simulations of continuous-interleaved-sampling cochlear implants.

Four-band and single-band noise-excited vocoders were used in acoustic simulations to investigate spectral and temporal cues to melodic pitch in the output of a cochlear implant speech processor. Noise carriers were modulated by amplitude envelopes extracted by half-wave rectification and low-pass filtering at 32 or 400 Hz. The four-band, but not the single-band processors, may preserve spectral correlates of fundamental frequency (F0). Envelope smoothing at 400 Hz preserves temporal correlates of F0, which are eliminated with 32-Hz smoothing. Inputs to the processors were sawtooth frequency glides, in which spectral variation is completely determined by F0, or synthetic diphthongal vowel glides, whose spectral shape is dominated by varying formant resonances. Normal listeners labeled the direction of pitch movement of the processed stimuli. For processed sawtooth waves, purely temporal cues led to decreasing performance with increasing F0. With purely spectral cues, performance was above chance despite the limited spectral resolution of the processors. For processed diphthongs, performance with purely spectral cues was at chance, showing that spectral envelope changes due to formant movement obscured spectral cues to F0. Performance with temporal cues was poorer for diphthongs than for sawtooths, with very limited discrimination at higher F0. These data suggest that, for speech signals through a typical cochlear implant processor, spectral cues to pitch are likely to have limited utility, while temporal envelope cues may be useful only at low F0.

Adult↗

The whistles of Hawaiian spinner dolphins.

The characteristics of the whistles of Hawaiian spinner dolphins (Stenella longirostris) are considered by examining concurrently the whistle repertoire (whistle types) and the frequency of occurrence of each whistle type (whistle usage). Whistles were recorded off six islands in the Hawaiian Archipelago. In this study Hawaiian spinner dolphins emitted frequency modulated whistles that often sweep up in frequency (47% of the whistles were upsweeps). The frequency span of the fundamental component was mainly between 2 and 22 kHz (about 94% of the whistles) with an average mid-frequency of 12.9 kHz. The duration of spinner whistles was relatively short, mainly within a span of 0.05 to 1.28 s (about 94% of the whistles) with an average value of 0.49 s. The average maximum frequency of 15.9 kHz obtained by this study is consistent with the body length versus maximum frequency relationship obtained by Wang et al. (1995a) when using spinner dolphin adult body length measurements. When comparing the average values of whistle parameters obtained by this and other studies in the Island of Hawaii, statistically significant differences were found between studies. The reasons for these differences are not obvious. Some possibilities include differences in the upper frequency limit of the recording systems, different spinner groups being recorded, and observer differences in viewing spectrograms. Standardization in recording and analysis procedure is clearly needed.

Animals↗

Quantifying the intelligibility of speech in noise for non-native talkers.

The intelligibility of speech pronounced by non-native talkers is generally lower than speech pronounced by native talkers, especially under adverse conditions, such as high levels of background noise. The effect of foreign accent on speech intelligibility was investigated quantitatively through a series of experiments involving voices of 15 talkers, differing in language background, age of second-language (L2) acquisition and experience with the target language (Dutch). Overall speech intelligibility of L2 talkers in noise is predicted with a reasonable accuracy from accent ratings by native listeners, as well as from the self-ratings for proficiency of L2 talkers. For non-native speech, unlike native speech, the intelligibility of short messages (sentences) cannot be fully predicted by phoneme-based intelligibility tests. Although incorrect recognition of specific phonemes certainly occurs as a result of foreign accent, the effect of reduced phoneme recognition on the intelligibility of sentences may range from severe to virtually absent, depending on (for instance) the speech-to-noise ratio. Objective acoustic-phonetic analyses of accented speech were also carried out, but satisfactory overall predictions of speech intelligibility could not be obtained with relatively simple acoustic-phonetic measures.

Adult↗

Effect of modulator asynchrony of sinusoidal and noise modulators on frequency and amplitude modulation detection interference.

The effect on modulation detection interference (MDI) of timing of gating of the modulation of target and interferer, with synchronously gated carriers, was investigated in three experiments. In a two-interval, two-alternative forced choice adaptive procedure, listeners had to detect 15 Hz sinusoidal amplitude modulation (AM) or frequency modulation (FM) imposed for 200 ms in the temporal center of a 600 ms target sinusoidal carrier. In the first experiment, 15 Hz sinusoidal FM was imposed in phase on both target and interferer carriers. Thresholds were lower for nonoverlapping than for synchronous modulation of target and interferer, but MDI still occurred for the former. Thresholds were significantly higher when the modulators were gated synchronously than when the interferer modulator was gated on before and off after that of the target. This contrasts with the findings of Oxenham and Dau [J. Acoust. Soc. Am. 110, 402-408 (2001)], who reported no effect of modulation asynchrony on AM detection thresholds, using a narrowband noise modulator. Using FM, experiment 2 showed that for temporally overlapping modulation of target and interferer, modulator asynchrony had no significant effect when the interferer was modulated by a narrowband noise. Experiment 3 showed that, for AM, synchronous gating of modulation of the target and interferer produced lower thresholds than asynchronous gating, especially for sinusoidal modulation of the interferer. Results are discussed in terms of specific cues available for periodic modulation, and differences between perceptual grouping on the basis of common AM and FM.

Acoustic Stimulation↗

Within-ear and across-ear interference in a cocktail-party listening task.

Although many researchers have shown that listeners are able to selectively attend to a target speech signal when a masking talker is present in the same ear as the target speech or when a masking talker is present in a different ear than the target speech, little is known about selective auditory attention in tasks with a target talker in one ear and independent masking talkers in both ears at the same time. In this series of experiments, listeners were asked to respond to a target speech signal spoken by one of two competing talkers in their right (target) ear while ignoring a simultaneous masking sound in their left (unattended) ear. When the masking sound in the unattended ear was noise, listeners were able to segregate the competing talkers in the target ear nearly as well as they could with no sound in the unattended ear. When the masking sound in the unattended ear was speech, however, speech segregation in the target ear was substantially worse than with no sound in the unattended ear. When the masking sound in the unattended ear was time-reversed speech, speech segregation was degraded only when the target speech was presented at a lower level than the masking speech in the target ear. These results show that within-ear and across-ear speech segregation are closely related processes that cannot be performed simultaneously when the interfering sound in the unattended ear is qualitatively similar to speech.

Adult↗

Susceptibility to acoustic trauma in young and aged gerbils.

The effect of age on susceptibility to noise-induced hearing loss (NIHL), the effect of gender on the interaction of age-related hearing loss (ARHL) and NIHL, and the relative contributions of ARHL and NIHL to total hearing loss are poorly understood. The issues are difficult to resolve empirically in human subjects because of lack of control over extrinsic variables and for ethical reasons. Accordingly, these issues were examined in a well-studied animal model of both ARHL and NIHL, the Mongolian gerbil. Animals were exposed to an intense tone (3.5 kHz, 113 dB SPL, 1 h) either as young adults (6-8 months) or near the end of the average lifespan of the species (34-38 months). Hearing thresholds were determined with the auditory brainstem response (ABR). ARHL was approximately 5-10 dB, with slightly more observed in males at 16 kHz (p<0.05). NIHL of approximately 15-20 dB was similar for the young and old groups, suggesting no differences in susceptibility as a function of age. There were no gender differences in NIHL. The relative contributions of ARHL and NIHL to total hearing loss in aged, noise-exposed gerbils were predicted by an addition of ARHL and NIHL in dB, similar to an international standard on hearing loss allocation, ISO-1999 [Determination of Occupational Noise Exposure and Estimation of Noise-Induced Hearing Impairment (1990)]. Previous evaluations of ISO-1999 using the gerbil animal model concluded that addition of ARHL and NIHL in dB overpredicts total hearing loss. However, in these studies, ARHL was large and nearly equal to NIHL. In the current study, where ARHL was much less than NIHL, addition of the two factors in dB, as recommended by ISO-1999, results in fairly accurate predictions of total hearing loss.

Age Factors↗

Evidence of upward spread of suppression in DPOAE measurements.

Measurements of DPOAE level in the presence of a suppressor were used to describe a pattern that is qualitatively similar to population studies in the auditory nerve and to behavioral studies of upward spread of masking. DPOAEs were measured in the presence of a suppressor (f3) fixed at either 2.1 or 4.2 kHz, and set to each of seven levels (L3) from 20 to 80 dB SPL. In the presence of a fixed f3 and L3 combination, f2 was varied from about 1 oct below to at least 1/2 oct above f3, while L2 was set to each of 6 values (20-70 dB SPL). L1 was set according to the equation L1 = 0.4L2 + 39 [Janssen et al., J. Acoust. Soc. Am. 103, 3418-3430 (1998)]. At each L2, L1 combination, DPOAE level was measured in a control condition in which no suppressor was presented. Data were converted into decrements (the amount of suppression, in dB) by subtracting the DPOAE level in the presence of each suppressor from the DPOAE level in the corresponding control condition. Plots of DPOAE decrements as a function of f2 showed maximum suppression when f2 approximately = f3. As L3 increased, the suppressive effect spread more towards higher f2 frequencies, with less spread towards lower frequencies relative to f3. DPOAE decrement versus L3 functions had steeper slopes when f2 > f3, compared to the slopes when f2 < f3. These data are consistent with other findings that have shown that response growth for a characteristic place (CP) or frequency (CF) depends on the relation between CP or CF and driver frequency, with steeper slopes when driver frequency is less than CF and shallower slopes when driver frequency is greater than CF. For a fixed amount of suppression (3 dB), L3 and L2 varied nearly linearly for conditions in which f3 approximately = f2, but grew more rapidly for conditions in which f3 < f2, reflecting the basal spread of excitation to the suppressor. The present data are similar in form to the results observed in population studies from the auditory nerve of lower animals and in behavioral masking studies in humans.

Adult↗

A narrow band pattern-matching model of vowel perception.

The purpose of this paper is to propose and evaluate a new model of vowel perception which assumes that vowel identity is recognized by a template-matching process involving the comparison of narrow band input spectra with a set of smoothed spectral-shape templates that are learned through ordinary exposure to speech. In the present simulation of this process, the input spectra are computed over a sufficiently long window to resolve individual harmonics of voiced speech. Prior to template creation and pattern matching, the narrow band spectra are amplitude equalized by a spectrum-level normalization process, and the information-bearing spectral peaks are enhanced by a "flooring" procedure that zeroes out spectral values below a threshold function consisting of a center-weighted running average of spectral amplitudes. Templates for each vowel category are created simply by averaging the narrow band spectra of like vowels spoken by a panel of talkers. In the present implementation, separate templates are used for men, women, and children. The pattern matching is implemented with a simple city-block distance measure given by the sum of the channel-by-channel differences between the narrow band input spectrum (level-equalized and floored) and each vowel template. Spectral movement is taken into account by computing the distance measure at several points throughout the course of the vowel. The input spectrum is assigned to the vowel template that results in the smallest difference accumulated over the sequence of spectral slices. The model was evaluated using a large database consisting of 12 vowels in /hVd/ context spoken by 45 men, 48 women, and 46 children. The narrow band model classified vowels in this database with a degree of accuracy (91.4%) approaching that of human listeners.

Adult↗

Virtual pitch integration for asynchronous harmonics.

This experiment examined the generation of virtual pitch for harmonically related tones that do not overlap in time. The interval between successive tones was systematically varied in order to gauge the integration period for virtual pitch. A pitch discrimination task was employed, and both harmonic and nonharmonic tone series were tested. The results confirmed that a virtual pitch can be generated by a series of brief, harmonically related tones that are separated in time. Robust virtual pitch information can be derived for intervals between successive 40-ms tones of up to about 45 ms, consistent with a minimum estimate of integration period of about 210 ms. Beyond intertone intervals of 45 ms, performance becomes more variable and approaches an upper limit where discrimination of tone sequences can be undertaken on the basis of the individual frequency components. The individual differences observed in this experiment suggest that the ability to derive a salient virtual pitch varies across listeners.

Adult↗

Characterizing cochlear mechano-electric transduction with a nonlinear system identification technique: the influence of the middle ear.

Previously a third-order polynomial equation characterizing mechano-electric transduction was obtained from a nonlinear system identification procedure applied to an ear canal acoustic signal and cochlear microphonic (CM/AC). In this paper, we examine the influence of the linearity and frequency response of the intervening middle ear on the nonlinearity, frequency response, and coherence of the third-order polynomial model of mechano-electric transduction (MET). Ear canal sound pressure (AC), cochlear microphonics (CM), and stapes velocity (SV) were simultaneously recorded from Mongolian gerbils. Linear and nonlinear transfer and coherence functions relating stapes velocity to the acoustic signal (SV/AC), CM to the acoustic signal (CM/AC), and CM to the stapes velocity (CM/SV) were computed. The results showed that SV/AC was linear while CM/AC and CM/SV were not, indicating that the nonlinearity of CM/AC was not due to nonlinearity of the middle ear. The frequency response of the linear term of CM/AC was similar to that of ST/AC but differed from that of CM/SV while the cubic term of CM/AC was similar to that of CM/SV. This indicates that the frequency dependence of CM/AC was due to both the middle ear and frequency dependence of the inner ear. Finally the fit of the polynomial model of MET without the middle ear (CM/SV) did not improve from the fit including the middle ear (CM/AC). A cochlear model of the CM indicated that the lack of improvement was due to the limitations of a third-order polynomial equation characterizing the hair cell transducer function.

Animals↗

Spectro-temporal processing in the envelope-frequency domain.

The frequency selectivity for amplitude modulation applied to tonal carriers and the role of beats between modulators in modulation masking were studied. Beats between the masker and signal modulation as well as intrinsic envelope fluctuations of narrow-band-noise modulators are characterized by fluctuations in the "second-order" envelope (referred to as the "venelope" in the following). In experiment 1, masked threshold patterns (MTPs), representing signal modulation threshold as a function of masker-modulation frequency, were obtained for signal-modulation frequencies of 4, 16, and 64 Hz in the presence of a narrow-band-noise masker modulation, both applied to the same sinusoidal carrier. Carrier frequencies of 1.4, 2.8, and 5.5 kHz were used. The shape and relative bandwidth of the MTPs were found to be independent of the signal-modulation frequency and the carrier frequency. Experiment 2 investigated the extent to which the detection of beats between signal and masker modulation is involved in tone-in-noise (TN), noise-in-tone (NT), and tone-in-tone (TT) modulation masking, whereby the TN condition was similar to the one used in the first experiment. A signal-modulation frequency of 64 Hz, applied to a 2.8-kHz carrier, was tested. Thresholds in the NT condition were always lower than in the TN condition, analogous to the masking effects known from corresponding experiments in the audio-frequency domain. TT masking conditions generally produced the lowest thresholds and were strongly influenced by the detection of beats between the signal and the masker modulation. In experiment 3, TT masked-threshold patterns were obtained in the presence of an additional sinusoidal masker at the beat frequency. Signal-modulation frequencies of 32, 64, and 128 Hz, applied to a 2.8-kHz carrier, were used. It was found that the presence of an additional modulation at the beat frequency hampered the subject's ability to detect the envelope beats and raised thresholds up to a level comparable to that found in the TN condition. The results of the current study suggest that (i) venelope fluctuations play a similar role in modulation masking as envelope fluctuations do in spectral masking, and (ii) envelope and venelope fluctuations are processed by a common mechanism. To interpret the empirical findings, a general model structure for the processing of envelope and venelope fluctuations is proposed.

Adult↗

Empirical refinements applicable to the recording of fish sounds in small tanks.

Many underwater bioacoustical recording experiments (e.g., fish sound production during courtship or agonistic encounters) are usually conducted in a controlled laboratory environment of small-sized tanks. The effects of reverberation, resonance, and tank size on the characteristics of sound recorded inside small tanks have never been fully addressed, although these factors are known to influence the recordings. In this work, 5-cycle tone bursts of 1-kHz sound were used as a test signal to investigate the sound recorded in a 170-l rectangular glass tank at various depths and distances from a transducer. The dominant frequency, sound-pressure level, and power spectrum recorded in small tanks were significantly distorted compared to the original tone bursts. Due to resonance, the dominant frequency varied with water depth, and power spectrum level of the projected frequency decreased exponentially with increased distance between the hydrophone and the sound source; however, the resonant component was nearly uniform throughout the tank. Based on the empirical findings and theoretical calculation, a working protocol is presented that minimizes distortion in fish sound recordings in small tanks. To validate this approach, sounds produced by the croaking gourami (Trichopsis vittata) during staged agonistic encounters were recorded according to the proposed protocol in an 1800-l circular tank and in a 37-l rectangular tank to compare differences in acoustic characteristics associated with tank size and recording position. The findings underscore pitfalls associated with recording fish sounds in small tanks. Herein, an empirical solution to correct these distortions is provided.

Acoustics↗