Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sound Spectrography”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,603 records · Page 89Linked to original sources

Noise duration for a single overflight.

Overflights in national parks and preserves interfere with communication and sounds of nature. The percentage of time that an aircraft is audible, P, can be used as a noise metric. To calculate P the overflight time for a single aircraft, tau, has to be known. The method of tau calculation is based on the assumption that an aircraft is a point source and the noise propagation is governed by geometrical spreading, air absorption, and refraction. The atmosphere is characterized by the effective sound speed gradient. Analytical formulas for tau are derived for down- and crosswind flights.

Acceleration↗

Spectral shape discrimination by hearing-impaired and normal-hearing listeners.

The ability to discriminate between sounds with different spectral shapes was evaluated for normal-hearing and hearing-impaired listeners. Listeners discriminated between a standard stimulus and a signal stimulus in which half of the standard components were decreased in level and half were increased in level. In one condition, the standard stimulus was the sum of six equal-amplitude tones (equal-SPL), and in another the standard stimulus was the sum of six tones at equal sensation levels re: audiometric thresholds for individual subjects (equal-SL). Spectral weights were estimated in conditions where the amplitudes of the individual tones were perturbed slightly on every presentation. Sensitivity was similar in all conditions for normal-hearing and hearing-impaired listeners. The presence of perturbation and equal-SL components increased thresholds for both groups, but only small differences in weighting strategy were measured between the groups depending on whether the equal-SPL or equal-SL condition was tested. The average data suggest that normal-hearing listeners may rely more on the central components of the spectrum whereas hearing-impaired listeners may have been more likely to use the edges. However, individual weighting functions were quite variable, especially for the HI listeners, perhaps reflecting difficulty in processing changes in spectral shape due to hearing loss. Differences in weighting strategy without changes in sensitivity suggest that factors other than spectral weights, such as internal noise or difficulty encoding a reference stimulus, also may dominate performance.

Adult↗

The role of contrasting temporal amplitude patterns in the perception of speech.

Despite a lack of traditional speech features, novel sentences restricted to a narrow spectral slit can retain nearly perfect intelligibility [R. M. Warren et al., Percept. Psychophys. 57, 175-182 (1995)]. The current study employed 514 listeners to elucidate the cues allowing this high intelligibility, and to examine generally the use of narrow-band temporal speech patterns. When 1/3-octave sentences were processed to preserve the overall temporal pattern of amplitude fluctuation, but eliminate contrasting amplitude patterns within the band, sentence intelligibility dropped from values near 100% to values near zero (experiment 1). However, when a 1/3-octave speech band was partitioned to create a contrasting pair of independently amplitude-modulated 1/6-octave patterns, some intelligibility was restored (experiment 2). An additional experiment (3) showed that temporal patterns can also be integrated across wide frequency separations, or across the two ears. Despite the linguistic content of single temporal patterns, open-set intelligibility does not occur. Instead, a contrast between at least two temporal patterns is required for the comprehension of novel sentences and their component words. These contrasting patterns can reside together within a narrow range of frequencies, or they can be integrated across frequencies or ears. This view of speech perception, in which across-frequency changes in energy are seen as systematic changes in the temporal fluctuation patterns at two or more fixed loci, is more in line with the physiological encoding of complex signals.

Adolescent↗

Time-frequency model for echo-delay resolution in wideband biosonar.

A time/frequency model of the bat's auditory system was developed to examine the basis for the fine (approximately 2 micros) echo-delay resolution of big brown bats (Eptesicus fuscus), and its performance at resolving closely spaced FM sonar echoes in the bat's 20-100-kHz band at different signal-to-noise ratios was computed. The model uses parallel bandpass filters spaced over this band to generate envelopes that individually can have much lower bandwidth than the bat's ultrasonic sonar sounds and still achieve fine delay resolution. Because fine delay separations are inside the integration time of the model's filters (approximately 250-300 micros), resolving them means using interference patterns along the frequency dimension (spectral peaks and notches). The low bandwidth content of the filter outputs is suitable for relay of information to higher auditory areas that have intrinsically poor temporal response properties. If implemented in fully parallel analog-digital hardware, the model is computationally extremely efficient and would improve resolution in military and industrial sonar receivers.

Animals↗

Recovery from prior stimulation: masking of speech by interrupted noise for younger and older adults with normal hearing.

In a previous study [Dubno et al, J. Acoust. Soc. Am. 111, 2897-2907 (2002)], older subjects benefitted less than younger subjects from momentary improvements in signal-to-noise ratio when listening to speech in interrupted maskers. It has been hypothesized that the benefit derived from interrupted maskers may be related to recovery from forward masking, i.e., the recovery of a response to a suprathreshold signal from prior stimulation by a masker. The effect of interrupted maskers on speech recognition may be well suited to test hypotheses regarding recovery from prior stimulation, given that both involve the perception of signals following a masker. Here, younger and older adults with normal but not identical audiograms listened to nonsense syllables at moderate and high levels in a speech-shaped noise that was modulated by a 2-, 10-, 25-, or 50-Hz square wave. An additional low-level noise was always present that was shaped to produce equivalent masked thresholds for all subjects. To assess recovery from forward masking, forward-masked thresholds were measured at 0.5 and 4.0 kHz as a function of the delay between the speech-shaped masker and the signal. Speech recognition in interrupted noise was poorer for older than younger subjects. Small but consistent age-related differences were observed in the decrease in score with interrupted noise relative to the score without interrupted noise. Forward-masked thresholds of older subjects were higher than those of younger subjects, but there were no age-related differences in the amount of forward masking or in simultaneous masking. Negative correlations were observed between speech-recognition scores in interrupted noise and forward-masked thresholds. That is, the benefit derived from momentary improvements in speech audibility in an interrupted noise decreased as forward-masked thresholds increased. Stronger correlations with forward masking were observed for the higher frequency signal, for higher noise interruption rates, and when the signal-to-noise ratio was poor. Comparisons of speech-recognition scores at moderate and high levels for younger and older subjects were not consistent with the hypothesis of an age-related difference in the contribution of low-spontaneous-rate fibers to speech recognition in interrupted noise.

Adult↗

Modulation masking in cochlear implant listeners: envelope versus tonotopic components.

It is hypothesized that channel-interaction in cochlear implant listeners as measured in a modulation-masking experiment would be influenced by both the tonotopic overlap between masker and signal as well as an interaction between their envelopes. Two experiments were conducted to measure the effects of maskers with noisy and steady-state envelopes on modulation detection by four adult Nucleus-22 cochlear implant listeners, as a function of tonotopic distance between the masker and the signal. In the first experiment, we measured detection thresholds for a 50-Hz modulation in the envelope of a 500-Hz carrier pulse train, in the presence of a masker stimulating regions basal and apical to, as well as overlapping with, the signal. The maskers had two kinds of envelopes: (i) amplitude-modulated by flat-spectrum noise (NAM) and (ii) steady-state (SS(peak)) at a level corresponding to the maximum of the noise fluctuation range. In general, modulation thresholds obtained in the presence of the NAM maskers significantly exceeded thresholds obtained with the corresponding SS(peak) maskers. The ratio p of the threshold modulation depth m obtained with the NAM masker to that obtained with the SS(peak) masker was defined as a conservative index of "envelope masking." In the second experiment, p was determined for two different tasks: the detection of modulation at 20 Hz and steady-state intensity increment detection. Compared to the 50-Hz modulation detection results, the ratio p was reduced for the 20-Hz modulation detection task and even more so for the steady-state increment detection task. It is concluded that channel-interaction can be significantly increased in cochlear implant listeners when dynamic stimuli are used in place of steady-state stimuli.

Adult↗

Modal analysis of a violin octet.

Experimental modal analysis of a complete Hutchins-Schelleng violin octet, combined with cavity mode analysis and room-averaged acoustic analysis, gives a highly detailed characterization of the dynamics for this historic group of instruments. All the "signature" modes in the open string pitch region--cavity modes A0 ("main air") and A1 (lowest longitudinal), C-bout "rhomboid," the first corpus bending modes B1- and B1+ (comprising the "main wood")--were observed across the octet. A0 was always the lowest dominant radiator, below all corpus modes. A1 contributed significant acoustic output only for larger instruments, but was the dominant contributor for the large bass in the "main wood" region. Acoustic results indicate either B1- or B1+ can be the major radiator. Damping results indicate that B1 modes overall radiate approximately 28% of their vibrational energy. "Doublet" B1 modes from substructure couplings were observed for three instruments. "A0-B0" coupling was not significant for the largest instruments. The original flat-plate-based scaling of the "main wood" resonance was generally successful across the octet, although that for the "main air" was not.

Equipment Design↗

Diversity in noise-induced temporary hearing loss in otophysine fishes.

The effects of intense white noise (158 dB re 1 microPa for 12 and 24 h) on the hearing abilities of two otophysine fish species--the nonvocal goldfish Carassius auramus and the vocalizing catfish Pimelodus pictus--were investigated in relation to noise exposure duration. Hearing sensitivity was determined utilizing the auditory brainstem response (ABR) recording technique. Measurements in the frequency range between 0.2 and 4.0 kHz were conducted prior and directly after noise exposure as well as after 3, 7, and 14 days of recovery. Both species showed a significant loss of sensitivity (up to 26 dB in C. auratus and 32 dB in P. pictus) immediately after noise exposure, with the greatest hearing loss in the range of their most sensitive frequencies. Hearing loss differed between both species, and was more pronounced in the catfish. Exposure duration had no influence on hearing loss. Hearing thresholds of C. auratus recovered within three days, whereas those of P. pictus only returned to their initial values within 14 days after exposure in all but one frequency. The results indicate that hearing specialists are affected differently by noise exposure and that acoustic communication might be restricted in noisy habitats.

Acoustic Stimulation↗

Segmental intelligibility of four currently used text-to-speech synthesis methods.

The study investigated the segmental intelligibility of four currently available text-to-speech (TTS) products under 0-dB and 5-dB signal-to-noise ratios. The products were IBM ViaVoice version 5.1, which uses formant coding, Festival version 1.4.2, a diphone-based LPC TTS product, AT&T Next-Gen, a half-phone-based TTS product that uses harmonic-plus-noise method for synthesis, and FlexVoice2, a hybrid TTS product that combines concatenative and formant coding techniques. Overall, concatenative techniques were more intelligible than formant or hybrid techniques, with formant coding slightly better at modeling vowels and concatenative techniques marginally better at synthesizing consonants. No TTS product was better at resisting noise interference than others, although all were more intelligible at 5 dB than at 0-dB SNR. The better TTS products in this study were, on the average, 22% less intelligible and had about 3 times more phoneme errors than human voice under comparable listening conditions. The hybrid TTS technology of FlexVoice had the lowest intelligibility and highest error rates. There were discernible patterns of errors for stops, fricatives, and nasals. Unrestricted TTS output--e-mail messages, news reports, and so on--under high noise conditions prevalent in automobiles, airports, etc. will likely challenge the listeners.

Adolescent↗

Speech recognition under conditions of frequency-place compression and expansion.

In normal acoustic hearing the mapping of acoustic frequency information onto the appropriate cochlear place is a natural biological function, but in cochlear implants it is controlled by the speech processor. The cochlear tonotopic range of the implant is determined by the length and insertion depth of the electrode array. Conventional cochlear implant electrode arrays are designed for an insertion of 25 mm inside the round window and the active electrodes occupy 16 mm, which would place the electrodes in a cochlear region corresponding to an acoustic frequency range of 500-6000 Hz. However, some implant speech processors map an acoustic frequency range from 150 to 10000 Hz onto these electrodes. While this mapping preserves the entire range of acoustic frequency information, it also results in a compression of the tonotopic pattern of speech information delivered to the brain. The present study measured the effects of such a compression of frequency-to-place mapping on speech recognition using acoustic simulations. Also measured were the effects of an expansion of the frequency-to-place mapping, which produces an expanded representation of speech in the cochlea. Such an expanded representation might improve speech recognition by improving the relative spatial (tonotopic) resolution, like an "acoustic fovea." Phoneme and sentence recognition was measured as a function of linear (in terms of cochlear distance) frequency-place compression and expansion. These conditions were presented to normal-hearing listeners using a noise-band vocoder, simulating cochlear implant electrodes with different insertion depths and different number of electrode channels. The cochlear tonotopic range was held constant by employing the same noise carrier bands for each condition, while the analysis frequency range was either compressed or expanded relative to the carrier frequency range. For each condition, the result was compared to that of the perfect frequency-place match, where the carrier and the analysis bands were perfectly matched. Speech recognition in the matched conditions was generally better than any condition of frequency-place expansion and compression, even when the matched condition eliminated a considerable amount of acoustic information. This result suggests that speech recognition, at least without training, is dependent on the mapping of acoustic frequency information onto the appropriate cochlear place.

Adult↗

Speech transmission index from running speech: a neural network approach.

Speech transmission index (STI) is an important objective parameter concerning speech intelligibility for sound transmission channels. It is normally measured with specific test signals to ensure high accuracy and good repeatability. Measurement with running speech was previously proposed, but accuracy is compromised and hence applications limited. A new approach that uses artificial neural networks to accurately extract the STI from received running speech is developed in this paper. Neural networks are trained on a large set of transmitted speech examples with prior knowledge of the transmission channels' STIs. The networks perform complicated nonlinear function mappings and spectral feature memorization to enable accurate objective parameter extraction from transmitted speech. Validations via simulations demonstrate the feasibility of this new method on a one-net-one-speech extract basis. In this case, accuracy is comparable with normal measurement methods. This provides an alternative to standard measurement techniques, and it is intended that the neural network method can facilitate occupied room acoustic measurements.

Humans↗

A practical method of predicting the loudness of complex electrical stimuli.

The output of speech processors for multiple-electrode cochlear implants consists of current waveforms with complex temporal and spatial patterns. The majority of existing processors output sequential biphasic current pulses. This paper describes a practical method of calculating loudness estimates for such stimuli, in addition to the relative loudness contributions from different cochlear regions. The method can be used either to manipulate the loudness or levels in existing processing strategies, or to control intensity cues in novel sound processing strategies. The method is based on a loudness model described by McKay et al [J. Acoust. Soc. Am. 110, 1514-1524 (2001)] with the addition of the simplifying approximation that current pulses falling within a temporal integration window of several milliseconds' duration contribute independently to the overall loudness of the stimulus. Three experiments were carried out with six implantees who use the CI24M device manufactured by Cochlear Ltd. The first experiment validated the simplifying assumption, and allowed loudness growth functions to be calculated for use in the loudness prediction method. The following experiments confirmed the accuracy of the method using multiple-electrode stimuli with various patterns of electrode locations and current levels.

Adult↗

Variation in chick-a-dee calls of a Carolina chickadee population, Poecile carolinensis: identity and redundancy within note types.

Chick-a-dee calls of chickadee species are structurally complex because calls possess a rudimentary syntax governing the ordering of their different note types. Chick-a-dee calls were recorded in an aviary from female and male birds from two field sites. This paper reports sources of variation of acoustical parameters of notes in these calls. There were significant sex and microgeographic differences in some of the measured parameters of the notes in the calls. In addition, the syntax of the call itself influenced characteristics of each of the notes. For example, calls with many introductory notes began with a note of higher frequency and longer duration, relative to calls with few introductory notes. Furthermore, the number of introductory notes influenced frequency and duration components of notes later in the call. Thus, single notes are predictive of the note composition of the signaler's call. This suggests that a receiver might gain the meaning in the call even if it hears only part of the call. Further, single notes within these complex calls can contain information enabling receivers to predict the sex of the signaler, and whether it is from the local population.

Animal Communication↗

Investigations of the precedence effect in budgerigars: the perceived location of auditory images.

The perceived location of auditory images has been recently studied in budgerigars [Dent and Dooling, J. Acoust. Soc. Am. 113, 2146-2158 (2003)]. Those results suggested that budgerigars (Melopsittacus undulatus) perceive precedence effect stimuli in a manner similar to humans and other animals. Here we extend those experiments to include the effects of intensity on the perceived location of auditory images and the perceived location of paired stimuli from multiple locations in space. We measured the abilities of budgerigars to discriminate between paired stimuli separated in time, intensity, and/or location. Increasing the intensity of a lag stimulus disrupted localization dominance. Budgerigars also perceived simultaneously presented (away from the midline) stimuli as very similar to a single sound presented from the midline, much like the phantom image reported in humans. The perception of paired stimuli from one side of the head versus two sides of the head was also examined and showed that the spatial cues available in these stimuli are important and that echoes are not perceptually inaccessible during localization dominance conditions. The results from these experiments add further data showing the precedence effect in budgerigars is similar to that found in humans and other animals.

Acoustic Stimulation↗

Transient emission suppression tuning curve attributes in relation to psychoacoustic threshold.

Ipsilateral suppression characteristics of transiently evoked otoacoustic emissions (TEOAEs) are described in relation to psychoacoustic threshold at 4000 Hz and the presence or absence of spontaneous otoacoustic emissions in 41 adults with normal hearing. TEOAE amplitudes were measured in response to 4000-Hz tonebursts presented in linear blocks at 40 and 50 dB SPL while puretone suppressors were introduced at a variety of frequencies and levels ipsilateral to and simultaneously with the tonebursts. Suppressors close to the toneburst frequency were most effective in decreasing the amplitude of the TEOAEs, while those more remote in frequency required significantly greater intensity for a similar amount of suppression. Consequently, characteristic tuning curve shapes were obtained. Tuning-curve tip levels were closely associated with the level of the toneburst and tip frequencies occurred at or above the toneburst frequency. Tuning-curve widths (Q10), however, varied significantly across subjects with similar psychoacoustic thresholds in quiet determined by a two-alternative forced-choice method. The results suggest that a portion of that variability may be explained by the presence or absence of spontaneous otoacoustic emissions in an individual ear.

Acoustic Stimulation↗

Amplitude and phase of distortion product otoacoustic emissions in the guinea pig in an (f1 ,f2) area study.

Lower sideband distortion product otoacoustic emissions (DPOAEs), measured in the ear canal upon stimulation with two continuous pure tones, are the result of interfering contributions from two different mechanisms, the nonlinear distortion component and the linear reflection component. The two contributors have been shown to have a different amplitude and, in particular, a different phase behavior as a function of the stimulus frequencies. The dominance of either component was investigated in an extensive (f1 ,f2) area study of DPOAE amplitude and phase in the guinea pig, which allows for both qualitative and quantitative analysis of isophase contours. Making a minimum of additional assumptions, simple relations between the direction of constant phase in the (f ,f2) plane and the group delays in f1-sweep, f2-sweep, and fixed f2/f1 paradigms can be derived, both for distortion (wave-fixed) and reflection (place-fixed) components. The experimental data indicate the presence of both components in the lower sideband DPOAEs, with the reflection component as the dominant contributor for low f2/f1 ratios and the distortion component for intermediate ratios. At high ratios the behavior cannot be explained by dominance of either component.

Acoustic Stimulation↗

Functional differences between vowel onsets and offsets in temporal perception of speech: local-change detection and speaking-rate discrimination.

To provide a perceptual framework for the objective evaluation of durational rules in speech synthesis, two experiments were conducted to investigate the differences between vowel (V) onsets and V-offsets in their functions of marking the perceived temporal structure of speech. The first experiment measured the detectability of temporal modifications given in four-mora (CVCVCVCV) Japanese words. In the V-onset condition, the inter-onset intervals of vowels were uniformly changed (either expanded or reduced) while their inter-offset intervals were preserved. In the V-offset condition, this was reversed. These manipulations did not change the duration of the entire word. Each of the modified words was paired with its unmodified counterpart, and the pair was given to listeners, who were asked to rate the difference between the paired words. The results show that there were no significant differences in the listeners' abilities to detect the temporal modification between the V-onset and V-offset conditions. In the second experiment, the listeners were asked to estimate the differences they perceived in speaking rates for the same stimulus set as that of the first experiment. Interestingly, the results show a clear difference in the listeners' performance between the V-onset and V-offset conditions. Specifically, changing the V-onset intervals changed the perceived speaking rates, which showed a linear relation (r = -0.9) despite the fact that the duration of the entire word remained unchanged. In contrast, modifying the V-offset intervals produced no clear relation with the perceived speaking rates. The second experiment also showed that the listeners performed well in speaking rate discrimination (3.5%-5% in the change ratio). These results are discussed in relation to the differences in the listeners' temporal processing range (local or global) between the two experiments.

Adult↗