Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sound Spectrography”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,621 records · Page 90Linked to original sources

Acoustic correlates of caller identity and affect intensity in the vowel-like grunt vocalizations of baboons.

Comparative, production-based research on animal vocalizations can allow assessments of continuity in vocal communication processes across species, including humans, and may aid in the development of general frameworks relating specific constitutional attributes of callers to acoustic-structural details of their vocal output. Analyses were undertaken on vowel-like baboon grunts to examine variation attributable to caller identity and the intensity of the affective state underlying call production. Six hundred six grunts from eight adult females were analyzed. Grunts derived from 128 bouts of calling in two behavioral contexts: concerted group movements and social interactions involving mothers and their young infants. Each context was subdivided into a high- and low-arousal condition. Thirteen acoustic features variously predicted to reflect variation in either caller identity or arousal intensity were measured for each grunt bout, including tempo-, source- and filter-related features. Grunt bouts were highly individually distinctive, differing in a variety of acoustic dimensions but with some indication that filter-related features contributed disproportionately to individual distinctiveness. In contrast, variation according to arousal condition was associated primarily with tempo- and source-related features, many matching those identified as vehicles of affect expression in other nonhuman primate species and in human speech and other nonverbal vocal signals.

Affect↗

Patterns in the vocalizations of male harbor seals.

Comparative analyses of the roar vocalization of male harbor seals from ten sites throughout their distribution showed that vocal variation occurs at the oceanic, regional, population, and subpopulation level. Genetic barriers based on the physical distance between harbor seal populations present a likely explanation for some of the observed vocal variation. However, site-specific vocal variations were present between genetically mixed subpopulations in California. A tree-based classification analysis grouped Scottish populations together with eastern Pacific sites, rather than amongst Atlantic sites as would be expected if variation was based purely on genetics. Lastly, within the classification tree no individual vocal parameter was consistently responsible for consecutive splits between geographic sites. Combined, these factors suggest that site-specific variation influences the development of vocal structure in harbor seals and these factors may provide evidence for the occurrence of vocal dialects.

Animal Communication↗

A nonlinear filter-bank model of the guinea-pig cochlear nerve: rate responses.

The aim of this study is to produce a functional model of the auditory nerve (AN) response of the guinea-pig that reproduces a wide range of important responses to auditory stimulation. The model is intended for use as an input to larger scale models of auditory processing in the brain-stem. A dual-resonance nonlinear filter architecture is used to reproduce the mechanical tuning of the cochlea. Transduction to the activity on the AN is accomplished with a recently proposed model of the inner-hair-cell. Together, these models have been shown to be able to reproduce the response of high-, medium-, and low-spontaneous rate fibers from the guinea-pig AN at high best frequencies (BFs). In this study we generate parameters that allow us to fit the AN model to data from a wide range of BFs. By varying the characteristics of the mechanical filtering as a function of the BF it was possible to reproduce the BF dependence of frequency-threshold tuning curves, AN rate-intensity functions at and away from BF, compression of the basilar membrane at BF as inferred from AN responses, and AN iso-intensity functions. The model is a convenient computational tool for the simulation of the range of nonlinear tuning and rate-responses found across the length of the guinea-pig cochlear nerve.

Animals↗

Enhancing interaural-delay-based extents of laterality at high frequencies by using "transposed stimuli".

An acoustic pointing task was used to determine whether interaural temporal disparities (ITDs) conveyed by high-frequency "transposed" stimuli would produce larger extents of laterality than ITDs conveyed by bands of high-frequency Gaussian noise. The envelopes of transposed stimuli are designed to provide high-frequency channels with information similar to that conveyed by the waveforms of low-frequency stimuli. Lateralization was measured for low-frequency Gaussian noises, the same noises transposed to 4 kHz, and high-frequency Gaussian bands of noise centered at 4 kHz. Extents of laterality obtained with the transposed stimuli were greater than those obtained with bands of Gaussian noise centered at 4 kHz and, in some cases, were equivalent to those obtained with low-frequency stimuli. In a second experiment, the general effects on lateral position produced by imposed combinations of bandwidth, ITD, and interaural phase disparities (IPDs) on low-frequency stimuli remained when those stimuli were transposed to 4 kHz. Overall, the data were fairly well accounted for by a model that computes the cross-correlation subsequent to known stages of peripheral auditory processing augmented by low-pass filtering of the envelopes within the high-frequency channels of each ear.

Adult↗

Temporary threshold shifts and recovery following noise exposure in the Atlantic bottlenosed dolphin (Tursiops truncatus).

Behaviorally determined hearing thresholds for a 7.5-kHz tone for an Atlantic bottlenosed dolphin (Tursiops truncatus) were obtained following exposure to fatiguing low-frequency octave band noise. The fatiguing stimulus ranged from 4 to 11 kHz and was gradually increased in intensity to 179 dB re 1 microPa and in duration to 55 min. Exposures occurred no more frequently than once per week. Measured temporary threshold shifts averaged 11 dB. Threshold determination took at least 20 min. Recovery was examined 360, 180, 90, and 45 min following exposure and was essentially complete within 45 min.

Acoustic Stimulation↗

On the importance of early reflections for speech in rooms.

This paper presents the results of new studies based on speech intelligibility tests in simulated sound fields and analyses of impulse response measurements in rooms used for speech communication. The speech intelligibility test results confirm the importance of early reflections for achieving good conditions for speech in rooms. The addition of early reflections increased the effective signal-to-noise ratio and related speech intelligibility scores for both impaired and nonimpaired listeners. The new results also show that for common conditions where the direct sound is reduced, it is only possible to understand speech because of the presence of early reflections. Analyses of measured impulse responses in rooms intended for speech show that early reflections can increase the effective signal-to-noise ratio by up to 9 dB. A room acoustics computer model is used to demonstrate that the relative importance of early reflections can be influenced by the room acoustics design.

Adult↗

Perceptual weights in auditory level discrimination.

Perceptual weights in level discrimination (also called intensity discrimination) were determined for 3-, 7-, 15-, and 24-component tone complexes with flat spectral envelopes using a correlational paradigm. Each frequency component was randomly and independently perturbed in level oneach presentation. For the target interval, frequency-component levels were additionally increased by the level increment to be detected, deltaL [= 201og10((p + deltap)/p), where p is pressure]. Weights were calculated from the across-trial correlation between the level perturbations for each frequency component and the interval chosen by the listener. Two conditions were investigated: (1) deltaL was equal across frequency components, and (2) deltaL increased progressively across frequency components. For both conditions, data for four listeners usually showed the greatest weight for the highest frequency component. The two-to-four highest frequency components generally were most important for level discrimination. The effect of increasing deltaL progressively with frequency was small and inconsistent. Additional measurements showed that flanking noise maskers designed to mask spread of excitation caused only small and generally unsystematic changes to the weights. Overall, these results indicate that listeners combine information across a wide range of auditory channels to arrive at a decision for level discrimination, but the weighting of channels appears to be suboptimal.

Adult↗

Children's detection of pure-tone signals: informational masking with contralateral maskers.

When normal-hearing adults and children are required to detect a 1000-Hz tone in a random-frequency multitone masker, masking is often observed in excess of that predicted by traditional auditory filter models. The excess masking is called informational masking. Though individual differences in the effect are large, the amount of informational masking is typically much greater in young children than in adults [Oh et al., J. Acoust. Soc. Am. 109, 2888-2895 (2001)]. One factor that reduces informational masking in adults is spatial separation of the target tone and masker. The present study was undertaken to determine whether or not a similar effect of spatial separation is observed in children. An extreme case of spatial separation was used in which the target tone was presented to one ear and the random multitone masker to the other ear. This condition resulted in nearly complete elimination of masking in adults. In young children, however, presenting the masker to the nontarget ear typically produced only a slight decrease in overall masking and no change in informational masking. The results for children are interpreted in terms of a model that gives equal weight to the auditory filter outputs from each ear.

Adolescent↗

Effects of speaking rate on second formant trajectories of selected vocalic nuclei.

The effect of speaking rate variations on second formant (F2) trajectories was investigated for a continuum of rates. F2 trajectories for the schwa preceding a voiced bilabial stop, and one of three target vocalic nuclei following the stop, were generated for utterances of the form "Put a bV here, where V was /i/,/ae/ or /oI/. Discrete spectral measures at the vowel-consonant and consonant-vowel interfaces, as well as vowel target values, were examined as potential parameters of rate variation; several different whole-trajectory analyses were also explored. Results suggested that a discrete measure at the vowel consonant (schwa-consonant) interface, the F2off value, was in many cases a good index of rate variation, provided the rates were not unusually slow (vowel durations less than 200 ms). The relationship of the spectral measure at the consonant-vowel interface, F2 onset, as well as that of the "target" for this vowel, was less clearly related to rate variation. Whole-trajectory analyses indicated that the rate effect cannot be captured by linear compressions and expansions of some prototype trajectory. Moreover, the effect of rate manipulation on formant trajectories interacts with speaker and vocalic nucleus type, making it difficult to specify general rules for these effects. However, there is evidence that a small number of speaker strategies may emerge from a careful qualitative and quantitative analysis of whole formant trajectories. Results are discussed in terms of models of speech production and a group of speech disorders that is usually associated with anomalies of speaking rate, and hence of formant frequency trajectories.

Adult↗

Pitch discrimination of diotic and dichotic tone complexes: harmonic resolvability or harmonic number?

Three experiments investigated the relationship between harmonic number, harmonic resolvability, and the perception of harmonic complexes. Complexes with successive equal-amplitude sine- or random-phase harmonic components of a 100- or 200-Hz fundamental frequency (f0) were presented dichotically, with even and odd components to opposite ears, or diotically, with all harmonics presented to both ears. Experiment 1 measured performance in discriminating a 3.5%-5% frequency difference between a component of a harmonic complex and a pure tone in isolation. Listeners achieved at least 75% correct for approximately the first 10 and 20 individual harmonics in the diotic and dichotic conditions, respectively, verifying that only processes before the binaural combination of information limit frequency selectivity. Experiment 2 measured fundamental frequency difference limens (f0 DLs) as a function of the average lowest harmonic number. Similar results at both f0's provide further evidence that harmonic number, not absolute frequency, underlies the order-of-magnitude increase observed in f0 DLs when only harmonics above about the 10th are presented. Similar results under diotic and dichotic conditions indicate that the auditory system, in performing f0 discrimination, is unable to utilize the additional peripherally resolved harmonics in the dichotic case. In experiment 3, dichotic complexes containing harmonics below the 12th, or only above the 15th, elicited pitches of the f0 and twice the f0, respectively. Together, experiments 2 and 3 suggest that harmonic number, regardless of peripheral resolvability, governs the transition between two different pitch percepts, one based on the frequencies of individual resolved harmonics and the other based on the periodicity of the temporal envelope.

Adolescent↗

Variation in humpback whale (Megaptera novaeangliae) song length in relation to low-frequency sound broadcasts.

Humpback whale song lengths were measured from recordings made off the west coast of the island of Hawai'i in March 1998 in relation to acoustic broadcasts ("pings") from the U.S. Navy SURTASS Low Frequency Active sonar system. Generalized additive models were used to investigate the relationships between song length and time of year, time of day, and broadcast factors. There were significant seasonal and diurnal effects. The seasonal factor was associated with changes in the density of whales sighted near shore. The diurnal factor was associated with changes in surface social activity. Songs that ended within a few minutes of the most recent ping tended to be longer than songs sung during control periods. Many songs that were overlapped by pings, and songs that ended several minutes after the most recent ping, did not differ from songs sung in control periods. The longest songs were sung between 1 and 2 h after the last ping. Humpbacks responded to louder broadcasts with longer songs. The fraction of variation in song length that could be attributed to broadcast factors was low. Much of the variation in humpback song length remains unexplained.

Animal Communication↗

Whole-lung resonance in a bottlenose dolphin (Tursiops truncatus) and white whale (Delphinapterus leucas).

An acoustic backscatter technique was used to estimate in vivo whole-lung resonant frequencies in a bottlenose dolphin (Tursiops truncatus) and white whale (Delphinapterus leucas). Subjects were trained to submerge and position themselves near an underwater sound projector and a receiving hydrophone. Acoustic pressure measurements were made near the thorax while the subject was insonified with pure tones at frequencies from 16 to 100 Hz. Whole-lung resonant frequencies were estimated by comparing pressures measured near the subject's thorax to those measured from the same location without the subject present. Experimentally measured resonant frequencies for the white whale and dolphin lungs were 30 and 36 Hz, respectively. These values were significantly higher than those predicted using a free-spherical air bubble model. Experimentally measured damping ratios and quality factors at resonance were 0.20 and 2.5, respectively, for the white whale, and 0.16 and 3.1, respectively, for the dolphin.

Acoustic Stimulation↗

Mammalian spontaneous otoacoustic emissions are amplitude-stabilized cochlear standing waves.

Mammalian spontaneous otoacoustic emissions (SOAEs) have been suggested to arise by three different mechanisms. The local-oscillator model, dating back to the work of Thomas Gold, supposes that SOAEs arise through the local, autonomous oscillation of some cellular constituent of the organ of Corti (e.g., the "active process" underlying the cochlear amplifier). Two other models, by contrast, both suppose that SOAEs are a global collective phenomenon--cochlear standing waves created by multiple internal reflection--but differ on the nature of the proposed power source: Whereas the "passive" standing-wave model supposes that SOAEs are biological noise, passively amplified by cochlear standing-wave resonances acting as narrow-band nonlinear filters, the "active" standing-wave model supposes that standing-wave amplitudes are actively maintained by coherent wave amplification within the cochlea. Quantitative tests of key predictions that distinguish the local-oscillator and global standing-wave models are presented and shown to support the global standing-wave model. In addition to predicting the existence of multiple emissions with a characteristic minimum frequency spacing, the global standing-wave model accurately predicts the mean value of this spacing, its standard deviation, and its power-law dependence on SOAE frequency. Furthermore, the global standing-wave model accounts for the magnitude, sign, and frequency dependence of changes in SOAE frequency that result from modulations in middle-ear stiffness. Although some of these SOAE characteristics may be replicable through artful ad hoc adjustment of local-oscillator models, they all arise quite naturally in the standing-wave framework. Finally, the statistics of SOAE time waveforms demonstrate that SOAEs are coherent, amplitude-stabilized signals, as predicted by the active standing-wave model. Taken together, the results imply that SOAEs are amplitude-stabilized standing waves produced by the cochlea acting as a biological, hydromechanical analog of a laser oscillator. Contrary to recent claims, spontaneous emission of sound from the ear does not require the autonomous mechanical oscillation of its cellular constituents.

Animals↗

Distortion product otoacoustic emission suppression tuning curves in normal-hearing and hearing-impaired human ears.

Distortion product otoacoustic emission (DPOAE) suppression measurements were made in 20 subjects with normal hearing and 21 subjects with mild-to-moderate hearing loss. The probe consisted of two primary tones (f2, f1), with f2 held constant at 4 kHz and f2/f1 = 1.22. Primary levels (L1, L2) were set according to the equation L1 = 0.4 L2 + 39 dB [Kummer et al., J. Acoust. Soc. Am. 103, 3431-3444 (1998)], with L2 ranging from 20 to 70 dB SPL (normal-hearing subjects) and 50-70 dB SPL (subjects with hearing loss). Responses elicited by the probe were suppressed by a third tone (f3), varying in frequency from 1 octave below to 1/2 octave above f2. Suppressor level (L3) varied from 5 to 85 dB SPL. Responses in the presence of the suppressor were subtracted from the unsuppressed condition in order to convert the data into decrements (amount of suppression). The slopes of the decrement versus L3 functions were less steep for lower frequency suppressors and more steep for higher frequency suppressors in impaired ears. Suppression tuning curves, constructed by selecting the L3 that resulted in 3 dB of suppression as a function of f3, resulted in tuning curves that were similar in appearance for normal and impaired ears. Although variable, Q10 and Q(ERB) were slightly larger in impaired ears regardless of whether the comparisons were made at equivalent SPL or equivalent sensation levels (SL). Larger tip-to-tail differences were observed in ears with normal hearing when compared at either the same SPL or the same SL, with a much larger effect at similar SL. These results are consistent with the view that subjects with normal hearing and mild-to-moderate hearing loss have similar tuning around a frequency for which the hearing loss exists, but reduced cochlear-amplifier gain.

Adolescent↗

Pitch of amplitude-modulated irregular-rate stimuli in acoustic and electric hearing.

The pitch of stimuli was studied under conditions where place-of-excitation was held constant, and where pitch was therefore derived from "purely temporal" cues. In experiment 1, the acoustical and electrical pulse trains consisted of pulses whose amplitudes alternated between a high and a low value, and whose interpulse intervals alternated between 4 and 6 ms. The attenuated pulses occurred after the 4-ms intervals in condition A, and after the 6-ms intervals in condition B. For both normal-hearing subjects and cochlear implantees, the period of an isochronous pulse train equal in pitch to this "4-6" stimulus increased from near 6 ms at the smallest modulation depth to nearly 10 ms at the largest depth. Additionally, the modulated pulse trains in condition A were perceived as being lower in pitch than those in condition B. Data are interpreted in terms of increased refractoriness in condition A, where the larger pulses are more closely followed by the smaller ones than in condition B. Consistent with this conclusion, the A-B difference was reduced at longer interpulse intervals. These findings provide a measure of supra-threshold effects of refractoriness on pitch perception, and increase our understanding of coding of temporal information in cochlear implant speech processing schemes.

Acoustic Stimulation↗

Perceived naturalness of spectrally distorted speech and music.

We determined how the perceived naturalness of music and speech (male and female talkers) signals was affected by various forms of linear filtering, some of which were intended to mimic the spectral "distortions" introduced by transducers such as microphones, loudspeakers, and earphones. The filters introduced spectral tilts and ripples of various types, variations in upper and lower cutoff frequency, and combinations of these. All of the differently filtered signals (168 conditions) were intermixed in random order within one block of trials. Levels were adjusted to give approximately equal loudness in all conditions. Listeners were required to judge the perceptual quality (naturalness) of the filtered signals on a scale from 1 to 10. For spectral ripples, perceived quality decreased with increasing ripple density up to 0.2 ripple/ERB(N) and with increasing ripple depth. Spectral tilts also degraded quality, and the effects were similar for positive and negative tilts. Ripples and/or tilts degraded quality more when they extended over a wide frequency range (87-6981 Hz) than when they extended over subranges. Low- and mid-frequency ranges were roughly equally important for music, but the mid-range was most important for speech. For music, the highest quality was obtained for the broadband signal (55-16,854 Hz). Increasing the lower cutoff frequency from 55 Hz resulted in a clear degradation of quality. There was also a distinct degradation as the upper cutoff frequency was decreased from 16,845 Hz. For speech, there was a marked degradation when the lower cutoff frequency was increased from 123 to 208 Hz and when the upper cutoff frequency was decreased from 10,869 Hz. Typical telephone bandwidth (313 to 3547 Hz) gave very poor quality.

Adolescent↗