Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sound Localization”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 829 records · Page 46Linked to original sources

Reverberation and frequency attenuation in forests--implications for acoustic communication in animals.

Rates of reverberative decay and frequency attenuation are measured within two Australian forests. In particular, their dependence on the distance between a source and receiver, and the relative heights of both, is examined. Distance is always the most influential of these factors. The structurally denser of the forests exhibits much slower reverberative decay, although the frequency dependence of reverberation is qualitatively similar in the two forests. There exists a central range of frequencies between 1 and 3 kHz within which reverberation varies relatively little with distance. Attenuation is much greater within the structurally denser forest, and in both forests it generally increases with increasing frequency and distance, although patterns of variation differ between the two forests. Increasing the source height generally reduces reverberation, while increasing the receiver height generally reduces attenuation. These findings have considerable implications for acoustic communication between inhabitants of these forests, particularly for the perching behaviors of birds. Furthermore, this work indicates the ease with which the general acoustic properties of forests can be measured and compared.

Animal Communication↗

Directionality of dog vocalizations.

The directionality patterns of sound emission in domestic dogs were measured in an anechoic environment using a microphone array. Mainly long-distance signals from four dogs were investigated. The radiation pattern of the signals differed clearly from an omnidirectional one with average differences in sound-pressure level between the frontal and rear position of 3-7 dB depending from the individual. Frequency dependence of directionality was shown for the range from 250 to 3200 Hz. The results indicate that when studying acoustic communication in mammals, more attention should be paid to the directionality pattern of sound emission.

Acoustic Stimulation↗

The role of head-induced interaural time and level differences in the speech reception threshold for multiple interfering sound sources.

Three experiments investigated the roles of interaural time differences (ITDs) and level differences (ILDs) in spatial unmasking in multi-source environments. In experiment 1, speech reception thresholds (SRTs) were measured in virtual-acoustic simulations of an anechoic environment with three interfering sound sources of either speech or noise. The target source lay directly ahead, while three interfering sources were (1) all at the target's location (0 degrees,0 degrees,0 degrees), (2) at locations distributed across both hemifields (-30 degrees,60 degrees,90 degrees), (3) at locations in the same hemifield (30 degrees,60 degrees,90 degrees), or (4) co-located in one hemifield (90 degrees,90 degrees,90 degrees). Sounds were convolved with head-related impulse responses (HRIRs) that were manipulated to remove individual binaural cues. Three conditions used HRIRs with (1) both ILDs and ITDs, (2) only ILDs, and (3) only ITDs. The ITD-only condition produced the same pattern of results across spatial configurations as the combined cues, but with smaller differences between spatial configurations. The ILD-only condition yielded similar SRTs for the (-30 degrees,60 degrees,90 degrees) and (0 degrees,0 degrees,0 degrees) configurations, as expected for best-ear listening. In experiment 2, pure-tone BMLDs were measured at third-octave frequencies against the ITD-only, speech-shaped noise interferers of experiment 1. These BMLDs were 4-8 dB at low frequencies for all spatial configurations. In experiment 3, SRTs were measured for speech in diotic, speech-shaped noise. Noises were filtered to reduce the spectrum level at each frequency according to the BMLDs measured in experiment 2. SRTs were as low or lower than those of the corresponding ITD-only conditions from experiment 1. Thus, an explanation of speech understanding in complex listening environments based on the combination of best-ear listening and binaural unmasking (without involving sound-localization) cannot be excluded.

Analysis of Variance↗

A numerical study of the role of the tragus in the big brown bat.

A comprehensive characterization of the spatial sensitivity of an outer ear from a big brown bat (Eptesicus fuscus) has been obtained using numerical methods and visualization techniques. Pinna shape information was acquired through x-ray microtomography. It was used to set up a finite-element model of diffraction from which directivities were predicted by virtue of forward wave-field projections based on a Kirchhoff integral formulation. Digital shape manipulation was used to study the role of the tragus in detailed numerical experiments. The relative position between tragus and pinna aperture was found to control the strength of an extensive asymmetric sidelobe which points in a frequency-dependent direction. An upright tragus position resulted in the strongest sidelobe sensitivity. Using a bootstrap validation paradigm, the results were found to be robust against small perturbations of the finite-element mesh boundaries. Furthermore, it was established that a major aspect of the tragus effect (position dependence) can be studied in a simple shape model, an obliquely truncated horn augmented by a flap representing the tragus. In the simulated wave field around the outer-ear structure, strong correlates of the tragus rotation were identified, which provide a direct link to the underlying physical mechanism.

Animals↗

Perception of pitch location within a speaker's F0 range.

Fundamental frequency (F0) is used for many purposes in speech, but its linguistic significance is based on its relation to the speaker's range, not its absolute value. While it may be that listeners can gauge a specific pitch relative to a speaker's range by recognizing it from experience, whether they can do the same for an unfamiliar voice is an open question. The present experiment explored that question. Twenty native speakers of English (10 male, 10 female) produced the vowel /a/ with a spoken (not sung) voice quality at varying pitches within their own ranges. Listeners then judged, without familiarization or context, where each isolated F0 lay within each speaker's range. Correlations were high both for the entire range (0.721) and for the range minus the extremes (0.609). Correlations were somewhat higher when the F0s were related to the range of all the speakers, either separated by sex (0.830) or pooled (0.848), but several factors discussed here may help account for this pattern. Regardless, the present data provide strong support for the hypothesis that listeners are able to locate an F0 reliably within a range without external context or prior exposure to a speaker's voice.

Adult↗

The effect of spatial separation on informational masking of speech in normal-hearing and hearing-impaired listeners.

The ability to understand speech in a multi-source environment containing informational masking may depend on the perceptual arrangement of signal and masker objects in space. In normal-hearing listeners, Arbogast et al. [J. Acoust. Soc. Am. 112, 2086-2098 (2002)] found an 18-dB spatial release from a primarily informational masker, compared to 7 dB for a primarily energetic masker. This article extends the earlier work to include the study of listeners with sensorineural hearing loss. Listeners performed closed-set speech recognition in two spatial conditions: 0 degrees and 90 degrees separation between signal and masker. Three maskers were tested: (1) the different-band sentence masker was designed to be primarily informational; (2) the different-band noise masker was a control for the different-band sentence; and (3) the same-band noise masker was designed to be primarily energetic. The spatial release from the different-band sentence was larger than for the other maskers, but was smaller (10 dB) for the hearing-impaired group than for the normal-hearing group (15 dB). The smaller benefit for the hearing-impaired listeners can be partially explained by masker sensation level. However, the results suggest that hearing-impaired listeners can use the perceptual effect of spatial separation to improve speech recognition in the presence of a primarily informational masker.

Adult↗

A Speech Intelligibility Index-based approach to predict the speech reception threshold for sentences in fluctuating noise for normal-hearing listeners.

The SII model in its present form (ANSI S3.5-1997, American National Standards Institute, New York) can accurately describe intelligibility for speech in stationary noise but fails to do so for nonstationary noise maskers. Here, an extension to the SII model is proposed with the aim to predict the speech intelligibility in both stationary and fluctuating noise. The basic principle of the present approach is that both speech and noise signal are partitioned into small time frames. Within each time frame the conventional SII is determined, yielding the speech information available to the listener at that time frame. Next, the SII values of these time frames are averaged, resulting in the SII for that particular condition. Using speech reception threshold (SRT) data from the literature, the extension to the present SII model can give a good account for SRTs in stationary noise, fluctuating speech noise, interrupted noise, and multiple-talker noise. The predictions for sinusoidally intensity modulated (SIM) noise and real speech or speech-like maskers are better than with the original SII model, but are still not accurate. For the latter type of maskers, informational masking may play a role.

Attention↗

Detection of high-frequency spectral notches as a function of level.

High-frequency spectral notches are important cues for sound localization. Our ability to detect them must depend on their representation as auditory nerve (AN) rate profiles. Because of the low threshold and the narrow dynamic range of most AN fibers, these rate profiles deteriorate at high levels. The system may compensate by using onset rate profiles whose dynamic range is wider, or by using low-spontaneous-rate fibers, whose threshold is higher. To test these hypotheses, the threshold notch depth necessary to discriminate between a flat spectrum broadband noise and a similar noise with a spectral notch centered at 8 kHz was measured at levels from 32 to 100 dB SPL. The importance of the onset rate-profile representation of the notch was estimated by varying the stimulus duration and its rise time. For a large proportion of listeners, threshold notch depth varied nonmonotonically with level, increasing for levels up to 70-80 dB SPL and decreasing thereafter. The nonmonotonic aspect of the function was independent of notch bandwidth and stimulus duration. Thresholds were independent of stimulus rise time but increased for the shorter noise bursts. Results are discussed in terms of the ability of the AN to convey spectral notch information at different levels.

Acoustic Stimulation↗

The influence of spectral, temporal, and interaural stimulus variations on the precedence effect.

The precedence effect describes phenomena that are believed to aid localization of sounds in reverberant environments. These phenomena relate to the emphasis given to the first-arriving or preceding sound. In this paper, experiments are described which study precedence using stimulus parametrizations spanning temporal, spectral, and interaural dimensions. Subjects report the sidedness of headphone stimuli comprising a source and single reflection placed symmetrically with respect to midline. Most of the experiments use long-duration noises with the onset and offset time-of-arrival differences windowed out from the combined lead and lag stimulus, thus requiring the subject to lateralize using cues in the ongoing portion of the stimuli where the lead and lag overlap completely. A similar experiment using click stimuli is included for comparison. The influence of spectral content is studied by varying either the bandwidth or the center frequency. Dependence on interaural cues is investigated by using either ITDs or IIDs to induce laterality in the individual lead and lag components. Results indicate that precedence continues into the ongoing portion of long-duration stimuli and is robust to the removal of initial onsets, to reduction of bandwidth, and to the choice of interaural cue used to induce laterality in the lead and lag.

Acoustic Stimulation↗

Sound-producing sources as objects of perception: rate normalization and nonspeech perception.

In a variety of experiments and paradigms, researchers have attempted to determine whether or not speech perception is specialized by comparing perception of speech syllables to perception of nonspeech analogs. While nonspeech analogs appear optimal as comparisons to speech because they are acoustically similar without being recognized as speechlike, it is argued that the comparison they offer is confounded and uninterpretable. Two experiments are designed to show that, in auditory perception generally where acoustic signals are causal consequences of mechanical events, perceptual experiences are of the mechanical events themselves, not of the acoustic signal. This has two consequences. One is that there is a confounding in comparisons of speech with sine wave analogs that, whereas the one perceived as speech also has a definite causal source, the other, perceived as nonspeech, has an indeterminate or ambiguous source. A second is that response patterns in classification tasks such as those used in the literature comparing speech to nonspeech will reflect properties of the perceived sound-producing event; they will not provide a clear window on auditory system processes used to recover event properties. Experiment 3 is designed to show that perception of many acoustic-signal-producing events can appear to be special by the logic of speech-sine wave comparisons--even events that cannot plausibly be supposed to involve a specialization.

Adult↗

Modeling the perception of concurrent vowels: vowels with different fundamental frequencies.

If two vowels with different fundamental frequencies (fo's) are presented simultaneously and monaurally, listeners often hear two talkers producing different vowels on different pitches. This paper describes the evaluation of four computational models of the auditory and perceptual processes which may underlie this ability. Each model involves four stages: (i) frequency analysis using an "auditory" filter bank, (ii) determination of the pitches present in the stimulus, (iii) segregation of the competing speech sources by grouping energy associated with each pitch to create two derived spectral patterns, and (iv) classification of the derived spectral patterns to predict the probabilities of listeners' vowel-identification responses. The "place" models carry out the operations of pitch determination and spectral segregation by analyzing the distribution of rms levels across the channels of the filter bank. The "place-time" models carry out these operations by analyzing the periodicities in the waveforms in each channel. In their "linear" versions, the place and place-time models operate directly on the waveforms emerging from the filters. In their "nonlinear" versions, analogous operations are applied to the output of an additional stage which applied a compressive nonlinearity to the filtered waveforms. Compared to the other three models, the nonlinear place-time model provides the most accurate estimates of the fo's of paris of concurrent synthetic vowels and comes closest to predicting the identification responses of listeners to such stimuli. Although the model has several limitations, the results are compatible with the idea that a place-time analysis is used to segregate competing sound sources.

Attention↗

The emission pattern of vocalizations and directionality of the sonar system in the echolocating bat, Pteronotus parnelli.

The radiation patterns of the first three harmonics (approx. 30, 60, 90 kHz) of the mustached bat biosonar signal were measured from vocalizations elicited by cortical microstimulation. The primary foci of the acoustic beam patterns were in front of the mouth but somewhat below the horizontal plane. The prominent second and third harmonics showed sharp cutoffs between 20 degrees and 30 degrees lateral to the midline. Sidelobes were found, suggesting the influence of some vocal tract interference. When compared with previously measured estimates of the directionality of the auditory system, the vocal emission patterns are roughly complementary: Regions of maximum auditory sensitivity are found in areas of submaximal power for the sonar pulse beam pattern. The result is that, for the two most important harmonics, the "biosonar system" (i.e., vocal beam pattern plus receiver directionality) has a broader and more uniform directionality than either component alone. Therefore, within a limited region of space, echo amplitude will vary less as a function of angular displacement. This reduces the confounding influences of absolute sound pressure level on interaural intensity differences.

Animals↗

Auditory space expansion via linear filtering.

A signal-processing algorithm that modifies the interaural time delays associated with directional sources is described. Signals received at two microphones are processed by four linear filters arranged in a lattice configuration to produce two outputs, one for each ear. Since the processing is linear, the method is equally applicable to single or multiple directional sources. The filters are designed to minimize the average squared error between a user specified desired space warping function and the actual warping function that they implement. Two classes of filters are considered: filters whose frequency response is unconstrained and filters constrained to be causal with finite impulse response. In both cases the solution of the least-squares problem is given and properties of the actual space warping function are examined. Perceptual experiments and analysis of acoustic waveforms are utilized to demonstrate the effectiveness of the algorithm. Extension of this method for utilizing more than two microphones is described.

Algorithms↗

A computational model of echo processing and acoustic imaging in frequency-modulated echolocating bats: the spectrogram correlation and transformation receiver.

The spectrogram correlation and transformation (SCAT) model of the sonar receiver in the big brown bat (Eptesicus fuscus) consists of a cochlear component for encoding the bat's frequency modulated (FM) sonar transmissions and multiple FM echoes in a spectrogram format, followed by two parallel pathways for processing temporal and spectral information in sonar echoes to reconstruct the absolute range and fine range structure of multiple targets from echo spectrograms. The outputs of computations taking place along these parallel pathways converge to be displayed along a computed image dimension of echo delay or target range. The resulting image depicts the location of various reflecting sources in different targets along the range axis. This series of transforms is equivalent to simultaneous, parallel forward and inverse transforms on sonar echoes, yielding the impulse responses of targets by deconvolution of the spectrograms. The performance of the model accurately reproduces the images perceived by Eptesicus in a variety of behavioral experiments on two-glint resolution in range, echo phase sensitivity, amplitude-latency trading of range estimates, dissociation of time- and frequency-domain image components, and ranging accuracy in noise.

Animals↗

Pressure transfer function and absorption cross section from the diffuse field to the human infant ear canal.

The diffuse-field pressure transfer function from a reverberant field to the ear canal of human infants, ages 1, 3, 6, 12, and 24 months, has been measured from 125-10700 Hz. The source was a loudspeaker using pink noise, and the diffuse-field pressure and the ear-canal pressure were simultaneously measured using a spatial averaging technique in a reverberant room. The results in most subjects show a two-peak structure in the 2-6-kHz range, corresponding to the ear-canal and concha resonances. The ear-canal resonance frequency decreases from 4.4 kHz at age 1 month to 2.9 kHz at age 24 months. The concha resonance frequency decreases from 5.5 kHz at age 1 month to 4.5 kHz at age 24 months. Below 2 kHz, the diffuse-field transfer function shows effects due to the torsos of the infant and parent, and varies with how the infant is held. Comparisons are reported of the diffuse-field absorption cross section for infants relative to adults. This quantity is a measure of power absorbed by the middle ear from a diffuse sound field, and large differences are observed in infants relative to adults. The radiation efficiencies of the infant and the adult ear are small at low frequencies, near unity at midfrequencies, and decrease at higher frequencies. The process of ear-canal development is not yet complete at age 24 months. The results have implications for experiments on hearing in infants.

Age Factors↗

Marine mammal call discrimination using artificial neural networks.

Recent work has applied a linear spectrogram correlator filter (SCF) to detect bowhead whale (Balaena mysticetus) song notes, outperforming both a time-series-matched filter and a hidden Markov model. The method relies on an empirical weighting matrix. An artificial neural net (ANN) may be better yet, since it offers two advantages; (i) the equivalent weighting matrix is determined by training and can converge to a more optimal solution and (ii) an ANN is a nonlinear estimator and can embody more sophisticated responses. A three-layer feed-forward ANN is ideally suited to this application and has been implemented on 1475 sounds, of which 54% were used for training and 46% kept as "unseen" test data. The trained ANN error rate was 1.5%, a twofold improvement over previous methods. It is shown that ANN hidden neurons can be interrogated to reveal the operating paradigm developed during training. The function of each of these neurons can be determined in terms of spectrographic features of the training calls. Furthermore, the operating paradigm can be controlled and training time reduced by assigning specific recognition tasks to hidden neurons prior to training, rather than initiating training with randomized weights. The ANN is compared to the SCF and the role of the "hidden" neurons and equivalent weighting matrices are discussed.

Animals↗

Physiological studies of the precedence effect in the inferior colliculus of the kitten.

The precedence effect (PE) is a perceptual phenomenon that reflects listeners' ability to suppress echoes in reverberant environments. The PE is not present at birth and appears only several months postnatal. Recent physiological studies have demonstrated correlates of the PE in the central nucleus of the inferior colliculus (ICC) of adult animals. The present study extended the same techniques to search for similar correlates in the ICC of kittens during the first postnatal month. Stimuli consisted of pairs of clicks or noise bursts presented from different locations in free field or with different inter-aural differences in time (ITD) under headphones, with an inter-stimulus-delay (ISD) between their onsets. Results suggest that a physiological correlate of the PE, i.e. suppression of responses to the second source, is present as early as 8 days postnatal, and occurs at similar ISDs to those recorded in adult cats. Suppression in kitten neurons varies with stimulus level, duration, and azimuthal position, in a similar manner to that in adult neurons. The age at which correlates of the PE in the kitten can be found precedes the age at which kittens can localize sound sources effectively, and presumably before the age at which they would demonstrate the PE behaviorally. Thus, the neural mechanisms that might be involved in the first stages of processing PE stimuli may be in place well before the behavioral correlate develops.

Animals↗

Auditory-visual spatial integration: a new psychophysical approach using laser pointing to acoustic targets.

The alignment of auditory and visual spatial perception was investigated in four experiments, employing a method of laser pointing toward acoustic targets in combination with various tasks of visual fixation in six subjects. Subjects had to fixate either a target LED or a laser spot projected on a screen in a dark, anechoic room and, while doing so, direct the laser beam toward the perceived azimuthal position of the sound stimulus (bandpass-filtered noise; bandwidth 1-3 kHz; 70 dB sound pressure level, duration 10 s). The sound was produced by one of nine loudspeakers, located behind the acoustically transparent screen between 22 degrees to the left and 22 degrees to the right of straight ahead. Systematic divergences between sound azimuth and laser adjustment were found, depending on the instructions given to the subjects. The eccentricity of acoustic targets was generally overestimated by up to 10.4 degrees with an only slight influence of gaze direction on this effect. When the sound source was straight ahead, gaze direction had a substantial influence in that the laser adjustments deviated by up to 5.6 degrees from sound azimuth, toward the side to which the gaze was directed. This effect of eye position decreased with increasing eccentricity of the sound. These results can be explained by the interactive effects of four distinct factors: the lateral overestimation of the auditory eccentricity, the effect of eye position on sound localization, the effect of the retinal eccentricity on visual localization, and the extraretinal effect of eye position on visual localization.

Acoustics↗