Cognitive restoration of reversed speech.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to K Saberi.
Explore the source record for details and available documents.
Detection performance for a masked auditory signal of fixed frequency can be substantially degraded if there is uncertainty about the frequency content of the masker. A quasimolecular psychophysical approach was used to examine response strategies in masker-uncertainty conditions, and to investigate the influence of uncertainty when the number of different masker samples was limited to ten or fewer. The task of the four listeners was to detect a 1000-Hz signal that was presented simultaneously with one of ten ten-tone masker samples. The masker sample was either fixed throughout a block of two-interval forced-choice trials or was randomized across or within trials. The primary results showed that: (1) When the signal level was low and the masker sample differed between the two intervals of a trial, most listeners based their responses more on the presence of specific masker samples than on the signal. (2) The detrimental effect of masker uncertainty was clearly evident when only four maskers were randomly presented, and grew as the size of the masker set was increased from two to ten. (3) The slopes of psychometric functions measured with the same masker samples differed among the fixed and two random-masker conditions. (4) There were large differences in the influence of masker uncertainty across masker samples and listeners. These data demonstrate the great susceptibility of human listeners to the influence of masker uncertainty and the ability of quasimolecular investigations to reveal important aspects of behavior in uncertainty condition.
Owls and other animals, including humans, use the difference in arrival time of sounds between the ears to determine the direction of a sound source in the horizontal plane. When an interaural time difference (ITD) is conveyed by a narrowband signal such as a tone, human beings may fail to derive the direction represented by that ITD. This is because they cannot distinguish the true ITD contained in the signal from its phase equivalents that are ITD +/- nT, where T is the period of the stimulus tone and n is an integer. This uncertainty is called phase-ambiguity. All ITD-sensitive neurons in birds and mammals respond to an ITD and its phase equivalents when the ITD is contained in narrowband signals. It is not known, however, if these animals show phase-ambiguity in the localization of narrowband signals. The present work shows that barn owls (Tyto alba) experience phase-ambiguity in the localization of tones delivered by earphones. We used sound-induced head-turning responses to measure the sound-source directions perceived by two owls. In both owls, head-turning angles varied as a sinusoidal function of ITD. One owl always pointed to the direction represented by the smaller of the two ITDs, whereas a second owl always chose the direction represented by the larger ITD (i.e., ITD - T).
The detection of interaural time differences (ITDs) for sound localization critically depends on the similarity between the left and right ear signals (interaural correlation). We show that, like humans, owls can localize phantom sound sources well until the correlation declines to a very low value, below which their performance rapidly deteriorates. Decreasing interaural correlation also causes the response of the owl's tectal auditory neurons to decline nonlinearly, with a rapid drop followed by a more gradual reduction. A detection-theoretic analysis of the statistical properties of neuronal responses could account for the variance of behavioral responses as interaural correlation is decreased. Finally, cross-correlation analysis suggests that low interaural correlations cause misalignment of cross-correlation peaks across different frequencies, contributing heavily to the nonlinear decline in neural and ultimately behavioral performance.
Interaural-delay sensitivity to high-frequency (> or = 3 kHz) sinusoidal-frequency-modulated (SFM) tones is examined for rates from 25 to 800 Hz and depths of -12 to 18 dB. Comparison is made to thresholds obtained for sinusoidal-amplitude-modulated (SAM) tones for the same observers and modulation rates. Both SAM and SFM threshold-by-rate functions are U-shaped with optimum sensitivity to SFM tones occurring at higher rates (fm = 200-400 Hz) compared to those for SAM tones (fm = 100-200 Hz). Effects of modulation depth were examined for rates from 50 to 300 Hz. In all cases thresholds improved considerably with increasing modulation depth. It is also shown that a hybrid dichotic signal composed of an SFM tone presented to one ear and an SAM tone to the other, can perceptually fuse and be lateralized, with the contingency that both stimuli have equal modulation rates but not necessarily equal carrier frequencies. Using bandpass noise to restrict off-frequency listening, it was shown that for this stimulus, observers can use information from filters either below or above the carrier frequency. Consistent with FM-to-AM conversion from cochlear bandpass filtering, several important differences between the SAM- and SFM-tone data can be predicted from a nonstationary stochastic model of binaural interaction whose parameters are uniquely determined from the SAM-tone data.
This is a brief report on the use of maximum-likelihood (ML) estimators in auditory psychophysics. Slope parameters of psychometric functions are characterized for three nonintensive auditory tasks: forced-choice discrimination of interaural time differences (delta ITD), frequency (delta f), and duration (delta t). Using these slope estimates, the ML method is implemented and threshold estimates are obtained for the three tasks and compared with previously published data. delta ITD thresholds were additionally measured for human observers by means of two other psychophysical procedures: the constant-stimuli (CS) and the 2-down 1-up methods (Wetherill & Levitt, 1965). Standard errors were smallest for the ML method. Finally, simulations showed ML estimates to be more efficient than the CS and k-down 1-up procedures for k = 2 to 5. For up-down procedures, efficiency was highest for k values of 3 and 4. The entropy (Shannon, 1949) of ML estimates was the smallest of the simulated procedures, but poorer than ideal by 0.5 bits.
In humans, the lateral movement of an acoustic source produces dynamic changes in the relative sound-pressure level and time of arrival of the acoustic wave at the 2 ears. The dynamic nature of these cues is assumed to play an important role in the perception of lateral motion. A phenomenon of auditory motion is reported whose lateral direction and relative velocity may be specified while interaural differences are kept constant. The stimulus producing this percept is a narrowband wave-form whose instantaneous bandwidth is a cosine function of time. This phenomenon is predicted from a model of cross-correlation that estimates the running position of an image from a weighted combination of 2 variables: (a) magnitude of interaural delay, with smaller delays receiving more weight, and (b) consistency of interaural information across frequency.
One class of adaptive psychophysical procedures was studied, using simulated and human observers. These procedures are those which require an increase in stimulus intensity after an incorrect response, and a decrease after k successive correct responses. This paper analyzes how step size and the value of k affect the mean and standard deviation of threshold estimates based on a k-down 1-up adaptive procedure. Computer simulations are used to study the bias in threshold estimates, which are most evident when larger step size and small values of k are used. The adaptive procedure can be characterized by a function called the imbalance of the track, the relative probability of adjusting the stimulus either up or down at equal stimulus distances from the equilibrium point. These imbalance functions can be used to understand the threshold biases obtained in the computer simulations. The computer simulations also show that the average number of reversals obtained per trial is dependent on different values of k, but are largely independent of step size. The standard error of the threshold estimates, however, varies systematically with step size, but are nearly independent of k. Finally, we compare the stability of threshold estimates for human listeners using two very different sets of parameters: a very large step size (approximately half the range of the psychometric function) with k = 4, and the conventional k = 3 with an initial 4-dB and a final 2-dB step size.
Onset dominance in sound localization was examined by estimating observer weighting of interaural delays for each click of a train of high-frequency filtered clicks. The interaural delay of each click was a normal deviate that was sampled independently on each trial of a single-interval design. In Experiment 1, observer weights were derived for trains of n = 2, 4, 8, or 16 clicks as a function of interclick interval (ICI = 1.8, 3.0, or 12.0 msec). For small n and short ICI (1.8 msec), the ratio of onset weight to remaining weights was as large as 10. As ICI increased, the relative onset weight was reduced. For large n and all ICIs, the ongoing train was weighted more heavily than the onset. This diminishing relative onset weight with increasing ICI and n is consistent with optimum distribution of weights among components. Efficiency of weight distribution is near ideal when ICI = 12 msec and n = 2 and very poor for shorter ICIs and larger ns. Further experiments showed that: (1) onset dominance involves both within- and between-frequency-channel mechanisms, and (2) the stimulus configuration (ICI, n, frequency content, and temporal gaps) affects weighting functions in a complex way not explained by cross-correlation analysis or contralateral inhibition (Lindemann, 1986a, 1986b).
Most naturally occurring sounds are modulated in amplitude or frequency; important examples include animal vocalizations and species-specific communication signals in mammals, insects, reptiles, birds and amphibians. Deciphering the information from amplitude-modulated (AM) sounds is a well-understood process, requiring a phase locking of primary auditory afferents to the modulation envelopes. The mechanism for decoding frequency modulation (FM) is not as clear because the FM envelope is flat (Fig. 1). One biological solution is to monitor amplitude fluctuations in frequency-tuned cochlear filters as the instantaneous frequency of the FM sweeps through the passband of these filters. This view postulates an FM-to-AM transduction whereby a change in frequency is transmitted as a change in amplitude. This is an appealing idea because, if such transduction occurs early in the auditory pathway, it provides a neurally economical solution to how the auditory system encodes these important sounds. Here we illustrate that an FM and AM sound must be transformed into a common neural code in the brain stem. Observers can accurately determine if the phase of an FM presented to one ear is leading or lagging, by only a fraction of a millisecond, the phase of an AM presented to the other ear. A single intracranial image is perceived, the spatial position of which is a function of this phase difference.
Explore the source record for details and available documents.
This study examines the ability to lateralize a complex signal characterized by correlated temporal activity across widely separated frequency regions. The high-frequency complex consisted of two narrow-band stimuli. The two stimuli had common interaural delays but different carriers centered on nonoverlapping critical bands. Two basic conditions were examined: The narrow-band stimuli had temporal envelopes which were (1) identical or (2) different. In the first experiment, narrow bands of noise were used which either had identical temporal envelopes (comodulated) or statistically independent envelopes (CFs = 2550 and 3350 Hz). In the second experiment, two sinusoidally amplitude-modulated (SAM) tones were used whose modulators either had the same starting phase or a different starting phase (CFs = 2550 and 4000 Hz). Results of the first experiment showed that for bandwidths narrower than 300 Hz, comodulated bands produced significantly lower interaural-delay thresholds compared to independent bands. Results of the second experiment showed that when the two SAM tones (100-Hz modulation rate) had the same modulator starting phase, interaural-delay thresholds were lowest.
Free-field release from masking was studied as a function of the spatial separation of a signal and masker in a two-interval, forced-choice (2IFC) adaptive paradigm. The signal was a 250-ms train of clicks (100/s) generated by filtering 50-microseconds pulses with a TDH-49 speaker (0.9 to 9.0 kHz). The masker was continuous broadband (0.7 to 11 kHz) white noise presented at a level of 44 dBA measured at the position of the subject's head. In experiment I, masked and absolute thresholds were measured for 36 signal source locations (10 degree increments) along the horizontal plane as a function of seven masking source locations (30 degree increments). In experiment II, both absolute and masked thresholds were measured for seven signal locations along three vertical planes located at azimuthal rotations of 0 degrees (median vertical plane), 45 degrees, and 90 degrees. In experiment III, monaural absolute and masked thresholds were measured for various signal-masker configurations. Masking-level differences (MLDs) were computed relative to the condition where the signal and mask were in front of the subjects after using absolute thresholds to account for differences in the signal's sound-pressure level (SPL) due to direction. Maximum MLDs were 15 dB along the horizontal plane, 8 dB along the vertical, and 9 dB under monaural conditions.
Visual search performance was examined in a two-alternative, forced-choice paradigm. The task involved locating and identifying which of two visual targets was present on a trial. The location of the targets varied relative to the subject's initial fixation point from 0 to 14.8 deg. The visual targets were either presented concurrently with a sound located at the same position as the visual target or were presented in silence. Both the number of distractor visual figures (0-63) present in the field during the search (Experiments 1 and 2) and the distinctness of the visual target relative to the distractors (Experiment 2) were considered. Under all conditions, visual search latencies were reduced when spatially correlated sounds were present. Aurally guided search was particularly enhanced when the visual target was located in the peripheral regions of the central visual field and when a larger number of distractor images (63) were present. Similar results were obtained under conditions in which the target was visually enhanced. These results indicate that spatially correlated sounds may have considerable utility in high-information environments (e.g., piloting an aircraft).
Minimum audible angle (MAA) thresholds were obtained for four subjects in a two-alternative, forced-choice, three up/one down, adaptive paradigm as a function of the orientation of the array of sources. With sources distributed on the horizontal plane, the mean MAA threshold was 0.97 degrees. With the sources distributed on the vertical plane (array rotated 90 degrees), the mean MAA threshold was 3.65 degrees. Performance in both conditions was well in line with previous experiments of this type. Tests were also conducted with sources distributed on oblique planes. As the array was rotated from 10 degrees-60 degrees from the horizontal plane, relatively little change in the MAA threshold was observed; the mean MAA thresholds ranged from 0.78 degrees to 1.06 degrees. Only when the array was nearly vertical (80 degrees) was there any appreciable loss in spatial resolution; the MAA threshold had increased to 1.8 degrees. The relevance of these results to research on auditory localization under natural listening conditions, especially in the presence of head movements, is also discussed.
Interaural differences of time (IDT) thresholds were measured with 600-microseconds transients. The initial experiment was a successful replication of previous experiments that have obtained the precedence effect in lateralization paradigms (e.g., Yost and Soderquist, 1984). When a dichotic click followed a diotic click with an interclick interval (ICI) less than 1 ms or larger than 5 ms, IDT thresholds were generally less than 40 microseconds. For ICIs between 1 to 5 ms, IDT thresholds increased to approximately 220 microseconds. Poorest performance was observed for ICIs of 1.75 to 2.35 ms. During the course of conducting a series of planned experiments on this effect, a substantial drop in IDT thresholds was observed across the ICIs of maximum interest (1 to 5 ms). The precedence effect, which we had replicated in our initial experiment, essentially "disappeared" when the subjects were given sufficient practice on the lateralization task. A number of conditions were explored in an unsuccessful attempt to recover the precedence effect in these experienced subjects. The implications of these results are discussed.
Auditory resolution of moving sound sources was determined in a simulated motion paradigm for sources traveling along horizontal, vertical, or oblique orientations in the subjects's frontal plane. With motion restricted to the horizontal orientation, minimum audible movement angles (MAMA) ranged from about 1.7 degrees at the lowest velocity (1.8 degrees/s) to roughly 10 degrees at the highest velocity (320 degrees/s). With the sound moving along an oblique orientation (rotated 45 degrees relative to the horizontal) MAMAs generally matched those of the horizontal condition. When motion was restricted to the vertical, MAMAs were substantially larger at all velocities (often exceeding 8 degrees). Subsequent tests indicated that MAMAs are a U-shaped function of velocity, with optimum resolution obtained at about 2 degrees/s for the horizontal (and oblique) and 7-11 degrees/s for the vertical orientation. Additional tests conducted at a fixed velocity of 1.8 degrees/s along oblique orientations of 80 degrees and 87 degrees indicated that even a small deviation from the vertical had a significant impact on MAMAs. A displacement of 10 degrees from the vertical orientation (a slope of 80 degrees) was sufficient to reduce thresholds (obtained at a velocity of 1.8 degrees/s) from about 11 degrees to approximately 2 degrees (a fivefold increase in acuity). These results are in good agreement with our previous study of minimum audible angles long oblique planes [Perrott and Saberi, J. Acoust. Soc. Am. 87, 1728-1731 (1990)].(ABSTRACT TRUNCATED AT 250 WORDS)
In Experiments 1 and 2, the time to locate and identify a visual target (visual search performance in a two-alternative forced-choice paradigm) was measured as a function of the location of the target relative to the subject's initial line of gaze. In Experiment 1, tests were conducted within a 260 degree region on the horizontal plane at a fixed elevation (eye level). In Experiment 2, the position of the target was varied in both the horizontal (260 degrees) and the vertical (+/- 46 degrees from the initial line of gaze) planes. In both experiments, and for all locations tested, the time required to conduct a visual search was reduced substantially (175-1,200 msec) when a 10-Hz click train was presented from the same location as that occupied by the visual target. Significant differences in latencies were still evident when the visual target was located within 10 degrees of the initial line of gaze (central visual field). In Experiment 3, we examined head and eye movements that occur as subjects attempt to locate a sound source. Concurrent movements of the head and eyes are commonly encountered during auditorily directed search behavior. In over half of the trials, eyelid closures were apparent as the subjects attempted to orient themselves toward the sound source. The results from these experiments support the hypothesis that the auditory spatial channel has a significant role in regulating visual gaze.