Intraspecific and geographic variation of West Indian manatee (Trichechus manatus spp.) vocalizations.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
This study expands the limited understanding of pinniped aerial auditory masking and includes measurements at some of the relatively low frequencies predominant in many pinniped vocalizations. Behavioral techniques were used to obtain aerial critical ratios (CRs) within a hemianechoic chamber for a northern elephant seal (Mirounga angustirostris), a harbor seal (Phoca vitulina), and a California sea lion (Zalophus californianus). Simultaneous, octave-band noise maskers centered at seven test frequencies (0.2-8.0 kHz) were used to determine aerial CRs. Narrower and variable bandwidth masking noise was also used in order to obtain direct critical bandwidths (CBWs). The aerial CRs are very similar in magnitude and in frequency-specific differences (increasing gradually with test frequency) to underwater CRs for these subjects, demonstrating that pinniped cochlear processes are similar both in air and water. While, like most mammals, these pinniped subjects apparently lack specialization for enhanced detection of specific frequencies over masking noise, they consistently detect signals across a wide range of frequencies at relatively low signal-to-noise ratios. Direct CBWs are 3.2 to 14.2 times wider than estimated based on aerial CRs. The combined masking data are significant in terms of assessing aerial anthropogenic noise impacts, effective aerial communicative ranges, and amphibious aspects of pinniped cochlear mechanics.
Fundamental frequency (F0) extraction is often used in voice quality analysis. In pathological voices with a high degree of instability in F0, it is common for F0 extraction algorithms to fail. In such cases, the faulty F0 values might spoil the possibilities for further data analysis. This paper presents the correlogram, a new method of displaying periodicity. The correlogram is based on the waveform-matching techniques often used in F0 extraction programs, but with no mechanism to select an actual F0 value. Instead, several candidates for F0 are shown as dark bands. The result is presented as a 3D plot with time on the x axis, correlation delay inverted to frequency on the y axis, and correlation on the z axis. The z axis is represented in a gray scale as in a spectrogram. Delays corresponding to integer multiples of the period time will receive high correlation, thus resulting in candidates at F0, F0/2, F0/3, etc. While the correlogram adds little to F0 analysis of normal voices, it is useful for analysis of pathological voices since it illustrates the full complexity of the periodicity in the voice signal. Also, in combination with manual tracing, the correlogram can be used for semimanual F0 extraction. If so, F0 extraction can be performed on many voices that cause problems for conventional F0 extractors. To demonstrate the properties of the method it is applied to synthetic and natural voices, among them six pathological voices, which are characterized by roughness, vocal fry, gratings/scrape, hypofunctional breathiness and voice breaks, or combinations of these.
Sounds of blue whales were recorded from U.S. Navy hydrophone arrays in the North Atlantic. The most common signals were long, patterned sequences of very-low-frequency sounds in the 15-20 Hz band. Sounds within a sequence were hierarchically organized into phrases consisting of one or two different sound types. Sequences were typically composed of two-part phrases repeated every 73 s: a constant-frequency tonal "A" part lasting approximately 8 s, followed 5 s later by a frequency-modulated "B" part lasting approximately 11 s. A common sequence variant consisted only of repetitions of part A. Sequences were separated by silent periods averaging just over four minutes. Two other sound types are described: a 2-5 s tone at 9 Hz, and a 5-7 s inflected tone that swept up in frequency to ca. 70 Hz and then rapidly down to 25 Hz. The general characteristics of repeated sequences of simple combinations of long-duration, very-low-frequency sound units repeated every 1-2 min are typical of blue whale sounds recorded in other parts of the world. However, the specific frequency, duration, and repetition interval features of these North Atlantic sounds are different than those reported from other regions, lending further support to the notion that geographically separate blue whale populations have distinctive acoustic displays.
Time-frequency representations (TFRs) of otoacoustic emissions (OAEs) provide information simultaneously in time and frequency that may be obscured in waveform or spectral analyses. TFRs were applied to transient-evoked stimulus-frequency (SF) and distortion-product (DP) OAEs to test cochlear model predictions. SFOAEs and DPOAEs were elicited in 18 normal-hearing subjects using gated tones and tone pips. Synchronous spontaneous (SS) OAEs were measured to assess their contributions to SFOAEs and DPOAEs. A common form of TFR of measured OAEs was a collection of frequency-specific components often aligned with SSOAE sites, with each component characterized by one or more brief segments or a single long-duration segment. The spectral envelope of evoked OAEs differed from that of the evoking stimulus. Strong emission regions or cochlear "hot spots" were detected, and sometimes accounted for OAE energy observed outside the stimulus bandwidth. Contributions of hot spots and multiple internal reflections to the OAE, and differences between measured and predicted OAE spectra, increased as stimulus level decreased, consistent with level-dependent changes in the estimated cochlear reflectance. Suppression and frequency-pulling effects between components were observed. A recursive formulation was described for the linear coherent reflection emission theory [Zweig and Shera, J. Acoust. Soc. Am. 98, 2018-2047 (1995)] that is well suited for time-domain calculations.
Explore the source record for details and available documents.
Stretching or compressing an outer hair cell alters its membrane potential and, conversely, changing the electrical potential alters its length. This bi-directional energy conversion takes place in the cell's lateral wall and resembles the direct and converse piezoelectric effects both qualitatively and quantitatively. A piezoelectric model of the lateral wall has been developed that is based on the electrical and material parameters of the lateral wall. An equivalent circuit for the outer hair cell that includes piezoelectricity shows a greater admittance at high frequencies than one containing only membrane resistance and capacitance. The model also predicts resonance at ultrasonic frequencies that is inversely proportional to cell length. These features suggest all mammals use outer hair cell piezoelectricity to support the high-frequency receptor potentials that drive electromotility. It is also possible that members of some mammalian orders use outer hair cell piezoelectric resonance in detecting species-specific vocalizations.
Efforts to study the social acoustic signaling behavior of delphinids have traditionally been restricted to audio-range (<20 kHz) analyses. To explore the occurrence of communication signals at ultrasonic frequencies, broadband recordings of whistles and burst pulses were obtained from two commonly studied species of delphinids, the Hawaiian spinner dolphin (Stenella longirostris) and the Atlantic spotted dolphin (Stenella frontalis). Signals were quantitatively analyzed to establish their full bandwidth, to identify distinguishing characteristics between each species, and to determine how often they occur beyond the range of human hearing. Fundamental whistle contours were found to extend beyond 20 kHz only rarely among spotted dolphins, but with some regularity in spinner dolphins. Harmonics were present in the majority of whistles and varied considerably in their number, occurrence, and amplitude. Many whistles had harmonics that extended past 50 kHz and some reached as high as 100 kHz. The relative amplitude of harmonics and the high hearing sensitivity of dolphins to equivalent frequencies suggest that harmonics are biologically relevant spectral features. The burst pulses of both species were found to be predominantly ultrasonic, often with little or no energy below 20 kHz. The findings presented reveal that the social signals produced by spinner and spotted dolphins span the full range of their hearing sensitivity, are spectrally quite varied, and in the case of burst pulses are probably produced more frequently than reported by audio-range analyses.
A behavioral response paradigm was used to measure underwater hearing thresholds in two California sea lions (Zalophus californianus) before and after exposure to underwater impulses from an arc-gap transducer. Preexposure and postexposure hearing thresholds were compared to determine if the subjects experienced temporary shifts in their masked hearing thresholds (MTTS). Hearing thresholds were measured at 1 and 10 kHz. Exposures consisted of single underwater impulses produced by an arc-gap transducer referred to as a "pulsed power device" (PPD). The electrical charge of the PPD was varied from 1.32 to 2.77 kJ; the distance between the subject and the PPD was varied over the range 3.4 to 25 m. No MTTS was observed in either subject at the highest received levels: peak pressures of approximately 6.8 and 14 kPa, rms pressures of approximately 178 and 183 dB re: 1 microPa, and total energy fluxes of 161 and 163 dB re: 1 microPa2s for the two subjects. Behavioral reactions to the tests were observed in both subjects. These reactions primarily consisted of temporary avoidance of the site where exposure to the PPD impulse had previously occurred.
In a psychophysical task with echoes that jitter in delay, big brown bats can detect changes as small as 10-20 ns at an echo signal-to-noise ratio of approximately 49 dB and 40 ns at approximately 36 dB. This performance is possible to achieve with ideal coherent processing of the wideband echoes, but it is widely assumed that the bat's peripheral auditory system is incapable of encoding signal waveforms to represent delay with the requisite precision or phase at ultrasonic frequencies. This assumption was examined by modeling inner-ear transduction with a bank of parallel bandpass filters followed by low-pass smoothing. Several versions of the filterbank model were tested to learn how the smoothing filters, which are the most critical parameter for controlling the coherence of the representation, affect replication of the bat's performance. When tested at a signal-to-noise ratio of 36 dB, the model achieved a delay acuity of 83 ns using a second-order smoothing filter with a cutoff frequency of 8 kHz. The same model achieved a delay acuity of 17 ns when tested with a signal-to-noise ratio of 50 dB. Jitter detection thresholds were an order of magnitude worse than the bat for fifth-order smoothing or for lower cutoff frequencies. Most surprising is that effectively coherent reception is possible with filter cutoff frequencies well below any of the ultrasonic frequencies contained in the bat's sonar sounds. The results suggest that only a modest rise in the frequency response of smoothing in the bat's inner ear can confer full phase sensitivity on subsequent processing and account for the bat's fine acuity or delay.
The West Indian manatee (trichechus manatus latirostris) has become endangered partly because of a growing number of collisions with boats. A system to warn boaters of the presence of manatees, that can signal to boaters that manatees are present in the immediate vicinity, could potentially reduce these boat collisions. In order to identify the presence of manatees, acoustic methods are employed. Within this paper, three different detection algorithms are used to detect the calls of the West Indian manatee. The detection systems are tested in the laboratory using simulated manatee vocalizations from an audio compact disk. The detection method that provides the best overall performance is able to correctly identify approximately 96% of the manatee vocalizations. However, the system also results in a false alarm rate of approximately 16%. The results of this work may ultimately lead to the development of a manatee warning system that can warn boaters of the presence of manatees.
The relationship between musical training and informational masking was studied for 24 young adult listeners with normal hearing. The listeners were divided into two groups based on musical training. In one group, the listeners had little or no musical training; the other group was comprised of highly trained, currently active musicians. The hypothesis was that musicians may be less susceptible to informational masking, which is thought to reflect central, rather than peripheral, limitations on the processing of sound. Masked thresholds were measured in two conditions, similar to those used by Kidd et al. [J. Acoust. Soc. Am. 95, 3475-3480 (1994)]. In both conditions the signal was comprised of a series of repeated tone bursts at 1 kHz. The masker was comprised of a series of multitone bursts, gated with the signal. In one condition the frequencies of the masker were selected randomly for each burst; in the other condition the masker frequencies were selected randomly for the first burst of each interval and then remained constant throughout the interval. The difference in thresholds between the two conditions was taken as a measure of informational masking. Frequency selectivity, using the notched-noise method, was also estimated in the two groups. The results showed no difference in frequency selectivity between the two groups, but showed a large and significant difference in the amount of informational masking between musically trained and untrained listeners. This informational masking task, which requires no knowledge specific to musical training (such as note or interval names) and is generally not susceptible to systematic short- or medium-term training effects, may provide a basis for further studies of analytic listening abilities in different populations.
Previous data on the masking level difference (MLD) have suggested that NoSpi detection for a long-duration signal is dominated by signal energy occurring in masker envelope minima. This finding was expanded upon using a brief 500-Hz tonal signal that coincided with either the envelope maximum or minimum of a narrow-band Gaussian noise masker centered at 500 Hz, and data were collected at a range of masker levels. Experiment 1 employed a typical MLD stimulus, consisting of a 30-ms signal and a 50-Hz-wide masker with abrupt spectral edges, and experiment 2 used stimuli generated to eliminate possible spectral cues. Results were quite similar for the two types of stimuli. At the highest masker level the MLD for signals coinciding with masker envelope minima was substantially larger than that for signals coinciding with envelope maxima, a result that was primarily due to decreased NoSpi thresholds in masker minima. For most observers this effect was greatly reduced or eliminated at the lowest masker level. These level effects are broadly consistent with the presence of physiological background noise and with a level-dependent binaural temporal window. Comparison of these results with predictions of a published model suggest that basilar-membrane compression alone does not account for this level effect.
The binaural interaction component (BIC=sum of monaural-true binaural) of the auditory brainstem response appears to reflect central binaural fusion/lateralization processes. Auditory middle-latency responses (AMLRs) are more robust and may reflect more completely such binaural processing. The AMLR also demonstrates such binaural interaction. The fusion of dichotically presented tones with an interaural frequency difference (IFD) offers another test of the extent to which electrophysiological and psychoacoustical measures agree. The effect of IFDs on both the BIC of the AMLR and a psychoacoustical measure of binaural fusion thus were examined. The perception of 20-ms tone bursts at/near 500 Hz with increasing IFDs showed, first, a deviated sound image from the center of the head, followed by clearly separate pitch percepts in each ear. Thresholds of detection of sound deviation and separation (i.e., nonfusion) were found to be 57 and 209 Hz, respectively. However, magnitudes of BICs of the AMLR were found to remain nearly. constant for IFDs up to the 400-Hz (limit of range tested), suggesting that the AMLR-BIC does not provide an objective index of this aspect of binaural processing, at least not under the conditions examined. The nature of lateralization due to IFDs and the concept of critical bands for binaural fusion are also discussed. Further research appears warranted to investigate the significance of the lack of effect of IFDs on the AMLR-BIC. Finally, the IFD paradigm itself would seem useful in that it permits determination of the limit for nonfusion of sounds presented binaurally, a limit not accessible via more conventional paradigms involving interaural time, phase, or intensity differences.
The gammatone filter was imported from auditory physiology to provide a time-domain version of the roex auditory filter and enable the development of a realistic auditory filterbank for models of auditory perception [Patterson et al., J. Acoust. Soc. Am. 98, 1890-1894 (1995)]. The gammachirp auditory filter was developed to extend the domain of the gammatone auditory filter and simulate the changes in filter shape that occur with changes in stimulus level. Initially, the gammachirp filter was limited to center frequencies in the 2.0-kHz region where there were sufficient "notched-noise" masking data to define its parameters accurately. Recently, however, the range of the masking data has been extended in two massive studies. This paper reports how a compressive version of the gammachirp auditory filter was fitted to these new data sets to define the filter parameters over the extended frequency range. The results show that the shape of the filter can be specified for the entire domain of the data using just six constants (center frequencies from 0.25 to 6.0 kHz and levels from 30 to 80 dB SPL). The compressive, gammachirp auditory filter also has the advantage of being consistent with physiological studies of cochlear filtering insofar as the compression of the filter is mainly limited to the passband and the form of the chirp in the impulse response is largely independent of level.
Vocal vibrato and tremor are characterized by oscillations in voice fundamental frequency (F0). These oscillations may be sustained by a control loop within the auditory system. One component of the control loop is the pitch-shift reflex (PSR). The PSR is a closed loop negative feedback reflex that is triggered in response to discrepancies between intended and perceived pitch with a latency of approximately 100 ms. Consecutive compensatory reflexive responses lead to oscillations in pitch every approximately 200 ms, resulting in approximately 5-Hz modulation of F0. Pitch-shift reflexes were elicited experimentally in six subjects while they sustained /u/ vowels at a comfortable pitch and loudness. Auditory feedback was sinusoidally modulated at discrete integer frequencies (1 to 10 Hz) with +/- 25 cents amplitude. Modulated auditory feedback induced oscillations in voice F0 output of all subjects at rates consistent with vocal vibrato and tremor. Transfer functions revealed peak gains at 4 to 7 Hz in all subjects, with an average peak gain at 5 Hz. These gains occurred in the modulation frequency region where the voice output and auditory feedback signals were in phase. A control loop in the auditory system may sustain vocal vibrato and tremorlike oscillations in voice F0.
Listeners' auditory discrimination of vowel sounds depends in part on the order in which stimuli are presented. Such presentation order effects have been argued to be language independent, and to result from psychophysical (not speech- or language-specific) factors such as the decay of memory traces over time or increased weighting of later-occurring stimuli. In the present study, native Cantonese speakers' discrimination of a linguistic tone continuum is shown to exhibit order of presentation effects similar to those shown for vowels in previous studies. When presented with two successive syllables differing in fundamental frequency by approximately 4 Hz, listeners were significantly more sensitive to this difference when the first syllable was higher in frequency than the second. However, American English-speaking listeners with no experience listening to Cantonese showed no such contrast effect when tested in the same manner using the same stimuli. Neither English nor Cantonese listeners showed any order of presentation effects in the discrimination of a nonspeech continuum in which tokens had the same fundamental frequencies as the Cantonese speech tokens but had a qualitatively non-speech-like timbre. These results suggest that tone presentation order effects, unlike vowel effects, may be language specific, possibly resulting from the need to compensate for utterance-related pitch declination when evaluating fundamental frequency for tone identification.
Experiments were conducted with a single, bilateral cochlear implant user to examine interaural level and time-delay cues that putatively underlie the design and efficacy of bilateral implant systems. The subject's two implants were of different types but custom equipment allowed presentation of controlled bilateral stimuli, particularly those with specified interaural time difference (ITD) and interaural level difference (ILD) cues. A lateralization task was used to measure the effect of these cues on the perceived location of the sensations elicited. For trains of fixed-amplitude, biphasic current pulses at 100 pps, the subject demonstrated sensitivity to an ITD of 300 micros, providing evidence of access to binaural information. The choice of bilateral electrode pair greatly influenced ITD sensitivity, suggesting that electrode pairings are likely to be an important consideration in the effort to provide binaural advantages. The selection of bilateral electrode pairs showing sensitivity to ITD was partially aided by comparisons of the pitch elicited by individual electrodes in each ear (when stimulated alone with fixed-amplitude current pulses at 813 pps): specifically, interaural electrodes with similar pitches were more likely (but not certain) to show ITD sensitivity. Significant changes in lateral position occurred with specific electrode pairs. With five bilateral electrode pairs of 14 tested, ITDs of 300 and 600 micros moved an auditory image significantly from right to left. With these same pairs, ILD changes of approximately 11% of the dynamic range (in microApp) moved an auditory image from the far left to the far right-significantly farther than the nine pairs not showing significant ITD sensitivity. However, even these nine pairs did show response changes as a function of the interaural (or confounding monaural) level cue. Overall, insofar as the access to bilateral cues demonstrated herein generalizes to other subjects, it provides hope that the normal binaural advantages for speech recognition and sound localization can be made available to bilateral implant users.