Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sound Spectrography”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,369 records · Page 76Linked to original sources

Source ranging with minimal environmental information using a virtual receiver and waveguide invariant theory.

A method is presented for estimating the range of an unknown broadband acoustic source in a waveguide, using a vertical array and a signal sample from another broadband source at a known location relative to the array. The method requires no modeling of the acoustic field, and little to no environmental information for flat bathymetries. Waveguide invariant theory [e.g., D'Spain and Kuperman, J. Acoust. Soc. Am 106, 2454-2468 (1999)] is applied to the "virtual receiver" [Siderius et al., J. Acoust. Soc. Am. 102, 3439-3449 (1997)] to create a "virtual aperture" (VA). In effect, the method effectively converts a source at known range r(g) into a continuum of receivers lying between ranges (1 +/- alpha/beta)*r(g), where beta is a scalar parameter called the acoustic invariant, and alpha approximately 0.1. This effective displacement is achieved by correlating the known source field, measured at frequency component omega, with the unknown source field, measured at frequency component omega + omegas. When the VA output is plotted as a function of omega and omegas, the slope of the resulting correlation contours yields the unknown source range. The concept is illustrated via both simulation and analysis of data collected from a pseudo-random noise source with 75-150-Hz bandwidth during SWellEx-3, a shallow water experiment conducted off the San Diego coast. The virtual aperture can be reformulated for range-dependent environments, if adiabatic propagation assumptions are valid, and if the bathymetry surrounding the array is known.

Acoustics↗

Sounds produced by Australian Irrawaddy dolphins, Orcaella brevirostris.

Sounds produced by Irrawaddy dolphins, Orcaella brevirostris, were recorded in coastal waters off northern Australia. They exhibit a varied repertoire, consisting of broadband clicks, pulsed sounds and whistles. Broad-band clicks, "creaks" and "buzz" sounds were recorded during foraging, while "squeaks" were recorded only during socializing. Both whistle types were recorded during foraging and socializing. The sounds produced by Irrawaddy dolphins do not resemble those of their nearest taxonomic relative, the killer whale, Orcinus orca. Pulsed sounds appear to resemble those produced by Sotalia and nonwhistling delphinids (e.g., Cephalorhynchus spp.). Irrawaddy dolphins exhibit a vocal repertoire that could reflect the acoustic specialization of this species to its environment.

Animal Communication↗

Behavioral responses of humpback whales (Megaptera novaeangliae) to full-scale ATOC signals.

Loud (195 dB re 1 microPa at 1 m) 75-Hz signals were broadcast with an ATOC projector to measure ocean temperature. Respiratory and movement behaviors of humpback whales off North Kauai, Hawaii, were examined for potential changes in response to these transmissions and to vessels. Few vessel effects were observed, but there were fewer vessels operating during this study than in previous years. No overt responses to ATOC were observed for received levels of 98-109 dB re 1 microPa. An analysis of covariance, using the no-sound behavioral rate as a covariate to control for interpod variation, found that the distance and time between successive surfacings of humpbacks increased slightly with an increase in estimated received ATOC sound level. These responses are very similar to those observed in response to scaled-amplitude playbacks of ATOC signals [Frankel and Clark, Can. J. Zool. 76, 521-535 (1998)]. These similar results were obtained with different sound projectors, in different years and locations, and at different ranges creating a different sound field. The repeatability of the findings for these two different studies indicates that these effects, while small, are robust. This suggests that at least for the ATOC signal, the received sound level is a good predictor of response.

Animals↗

A conjugated infinite element method for half-space acoustic problems.

Many acoustic problems (especially in environmental acoustics) involve half-space domains bounded by a plane subjected to normal admittance boundary conditions. In the "low" frequency domain, the numerical treatment of such problems usually relies on boundary element methods based on a particular Green's function suited for the half-(admittance) plane. In the present paper, an alternative hybrid finite/infinite element scheme is proposed. The method relies on a direct treatment of nonhomogeneous boundary conditions along infinite element edges (or faces). The procedure is validated through comparisons with an available reference solution.

Acoustics↗

Chinese dialect identification using segmental and prosodic features.

Several approaches to Chinese dialect identification based on segmental and prosodic features of speech are described in this paper. When using segmental information only, the system performs phonotactic analysis after speech utterances have been tokenized into sequences of broad phonetic classes. The second scheme comprises prosodic models which are trained to capture tone sequence information for individual dialects. Also proposed is a novel approach that examines differences between Chinese dialects at broad phonetic and prosodic levels. These algorithms were evaluated via a multispeaker read-speech mode. Simulation results indicate that the combined use of segmental and prosodic features allows the proposed system to discriminate among three major Chinese dialects spoken in Taiwan with 93.0% accuracy.

Adolescent↗

Analysis of acoustic communication by ants.

An analysis is presented of acoustic communication by ants, based on near-field theory and on data obtained from the black imported fire ant Solenopsis richteri and other sources. Generally ant stridulatory sounds are barely audible, but they occur continuously in ant colonies. Because ants appear unresponsive to airborne sound, myrmecologists have concluded that stridulatory signals are transmitted through the substrate. However, transmission through the substrate is unlikely, for reasons given in the paper. Apparently ants communicate mainly through the air, and the acoustic receptors are hairlike sensilla on the antennae that respond to particle sound velocity. This may seem inconsistent with the fact that ants are unresponsive to airborne sound (on a scale of meters), but the inconsistency can be resolved if acoustic communication occurs within the near field, on a scale of about 100 mm. In the near field, the particle sound velocity is significantly enhanced and has a steep gradient. These features can be used to exclude extraneous sound, and to determine the direction and distance of a near-field source. Additionally, we observed that the tracheal air sacs of S. richteri can expand within the gaster, possibly amplifying the radiation of stridulatory sound.

Animal Communication↗

Localization of multiple sound sources with two microphones.

This paper presents a two-microphone technique for localization of multiple sound sources. Its fundamental structure is adopted from a binaural signal-processing scheme employed in biological systems for the localization of sources using interaural time differences (ITD). The two input signals are transformed to the frequency domain and analyzed for coincidences along left/right-channel delay-line pairs. The coincidence information is enhanced by a nonlinear operation followed by a temporal integration. The azimuths of the sound sources are estimated by integrating the coincidence locations across the broadband of frequencies in speech signals (the "direct" method). Further improvement is achieved by using a novel "stencil" filter pattern recognition procedure. This includes coincidences due to phase delays of greater than 2pi, which are generally regarded as ambiguous information. It is demonstrated that the stencil method can greatly enhance localization of lateral sources over the direct method. Also discussed and analyzed are two limitations involved in both methods, namely missed and artifactual sound sources. Anechoic chamber tests as well as computer simulation experiments showed that the signal-processing system generally worked well in detecting the spatial azimuths of four or six simultaneously competing sound sources.

Adult↗

Nonlinear interactions that could explain distortion product interference response areas.

Suppression and/or enhancement of third- and fifth-order distortion products by a third tone that can have a frequency more than an octave above and a level more than 40 dB below the primary tones have recently been measured by Martin et al. [Hear. Res. 136, 105-123 (1999)]. Contours of iso-suppression and iso-enhancement that are plotted as a function of third-tone frequency and level are called interference response areas. After ruling out order aliasing, two possible mechanisms for this effect have been developed, a harmonic mechanism and a catalyst mechanism. The harmonic mechanism produces distortion products by mixing a harmonic of one of the primary tones with the other primary tone. The catalyst mechanism produces distortion products by mixing one or more intermediate distortion products that are produced by the third tone with one or more of the input tones. The harmonic mechanism does not need a third tone and the catalyst mechanism does. Because the basilar membrane frequency response is predicted to affect each of these mechanisms differently, it is concluded that the catalyst mechanism will be dominant in the high-frequency regions of the cochlea and the harmonic mechanism will have significant strength in the low-frequency regions of the cochlea. The mechanisms are dependent on the existence of both even- and odd-order distortion, and significant even- and odd-order distortion have been measured in the experimental animals. Furthermore, the nonlinear part of the cochlear mechanical response must be well into saturation when input tones are 50 or more dB SPL.

Acoustics↗

Acoustic noise during functional magnetic resonance imaging.

Functional magnetic resonance imaging (fMRI) enables sites of brain activation to be localized in human subjects. For studies of the auditory system, acoustic noise generated during fMRI can interfere with assessments of this activation by introducing uncontrolled extraneous sounds. As a first step toward reducing the noise during fMRI, this paper describes the temporal and spectral characteristics of the noise present under typical fMRI study conditions for two imagers with different static magnetic field strengths. Peak noise levels were 123 and 138 dB re 20 microPa in a 1.5-tesla (T) and a 3-T imager, respectively. The noise spectrum (calculated over a 10-ms window coinciding with the highest-amplitude noise) showed a prominent maximum at 1 kHz for the 1.5-T imager (115 dB SPL) and at 1.4 kHz for the 3-T imager (131 dB SPL). The frequency content and timing of the most intense noise components indicated that the noise was primarily attributable to the readout gradients in the imaging pulse sequence. The noise persisted above background levels for 300-500 ms after gradient activity ceased, indicating that resonating structures in the imager or noise reverberating in the imager room were also factors. The gradient noise waveform was highly repeatable. In addition, the coolant pump for the imager's permanent magnet and the room air-handling system were sources of ongoing noise lower in both level and frequency than gradient coil noise. Knowledge of the sources and characteristics of the noise enabled the examination of general approaches to noise control that could be applied to reduce the unwanted noise during fMRI sessions.

Acoustics↗

Monaural and binaural detection of sinusoidal phase modulation of a 500-Hz tone.

The detectability of phase modulation was measured for three subjects in two-alternative temporal forced-choice experiments. In experiment 1, the detectability of sinusoidal phase modulation in a 1500-ms burst of an 80-dB (SPL), 500-Hz sinusoidal carrier presented to the left ear (monaural condition) was measured. The experiment was repeated with an 80-dB, 500-Hz static (unmodulated) tone at the right ear (dichotic condition). At a modulation rate of 1 Hz, subjects were an order of magnitude more sensitive to phase modulation in the dichotic condition than in the monaural condition. The dichotic advantage decreased monotonically with increasing modulation rate. Subjects ceased to detect movement in the dichotic stimulus above 10 Hz, but a dichotic advantage remained up to a modulation rate of 40 Hz. Thus, although sound movement detection is sluggish, detection of internal phase modulation is not. In experiment 2, thresholds for detecting 2-Hz phase modulation were measured in the dichotic condition as a function of the level of the pure tone in the right ear. The dichotic advantage persisted even when the level of the pure tone was reduced by 50 dB or more. The findings demonstrate a large dichotic advantage which persists to high modulation rates and which depends very little on interaural level differences.

Adult↗

Localization of brief sounds: effects of level and background noise.

Listeners show systematic errors in vertical-plane localization of wide-band sounds when tested with brief-duration stimuli at high intensities, but long-duration sounds at any comfortable level do not produce such errors. Improvements in high-level sound localization associated with increased stimulus duration might result from temporal integration or from adaptation that might allow reliable processing of later portions of the stimulus. Free-field localization judgments were obtained for clicks and for 3- and 100-ms noise bursts presented at sensation levels from 30 to 55 dB. For the brief (clicks and 3-ms) stimuli, listeners showed compression of elevation judgments and increased rates and unusual patterns of front/back confusion at sensation levels higher than 40-45 dB. At lower sensation levels, brief sounds were localized accurately. The localization task was repeated using 3-ms noise burst targets in a background of spatially diffuse, wide-band noise intended to pre-adapt the system prior to the target onset. For high-level targets, the addition of background noise afforded mild release from the elevation compression effect. Finally, a train of identical, high-level, 3-ms bursts was found to be localized more accurately than a single burst. These results support the adaptation hypothesis.

Adolescent↗

Application of a finite-element model to low-frequency sound insulation in dwellings.

The sound transmission between adjacent rooms has been modeled using a finite-element method. Predicted sound-level difference gave good agreement with experimental data using a full-scale and a quarter-scale model. Results show that the sound insulation characteristics of a party wall at low frequencies strongly depend on the modal characteristics of the sound field of both rooms and of the partition. The effect of three edge conditions of the separating wall on the sound-level difference at low frequencies was examined: simply supported, clamped, and a combination of clamped and simply supported. It is demonstrated that a clamped partition provides greater sound-level difference at low frequencies than a simply supported. It also is confirmed that the sound-pressure level difference is lower in equal room than in unequal room configurations.

Acoustics↗

On the relationships between the fixed-f1, fixed-f2, and fixed-ratio phase derivatives of the 2f1-f2 distortion product otoacoustic emission.

For primary frequency ratios, f2/f1, in the range 1.1-1.3, the fixed-f1 ("f2-sweep") phase derivative of the 2f1-f2 distortion product otoacoustic emission (DPOAE) is larger than the fixed-f2("f1-sweep") one. It has been proposed by some researchers that part or all of the difference between these delays may be attributed to the so-called cochlear filter "build-up" or response time in the DPOAE generation region around the f2 tonotopic site. The analysis of an approximate theoretical expression for the DPOAE signal [Talmadge et al., J. Acoust. Soc. Am. 104, 1517-1543 (1998)] shows that the contributions to the phase derivatives associated with the cochlear filter response is small. It is also shown that the difference between the phase derivatives can be qualitatively accounted for by assuming the approximate scale invariance of cochlear mechanics. The effects of DPOAE fine structure on the phase derivative are also explored, and it is found that the interpretation of the phase derivative in terms of the phase variation of a single DPOAE component can be quite problematic.

Humans↗

Effects of the salience of pitch and periodicity information on the intelligibility of four-channel vocoded speech: implications for cochlear implants.

Recent simulations of continuous interleaved sampling (CIS) cochlear implant speech processors have used acoustic stimulation that provides only weak cues to pitch, periodicity, and aperiodicity, although these are regarded as important perceptual factors of speech. Four-channel vocoders simulating CIS processors have been constructed, in which the salience of speech-derived periodicity and pitch information was manipulated. The highest salience of pitch and periodicity was provided by an explicit encoding, using a pulse carrier following fundamental frequency for voiced speech, and a noise carrier during voiceless speech. Other processors included noise-excited vocoders with envelope cutoff frequencies of 32 and 400 Hz. The use of a pulse carrier following fundamental frequency gave substantially higher performance in identification of frequency glides than did vocoders using envelope-modulated noise carriers. The perception of consonant voicing information was improved by processors that preserved periodicity, and connected discourse tracking rates were slightly faster with noise carriers modulated by envelopes with a cutoff frequency of 400 Hz compared to 32 Hz. However, consonant and vowel identification, sentence intelligibility, and connected discourse tracking rates were generally similar through all of the processors. For these speech tasks, pitch and periodicity beyond the weak information available from 400 Hz envelope-modulated noise did not contribute substantially to performance.

Cochlear Implants↗

Spontaneous speech recognition using a statistical coarticulatory model for the vocal-tract-resonance dynamics.

A statistical coarticulatory model is presented for spontaneous speech recognition, where knowledge of the dynamic, target-directed behavior in the vocal tract resonance is incorporated into the model design, training, and in likelihood computation. The principal advantage of the new model over the conventional HMM is the use of a compact, internal structure that parsimoniously represents long-span context dependence in the observable domain of speech acoustics without using additional, context-dependent model parameters. The new model is formulated mathematically as a constrained, nonstationary, and nonlinear dynamic system, for which a version of the generalized EM algorithm is developed and implemented for automatically learning the compact set of model parameters. A series of experiments for speech recognition and model synthesis using spontaneous speech data from the Switchboard corpus are reported. The promise of the new model is demonstrated by showing its consistently superior performance over a state-of-the-art benchmark HMM system under controlled experimental conditions. Experiments on model synthesis and analysis shed insight into the mechanism underlying such superiority in terms of the target-directed behavior and of the long-span context-dependence property, both inherent in the designed structure of the new dynamic model of speech.

Algorithms↗

Echolocation behavior of big brown bats, Eptesicus fuscus, in the field and the laboratory.

Echolocation signals were recorded from big brown bats, Eptesicus fuscus, flying in the field and the laboratory. In open field areas the interpulse intervals (IPI) of search signals were either around 134 ms or twice that value, 270 ms. At long IPI's the signals were of long duration (14 to 18-20 ms), narrow bandwidth, and low frequency, sweeping down to a minimum frequency (Fmin) of 22-25 kHz. At short IPI's the signals were shorter (6-13 ms), of higher frequency, and broader bandwidth. In wooded areas only short (6-11 ms) relatively broadband search signals were emitted at a higher rate (avg. IPI= 122 ms) with higher Fmin (27-30 kHz). In the laboratory the IPI was even shorter (88 ms), the duration was 3-5 ms, and the Fmin 30- 35 kHz, resembling approach phase signals of field recordings. Excluding terminal phase signals, all signals from all areas showed a negative correlation between signal duration and Fmin, i.e., the shorter the signal, the higher was Fmin. This correlation was reversed in the terminal phase of insect capture sequences, where Fmin decreased with decreasing signal duration. Overall, the signals recorded in the field were longer, with longer IPI's and greater variability in bandwidth than signals recorded in the laboratory.

Animals↗

The influence of interaural stimulus uncertainty on binaural signal detection.

This paper investigated the influence of stimulus uncertainty in binaural detection experiments and the predictions of several binaural models for such conditions. Masked thresholds of a 500-Hz sinusoid were measured in an NrhoSpi condition for both running and frozen-noise maskers using a three interval, forced-choice (3IFC) procedure. The nominal masker correlation varied between 0.64 and 1, and the bandwidth of the masker was either 10, 100, or 1,000 Hz. The running-noise thresholds were expected to be higher than the frozen-noise thresholds because of stimulus uncertainty in the running-noise conditions. For an interaural correlation close to +1, no difference between frozen-noise and running-noise thresholds was expected for all values of the masker bandwidth. These expectations were supported by the experimental data: for interaural correlations less than 1.0, substantial differences between frozen and running-noise conditions were observed for bandwidths of 10 and 100 Hz. Two additional conditions were tested to further investigate the influence of stimulus uncertainty. In the first condition a different masker sample was chosen on each trial, but the correlation of the masker was forced to a fixed value. In the second condition one of two independent frozen-noise maskers was randomly chosen on each trial. Results from these experiments emphasized the influence of stimulus uncertainty in binaural detection tasks: if the degree of uncertainty in binaural cues was reduced, thresholds decreased towards thresholds in the conditions without any stimulus uncertainty. In the analysis of the data, stimulus uncertainty was expressed in terms of three theories of binaural processing: the interaural correlation, the EC theory, and a model based on the processing of interaural intensity differences (IIDs) and interaural time differences (ITDs). This analysis revealed that none of the theories tested could quantitatively account for the observed thresholds. In addition, it was found that, in conditions with stimulus uncertainty, predictions based on correlation differ from those based on the EC theory.

Attention↗

Independence of frequency channels in auditory temporal gap detection.

The ability of listeners to detect a temporal gap in a 1600-Hz-wide noiseband (target) was studied as a function of the absence and presence of concurrent stimulation by a second 1600-Hz-wide noiseband (distractor) with a nonoverlapping spectrum. Gap detection thresholds for single noisebands centered on 1.0, 2.0, 4.0, and 5.0 kHz were in the range from 4 to 6 ms, and were comparable to those described in previous studies. Gap thresholds for the same target noisebands were only modestly improved by the presence of a synchronously gated gap in a second frequency band. Gap thresholds were unaffected by the presence of a continuous distractor that was either proximate or remote from the target frequency band. Gap thresholds for the target noiseband were elevated if the distractor noiseband also contained a gap which "roved" in time in temporal proximity to the target gap. This effect was most marked in inexperienced listeners. Between-channel gap thresholds, obtained using leading and trailing markers that differed in frequency, were high in all listeners, again consistent with previous findings. The data are discussed in terms of the levels of the auditory perceptual processing stream at which the listener can voluntarily access auditory events in distinct frequency channels.

Adult↗