Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sound Spectrography”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,549 records · Page 86Linked to original sources

Multiple internal reflections in the cochlea and their effect on DPOAE fine structure.

In recent years, evidence has accumulated in support of a two-source model of distortion product otoacoustic emissions (DPOAEs). According to such models DPOAEs recorded in the ear canal are associated with two separate sources of cochlear origin. It is the interference between the contributions from the two sources that gives rise to the DPOAE fine structure (a pseudoperiodic change in DPOAE level or group delay with frequency). Multiple internal reflections between the base of the cochlea (oval window) and the DP tonotopic place can add additional significant components for certain stimulus conditions and thus modify the DPOAE fine structure. DPOAEs, at frequency increments between 4 and 8 Hz, were recorded at fixed f2/f1 ratios of 1.053, 1.065, 1.08, 1.11, 1.14, 1.18, 1.22, 1.26, 1.30, 1.32, 1.34, and 1.36 from four subjects. The resulting patterns of DPOAE amplitude and group delay (the negative of the slope of phase) revealed several previously unreported patterns in addition to the commonly reported log sine variation with frequency. These observed "exotic" patterns are predicted in computational simulations when multiple internal reflections are included. An inverse FFT algorithm was used to convert DPOAE data from the frequency to the "time" domain. Comparison of data in the time and frequency domains confirmed the occurrence of these "exotic" patterns in conjunction with the presence of multiple internal reflections. Multiple internal reflections were observed more commonly for high primary ratios (f2/f1 > or = 1.3). These results indicate that a full interpretation of the DPOAE level and phase (group delay) must include not only the two generation sources, but also multiple internal reflections.

Acoustic Stimulation↗

Echolocation signals of wild Atlantic spotted dolphin (Stenella frontalis).

An array of four hydrophones arranged in a symmetrical star configuration was used to measure the echolocation signals of the Atlantic spotted dolphin (Stenella frontalis) in the Bahamas. The spacing between the center hydrophone and the other hydrophones was 45.7 cm. A video camera was attached to the array and a video tape recorder was time synchronized with the computer used to digitize the acoustic signals. The echolocation signals had bi-modal frequency spectra with a low-frequency peak between 40 and 50 kHz and a high-frequency peak between 110 and 130 kHz. The low-frequency peak was dominant when the signal the source level was low and the high-frequency peak dominated when the source level was high. Peak-to-peak source levels as high as 210 dB re 1 microPa were measured. The source level varied in amplitude approximately as a function of the one-way transmission loss for signals traveling from the animals to the array. The characteristics of the signals were similar to those of captive Tursiops truncatus, Delphinapterus leucas and Pseudorca crassidens measured in open waters under controlled conditions.

Animals↗

Antimasking aspects of harp seal (Pagophilus groenlandicus) underwater vocalizations.

Underwater sounds are very important in social communication of harp seals (Pagophilus groenlandicus) because they are the main means of long- and short-distance communication. Individual harp seals must try to avoid being masked and emit only those calls that will benefit them. Underwater vocalizations of harp seals were recorded during the breeding season. The physical characteristics associated with antimasking attributes of 16 call types were examined. Rising frequency or increasing amplitude within calls were not common. Most of the calls ended abruptly (range 145-966 dB/s), but call onset was more gradual. At high calling rates (95.1-135 calls/min) there were significantly more calls overlapping temporally than at medium (75.1-95 calls/min) or low (35-75 calls/min) calling rates, but even at the highest calling rates, 79.1% of the calls were not overlapped. When 2, 3, or 4 calls overlapped, there were significantly fewer frequency separations of less than 1/3 octave than would be expected by chance. This is important because sounds that are separated by less than 1/3 octave likely mask each other. When 2-4 calls are occurring simultaneously, only 4.5% to 14.2% are masked by virtue of being within 1/3 octave from their nearest neighbor. None of the overlappping calls was of the same type. This suggests that the seals are actively listening to each other's calls and are not randomly using the different call types. Harp seals use frequency and temporal separation in conjunction with a wide vocal repertoire to avoid masking each other.

Animal Communication↗

An auditory-periphery model of the effects of acoustic trauma on auditory nerve responses.

Acoustic trauma degrades the auditory nerve's tonotopic representation of acoustic stimuli. Recent physiological studies have quantified the degradation in responses to the vowel /E/ and have investigated amplification schemes designed to restore a more correct tonotopic representation than is achieved with conventional hearing aids. However, it is difficult from the data to quantify how much different aspects of the cochlear pathology contribute to the impaired responses. Furthermore, extensive experimental testing of potential hearing aids is infeasible. Here, both of these concerns are addressed by developing models of the normal and impaired auditory peripheries that are tested against a wide range of physiological data. The effects of both outer and inner hair cell status on model predictions of the vowel data were investigated. The modeling results indicate that impairment of both outer and inner hair cells contribute to degradation in the tonotopic representation of the formant frequencies in the auditory nerve. Additionally, the model is able to predict the effects of frequency-shaping amplification on auditory nerve responses, indicating the model's potential suitability for more rapid development and testing of hearing aid schemes.

Animals↗

Effect of current stimulus on in vivo cochlear mechanics.

In this paper, the influence of direct current stimulation on the acoustic impulse response of the basilar membrane (BM) is studied. A positive current applied in the scala vestibuli relative to a ground electrode in the scala tympani is found to enhance gain and increase the best frequency at a given location on the BM. An opposite effect is found for a negative current. Also, the amplitude of low-frequency cochlear microphonic at high sound levels is found to change with the concurrent application of direct current stimulus. BM vibrations in response to pure tone acoustic excitation are found to possess harmonics whose levels relative to the fundamental increase with the application of positive current and decrease with the application of negative current. A model for outer hair cell activity that couples changes in length and stiffness to transmembrane potential is used to interpret the results of these experiments and others in the literature. The importance of the in vivo mechanical and electrical loading is emphasized. Simulation results show the somewhat paradoxical finding that for outer hair cells under tension, hyperpolarization causes shortening of the cell length due to the dominance of voltage dependent stiffness changes.

Acoustic Stimulation↗

Spectral models of additive and modulation noise in speech and phonatory excitation signals.

The article presents spectral models of additive and modulation noise in speech. The purpose is to learn about the causes of noise in the spectra of normal and disordered voices and to gauge whether the spectral properties of the perturbations of the phonatory excitation signal can be inferred from the spectral properties of the speech signal. The approach to modeling consists of deducing the Fourier series of the perturbed speech, assuming that the Fourier series of the noise and of the clean monocycle-periodic excitation are known. The models explain published data, take into account the effects of supraglottal tremor, demonstrate the modulation distortion owing to vocal tract filtering, establish conditions under which noise cues of different speech signals may be compared, and predict the impossibility of inferring the spectral properties of the frequency modulating noise from the spectral properties of the frequency modulation noise (e.g., phonatory jitter and frequency tremor). The general conclusion is that only phonatory frequency modulation noise is spectrally relevant. Other types of noise in speech are either epiphenomenal, or their spectral effects are masked by the spectral effects of frequency modulation noise.

Fourier Analysis↗

Effects of prosodic boundary on /aC/ sequences: acoustic results.

This study presents various acoustic measures used to examine the sequence /a # C/, where "#" represents different prosodic boundaries in French. The 6 consonants studied are /b d g f s S/ (3 stops and 3 fricatives). The prosodic units investigated are the utterance, the intonational phrase, the accentual phrase, and the word. It is found that vowel target values, formant transitions into the stop consonant, and the rate of change in spectral tilt into the fricative, are affected by the strength of the prosodic boundary. F1 becomes higher for /a/ the stronger the prosodic boundary, with the exception of one speaker's utterance data, which show the effects of articulatory declension at the utterance level. Various effects of the stop consonant context are observed, the most notable being a tendency for the vowel /a/ to be displaced in the direction of the F2 consonant "locus" for /d/ (the F2 consonant values for which remain relatively stable across prosodic boundaries) and for /g/ (the F2 consonant values for which are displaced in the direction of the velar locus in weaker prosodic boundaries, together with those of the vowel). Velocity of formant transition may be affected by prosodic boundary (with greater velocity at weaker boundaries), though results are not consistent across speakers. There is also a tendency for the rate of change in spectral tilt moving from the vowel to the fricative to be affected by the presence of a prosodic boundary, with a greater rate of change at the weaker prosodic boundaries. It is suggested that spectral cues, in addition to duration, amplitude, and F0 cues, may alert listeners to the presence of a prosodic boundary.

Adult↗

Unfolding of phonetic information over time: a database of Dutch diphone perception.

We present the results of a large-scale study on speech perception, assessing the number and type of perceptual hypotheses which listeners entertain about possible phoneme sequences in their language. Dutch listeners were asked to identify gated fragments of all 1179 diphones of Dutch, providing a total of 488,520 phoneme categorizations. The results manifest orderly uptake of acoustic information in the signal. Differences across phonemes in the rate at which fully correct recognition was achieved arose as a result of whether or not potential confusions could occur with other phonemes of the language (long with short vowels, affricates with their initial components, etc.). These data can be used to improve models of how acoustic-phonetic information is mapped onto the mental lexicon during speech comprehension.

Adult↗

The synergy between speech production and perception.

Speech intelligibility is known to be relatively unaffected by certain deformations of the acoustic spectrum. These include translations, stretching or contracting dilations, and shearing of the spectrum (represented along the logarithmic frequency axis). It is argued here that such robustness reflects a synergy between vocal production and auditory perception. Thus, on the one hand, it is shown that these spectral distortions are produced by common and unavoidable variations among different speakers pertaining to the length, cross-sectional profile, and losses of their vocal tracts. On the other hand, it is argued that these spectral changes leave the auditory cortical representation of the spectrum largely unchanged except for translations along one of its representational axes. These assertions are supported by analyses of production and perception models. On the production side, a simplified sinusoidal model of the vocal tract is developed which analytically relates a few "articulatory" parameters, such as the extent and location of the vocal tract constriction, to the spectral peaks of the acoustic spectra synthesized from it. The model is evaluated by comparing the identification of synthesized sustained vowels to labeled natural vowels extracted from the TIMIT corpus. On the perception side a "multiscale" model of sound processing is utilized to elucidate the effects of the deformations on the representation of the acoustic spectrum in the primary auditory cortex. Finally, the implications of these results for the perception of generally identifiable classes of sound sources beyond the specific case of speech and the vocal tract are discussed.

Auditory Cortex↗

Measuring hearing in the harbor seal (Phoca vitulina): comparison of behavioral and auditory brainstem response techniques.

Auditory brainstem response (ABR) and standard behavioral methods were compared by measuring in-air audiograms for an adult female harbor seal (Phoca vitulina). Behavioral audiograms were obtained using two techniques: the method of constant stimuli and the staircase method. Sensitivity was tested from 0.250 to 30 kHz. The seal showed good sensitivity from 6 to 12 kHz [best sensitivity 8.1 dB (re 20 microPa2 x s) RMS at 8 kHz]. The staircase method yielded thresholds that were lower by 10 dB on average than the method of constant stimuli. ABRs were recorded at 2, 4, 8, 16, and 22 kHz and showed a similar best range (8-16 kHz). ABR thresholds averaged 5.7 dB higher than behavioral thresholds at 2, 4, and 8 kHz. ABRs were at least 7 dB lower at 16 kHz, and approximately 3 dB higher at 22 kHz. The better sensitivity of ABRs at higher frequencies could have reflected differences in the seal's behavior during ABR testing and/or bandwidth characteristics of test stimuli. These results agree with comparisons of ABR and behavioral methods performed in other recent studies and indicate that ABR methods represent a good alternative for estimating hearing range and sensitivity in pinnipeds, particularly when time is a critical factor and animals are untrained.

Animals↗

Echolocation in the Risso's dolphin, Grampus griseus.

The Risso's dolphin (Grampus griseus) is an exclusively cephalopod-consuming delphinid with a distinctive vertical indentation along its forehead. To investigate whether or not the species echolocates, a female Risso's dolphin was trained to discriminate an aluminum cylinder from a nylon sphere (experiment 1) or an aluminum sphere (experiment 2) while wearing eyecups and free swimming in an open-water pen in Kaneohe Bay, Hawaii. The dolphin completed the task with little difficulty despite being blindfolded. Clicks emitted by the dolphin were acquired at average amplitudes of 192.6 dB re 1 microPa, with estimated sources levels up to 216 dB re 1 microPa-1 m. Clicks were acquired with peak frequencies as high as 104.7 kHz (Mf(p) = 47.9 kHz), center frequencies as high as 85.7 kHz (Mf(0) = 56.5 kHz), 3-dB bandwidths up to 94.1 kHz (M(BW) = 39.7 kHz), and root-mean-square bandwidths up to 32.8 kHz (M(RMS) = 23.3 kHz). Click durations were between 40 and 70 micros. The data establish that the Risso's dolphin echolocates, and that, aside from slightly lower amplitudes and frequencies, the clicks emitted by the dolphin were similar to those emitted by other echolocating odontocetes. The particular acoustic and behavioral findings in the study are discussed with respect to the possible direction of the sonar transmission beam of the species.

Animals↗

Individual talker differences in voice-onset-time.

Individual talkers differ in the acoustic properties of their speech, and at least some of these differences are in acoustic properties relevant for phonetic perception. Recent findings from studies of speech perception have shown that listeners can exploit such differences to facilitate both the recognition of talkers' voices and the recognition of words spoken by familiar talkers. These findings motivate the current study, whose aim is to examine individual talker variation in a particular phonetically-relevant acoustic property, voice-onset-time (VOT). VOT is a temporal property that robustly specifies voicing in stop consonants. From the broad literature involving VOT, it appears that individual talkers differ from one another in their VOT productions. The current study confirmed this finding for eight talkers producing monosyllabic words beginning with voiceless stop consonants. Moreover, when differences in VOT due to variability in speaking rate across the talkers were factored out using hierarchical linear modeling, individual talkers still differed from one another in VOT, though these differences were attenuated. These findings provide evidence that VOT varies systematically from talker to talker and may therefore be one phonetically-relevant acoustic property underlying listeners' capacity to benefit from talker-specific experience.

Adolescent↗

The influence of flight speed on the ranging performance of bats using frequency modulated echolocation pulses.

Many species of bat use ultrasonic frequency modulated (FM) pulses to measure the distance to objects by timing the emission and reception of each pulse. Echolocation is mainly used in flight. Since the flight speed of bats often exceeds 1% of the speed of sound, Doppler effects will lead to compression of the time between emission and reception as well as an elevation of the echo frequencies, resulting in a distortion of the perceived range. This paper describes the consequences of these Doppler effects on the ranging performance of bats using different pulse designs. The consequences of Doppler effects on ranging performance described in this paper assume bats to have a very accurate ranging resolution, which is feasible with a filterbank receiver. By modeling two receiver types, it was first established that the effects of Doppler compression are virtually independent of the receiver type. Then, used a cross-correlation model was used to investigate the effect of flight speed on Doppler tolerance and range-Doppler coupling separately. This paper further shows how pulse duration, bandwidth, function type, and harmonics influence Doppler tolerance and range-Doppler coupling. The influence of each signal parameter is illustrated using calls of several bat species. It is argued that range-Doppler coupling is a significant source of error in bat echolocation, and various strategies bats could employ to deal with this problem, including the use of range rate information are discussed.

Acceleration↗

Effect of amplitude modulation coherence for masked speech signals filtered into narrow bands.

Introduction of masker amplitude modulation (AM) can improve signal detection in a number of paradigms. In some cases this advantage depends on the coherence of modulation across a relatively wide frequency range. In the experiments described below, observers were asked to identify masked spondee words produced by a single male talker. The target spondees and masking noise were filtered into nine narrow bands, and the coherence of AM of either the speech signal or noise masker was manipulated. Inherent modulation of the masker bands was manipulated via assignment of real and imaginary values to the associated components of each band in the frequency domain, and AM of speech bands was achieved via multiplication with envelopes extracted from these maskers. Responses were based on two alternatives, four alternatives, or open response sets. The effect of masker AM coherence was highly dependent upon the size of the response set: coherent AM was associated with better thresholds in a two-alternative response set, but poorer thresholds in an open response set. Results with AM speech did not depend critically upon the across-frequency temporal synchrony of AM imposed on the speech material.

Adult↗

Laryngeal biomechanics and vocal communication in the squirrel monkey (Saimiri boliviensis).

The larynges of eight squirrel monkeys were harvested, dissected, mounted on a pseudotracheal tube, and phonated using compressed air. Patterns of vocal fold oscillation were compared with sound spectrograms of calls recorded from monkeys in our colony. Four different regimes of vocal fold activation were identified. Regime 1 resembled typical human vowel production, with regular vocal-fold vibration, a prominent fundamental frequency, and an accompanying series of harmonic overtones. This regime is likely to give rise to squirrel monkey "cackles," as well as a variety of other harmonically structured calls. In regime 2, the pattern of vibrations exhibited the presence of two or more unrelated frequencies (biphonation). This regime of glottal activity resembled the biphonation observed in many exemplars of "twitter" and "kecker" calls. The vocal folds oscillated continuously in regime 3, but produced glottal pulses whose amplitudes waxed and waned rhythmically. This phenomenon resulted in the percept of a series of discrete pulses, and may give rise to "errs," "churrs," and other calls composed of a rapid sequence of acoustic elements. In regime 4, the period of each oscillation was quasi-irregular. Shrieks and other broadband calls or call elements that lack an apparent fundamental frequency may be produced in this manner.

Animals↗