Sounds produced by individual white whales, Delphinapterus leucas, from Svalbard during capture.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Modifying the vocal tract alters a speaker's previously learned acoustic-articulatory relationship. This study investigated the contribution of auditory feedback to the process of adapting to vocal-tract modifications. Subjects said the word /tas/ while wearing a dental prosthesis that extended the length of their maxillary incisor teeth. The prosthesis affected /s/ productions and the subjects were asked to learn to produce "normal" /s/'s. They alternately received normal auditory feedback and noise that masked their natural feedback during productions. Acoustic analysis of the speakers' /s/ productions showed that the distribution of energy across the spectra moved toward that of normal, unperturbed production with increased experience with the prosthesis. However, the acoustic analysis did not show any significant differences in learning dependent on auditory feedback. By contrast, when naive listeners were asked to rate the quality of the speakers' utterances, productions made when auditory feedback was available were evaluated to be closer to the subjects' normal productions than when feedback was masked. The perceptual analysis showed that speakers were able to use auditory information to partially compensate for the vocal-tract modification. Furthermore, utterances produced during the masked conditions also improved over a session, demonstrating that the compensatory articulations were learned and available after auditory feedback was removed.
Bottlenose dolphins (Tursiops truncatus) detect and discriminate underwater objects by interrogating the environment with their native echolocation capabilities. Study of dolphins' ability to detect complex (multihighlight) signals in noise suggest echolocation object detection using an approximate 265-micros energy integration time window sensitive to the echo region of highest energy or containing the highlight with highest energy. Backscatter from many real objects contains multiple highlights, distributed over multiple integration windows and with varying amplitude relationships. This study used synthetic echoes with complex highlight structures to test whether high-amplitude initial highlights would interfere with discrimination of low-amplitude trailing highlights. A dolphin was trained to discriminate two-highlight synthetic echoes using differences in the center frequencies of the second highlights. The energy ratio (delta dB) and the timing relationship (delta T) between the first and second highlights were manipulated. An iso-sensitivity function was derived using a factorial design testing delta dB at -10, -15, -20, and -25 dB and delta T at 10, 20, 40, and 80 micros. The results suggest that the animal processed multiple echo highlights as separable analyzable features in the discrimination task, perhaps perceived through differences in spectral rippling across the duration of the echoes.
Training American listeners to perceive Mandarin tones has been shown to be effective, with trainees' identification improving by 21%. Improvement also generalized to new stimuli and new talkers, and was retained when tested six months after training [Y. Wang et al., J. Acoust. Soc. Am. 106, 3649-3658 (1999)]. The present study investigates whether the tone contrasts gained perceptually transferred to production. Before their perception pretest and after their post-test, the trainees were recorded producing a list of Mandarin words. Their productions were first judged by native Mandarin listeners in an identification task. Identification of trainees' post-test tone productions improved by 18% relative to their pretest productions, indicating significant tone production improvement after perceptual training. Acoustic analyses of the pre- and post-training productions further reveal the nature of the improvement, showing that post-training tone contours approximate native norms to a greater degree than pretraining tone contours. Furthermore, pitch height and pitch contour are not mastered in parallel, with the former being more resistant to improvement than the latter. These results are discussed in terms of the relationship between non-native tone perception and production as well as learning at the suprasegmental level.
An interrupted noise exposure of sufficient intensity, presented on a daily repeating cycle, produces a threshold shift (TS) following the first day of exposure. TSs measured on subsequent days of the exposure sequence have been shown to decrease relative to the initial TS. This reduction of TS, despite the continuing daily exposure regime, has been called a cochlear toughening effect and the exposures referred to as toughening exposures. Four groups of chinchillas were exposed to one of four different noises presented on an interrupted (6 h/day for 20 days) or noninterrupted (24 h/day for 5 days) schedule. The exposures had equivalent total energy, an overall level of 100 dB(A) SPL, and approximately the same flat, broadband long-term spectrum. The noises differed primarily in their temporal structures; two were Gaussian and two were non-Gausssian, nonstationary. Brainstem auditory evoked potentials were used to estimate hearing thresholds and surface preparation histology was used to determine sensory cell loss. The experimental results presented here show that: (1) Exposures to interrupted high-level, non-Gaussian signals produce a toughening effect comparable to that produced by an equivalent interrupted Gaussian noise. (2) Toughening, whether produced by Gaussian or non-Gaussian noise, results in reduced trauma compared to the equivalent uninterrupted noise, and (3) that both continuous and interrupted non-Gaussian exposures produce more trauma than do energy and spectrally equivalent Gaussian noises. Over the course of the 20-day exposure, the pattern of TS following each day's exposure could exhibit a variety of configurations. These results do not support the equal energy hypothesis as a unifying principal for estimating the potential of a noise exposure to produce hearing loss.
The study of mosque acoustics, with regard to acoustical characteristics, sound quality for speech intelligibility, and other applicable acoustic criteria, has been largely neglected. In this study a background as to why mosques are designed as they are and how mosque design is influenced by worship considerations is given. In the study the acoustical characteristics of typically constructed contemporary mosques in Saudi Arabia have been investigated, employing a well-known impulse response. Extensive field measurements were taken in 21 representative mosques of different sizes and architectural features in order to characterize their acoustical quality and to identify the impact of air conditioning, ceiling fans, and sound reinforcement systems on their acoustics. Objective room-acoustic indicators such as reverberation time (RT) and clarity (C50) were measured. Background noise (BN) was assessed with and without the operation of air conditioning and fans. The speech transmission index (STI) was also evaluated with and without the operation of existing sound reinforcement systems. The existence of acoustical deficiencies was confirmed and quantified. The study, in addition to describing mosque acoustics, compares design goals to results obtained in practice and suggests acoustical target values for mosque design. The results show that acoustical quality in the investigated mosques deviates from optimum conditions when unoccupied, but is much better in the occupied condition.
Many competing noises in real environments are modulated or fluctuating in level. Listeners with normal hearing are able to take advantage of temporal gaps in fluctuating maskers. Listeners with sensorineural hearing loss show less benefit from modulated maskers. Cochlear implant users may be more adversely affected by modulated maskers because of their limited spectral resolution and by their reliance on envelope-based signal-processing strategies of implant processors. The current study evaluated cochlear implant users' ability to understand sentences in the presence of modulated speech-shaped noise. Normal-hearing listeners served as a comparison group. Listeners repeated IEEE sentences in quiet, steady noise, and modulated noise maskers. Maskers were presented at varying signal-to-noise ratios (SNRs) at six modulation rates varying from 1 to 32 Hz. Results suggested that normal-hearing listeners obtain significant release from masking from modulated maskers, especially at 8-Hz masker modulation frequency. In contrast, cochlear implant users experience very little release from masking from modulated maskers. The data suggest, in fact, that they may show negative effects of modulated maskers at syllabic modulation rates (2-4 Hz). Similar patterns of results were obtained from implant listeners using three different devices with different speech-processor strategies. The lack of release from masking occurs in implant listeners independent of their device characteristics, and may be attributable to the nature of implant processing strategies and/or the lack of spectral detail in processed stimuli.
This study examined whether cochlear implant users must perceive differences along phonetic continua in the same way as do normal hearing listeners (i.e., sharp identification functions, poor within-category sensitivity, high between-category sensitivity) in order to recognize speech accurately. Adult postlingually deafened cochlear implant users, who were heterogeneous in terms of their implants and processing strategies, were tested on two phonetic perception tasks using a synthetic /da/-/ta/ continuum (phoneme identification and discrimination) and two speech recognition tasks using natural recordings from ten talkers (open-set word recognition and forced-choice /d/-/t/ recognition). Cochlear implant users tended to have identification boundaries and sensitivity peaks at voice onset times (VOT) that were longer than found for normal-hearing individuals. Sensitivity peak locations were significantly correlated with individual differences in cochlear implant performance; individuals who had a /d/-/t/ sensitivity peak near normal-hearing peak locations were most accurate at recognizing natural recordings of words and syllables. However, speech recognition was not strongly related to identification boundary locations or to overall levels of discrimination performance. The results suggest that perceptual sensitivity affects speech recognition accuracy, but that many cochlear implant users are able to accurately recognize speech without having typical normal-hearing patterns of phonetic perception.
The present study investigated the hypothesis that the cues for modulation rate discrimination for unresolved spectral components differ as a function of the spectral region occupied by the stimuli. Specifically, it was hypothesized that when components occupy relatively low spectral regions, phase locking both to the fine structure and to the envelope are useful cues. However, as the spectral region occupied by the components increases, phase locking to the fine structure becomes less robust, whereas phase locking to the envelope remains as a potentially strong cue. Observers were asked to detect a decrease in modulation rate for carrier frequencies between 1500 and 6000 Hz. Both amplitude-modulated (AM) and quasifrequency-modulated (QFM) tones were used in order to produce stimuli having strong and weak envelope cues, respectively. Although there were marked individual differences, the results showed an interaction between modulation type and spectral region, with AM and QFM performance being relatively similar at low spectral region, but with QFM showing a steeper reduction in performance as the spectral region of the carrier frequency increased. Overall, the data are consistent with an interpretation that pitch perception for unresolved components depends upon both fine structure and envelope cues, and that the relative importance of these cues depends upon the spectral region occupied by the stimuli.
The underwater hearing sensitivity of a striped dolphin was measured in a pool using standard psycho-acoustic techniques. The go/no-go response paradigm and up-down staircase psychometric method were used. Auditory sensitivity was measured by using 12 narrow-band frequency-modulated signals having center frequencies between 0.5 and 160 kHz. The 50% detection threshold was determined for each frequency. The resulting audiogram for this animal was U-shaped, with hearing capabilities from 0.5 to 160 kHz (8 1/3 oct). Maximum sensitivity (42 dB re 1 microPa) occurred at 64 kHz. The range of most sensitive hearing (defined as the frequency range with sensitivities within 10 dB of maximum sensitivity) was from 29 to 123 kHz (approximately 2 oct). The animal's hearing became less sensitive below 32 kHz and above 120 kHz. Sensitivity decreased by about 8 dB per octave below 1 kHz and fell sharply at a rate of about 390 dB per octave above 140 kHz.
The tissue mechanics governing vocal-fold closure and collision during phonation are modeled in order to evaluate the role of elastic forces in glottal closure and in the development of stresses that may be a risk factor for pathology development. The model is a nonlinear dynamic contact problem that incorporates a three-dimensional, linear elastic, finite-element representation of a single vocal fold, a rigid midline surface, and quasistatic air pressure boundary conditions. Qualitative behavior of the model agrees with observations of glottal closure during normal voice production. The predicted relationship between subglottal pressure and peak collision force agrees with published experimental measurements. Accurate predictions of tissue dynamics during collision suggest that elastic forces play an important role during glottal closure and are an important determinant of aerodynamic variables that are associated with voice quality. Model predictions of contact force between the vocal folds are directly proportional to compressive stress (r2 = 0.79), vertical shear stress (r2 = 0.69), and Von Mises stress (r2 = 0.83) in the tissue. These results guide the interpretation of experimental measurements by relating them to a quantity that is important in tissue damage.
A model for the generation of auditory brainstem responses (ABR) and frequency following responses (FFRs) is presented. The model is based on the concept introduced by Goldstein and Kiang [J. Acoust. Soc. Am. 30, 107-114 (1958)] that evoked potentials recorded at remote electrodes can theoretically be given by convolution of an elementary unit waveform (unitary response) with the instantaneous discharge rate function for the corresponding unit. In the present study, the nonlinear computational auditory-nerve model recently developed by Heinz et al. [ARLO 2(3), 91-96 (2001)] was used to calculate the instantaneous discharge rate ri(t) for fibers i in the frequency range from 0.1 and 10 kHz. The summed activity across frequency was convolved with a unitary response which is assumed to reflect contributions from different cell populations within the auditory brainstem, recorded at a given pair of electrodes on the scalp. Predicted potential patterns are compared with experimental data for a number of stimulus and level conditions. Clicks, chirps as defined in Dau et al. [J. Acoust. Soc. Am. 107, 1530-1540 (2000)], long-duration stimuli comprising the chirp, as well as tones and slowly varying tonal sweeps were considered. The results demonstrate the importance of considering the effects of the basilar-membrane traveling wave and auditory-nerve processing for the formation of ABR and FFR. Specifically, the results support the hypothesis that the FFR to low-frequency tones represents synchronized activity mainly stemming from mid- and high-frequency units at more basal sites, and not from units tuned to frequencies around the signal frequency.
Function words, especially frequently occurring ones such as (the, that, and, and of), vary widely in pronunciation. Understanding this variation is essential both for cognitive modeling of lexical production and for computer speech recognition and synthesis. This study investigates which factors affect the forms of function words, especially whether they have a fuller pronunciation (e.g., thi, thaet, aend, inverted-v v) or a more reduced or lenited pronunciation (e.g., thax, thixt, n, ax). It is based on over 8000 occurrences of the ten most frequent English function words in a 4-h sample from conversations from the Switchboard corpus. Ordinary linear and logistic regression models were used to examine variation in the length of the words, in the form of their vowel (basic, full, or reduced), and whether final obstruents were present or not. For all these measures, after controlling for segmental context, rate of speech, and other important factors, there are strong independent effects that made high-frequency monosyllabic function words more likely to be longer or have a fuller form (1) when neighboring disfluencies (such as filled pauses uh and um) indicate that the speaker was encountering problems in planning the utterance; (2) when the word is unexpected, i.e., less predictable in context; (3) when the word is either utterance initial or utterance final. Looking at the phenomenon in a different way, frequent function words are more likely to be shorter and to have less-full forms in fluent speech, in predictable positions or multiword collocations, and utterance internally. Also considered are other factors such as sex (women are more likely to use fuller forms, even after controlling for rate of speech, for example), and some of the differences among the ten function words in their response to the factors.
Cochlear nonlinearity was estimated over a wide range of center frequencies and levels in listeners with normal hearing, using a forward-masking method. For a fixed low-level probe, the masker level required to mask the probe was measured as a function of the masker-probe interval, to produce a temporal masking curve (TMC). TMCs were measured for probe frequencies of 500, 1000, 2000, 4000, and 8000 Hz, and for masker frequencies 0.5, 0.7, 0.9, 1.0 (on frequency), 1.1, and 1.6 times the probe frequency. Across the range of probe frequencies, the TMCs for on-frequency maskers showed two or three segments with clearly distinct slopes. If it is assumed that the rate of decay of the internal effect of the masker is constant across level and frequency, the variations in the slopes of the TMCs can be attributed to variations in cochlear compression. Compression-ratio estimates for on-frequency maskers were between 3:1 and 5:1 across the range of probe frequencies. Compression did not decrease at low frequencies. The slopes of the TMCs for the lowest frequency probe (500 Hz) did not change with masker frequency. This suggests that compression extends over a wide range of stimulus frequencies relative to characteristic frequency in the apical region of the cochlea.
The mammalian cochlea is a structure comprising a number of components connected by elastic elements. A mechanical system of this kind is expected to have multiple normal modes of oscillation and associated resonances. The guinea pig cochlear mechanics was probed using distortion components generated in the cochlea close to the place of overlap between two tones presented simultaneously. Otoacoustic emissions at frequencies of the distortion components were recorded in the ear canal. The phase behavior of the emissions reveals the presence of a nonlinear resonance at a frequency about a half octave below that of the high-frequency primary tone. The location of the resonance is level dependent and the resonance shifts to lower frequencies with increasing stimulus intensity. This resonance is thought to be associated with the tectorial membrane. The resonance tends to minimize input to the cochlear receptor cells at frequencies below the high-frequency primary and increases the dynamic load to the stereocilia of the receptor cells at the primary frequency when the tectorial membrane and reticular lamina move in counterphase.
Characteristics of distortion product otoacoustic emissions (DPOAEs) and auditory brainstem responses (ABRs) were measured in Mongolian gerbil before and after the introduction of two different auditory dysfunctions: (1) acoustic damage with a high-intensity tone, or (2) furosemide intoxication. The goal was to find emission parameters and measures that best differentiated between the two dysfunctions, e.g., at a given ABR threshold elevation. Emission input-output or "growth" functions were used (frequencies f1 and f2, f2/f1 = 1.21) with equal levels, L1 = L2, and unequal levels, with L1 = L2 + 20 dB. The best parametric choice was found to be unequal stimulus levels, and the best measure was found to be the change in the emission threshold level, delta x. The emission threshold was defined as the stimulus level required to reach a criterion emission amplitude, in this case -10 dB SPL. (The next best measure was the change in emission amplitude at high stimulus levels, specifically that measured at L1 x L2 = 90 x 70 dB SPL.) For an ABR threshold shift of 20 dB or more, there was essentially no overlap in the emission threshold measures for the two conditions, sound damage or furosemide. The dividing line between the two distributions increased slowly with the change in ABR threshold, delta ABR, and was given by delta x(t) = 0.6 delta ABR + 8 dB. For a given delta ABR, if the shift in emission threshold was more than the calculated dividing line value, delta x(t), the auditory dysfunction was due to acoustic damage, if less, it was due to furosemide.
The response at the surface of an isotropic viscoelastic medium to buried fundamental acoustic sources is studied theoretically, computationally and experimentally. Finite and infinitesimal monopole and dipole sources within the low audible frequency range (40-400 Hz) are considered. Analytical and numerical integral solutions that account for compression, shear and surface wave response to the buried sources are formulated and compared with numerical finite element simulations and experimental studies on finite dimension phantom models. It is found that at low audible frequencies, compression and shear wave propagation from point sources can both be significant, with shear wave effects becoming less significant as frequency increases. Additionally, it is shown that simple closed-form analytical approximations based on an infinite medium model agree well with numerically obtained "exact" half-space solutions for the frequency range and material of interest in this study. The focus here is on developing a better understanding of how biological soft tissue affects the transmission of vibro-acoustic energy from biological acoustic sources below the skin surface, whose typical spectral content is in the low audible frequency range. Examples include sound radiated from pulmonary, gastro-intestinal and cardiovascular system functions, such as breath sounds, bowel sounds and vascular bruits, respectively.
Five commonly used methods for determining the onset of voicing of syllable-initial stop consonants were compared. The speech and glottal activity of 16 native speakers of Cantonese with normal voice quality were investigated during the production of consonant vowel (CV) syllables in Cantonese. Syllables consisted of the initial consonants /ph/, /th/, /kh/, /p/, /t/, and /k/ followed by the vowel /a/. All syllables had a high level tone, and were all real words in Cantonese. Measurements of voicing onset were made based on the onset of periodicity in the acoustic waveform, and on spectrographic measures of the onset of a voicing bar (f0), the onset of the first formant (F1), second formant (F2), and third formant (F3). These measurements were then compared against the onset of glottal opening as determined by electroglottography. Both accuracy and variability of each measure were calculated. Results suggest that the presence of aspiration in a syllable decreased the accuracy and increased the variability of spectrogram-based measurements, but did not strongly affect measurements made from the acoustic waveform. Overall, the acoustic waveform provided the most accurate estimate of voicing onset; measurements made from the amplitude waveform were also the least variable of the five measures. These results can be explained as a consequence of differences in spectral tilt of the voicing source in breathy versus modal phonation.