Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sound Spectrography”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,387 records · Page 77Linked to original sources

Relationship between N1 evoked potential morphology and the perception of voicing.

Auditory evoked potential (AEP) correlates of the neural representation of stimuli along a /ga/-/ka/ and a /ba/-/pa/ continuum were examined to determine whether the voice-onset time (VOT)-related change in the N1 onset response from a single to double-peaked component is a reliable indicator of the perception of voiced and voiceless sounds. Behavioral identification results from ten subjects revealed a mean category boundary at a VOT of 46 ms for the /ga/-/ka/ continuum and at a VOT of 27.5 ms for the /ba/-/pa/ continuum. In the same subjects, electrophysiologic recordings revealed that a single N1 component was seen for stimuli with VOTs of 30 ms and less, and two components (N1' and N1) were seen for stimuli with VOTs of 40 ms and more for both continua. That is, the change in N1 morphology (from single to double-peaked) coincided with the change in perception from voiced to voiceless for stimuli from the /ba/-/pa/ continuum, but not for stimuli from the /ga/-/ka/ continuum. The results of this study show that N1 morphology does not reliably predict phonetic identification of stimuli varying in VOT. These findings also suggest that the previously reported appearance of a "double-peak" onset response in aggregate recordings from the auditory cortex does not indicate a cortical correlate of the perception of voicelessness.

Adult↗

Modeling the combined effects of basilar membrane nonlinearity and roughness on stimulus frequency otoacoustic emission fine structure.

A theoretical framework for describing the effects of nonlinear reflection on otoacoustic emission fine structure is presented. The following models of cochlear reflection are analyzed: weak nonlinearity, distributed roughness, and a combination of weak nonlinearity and distributed roughness. In particular, these models are examined in the context of stimulus frequency otoacoustic emissions (SFOAEs). In agreement with previous studies, it is concluded that only linear cochlear reflection can explain the underlying properties of cochlear fine structures. However, it is shown that nonlinearity can unexpectedly, in some cases, significantly modify the level and phase behaviors of the otoacoustic emission fine structure, and actually enhance the pattern of fine structures observed. The implications of these results on the stimulus level dependence of SFOAE fine structure are also explored.

Basilar Membrane↗

Channel weights for speech recognition in cochlear implant users.

The purpose of this study was to develop and validate a method of estimating the relative "weight" that a multichannel cochlear implant user places on individual channels, indicating its contribution to overall speech recognition. The correlational method as applied to speech recognition was used both with normal-hearing listeners and with cochlear implant users fitted with six-channel speech processors. Speech was divided into frequency bands corresponding to the bands of the processor and a randomly chosen level of corresponding filtered noise was added to each channel on each trial. Channels in which the signal-to-noise ratio was more highly correlated with performance have higher weights, and conversely, channels in which the correlations were smaller have lower weights. Normal-hearing listeners showed approximately equal weights across frequency bands. In contrast, cochlear implant users showed unequal weighting across bands, and varied from individual to individual with some channels apparently not contributing significantly to speech recognition. To validate these channel weights, individual channels were removed and speech recognition in quiet was tested. A strong correlation was found between the relative weight of the channel removed and the decrease in speech recognition, thus providing support for use of the correlational method for cochlear implant users.

Adult↗

Simultaneous effects on vowel duration in American English: a covariance structure modeling approach.

The powerful techniques of covariance structure modeling (CSM) long have been used to study complex behavioral phenomenon in the social and behavioral sciences. This study employed these same techniques to examine simultaneous effects on vowel duration in American English. Additionally, this study investigated whether a single population model of vowel duration fits observed data better than a dual population model where separate parameters are generated for syllables that carry large information loads and for syllables that specify linguistic relationships. For the single population model, intrinsic duration, phrase final position, lexical stress, post-vocalic consonant voicing, and position in word all were significant predictors of vowel duration. However, the dual population model, in which separate model parameters were generated for (1) monosyllabic content words and lexically stressed syllables and (2) monosyllabic function words and lexically unstressed syllables, fit the data better than the single population model. Intrinsic duration and phrase final position affected duration similarly for both the populations. On the other hand, the effects of post-vocalic consonant voicing and position in word, while significant predictors of vowel duration in content words and stressed syllables, were not significant predictors of vowel duration in function words or unstressed syllables. These results are not unexpected, based on previous research, and suggest that covariance structure analysis can be used as a complementary technique in linguistic and phonetic research.

Adult↗

Children's perception of speech in multitalker babble.

Children 5, 9, and 11 years of age and young adults attempted to identify the final word of sentences recorded by a female speaker. The sentences were presented in two levels of multitalker babble, and participants responded by selecting one of four pictures. In a low-noise condition, the signal-to-noise ratio (SNR) was adjusted for each age group to yield 85% correct performance. In a high-noise condition, the SNR was set 7 dB lower than the low-noise condition. Although children required more favorable SNRs than adults to achieve comparable performance in low noise, an equivalent decrease in SNR had comparable consequences for all age groups. Thus age-related differences on this task can be attributed primarily to sensory factors.

Adolescent↗

Interrelations among distortion-product phase-gradient delays: their connection to scaling symmetry and its breaking.

Distortion-product-otoacoustic-emission (DPOAE) phase-versus-frequency functions and corresponding phase-gradient delays have received considerable attention because of their potential for providing information about mechanisms of emission generation, cochlear wave latencies, and characteristics of cochlear tuning. The three measurement paradigms in common use (fixed-f1, fixed-f2, and fixed-f2/f1) yield significantly different delays, suggesting that they depend on qualitatively different aspects of cochlear mechanics. In this paper, theory and experiment are combined to demonstrate that simple phenomenological arguments, which make no detailed mechanistic assumptions concerning the underlying cochlear mechanics, predict relationships among the delays that are in good quantitative agreement with experimental data obtained in guinea pigs. To understand deviations between the simple theory and experiment, a general equation is found that relates the three delays for any deterministic model of DPOAE generation. Both model-independent and exact, the general relation provides a powerful consistency check on the measurements and a useful tool for organizing and understanding the structure in DPOAE phase data (e.g., for interpreting the relative magnitudes and intensity-dependencies of the three delays). Analysis of the general relation demonstrates that the success of the simple, phenomenological approach can be understood as a consequence of the mechanisms of emission generation and the approximate local scaling symmetry of cochlear mechanics. The general relation is used to quantify deviations from scaling manifest in the measured phase-gradient delays; the results indicate that deviations from scaling are typically small and that both linear and nonlinear mechanisms contribute significantly to these deviations. Intensity-dependent mechanisms contributing to deviations from scaling include cochlear-reflection and wave-interference effects associated with the mixing of distortion- and reflection-source emissions (as in DPOAE fine structure). Finally, the ratio of the fixed-f1 and fixed-f2 phase-gradient delays is shown to follow from the choice of experimental paradigm and, in the scaling limit, contains no information about cochlear physiology whatsoever. These results cast considerable doubt on the theoretical basis of recent attempts to use relative DPOAE phase-gradient delays to estimate the bandwidths of peripheral auditory filters.

Animals↗

Seismic properties of Asian elephant (Elephas maximus) vocalizations and locomotion.

Seismic and acoustic data were recorded simultaneously from Asian elephants (Elephas maximus) during periods of vocalizations and locomotion. Acoustic and seismic signals from rumbles were highly correlated at near and far distances and were in phase near the elephant and were out of phase at an increased distance from the elephant. Data analyses indicated that elephant generated signals associated with rumbles and "foot stomps" propagated at different velocities in the two media, the acoustic signals traveling at 309 m/s and the seismic signals at 248-264 m/s. Both types of signals had predominant frequencies in the range of 20 Hz. Seismic signal amplitudes considerably above background noise were recorded at 40 m from the generating elephants for both the rumble and the stomp. Seismic propagation models suggest that seismic waveforms from vocalizations are potentially detectable by instruments at distances of up to 16 km, and up to 32 km for locomotion generated signals. Thus, if detectable by elephants, these seismic signals could be useful for long distance communication.

Animal Communication↗

Some effects of duration on vowel recognition.

This study was designed to examine the role of duration in vowel perception by testing listeners on the identification of CVC syllables generated at different durations. Test signals consisted of synthesized versions of 300 utterances selected from a large, multitalker database of /hVd/ syllables [Hillenbrand et al., J. Acoust. Soc. Am. 97, 3099-3111 (1995)]. Four versions of each utterance were synthesized: (1) an original duration set (vowel duration matched to the original utterance), (2) a neutral duration set (duration fixed at 272 ms, the grand mean across all vowels), (3) a short duration set (duration fixed at 144 ms, two standard deviations below the mean), and (4) a long duration set (duration fixed at 400 ms, two standard deviations above the mean). Experiment 1 used a formant synthesizer, while a second experiment was an exact replication using a sinusoidal synthesis method that represented the original vowel spectrum more precisely than the formant synthesizer. Findings included (1) duration had a small overall effect on vowel identity since the great majority of signals were identified correctly at their original durations and at all three altered durations; (2) despite the relatively small average effect of duration, some vowels, especially [see text] and [see text], were significantly affected by duration; (3) some vowel contrasts that differ systematically in duration, such as [see text], and [see text], were minimally affected by duration; (4) a simple pattern recognition model appears to be capable of accounting for several features of the listening test results, especially the greater influence of duration on some vowels than others; and (5) because a formant synthesizer does an imperfect job of representing the fine details of the original vowel spectrum, results using the formant-synthesized signals led to a slight overestimate of the role of duration in vowel recognition, especially for the shortened vowels.

Attention↗

Whistles of boto, Inia geoffrensis, and tucuxi, Sotalia fluviatilis.

Whistles were recorded and analyzed from free-ranging single or mixed species groups of boto and tucuxi in the Peruvian Amazon, with sonograms presented. Analysis revealed whistles recorded falling into two discrete groups: a low-frequency group with maximum frequency below 5 kHz, and a high-frequency group with maximum frequencies above 8 kHz and usually above 10 kHz. Whistles in the two groups differed significantly in all five measured variables (beginning frequency, end frequency, minimum frequency, maximum frequency, and duration). Comparisons with published details of whistles by other platanistoid river dolphins and by oceanic dolphins suggest that the low-frequency whistles were produced by boto, the high-frequency whistles by tucuxi. Tape recordings obtained on three occasions when only one species was present tentatively support this conclusion, but it is emphasized that this is based on few data.

Animal Communication↗

Isolating the auditory system from acoustic noise during functional magnetic resonance imaging: examination of noise conduction through the ear canal, head, and body.

Approaches were examined for reducing acoustic noise levels heard by subjects during functional magnetic resonance imaging (fMRI), a technique for localizing brain activation in humans. Specifically, it was examined whether a device for isolating the head and ear canal from sound (a "helmet") could add to the isolation provided by conventional hearing protection devices (i.e., earmuffs and earplugs). Both subjective attenuation (the difference in hearing threshold with versus without isolation devices in place) and objective attenuation (difference in ear-canal sound pressure) were measured. In the frequency range of the most intense fMRI noise (1-1.4 kHz), a helmet, earmuffs, and earplugs used together attenuated perceived sound by 55-63 dB, whereas the attenuation provided by the conventional devices alone was substantially less: 30-37 dB for earmuffs, 25-28 dB for earplugs, and 39-41 dB for earmuffs and earplugs used together. The data enabled the clarification of the relative importance of ear canal, head, and body conduction routes to the cochlea under different conditions: At low frequencies (< or =500 Hz), the ear canal was the dominant route of sound conduction to the cochlea for all of the device combinations considered. At higher frequencies (>500 Hz), the ear canal was the dominant route when either earmuffs or earplugs were worn. However, the dominant route of sound conduction was through the head when both earmuffs and earplugs were worn, through both ear canal and body when a helmet and earmuffs were worn, and through the body when a helmet, earmuffs, and earplugs were worn. It is estimated that a helmet, earmuffs, and earplugs together will reduce the most intense fMRI noise levels experienced by a subject to 60-65 dB SPL. Even greater reductions in noise should be achievable by isolating the body from the surrounding noise field.

Adult↗

Data processing options and response scoring for OAE-based newborn hearing screening.

Scoring of click-evoked otoacoustic emissions (CEOAEs) is typically achieved by the evaluation of the reproducibility of the whole emission and/or within narrow bands. Screening outcomes are influenced not only by the specific combination of the subdivision scheme (i.e., the number, position, and bandwidth of the narrow bands) and the threshold used to determine pass and refer, but also by the accuracy with which the reproducibility is estimated. This study was designed to examine what factors affect the accuracy of the reproducibility estimate and how the accuracy of the reproducibility estimate together with the choice of the subdivision scheme/thresholds affect CEOAE scoring. Simulations with real CEOAEs corrupted with synthesized noise indicated that the reproducibility estimate is influenced by time-windowing and band-pass filtering: the longer the time-window or the broader the bandwidth of the filter, the more accurate the estimate. Quantitative figures on numerical scoring were given in terms of the referral rate and were derived from CEOAEs recorded in a clinical environment from more than 3400 newborns. The narrow bands were extracted according to 12 different subdivision schemes covering the 1.5-4-kHz range. The referral rate was found to depend on the subdivision scheme being used: (i) the worst results were obtained considering four narrow bands at 1.6-2.4-3.2-4 kHz; (ii) the best results were obtained considering two narrow bands at 2.25 and 3.75 kHz; (iii) bandwidths greater than 1 kHz resulted in the lowest referral rates. Also, scoring based on the extraction of four narrow bands produced the most unstable results, i.e., a small change in the threshold might cause even a great change in the referral rate.

Female↗

On the annoyance caused by impulse sounds produced by small, medium-large, and large firearms.

A laboratory study was designed in which the annoyance was investigated for 14 different impulse sound types produced by various firearms ranging in caliber from 7.62 to 155 mm. Sixteen subjects rated the annoyance for the simulated conditions of (1) being outdoors, and (2) being indoors with the windows closed. In the latter case, a representative outdoor-to-indoor reduction in sound level was applied. It was anticipated that the presumed additional annoyance caused by the "heaviness" of the impulse sounds might be predicted from the difference between the C-weighted sound exposure level (CSEL; LCE) and the A-weighted sound exposure level (ASEL; LAE). In the outdoor rating conditions, the annoyance was almost entirely determined by ASEL. The explained variance, r2, in the mean ratings by ASEL was 0.95. In the indoor rating conditions, however, the explained variance in the annoyance ratings by (outdoor) ASEL was significantly increased from r2 = 0.87 to r2= 0.97 by adding the product (LCE-LAE)(LAE-alpha) as a second variable. In combination with a 12-dB adjustment for small firearms, the present results showed that for the entire set of impulse sounds rated indoors with windows closed, the rating sound level, Lr, is given by Lr=LAE +12dB+beta(LCE-LAE)(LAE-alpha), with alpha=45dB and beta=0.015dB(-1). For the outdoor rating condition, the optimal parameter values were equal to alpha=57 dB and, again, beta=0.015 dB(-1). In validation studies, in which the effects of the present rating procedure will be compared to field data, it has to be determined to what extent the constants alpha and beta have to be adjusted.

Adolescent↗

Manipulating the "straightness" and "curvature" of patterns of interaural cross correlation affects listeners' sensitivity to changes in interaural delay.

The purpose of this study was to test the hypothesis that stimuli characterized by "straight" trajectories of their patterns of cross correlation foster greater sensitivity to changes in interaural temporal disparities (ITDs) than do stimuli characterized by more "curved" trajectories of their patterns of cross correlation. To do so, sensitivity to changes in ITD was measured, as a function of duration, using a set of "reference" stimuli that yielded differing relative amounts of straightness within their patterns of cross correlation while keeping the dominant trajectory at or near midline. The relative amounts of straightness were manipulated by employing specific combinations of bandwidth, ITD, and interaural phase disparity (IPD) of Gaussian noises centered at 500 Hz. The results were consistent with expectations in that the patterning of the threshold ITDs revealed increasingly poorer sensitivity as greater and greater curvature was imposed on the dominant, "midline," trajectory. The variations in threshold ITD across the stimulus conditions can be accounted for quite well quantitatively by assuming either that the listeners based their judgments on changes in the position of the most central peak of the cross-correlation function or that they based their judgments on changes in the centroid of a second-level cross-correlation function. In a second experiment, binaural detection was measured using a subset of the reference stimuli as maskers. As expected, sensitivity was poorest with the maskers characterized by the greatest curvature, which were also those having the lowest interaural correlation.

Adult↗

A masking level difference due to harmonicity.

The role of harmonicity in masking was studied by comparing the effect of harmonic and inharmonic maskers on the masked thresholds of noise probes using a three-alternative, forced-choice method. Harmonic maskers were created by selecting sets of partials from a harmonic series with an 88-Hz fundamental and 45 consecutive partials. Inharmonic maskers differed in that the partial frequencies were perturbed to nearby values that were not integer multiples of the fundamental frequency. Average simultaneous-masked thresholds were as much as 10 dB lower with the harmonic masker than with the inharmonic masker, and this difference was unaffected by masker level. It was reduced or eliminated when the harmonic partials were separated by more than 176 Hz, suggesting that the effect is related to the extent to which the harmonics are resolved by auditory filters. The threshold difference was not observed in a forward-masking experiment. Finally, an across-channel mechanism was implicated when the threshold difference was found between a harmonic masker flanked by harmonic bands and a harmonic masker flanked by inharmonic bands. A model developed to explain the observed difference recognizes that an auditory filter output envelope is modulated when the filter passes two or more sinusoids, and that the modulation rate depends on the differences among the input frequencies. For a harmonic masker, the frequency differences of adjacent partials are identical, and all auditory filters have the same dominant modulation rate. For an inharmonic masker, however, the frequency differences are not constant and the envelope modulation rate varies across filters. The model proposes that a lower variability facilitates detection of a probe-induced change in the variability, thus accounting for the masked threshold difference. The model was supported by significantly improved predictions of observed thresholds when the predictor variables included envelope modulation rate variance measured using simulated auditory filters.

Adult↗

Investigation of the relationship among three common measures of precedence: fusion, localization dominance, and discrimination suppression.

Listeners have a remarkable ability to localize and identify sound sources in reverberant environments. The term "precedence effect" (PE; also known as the "Haas effect," "law of the first wavefront," and "echo suppression") refers to a group of auditory phenomena that is thought to be related to this ability. Traditionally, three measures have been used to quantify the PE: (1) Fusion: at short delays (1-5 ms for clicks) the lead and lag perceptually fuse into one auditory event; (2) Localization dominance: the perceived location of the leading source dominates that of the lagging source; and (3) Discrimination suppression: at short delays, changes in the location or interaural parameters of the lag are difficult to discriminate compared with changes in characteristics of the lead. Little is known about the relation among these aspects of the PE, since they are rarely studied in the same listeners. In the present study, extensive measurements of these phenomena were made for six normal-hearing listeners using 1-ms noise bursts. The results suggest that, for clicks, fusion lasts 1-5 ms; by 5 ms most listeners hear two sounds on a majority of trials. However, localization dominance and discrimination suppression remain potent for delays of 10 ms or longer. Results are consistent with a simple model in which information from the lead and lag interacts perceptually and in which the strength of this interaction decreases with spatiotemporal separation of the lead and lag. At short delays, lead and lag both contribute to spatial perception, but the lead dominates (to the extent that only one position is ever heard). At the longest delays tested, two distinct sounds are perceived (as measured in a fusion task), but they are not always heard at independent spatial locations (as measured in a localization dominance task). These results suggest that directional cues from the lag are not necessarily salient for all conditions in which the lag is subjectively heard as a separate event.

Adult↗

The prediction of speech intelligibility in underground stations of rectangular cross section.

Long enclosures are spaces with nondiffuse sound fields, for which the classical theory of acoustics is not appropriate. Thus, the modeling of the sound field in a long enclosure is very different from the prediction of the behavior of sound in a diffuse space. Ray-tracing computer models have been developed for the prediction of the sound field in long enclosures, with particular reference to spaces such as underground stations which are generally long spaces of rectangular or curved cross section. This paper describes the development of a model for use in underground stations of rectangular cross section. The model predicts the sound-pressure level, early decay time, clarity index, and definition at receiver points along the enclosure. The model also calculates the value of the speech transmission index at individual points. Measurements of all parameters have been made in a station of rectangular cross section, and compared with the predicted values. The predictions of all parameters show good agreement with measurements at all frequencies, particularly in the far field of the sound source, and the trends in the behavior of the parameters along the enclosure have been correctly predicted.

Confined Spaces↗