Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sound Spectrography”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,495 records · Page 83Linked to original sources

Acoustic identification of female Steller sea lions (Eumetopias jubatus).

Steller sea lion (Eumetopias jubatus) mothers and pups establish and maintain contact with individually distinctive vocalizations. Our objective was to develop a robust neural network to classify females based on their mother-pup contact calls. We catalogued 573 contact calls from 25 females in 1998 and 1323 calls from 46 females in 1999. From this database, a subset of 26 females with sufficient samples of calls was selected for further study. Each female was identified visually by marking patterns, which provided the verification for acoustic identification. Average logarithmic spectra were extracted for each call, and standardized training and generalization datasets created for the neural network classifier. A family of backpropagation networks was generated to assess relative contribution of spectral input bandwidth, frequency resolution, and network architectural variables to classification accuracy. The network with best overall generalization accuracy (71%) used an input representation of 0-3 kHz of bandwidth at 10.77 Hz/bin frequency resolution, and a 2:1 hidden:output layer neural ratio. The network was analyzed to reveal which portions of the call spectra were most influential for identification of each female. Acoustical identification of distinctive female acoustic signatures has several potentially important conservation applications for this endangered species, such as rapid survey of females present on a rookery.

Animals↗

Mechanisms of modulation gap detection.

It has been postulated that the central auditory system contains an array of modulation filters, each responsive to a different range of modulation frequencies present at the outputs of the (peripheral) auditory filters. In the present experiments, we tested what we call the "dip hypothesis," that a gap in modulation is detected using the "dip" in the output of the modulation filter tuned to the modulator frequency. In experiment 1, the task was to detect a gap in the sinusoidal amplitude modulation imposed on a 4-kHz carrier. The modulator preceding the gap ended with a positive-going zero-crossing. There were three conditions, differing in the phase at which the modulator started at the end of the gap; zero-phase, at a positive-going zero-crossing; pi-phase, at a negative-going zero-crossing; and "preserved" phase, at the phase the modulator would have had if it had continued without interruption. Modulation frequencies were 5, 10, 20, and 40 Hz. Psychometric functions for detection of the gap were measured using a two-alternative forced-choice task. For the zero-phase and preserved-phase conditions, the detectability index, d', increased monotonically with increasing gap duration. For the pi-phase condition, performance was good (d' > 1) for small gap durations, and initially worsened with increasing gap duration, before improving again for longer gap durations. This is the pattern of results expected from the dip hypothesis, provided that the modulation filters have Q values of 2 or more. However, it is also possible that a rhythm cue was used to improve performance in the pi-phase condition for short gap durations; the introduction of the gap markedly disrupted the regular rhythm produced by the modulator peaks. In experiment 2, the rhythm cue was disrupted by varying the modulator period randomly around its nominal value, except for the modulator periods immediately before and after the gap. This markedly impaired performance, and resulted in psychometric functions that were very similar for the zero-phase and pi-phase conditions. This pattern of results is inconsistent with the dip hypothesis. For both experiments, modulation gap "thresholds" (d' approximately 1) were roughly constant when expressed as a proportion of the modulator period. Possible mechanisms of modulation gap detection are discussed and evaluated.

Attention↗

Effects of age and frequency disparity on gap discrimination.

Temporal discrimination was measured using a gap discrimination paradigm for three groups of listeners with normal hearing: (1) ages 18-30, (2) ages 40-52, and (3) ages 62-74 years. Normal hearing was defined as pure-tone thresholds < or = 25 dB HL from 250 to 6000 Hz and < or = 30 dB HL at 8000 Hz. Silent gaps were placed between 1/4-octave bands of noise centered at one of six frequencies. The noise band markers were paired so that the center frequency of the leading marker was fixed at 2000 Hz, and the center frequency of the trailing marker varied randomly across experimental runs. Gap duration discrimination was significantly poorer for older listeners than for young and middle-aged listeners, and the performance of the young and middle-aged listeners did not differ significantly. Age group differences were more apparent for the more frequency-disparate stimuli (2000-Hz leading marker followed by a 500-Hz trailing marker) than for the fixed-frequency stimuli (2000-Hz lead and 2000-Hz trail). The gap duration difference limens of the older listeners increased more rapidly with frequency disparity than those of the other listeners. Because age effects were more apparent for the more frequency-disparate conditions, and gap discrimination was not affected by differences in hearing sensitivity among listeners, it is suggested that gap discrimination depends upon temporal mechanisms that deteriorate with age and stimulus complexity but are unaffected by hearing loss.

Adolescent↗

An EMA study of VCV coarticulatory direction.

This study addresses three issues that are relevant to coarticulation theory in speech production: whether the degree of articulatory constraint model (DAC model) accounts for patterns of the directionality of tongue dorsum coarticulatory influences; the extent to which those patterns in tongue dorsum coarticulatory direction are similar to those for the tongue tip; and whether speech motor control and phonemic planning use a fixed or a context-dependent temporal window. Tongue dorsum and tongue tip movement data on vowel-to-vowel coarticulation are reported for Catalan VCV sequences with vowels /i/, /a/, and /u/, and consonants /p/, /n/, dark /l/, /s/, /S/, alveolopalatal /n/ and /k/. Electromidsagittal articulometry recordings were carried out for three speakers using the Carstens articulograph. Trajectory data are presented for the vertical dimension for the tongue dorsum, and for the horizontal dimension for tongue dorsum and tip. In agreement with predictions of the DAC model, results show that directionality patterns of tongue dorsum coarticulation can be accounted for to a large extent based on the articulatory requirements on consonantal production. While dorsals exhibit analogous trends in coarticulatory direction for all articulators and articulatory dimensions, this is mostly so for the tongue dorsum and tip along the horizontal dimension in the case of lingual fricatives and apicolaminal consonants. This finding results from different articulatory strategies: while dorsal consonants are implemented through homogeneous tongue body activation, the tongue tip and tongue dorsum act more independently for more anterior consonantal productions. Discontinuous coarticulatory effects reported in the present investigation suggest that phonemic planning is adaptative rather than context independent.

Biomechanical Phenomena↗

Temporary shift in masked hearing thresholds in odontocetes after exposure to single underwater impulses from a seismic watergun.

A behavioral response paradigm was used to measure masked underwater hearing thresholds in a bottlenose dolphin (Tursiops truncatus) and a white whale (Delphinapterus leucas) before and after exposure to single underwater impulsive sounds produced from a seismic watergun. Pre- and postexposure thresholds were compared to determine if a temporary shift in masked hearing thresholds (MTTS), defined as a 6-dB or larger increase in postexposure thresholds, occurred. Hearing thresholds were measured at 0.4, 4, and 30 kHz. MTTSs of 7 and 6 dB were observed in the white whale at 0.4 and 30 kHz, respectively, approximately 2 min following exposure to single impulses with peak pressures of 160 kPa, peak-to-peak pressures of 226 dB re 1 microPa, and total energy fluxes of 186 dB re 1 microPa2 x s. Thresholds returned to within 2 dB of the preexposure value approximately 4 min after exposure. No MTTS was observed in the dolphin at the highest exposure conditions: 207 kPa peak pressure, 228 dB re 1 microPa peak-to-peak pressure, and 188 dB re 1 microPa2 x s total energy flux.

Acoustic Stimulation↗

One source for distortion product otoacoustic emissions generated by low- and high-level primaries.

Distortion product otoacoustic emissions (DPOAE) elicited by tones below 60-70 dB sound pressure level (SPL) are significantly more sensitive to cochlear insults. The vulnerable, low-level DPOAE have been associated with the postulated active cochlear process, whereas the relatively robust high-level DPOAE component has been attributed to the passive, nonlinear macromechanical properties of the cochlea. However, it is proposed that the differences in the vulnerability of DPOAEs to high and low SPLs is a natural consequence of the way the cochlea responds to high and low SPLs. An active process boosts the basilar membrane (BM) vibrations, which are attenuated when the active process is impaired. However, at high SPLs the contribution of the active process to BM vibration is small compared with the dominating passive mechanical properties of the BM. Consequently, reduction of active cochlear amplification will have greatest effect on BM vibrations and DPOAEs at low SPLs. To distinguish between the "two sources" and the "single source" hypotheses we analyzed the level dependence of the notch and corresponding phase discontinuity in plots of DPOAE magnitude and phase as functions of the level of the primaries. In experiments where furosemide was used to reduce cochlear amplification, an upward shift of the notch supports the conclusion that both the low- and high-level DPOAEs are generated by a single source, namely a nonlinear amplifier with saturating I/O characteristic.

Animals↗

Benefit of modulated maskers for speech recognition by younger and older adults with normal hearing.

To assess age-related differences in benefit from masker modulation, younger and older adults with normal hearing but not identical audiograms listened to nonsense syllables in each of two maskers: (1) a steady-state noise shaped to match the long-term spectrum of the speech, and (2) this same noise modulated by a 10-Hz square wave, resulting in an interrupted noise. An additional low-level broadband noise was always present which was shaped to produce equivalent masked thresholds for all subjects. This minimized differences in speech audibility due to differences in quiet thresholds among subjects. An additional goal was to determine if age-related differences in benefit from modulation could be explained by differences in thresholds measured in simultaneous and forward maskers. Accordingly, thresholds for 350-ms pure tones were measured in quiet and in each masker; thresholds for 20-ms signals in forward and simultaneous masking were also measured at selected signal frequencies. To determine if benefit from modulated maskers varied with masker spectrum and to provide a comparison with previous studies, a subgroup of younger subjects also listened in steady-state and interrupted noise that was not spectrally shaped. Articulation index (AI) values were computed and speech-recognition scores were predicted for steady-state and interrupted noise; predicted benefit from modulation was also determined. Masked thresholds of older subjects were slightly higher than those of younger subjects; larger age-related threshold differences were observed for short-duration than for long-duration signals. In steady-state noise, speech recognition for older subjects was poorer than for younger subjects, which was partially attributable to older subjects' slightly higher thresholds in these maskers. In interrupted noise, although predicted benefit was larger for older than younger subjects, scores improved more for younger than for older subjects, particularly at the higher noise level. This may be related to age-related increases in thresholds in steady-state noise and in forward masking, especially at higher frequencies. Benefit of interrupted maskers was larger for unshaped than for speech-shaped noise, consistent with AI predictions.

Adult↗

Asymmetry of masking between complex tones and noise: the role of temporal structure and peripheral compression.

Thresholds for the detection of harmonic complex tones in noise were measured as a function of masker level. The rms level of the masker ranged from 40 to 70 dB SPL in 10-dB steps. The tones had a fundamental frequency (F0) of 62.5 or 250 Hz, and components were added in either cosine or random phase. The complex tones and the noise were bandpass filtered into the same frequency region, from the tenth harmonic up to 5 kHz. In a different condition, the roles of masker and signal were reversed, keeping all other parameters the same; subjects had to detect the noise in the presence of a harmonic tone masker. In both conditions, the masker was either gated synchronously with the 700-ms signal, or it started 400 ms before and stopped 200 ms after the signal. The results showed a large asymmetry in the effectiveness of masking between the tones and noise. Even though signal and masker had the same bandwidth, the noise was a more effective masker than the complex tone. The degree of asymmetry depended on F0, component phase, and the level of the masker. The maximum difference between masked thresholds for tone and noise was about 28 dB; this occurred when the F0 was 62.5 Hz, the components were in cosine phase, and the masker level was 70 dB SPL. In most conditions, the growth-of-masking functions had slopes close to 1 (on a dB versus dB scale). However, for the cosine-phase tone masker with an F0 of 62.5 Hz, a 10-dB increase in masker level led to an increase in masked threshold of the noise of only 3.7 dB, on average. We suggest that the results for this condition are strongly affected by the active mechanism in the cochlea.

Adult↗

The mid-level hump at 2 kHz.

Shortening the duration of a Gaussian-shaped 2-kHz tone-pip causes the intensity-difference limen (DL) to depart from the "near-miss to Weber's law" and swell into a mid-level hump [Nizami et al., J. Acoust. Soc. Am. 110, 2505-2515 (2001)]. For some subjects the size of this hump approaches or exceeds the size reported for longer tones under forward masking, suggesting that forward masking might make little difference to the DL for very brief probes. To test this hypothesis, DLs were determined over 30 to 90 dB SPL for a brief Gaussian-shaped 2-kHz tone-pip. DLs were obtained first without forward masking, then with the pip placed 10 or 100 ms after a 200-ms 2-kHz tone of 50 dB SPL, or 100 ms after a 200-ms 2-kHz tone of 70 dB SPL. DLs inflated significantly under all forward-masking conditions. DLs also enlarged under an 80 dB SPL forward masker at pip delays of 4, 10, 40, and 100 ms. The peaks of the humps obtained under forward masking clustered around a sensation level (SL) that was significantly lower than the average SL for the peaks of the humps obtained without forward masking. Overall, the results do not support the neuronal-recovery-rate model of Zeng et al. [Hear. Res. 55, 223-230 (1991)], but are not incompatible with the Carlyon and Beveridge hypothesis [J. Acoust. Soc. Am. 93, 2886-2895 (1993)] that nonsimultaneous maskers corrupt the memory trace evoked by the probe.

Adult↗

Word recognition in competing babble and the effects of age, temporal processing, and absolute sensitivity.

This study was designed to clarify whether speech understanding in a fluctuating background is related to temporal processing as measured by the detection of gaps in noise bursts. Fifty adults with normal hearing or mild high-frequency hearing loss served as subjects. Gap detection thresholds were obtained using a three-interval, forced-choice paradigm. A 150-ms noise burst was used as the gap carrier with the gap placed close to carrier onset. A high-frequency masker without a temporal gap was gated on and off with the noise bursts. A continuous white-noise floor was present in the background. Word scores for the subjects were obtained at a presentation level of 55 dB HL in competing babble levels of 50, 55, and 60 dB HL. A repeated measures analysis of covariance of the word scores examined the effects of age, absolute sensitivity, and temporal sensitivity. The results of the analysis indicated that word scores in competing babble decreased significantly with increases in babble level, age, and gap detection thresholds. The effects of absolute sensitivity on word scores in competing babble were not significant. These results suggest that age and temporal processing influence speech understanding in fluctuating backgrounds in adults with normal hearing or mild high-frequency hearing loss.

Adult↗

Early pitch-shift response is active in both steady and dynamic voice pitch control.

When air conducted auditory feedback pitch is experimentally shifted upward or downward during steady phonation, voice pitch changes in response. The first pitch change is an automatic deflection opposite in direction to the feedback shift. It appears to help stabilize voice pitch by counteracting unintended changes. But what happens during an intended pitch change? If the purpose of the first pitch-shift response is to stabilize voice pitch around a fixed target, it should be suppressed during voluntary pitch changes. Alternatively, if the pitch-shift response is a general process of voice control it should be modified during intended pitch changes to bring production in line with the desired output. Auditory feedback pitch was shifted during steady pitch and upward glissando vocalizations by thirty trained singers. Contrary to the "steady-specific" hypothesis, pitch-shift responses occurred during dynamic pitch vocalizations. Responses were comparable in direction, peak time, and slope, but had significantly longer latency and smaller magnitude than responses elicited during steady note phonation. Results indicate that the early pitch-shift response is a general component of voice control that serves to automatically bring phonation pitch into agreement with an intended target, whether that target is constant or changing in time.

Adolescent↗

Temporal pitch mechanisms in acoustic and electric hearing.

Two experiments investigated pitch perception for stimuli where the place of excitation was held constant. Experiment 1 used pulse trains in which the interpulse interval alternated between 4 and 6 ms. In experiment 1a these "4-6" pulse trains were bandpass filtered between 3900 and 5300 Hz and presented acoustically against a noise background to normal listeners. The rate of an isochronous pulse train (in which all the interpulse intervals were equal) was adjusted so that its pitch matched that of the "4-6" stimulus. The pitch matches were distributed unimodally, had a mean of 5.7 ms, and never corresponded to either 4 or to 10 ms (the period of the stimulus). In experiment 1b the pulse trains were presented both acoustically to normal listeners and electrically to users of the LAURA cochlear implant, via a single channel of their device. A forced-choice procedure was used to measure psychometric functions, in which subjects judged whether the 4-6 stimulus was higher or lower in pitch than isochronous pulse trains having periods of 3, 4, 5, 6, or 7 ms. For both groups of listeners, the point of subjective equality corresponded to a period of 5.6 to 5.7 ms. Experiment 1c confirmed that these psychometric functions were monotonic over the range 4-12 ms. In experiment 2, normal listeners adjusted the rate of an isochronous filtered pulse train to match the pitch of mixtures of pulse trains having rates of F1 and F2 Hz, passed through the same bandpass filter (3900-5400 Hz). The ratio F2/F1 was 1.29 and F1 was either 70, 92, 109, or 124 Hz. Matches were always close to F2 Hz. It is concluded that the results of both experiments are inconsistent with models of pitch perception which rely on higher-order intervals. Together with those of other published data on purely temporal pitch perception, the data are consistent with a model in which only first-order interpulse intervals contribute to pitch, and in which, over the range 0-12 ms, longer intervals receive higher weights than short intervals.

Adult↗

Comodulation masking release in consonant recognition.

Comodulation masking release (CMR) refers to an improvement in the detection threshold of a signal masked by noise with coherent amplitude fluctuation across frequency, as compared to noise without the envelope coherence. The present study tested whether such an advantage for signal detection would facilitate the identification of speech phonemes. Consonant identification of bandpass speech was measured under the following three masker conditions: (1) a single band of noise in the speech band ("on-frequency" masker); (2) two bands of noise, one in the on-frequency band and the other in the "flanking band," with coherence of temporal envelope fluctuation between the two bands (comodulation); and (3) two bands of noise (on-frequency band and flanking band), without the coherence of the envelopes (noncomodulation). A pilot experiment with a small number of consonant tokens was followed by the main experiment with 12 consonants and the following masking conditions: three frequency locations of the flanking band and two masker levels. Results showed that in all conditions, the comodulation condition provided higher identification scores than the noncomodulation condition, and the difference in score was 3.5% on average. No significant difference was observed between the on-frequency only condition and the comodulation condition, i.e., an "unmasking" effect by the addition of a comodulated flaking band was not observed. The positive effect of CMR on consonant recognition found in the present study endorses a "cued-listening" theory, rather than an envelope correlation theory, as a basis of CMR in a suprathreshold task.

Adult↗

The relative detectability for mice of gaps having different ramp durations at their onset and offset boundaries.

The effect on gap detectability of varying noise fall time (FT) and rise time (RT) of the gap boundary ramps was examined in mice using reflex modification audiometry, measuring inhibition of acoustic startle reflexes by variously shaped gaps just preceding reflex expression. In experiment 1 (n = 12) inhibition increased up to near-asymptotic values with longer FT (0, 1, 2, 3, 5, or 10 ms) and QT (quiet time, 0 to 13 ms), with a 2:1 trade-off between FT and QT. In experiment 2 (n = 24) inhibition increased for any RT above 0 ms (2, 3, 5, or 7 ms) if QT= 1 ms, but diminished with increased RT when QT = 3 or 8 ms. Enhanced detectability for subthreshold gaps by longer ramps results from their extending the apparent gap duration. The negative effect of increased RT for threshold gaps suggests the importance for gap detection of the stronger neural responses to sharp edges at the end of the gap shown previously in the mouse inferior colliculus. These effects are specific to gaps: inhibition for fixed (70-dB SPL) or varied level pulses (30 to 60 dB) was unaffected by varying the ramped edges (experiments 3 and 4, n = 9).

Animals↗

Rating, ranking, and understanding acoustical quality in university classrooms.

Nonoptimal classroom acoustical conditions directly affect speech perception and, thus, learning by students. Moreover, they may lead to voice problems for the instructor, who is forced to raise his/her voice when lecturing to compensate for poor acoustical conditions. The project applied previously developed simplified methods to predict speech intelligibility in occupied classrooms from measurements in unoccupied and occupied university classrooms. The methods were used to predict the speech intelligibility at various positions in 279 University of British Columbia (UBC) classrooms, when 70% occupied, and for four instructor voice levels. Classrooms were classified and rank ordered by acoustical quality, as determined by the room-average speech intelligibility. This information was used by UBC to prioritize classrooms for renovation. Here, the statistical results are reported to illustrate the range of acoustical qualities found at a typical university. Moreover, the variations of quality with relevant classroom acoustical parameters were studied to better understand the results. In particular, the factors leading to the best and worst conditions were studied. It was found that 81% of the 279 classrooms have "good," "very good," or "excellent" acoustical quality with a "typical" (average-male) instructor. However, 50 (18%) of the classrooms had "fair" or "poor" quality, and two had "bad" quality, due to high ventilation-noise levels. Most rooms were "very good" or "excellent" at the front, and "good" or "very good" at the back. Speech quality varied strongly with the instructor voice level. In the worst case considered, with a quiet female instructor, most of the classrooms were "bad" or "poor." Quality also varies with occupancy, with decreased occupancy resulting in decreased quality. The research showed that a new classroom acoustical design and renovation should focus on limiting background noise. They should promote high instructor speech levels at the back of the classrooms. This involves, in part, limiting the amount of sound absorption that is introduced into classrooms to control reverberation. Speech quality is not very sensitive to changes in reverberation, so controlling it for its own sake should not be a design priority.

Acoustics↗

Normalized amplitude quotient for parametrization of the glottal flow.

Normalized amplitude quotient (NAQ) is presented as a method to parametrize the glottal closing phase using two amplitude-domain measurements from waveforms estimated by inverse filtering. In this technique, the ratio between the amplitude of the ac flow and the negative peak amplitude of the flow derivative is first computed using the concept of equivalent rectangular pulse, a hypothetical signal located at the instant of the main excitation of the vocal tract. This ratio is then normalized with respect to the length of the fundamental period. Comparison between NAQ and its counterpart among the conventional time-domain parameters, the closing quotient, shows that the proposed parameter is more robust against distortion such as measurement noise that make the extraction of conventional time-based parameters of the glottal flow problematic. Experiments with breathy, normal, and pressed vowels indicate that NAQ is also able to separate the type of phonation effectively.

Adult↗