Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sound Spectrography”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,783 records · Page 99Linked to original sources

Hearing-aid automatic gain control adapting to two sound sources in the environment, using three time constants.

A hearing aid AGC algorithm is presented that uses a richer representation of the sound environment than previous algorithms. The proposed algorithm is designed to (1) adapt slowly (in approximately 10 s) between different listening environments, e.g., when the user leaves a single talker lecture for a multi-babble coffee-break; (2) switch rapidly (about 100 ms) between different dominant sound sources within one listening situation, such as the change from the user's own voice to a distant speaker's voice in a quiet conference room; (3) instantly reduce gain for strong transient sounds and then quickly return to the previous gain setting; and (4) not change the gain in silent pauses but instead keep the gain setting of the previous sound source. An acoustic evaluation showed that the algorithm worked as intended. The algorithm was evaluated together with a reference algorithm in a pilot field test. When evaluated by nine users in a set of speech recognition tests, the algorithm showed similar results to the reference algorithm.

Acoustics↗

Human temporal auditory acuity as assessed by envelope following responses.

Temporal auditory acuity, the ability to discriminate rapid changes in the envelope of a sound, is essential for speech comprehension. Human envelope following responses (EFRs) recorded from scalp electrodes were evaluated as an objective measurement of temporal processing in the auditory nervous system. The temporal auditory acuity of older and younger participants was measured behaviorally using both gap and modulation detection tasks. These findings were then related to EFRs evoked by white noise that was amplitude modulated (25% modulation depth) with a sweep of modulation frequencies from 20 to 600 Hz. The frequency at which the EFR was no longer detectable was significantly correlated with behavioral measurements of gap detection (r = -0.43), and with the maximum perceptible modulation frequency (r = 0.72). The EFR techniques investigated here might be developed into a clinically useful objective estimate of temporal auditory acuity for subjects who cannot provide reliable behavioral responses.

Acoustic Stimulation↗

Envelope-onset asynchrony as a cue to voicing in initial english consonants.

An acoustic cue for voicing is proposed based on the underlying processes associated with the production of the voicing contrast. This cue is based on the time asynchrony between the onsets of two amplitude-envelope signals derived from different bands of speech (i.e., envelopes derived from a lowpass-filtered band at 350 Hz and from a highpass-filtered band at 3000 Hz). Acoustic measurements made on the envelope signals of a set of 16 initial consonants represented through multiple tokens of C1VC2 syllables indicate that the onset-timing difference between the low- and high-frequency envelopes (Envelope-Onset Asynchrony or EOA) provides a reliable and robust cue for distinguishing voiced from voiceless consonants. This cue, which is simply derived in real-time, has applications to the design of sensory aids for persons with profound hearing impairments (e.g., as a supplement to lipreading), as well as to automatic speech recognition.

Adult↗

Analysis of speech-based Speech Transmission Index methods with implications for nonlinear operations.

The Speech Transmission Index (STI) is a physical metric that is well correlated with the intelligibility of speech degraded by additive noise and reverberation. The traditional STI uses modulated noise as a probe signal and is valid for assessing degradations that result from linear operations on the speech signal. Researchers have attempted to extend the STI to predict the intelligibility of nonlinearly processed speech by proposing variations that use speech as a probe signal. This work considers four previously proposed speech-based STI methods and four novel methods, studied under conditions of additive noise, reverberation, and two nonlinear operations (envelope thresholding and spectral subtraction). Analyzing intermediate metrics in the STI calculation reveals why some methods fail for nonlinear operations. Results indicate that none of the previously proposed methods is adequate for all of the conditions considered, while four proposed methods produce qualitatively reasonable results and warrant further study. The discussion considers the relevance of this work to predicting the intelligibility of cochlear-implant processed speech.

Auditory Threshold↗

The effect of recording and analysis bandwidth on acoustic identification of delphinid species.

Because many cetacean species produce characteristic calls that propagate well under water, acoustic techniques can be used to detect and identify them. The ability to identify cetaceans to species using acoustic methods varies and may be affected by recording and analysis bandwidth. To examine the effect of bandwidth on species identification, whistles were recorded from four delphinid species (Delphinus delphis, Stenella attenuata, S. coeruleoalba, and S. longirostris) in the eastern tropical Pacific ocean. Four spectrograms, each with a different upper frequency limit (20, 24, 30, and 40 kHz), were created for each whistle (n = 484). Eight variables (beginning, ending, minimum, and maximum frequency; duration; number of inflection points; number of steps; and presence/absence of harmonics) were measured from the fundamental frequency of each whistle. The whistle repertoires of all four species contained fundamental frequencies extending above 20 kHz. Overall correct classification using discriminant function analysis ranged from 30% for the 20-kHz upper frequency limit data to 37% for the 40-kHz upper frequency limit data. For the four species included in this study, an upper bandwidth limit of at least 24 kHz is required for an accurate representation of fundamental whistle contours.

Animals↗

Design and optimization of a noise reduction system for infrasonic measurements using elements with low acoustic impedance.

The implementation of the infrasound network of the International Monitoring System (IMS) for the enforcement of the Comprehensive Nuclear-Test-Ban Treaty (CTBT) increases the effort in the design of suitable noise reducer systems. In this paper we present a new design consisting of low impedance elements. The dimensioning and the optimization of this discrete mechanical system are based on numerical simulations, including a complete electroacoustical modeling and a realistic wind-noise model. The frequency response and the noise reduction obtained for a given wind speed are compared to statistical noise measurements in the [0.02-4] Hz frequency band. The effects of the constructive parameters-the length of the pipes, inner diameters, summing volume, and number of air inlets-are investigated through a parametric study. The studied system consists of 32 air inlets distributed along an overall diameter of 16 m. Its frequency response is flat up to 4 Hz. For a 2 m/s wind speed, the maximal noise reduction obtained is 15 dB between 0.5 and 4 Hz. At lower frequencies, the noise reduction is improved by the use of a system of larger diameter. The main drawback is the high-frequency limitation introduced by acoustical resonances inside the pipes.

Acoustic Impedance Tests↗

Drilling and operational sounds from an oil production island in the ice-covered Beaufort sea.

Recordings of sounds underwater and in air, and of iceborne vibrations, were obtained at Northstar Island, an artificial gravel island in the Beaufort Sea near Prudhoe Bay (Alaska). The aim was to document the levels, characteristics, and range dependence of sounds and vibrations produced by drilling and oil production during the winter, when the island was surrounded by shore-fast ice. Drilling produced the highest underwater broadband (10-10,000 Hz) levels (maximum= 124 dB re: 1 microPa at 1 km), and mainly affected 700-1400 Hz frequencies. In contrast, drilling did not increase broadband levels in air or ice relative to levels during other island activities. Production did not increase broadband levels for any of the sensors. In all media, broadband levels decreased by approximately 20 dB/tenfold change in distance. Background levels underwater were reached by 9.4 km during drilling and 3-4 km without. In the air and ice, background levels were reached 5-10 km and 2-10 km from Northstar, respectively, depending on the wind but irrespective of drilling. A comparison of the recorded sounds with harbor and ringed seal audiograms showed that Northstar sounds were probably audible to seals, at least intermittently, out to approximately 1.5 km in water and approximately 5 km in air.

Acoustics↗

Robustness of spatial average equalization: a statistical reverberation model approach.

Traditionally, multiple listener room equalization is performed to improve sound quality at all listeners, during audio playback, in a multiple listener environment (e.g., movie theaters, automobiles, etc.). A typical way of doing multiple listener equalization is through spatial averaging, where the room responses are averaged spatially between positions and an inverse equalization filter is found from the spatially averaged result. However, the equalization performance, will be affected if there is a mismatch between the position of the microphones (which are used for measuring the room responses for designing the equalization filter) and the actual center of listener head position (during playback). In this paper, we will present results on the effects of microphone-listener mismatch on spatial average equalization performance. The results indicate that, for the analyzed rectangular configuration, the region of effective equalization depends on (i) the distance of a listener from the source, (ii) the amount of mismatch between the responses, and (iii) the frequency of the audio signal. We also present some convergence analysis to interpret the results.

Architecture↗

Habitat-dependent ambient noise: consistent spectral profiles in two African forest types.

Many animal species use acoustic signals to attract mates, to defend territories, or to convey information that may contribute to their fitness in other ways. However, the natural environment is usually filled with competing sounds. Therefore, if ambient noise conditions are relatively constant, acoustic interference can drive evolutionary changes in animal signals. Furthermore, masking noise may cause acoustic divergence between populations of the same species if noise conditions differ consistently among habitats. In this study, ambient noise was sampled in a replicate set of sites in two habitat types in Cameroon: contiguous rainforest and ecotone forest patches north of the rainforest. The noise characteristics of the two forest types show significant and consistent differences. Multiple samples taken at two rainforest sites in different seasons vary little and remain distinct from those in ecotone forest. The rainforest recordings show many distinctive frequency bands, with a general increase in amplitude from low to high frequencies. Ecotone forest only shows a distinctive high-frequency band at some parts of the day. Habitat-dependent abiotic and biotic sound sources and to some extent habitat-dependent sound transmission are the likely causes of these habitat-dependent noise spectra.

Animals↗

An echolocation model for the restoration of an acoustic image from a single-emission echo.

Bats can form a fine acoustic image of an object using frequency-modulated echolocation sound. The acoustic image is an impulse response, known as a reflected-intensity distribution, which is composed of amplitude and phase spectra over a range of frequencies. However, bats detect only the amplitude spectrum due to the low-time resolution of their peripheral auditory system, and the frequency range of emission is restricted. It is therefore necessary to restore the acoustic image from limited information. The amplitude spectrum varies with the changes in the configuration of the reflected-intensity distribution, while the phase spectrum varies with the changes in its configuration and location. Here, by introducing some reasonable constraints, a method is proposed for restoring an acoustic image from the echo. The configuration is extrapolated from the amplitude spectrum of the restricted frequency range by using the continuity condition of the amplitude spectrum at the minimum frequency of the emission and the minimum phase condition. The determination of the location requires extracting the amplitude spectra, which vary with its location. For this purpose, the Gaussian chirplets with a carrier frequency compatible with bat emission sweep rates were used. The location is estimated from the temporal changes of the amplitude spectra.

Animals↗

The bat head-related transfer function reveals binaural cues for sound localization in azimuth and elevation.

Directional properties of the sound transformation at the ear of four intact echolocating bats, Eptesicus fuscus, were investigated via measurements of the head-related transfer function (HRTF). Contributions of external ear structures to directional features of the transfer functions were examined by remeasuring the HRTF in the absence of the pinna and tragus. The investigation mainly focused on the interactions between the spatial and the spectral features in the bat HRTF. The pinna provides gain and shapes these features over a large frequency band (20-90 kHz), and the tragus contributes gain and directionality at the high frequencies (60 to 90 kHz). Analysis of the spatial and spectral characteristics of the bat HRTF reveals that both interaural level differences (ILD) and monaural spectral features are subject to changes in sound source azimuth and elevation. Consequently, localization cues for horizontal and vertical components of the sound source location interact. Availability of multiple cues about sound source azimuth and elevation should enhance information to support reliable sound localization. These findings stress the importance of the acoustic information received at the two ears for sound localization of sonar target position in both azimuth and elevation.

Animals↗

Distortion product otoacoustic emission (2f1-f2) suppression in 3-month-old infants: evidence for postnatal maturation of human cochlear function?

The complete timeline for maturation of human cochlear function has not been defined. Distortion product otoacoustic emission (DPOAE)-based measures of cochlear function show non-adult-like responses from premature and term-born neonates at high f2 frequencies; however, older infants were not included in these studies. In the present experiment, previously collected DPOAE ipsilateral suppression data from premature neonates were combined with new data collected from adults, term-born neonates, and 3-month-old infants to further examine the time course for maturation of cochlear function. DPOAE suppression tuning curves (STC) and suppression growth patterns were measured in the three age groups at f2 = 6000 Hz, L1 = 65, L2 = 55 dB SPL, with an f2/f1 of 1.2. Results indicate that term-born neonates and 3-month-old infants have non-adult-like STC width, slope on the low-frequency flank, and tip features. However, the two infant groups are not significantly different from one another. Suppression growth patterns for low-frequency suppressor tones show a clear developmental progression. In general, the younger the infant, the more shallow and compressive the suppression growth for the lowest suppressor frequencies. These findings suggest a high-frequency postnatal immaturity in cochlear function as measured by DPOAE suppression. Results may have been influenced by noncochlear factors, such as middle-ear immaturity. These factors are reviewed and considered.

Acoustic Stimulation↗

Auditory processing of real and illusory changes in frequency modulation (FM) phase.

Auditory processing of frequency modulation (FM) was explored. In experiment 1, detection of a tau-radians modulator phase shift deteriorated as modulation rate increased from 2.5 to 20 Hz, for 1- and 6-kHz carriers. In experiment 2, listeners discriminated between two 1-kHz carriers, where, mid-way through, the 10-Hz frequency modulator had either a phase shift or increased in depth by deltaD% for half a modulator period. Discrimination was poorer for deltaD = 4% than for smaller or larger increases. These results are consistent with instantaneous frequency being smoothed by a time window with a total duration of about 110 ms. In experiment 3, the central 200-ms of a 1-s 1-kHz carrier modulated at 5 Hz was replaced by noise, or by a faster FM applied to a more intense 1-kHz carrier. Listeners heard the 5-Hz FM continue at the same depth throughout the stimulus. Experiments 4 and 5 showed that, after an FM tone had been interrupted by a 200-ms noise, listeners were insensitive to the phase at which the FM resumed. It is argued that the auditory system explicitly encodes the presence, and possibly the rate and depth, of FM in a way that does not preserve information on FM phase.

Auditory Threshold↗

Vocal tract filtering and the "coo" of doves.

Ring doves (Streptopelia risoria) produce a "coo" vocalization that is essentially a pure-tone sound at a frequency of about 600 Hz and with a duration of about 1.5 s. While making this vocalization, the dove inflates the upper part of its esophagus to form a thin-walled sac structure that radiates sound to the surroundings. It is a reasonable assumption that the combined influence of the trachea, glottis and inflated upper esophagus acts as an effective band-pass filter to eliminate higher harmonics generated by the vibrating syringeal valve. Calculations reported here indicate that this is indeed the case. The tracheal tube, terminated by a glottal constriction, is the initial resonant structure, and subsequent resonant filtering takes place through the action of the inflated esophageal sac. The inflated esophagus proves to be a more efficient sound radiating mechanism than an open beak. The action of this sac is only moderately affected by the degree of inflation, although an uninflated esophagus is inactive as a sound radiator. These conclusions are supported by measurements and observations that have been reported in a companion paper.

Airway Resistance↗

Distortion product otoacoustic emissions provide clues hearing mechanisms in the frog ear.

2f1-f2 and 2 f2-f1 distortion product otoacoustic emissions (DPOAEs) were recorded from both ears of male and female Rana pipiens pipiens and Rana catesbeiana. The input-output (I/O) curves obtained from the amphibian papilla (AP) of both frog species are analogous to I/O curves recorded from mammals suggesting that, similarly to the mammalian cochlea, there may be an amplification process present in the frog AP. DPOAE level dependence on L1-L2 is different from that in mammals and consistent with intermodulation distortion expectations. Therefore, if a mechanical structure in the frog inner ear is functioning analogously to the mammalian basilar membrane, it must be more broadly tuned. DPOAE audiograms were obtained for primary frequencies spanning the animals' hearing range and selected stimulus levels. The results confirm that DPOAEs are produced in both papillae, with R. catesbeiana producing stronger emissions than R. p. pipiens. Consistent with previously reported sexual dimorphism in the mammalian and anuran auditory systems, females of both species produce stronger emissions than males. Moreover, it appears that 2 f1-f2 in the frog is generated primarily at the DPOAE frequency place, while 2 f2-f1 is generated primarily at a frequency place around the primaries. Regardless of generation place, both emissions within the AP may be subject to the same filtering mechanism, possibly the tectorial membrane.

Acoustic Stimulation↗

Direct measurement of onset and offset phonation threshold pressure in normal subjects.

Phonation threshold pressures were directly measured in five normal subjects in a variety of voicing conditions. The effects of fundamental frequency, intensity, closure speed of the vocal folds, and laryngeal airway resistance on phonation threshold pressures were determined. Subglottic air pressures were measured using percutaneous puncture of the cricothyroid membrane. Both onset and offset of phonation were studied to see if a hysteresis effect produced lower offset pressures than onset pressures. Univariate analysis showed that phonation threshold pressure was influenced most strongly by fundamental frequency and intensity. Multiple linear regression showed that these two variables, as well as laryngeal airway resistance, most strongly predicted phonation threshold pressure. Two of the five subjects demonstrated a significant hysteresis effect, but one subject actually had higher offset pressures than onset pressures.

Adult↗

The temporal representation of speech in a nonlinear model of the guinea pig cochlea.

The temporal representation of speechlike stimuli in the auditory-nerve output of a guinea pig cochlea model is described. The model consists of a bank of dual resonance nonlinear filters that simulate the vibratory response of the basilar membrane followed by a model of the inner hair cell/auditory nerve complex. The model is evaluated by comparing its output with published physiological auditory nerve data in response to single and double vowels. The evaluation includes analyses of individual fibers, as well as ensemble responses over a wide range of best frequencies. In all cases the model response closely follows the patterns in the physiological data, particularly the tendency for the temporal firing pattern of each fiber to represent the frequency of a nearby formant of the speech sound. In the model this behavior is largely a consequence of filter shapes; nonlinear filtering has only a small contribution at low frequencies. The guinea pig cochlear model produces a useful simulation of the measured physiological response to simple speech sounds and is therefore suitable for use in more advanced applications including attempts to generalize these principles to the response of human auditory system, both normal and impaired.

Animals↗

Rapid adaptation to foreign-accented English.

This study explored the perceptual benefits of brief exposure to non-native speech. Native English listeners were exposed to English sentences produced by non-native speakers. Perceptual processing speed was tracked by measuring reaction times to visual probe words following each sentence. Three experiments using Spanish- and Chinese-accented speech indicate that processing speed is initially slower for accented speech than for native speech but that this deficit diminishes within one minute of exposure. Control conditions rule out explanations for the adaptation effect based on practice with the task and general strategies for dealing with difficult speech. Further results suggest that adaptation can occur within as few as two to four sentence-length utterances. The findings emphasize the flexibility of human speech processing and require models of spoken word recognition that can rapidly accommodate significant acoustic-phonetic deviations from native language speech patterns.

Adaptation, Psychological↗