Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sound Spectrography”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,513 records · Page 84Linked to original sources

The effects of spatial separation in distance on the informational and energetic masking of a nearby speech signal.

Although many studies have shown that intelligibility improves when a speech signal and an interfering sound source are spatially separated in azimuth, little is known about the effect that spatial separation in distance has on the perception of competing sound sources near the head. In this experiment, head-related transfer functions (HRTFs) were used to process stimuli in order to simulate a target talker and a masking sound located at different distances along the listener's interaural axis. One of the signals was always presented at a distance of 1 m, and the other signal was presented 1 m, 25 cm, or 12 cm from the center of the listener's head. The results show that distance separation has very different effects on speech segregation for different types of maskers. When speech-shaped noise was used as the masker, most of the intelligibility advantages of spatial separation could be accounted for by spectral differences in the target and masking signals at the ear with the higher signal-to-noise ratio (SNR). When a same-sex talker was used as the masker, the intelligibility advantages of spatial separation in distance were dominated by binaural effects that produced the same performance improvements as a 4-5-dB increase in the SNR of a diotic stimulus. These results suggest that distance-dependent changes in the interaural difference cues of nearby sources play a much larger role in the reduction of the informational masking produced by an interfering speech signal than in the reduction of the energetic masking produced by an interfering noise source.

Adult↗

Buildup and breakdown of echo suppression for stimuli presented over headphones-the effects of interaural time and level differences.

The current study investigates buildup and breakdown of echo suppression for stimuli presented over headphones. The stimuli consisted of pairs of 120-micros clicks. The leading click (lead) and the lagging click (lag) in each pair were lateralized on opposite sides of the midline by means of interaural level differences (ILDs) of +/-10 dB or interaural time differences (ITDs) of +/-300 micros. Echo threshold was measured with an adaptive one-interval, two-alternative, forced-choice procedure with a subjective decision criterion, in which listeners had to report whether they heard a single, fused auditory event on one side of the midline, or two separate events on both sides. In the control conditions, referred to as the "single" conditions, echo threshold was measured for a single click pair, the test pair, presented in isolation. In addition to the control conditions, two kinds of test conditions were investigated, in which the test pair was preceded by 12 identical conditioning pairs: in the "same" conditions, the interaural configuration (ILDs or ITDs) of the conditioning pairs was identical to that of the test pair; in the "switch" conditions, the interaural configuration of lead and lag was reversed between the conditioning pairs and the test pair, in order to produce a switch in the lateralizations of the stimuli between the conditioning train and the test pair. No matter whether the lateralization of the clicks was produced by ILDs or by ITDs, most listeners experienced a buildup of echo suppression in the "same" conditions, manifested by a prolongation of echo threshold relative to the respective "single" conditions. However, the breakdown of echo suppression was much stronger in the ILD-switch than in the ITD-switch conditions. In five out of six listeners, the ITD switch had hardly any effect on echo threshold, although the ITDs (+/-300 micros) produced roughly the same degree of lateral displacement as the ILDs (+/-10 dB). These results suggest that the dynamic processes in echo suppression operate differentially in pathways responsible for the processing of interaural time and level differences.

Adult↗

Broadband sound generation by confined turbulent jets.

Sound generation by confined stationary jets is of interest to the study of voice and speech production, among other applications. The generation of sound by low Mach number, confined, stationary circular jets was investigated. Experiments were performed using a quiet flow supply, muffler-terminated rigid uniform tubes, and acrylic orifice plates. A spectral decomposition method based on a linear source-filter model was used to decompose radiated nondimensional sound pressure spectra measured for various gas mixtures and mean flow velocities into the product of (1) a source spectral distribution function; (2) a function accounting for near field effects and radiation efficiency; and (3) an acoustic frequency response function. The acoustic frequency response function agreed, as expected, with the transfer function between the radiated acoustic pressure at one fixed location and the strength of an equivalent velocity source located at the orifice. The radiation efficiency function indicated a radiation efficiency of the order (kD)2 over the planar wave frequency range and (kD)4 at higher frequencies, where k is the wavenumber and D is the tube cross sectional dimension. This is consistent with theoretical predictions for the planar wave radiation efficiency of quadrupole sources in uniform rigid anechoic tubes. The effects of the Reynolds number, Re, on the source spectral distribution function were found to be insignificant over the range 2000 2.5. The influence of a reflective open tube termination on the source function spectral distribution was found to be insignificant, confirming the absence of a feedback mechanism.

Aircraft↗

Acoustic characteristics of twin jets.

Experiments were conducted to investigate the acoustic characteristics of underexpanded supersonic twin jets in different azimuthal measurement planes. Compared with two independent jets, the twin jets produced additional noise due to the enhanced mixing and entrainment. The larger pressure ratio for switching from the axisymmetric mode to the helical mode led to lower noise levels at 90 degrees than for two independent jets. For pressure ratios greater than 5.00, the noise reduction was due to cessation of screeching of the twin jets while screeching of a single jet was still detected. The apparent shielding phenomenon was measured for the screech helical mode. The screech tone intensities were attenuated largely due to the shielding effects. The noise reductions due to shielding were obtained over a wide range of pressure ratios relative to the sum of two independent jets.

Acoustics↗

Modulation frequency and modulation level owing to vocal microtremor.

Vocal microtremor designates a normal slow modulation of the vocal cycle lengths of speakers who do not suffer from pathological tremor of the limbs and whose voices are not perceived as tremulous. Vocal microtremor is therefore distinct from pathological vocal tremor. The objective is to report data about the modulation frequency and modulation level owing to vocal microtremor. The modulation data have been obtained for vowels [a], [i], and [u] sustained by normophonic and mildly dysphonic male and female speakers. The results are the following. First, modulation frequencies and relative modulation levels do not differ significantly for male and female speakers, normophonic and mildly dysphonic speakers, as well as for vowel timbres [a], [i], and [u]. Second, the typical interquartile intervals of the modulation frequency and modulation level are equal to 2.0-4.7 Hz and 0.4%-1.3%, respectively. Third, dissimilarities between data reported by different studies are due to different cutoff frequencies below which spectral peaks are considered not to contribute to vocal microtremor.

Data Interpretation, Statistical↗

Acoustic intensity, impedance and reflection coefficient in the human ear canal.

The sound power per unit cross-sectional area was determined in human ear canals using a new method based on measuring the pressure distribution (P) along the length of variable cross-section acoustic waveguides. The technique provides the pressure/power reflection coefficients (R/R) as well as the acoustic intensity of the nonplanar incident wave (I+, the acoustic input to the ear) and the nonplanar outgoing wave (I-, the acoustic output of the ear). Results were compared to the classical acoustic impedance (Z) and associated plane-wave power reflection coefficient (R(Z)). Performance of the method was investigated theoretically using horn equation simulations and evaluated experimentally using pressure data recorded in nonuniform waveguides. The method was applied in normal-hearing young adults to determine ear-canal position- and frequency-dependence of I(+/-), R, and R(Z) using random phase broadband stimuli (1-15 kHz; approximately 75 dB SPL). Reflection coefficient (R) measurements at two different locations within individual human ear canals exhibited a position dependence averaging deltaR approximately 0.1 (over 6 mm distance)--a difference consistent with predictions of inviscid acoustics in nonuniform waveguides. Since this position dependence was relatively small, an "optimized" position-independent reflection coefficient was defined to facilitate practical application and intersubject comparisons.

Acoustic Impedance Tests↗

Auditory temporal resolution in birds: discrimination of harmonic complexes.

The ability of three species of birds to discriminate among selected harmonic complexes with fundamental frequencies varying from 50 to 1000 Hz was examined in behavioral experiments. The stimuli were synthetic harmonic complexes with waveform shapes altered by component phase selection, holding spectral and intensive information constant. Birds were able to discriminate between waveforms with randomly selected component phases and those with all components in cosine phase, as well as between positive and negative Schroeder-phase waveforms with harmonic periods as short as 1-2 ms. By contrast, human listeners are unable to make these discriminations at periods less than about 3-4 ms. Electrophysiological measures, including cochlear microphonic and compound action potential measurements to the same stimuli used in behavioral tests, showed differences between birds and gerbils paralleling, but not completely accounting for, the psychophysical differences observed between birds and humans. It appears from these data that birds can hear the fine temporal structure in complex waveforms over very short periods. These data show birds are capable of more precise temporal resolution for complex sounds than is observed in humans and perhaps other mammals. Physiological data further show that at least part of the mechanisms underlying this high temporal resolving power resides at the peripheral level of the avian auditory system.

Animals↗

Auditory brainstem responses in adult budgerigars (Melopsittacus undulatus).

The auditory brainstem response (ABR) was recorded in adult budgerigars (Melopsittacus undulatus) in response to clicks and tones. The typical budgerigar ABR waveform showed two prominent peaks occurring within 4 ms of the stimulus onset. As sound-pressure levels increased, ABR peak latency decreased, and peak amplitude increased for all waves while interwave interval remained relatively constant. While ABR thresholds were about 30 dB higher than behavioral thresholds, the shape of the budgerigar audiogram derived from the ABR closely paralleled that of the behavioral audiogram. Based on the ABR, budgerigars hear best between 1000 and 5700 Hz with best sensitivity at 2860 Hz-the frequency corresponding to the peak frequency in budgerigar vocalizations. The latency of ABR peaks increased and amplitude decreased with increasing repetition rate. This rate-dependent latency increase is greater for wave 2 as indicated by the latency increase in the interwave interval. Generally, changes in the ABR to stimulation intensity, frequency, and repetition rate are comparable to what has been found in other vertebrates.

Acoustic Stimulation↗

A model cochlear partition involving longitudinal elasticity.

This paper addresses the issue of longitudinal stiffness within the cochlea. A one-dimensional model of the cochlear partition is presented in which the resonant sections are coupled by longitudinal elastic elements. These elements functionally represent the aggregate mechanical effect of the connective tissue that spans the length of the organ of Corti. With the plate-like morphology of the cochlear partition in mind, the contribution of longitudinal elasticity to partition dynamics is appreciable, though weak and nonlinear. If the elasticity is considered Hookian then the nonlinearity takes a cubic form. Numerical solutions are presented that demonstrate the compressive nature of the partial differential nonlinear equations and their ability to produce realistic cubic distortion product otoacoustic emissions. Within the framework of this model, some speculations can be made regarding the dynamical function of the phalangeal processes, the sharpness of active cochlear mechanics, and the propogation of pathology along the partition.

Basilar Membrane↗

Captive dolphins, Tursiops truncatus, develop signature whistles that match acoustic features of human-made model sounds.

This paper presents a cross-sectional study testing whether dolphins that are born in aquarium pools where they hear trainers' whistles develop whistles that are less frequency modulated than those of wild dolphins. Ten pairs of captive and wild dolphins were matched for age and sex. Twenty whistles were sampled from each dolphin. Several traditional acoustic features (total duration, duration minus any silent periods, etc.) were measured for each whistle, in addition to newly defined flatness parameters: total flatness ratio (percentage of whistle scored as unmodulated), and contiguous flatness ratio (duration of longest flat segment divided by total duration). The durations of wild dolphin whistles were found to be significantly longer, and the captive dolphins had whistles that were less frequency modulated and more like the trainers' whistles. Using a standard t-test, the captive dolphin had a significantly higher total flatness ratio in 9/10 matched pairs, and in 8/10 pairs the captive dolphin had significantly higher contiguous flatness ratios. These results suggest that captive-born dolphins can incorporate features of artificial acoustic models made by humans into their signature whistles.

Animal Communication↗

Rules for controlling low-dimensional vocal fold models with muscle activation.

A low-dimensional, self-oscillation model of the vocal folds is used to capture three primary modes of vibration, a shear mode and two compressional modes. The shear mode is implemented with either two vertical masses or a rotating plate, and the compressional modes are implemented with an additional bar mass between the vertically stacked masses and the lateral boundary. The combination of these elements allows for the anatomically important body-cover differentiation of vocal fold tissues. It also allows for reconciliation of lumped-element mechanics with continuum mechanics, but in this reconciliation the oscillation region is restricted to a nearly rectangular glottis (as in all low-dimensional models) and a small effective thickness of vibration (<3 mm). The model is controlled by normalized activation levels of the cricothyroid (CT), thyroarytenoid (TA), lateral cricoarytenoid (LCA), and posterior cricoarytenoid (PCA) muscles, and lung pressure. An empirically derived set of rules converts these muscle activities into physical quantities such as vocal fold strain, adduction, glottal convergence, mass, thickness, depth, and stiffness. Results show that oscillation regions in muscle activation control spaces are similar to those measured by other investigations on human subjects.

Biomechanical Phenomena↗

Learning to perceive speech: how fricative perception changes, and how it stays the same.

A part of becoming a mature perceiver involves learning what signal properties provide relevant information about objects and events in the environment. Regarding speech perception, evidence supports the position that allocation of attention to various signal properties changes as children gain experience with their native language, and so learn what information is relevant to recognizing phonetic structure in that language. However, one weakness in that work has been that data have largely come from experiments that all use similarly designed stimuli and show similar age-related differences in labeling. In this study, two perception experiments were conducted that used stimuli designed differently from past experiments, with different predictions. In experiment 1, adults and children (4, 6, and 8 years of age) labeled stimuli with natural /f/ and /[see text]/ noises and synthetic vocalic portions that had initial formant transitions varying in appropriateness for /f/ or /[see text]/. The prediction was that similar labeling patterns would be found for all listeners. In experiment 2, adults and children labeled stimuli with initial /s/-like and /[see text]/-like noises and synthetic vocalic portions that had initial formant transitions varying in appropriateness for /s/ or /[see text]/. The prediction was that, as found before, children would weight formant transitions more and fricative noises less than adults, but that this age-related difference would elicit different patterns of labeling from those found previously. Results largely matched predictions, and so further evidence was garnered for the position that children learn which properties of the speech signal provide relevant information about phonetic structure in their native language.

Adult↗

Limitations on rate discrimination.

We investigated the limits of temporal pitch processing under conditions where the place and rate of stimulation on the basilar membrane were independent. Stimuli were harmonic complexes passed through a fixed bandpass filter and resembled filtered pulse trains. The task was to detect a difference in F0. When the harmonics were filtered between 3900-5400 Hz, presented monaurally, and summed in sine phase, subjects could perform the task at all FOs studied. However, when the pulse rate was doubled by summing components in alternating phase, thresholds increased with increasing F0 until the task was impossible at F0 = 300 Hz (pulse rate=600 pps). Thresholds improved again at higher FOs, presumably because some harmonics became resolved. The F0 at which this breakdown occurred decreased when the complexes were filtered into a lower frequency region, and increased when they were filtered into a higher region. In the highest region tested (7800-10800 Hz), all listeners could detect an increase of less than about 20% re: a pulse rate of 600 pps for alternating-phase complexes. Presenting a copy of the standard (lower-F0) stimulus to the contralateral ear during all intervals of a forced-choice trial improved performance markedly under conditions where monaural rate discrimination was very poor. This showed that temporal information is present in the auditory nerve that is unavailable to the temporal pitch mechanism, but which is accessible when a binaural cue is available. The results are compared to the inability of most cochlear implantees to detect increases in the rate of electrical pulse trains above about 300 pps. It is concluded that this inability is unlikely to result entirely from a central pitch limitation, because, with analogous acoustic stimulation, normal listeners can perform the task at substantially higher rates.

Adolescent↗

Temporal weighting in sound localization.

The dynamics of sound localization were studied using a free-field direct localization task (pointing to sound sources) and an observer-weighting analysis that assessed the relative influence of each click in a click-train stimulus. In agreement with previous studies of the precedence effect and binaural adaptation, weighting functions showed increased influence of the onset click when the interclick interval (ICI) was short (<5 ms). For longer ICIs, all clicks in a train contributed roughly the same amount to listeners' localization responses. Finally, when a short gap was introduced in the middle of a train, the influence of the click immediately following the gap increased, in agreement with the "restarting" results obtained by Hafter and Buell [J. Acoust. Soc. Am. 88, 806-812 (1990)].

Acoustic Stimulation↗

Broadband transmission noise reduction of smart panels featuring piezoelectric shunt circuits and sound-absorbing material.

The possibility of a broadband noise reduction of piezoelectric smart panels is experimentally studied. A piezoelectric smart panel is basically a plate structure on which piezoelectric patches with electrical shunt circuits are mounted and sound-absorbing material is bonded on the surface of the structure. Sound-absorbing material can absorb the sound transmitted at the midfrequency region effectively while the use of piezoelectric shunt damping can reduce the transmission at resonance frequencies of the panel structure. To be able to reduce the sound transmission at low panel resonance frequencies, piezoelectric damping using the measured electrical impedance model is adopted. A resonant shunt circuit for piezoelectric shunt damping is composed of resistor and inductor in series, and they are determined by maximizing the dissipated energy through the circuit. The transmitted noise-reduction performance of smart panels is tested in an acoustic tunnel. The tunnel is a square cross-sectional tube and a loudspeaker is mounted at one side of the tube as a sound source. Panels are mounted in the middle of the tunnel and the transmitted sound pressure across panels is measured. When an absorbing material is bonded on a single plate, a remarkable transmitted noise reduction in the midfrequency region is observed except for the fundamental resonance frequency of the plate. By enabling the piezoelectric shunt damping, noise reduction is achieved at the resonance frequency as well. Piezoelectric smart panels incorporating passive absorbing material and piezoelectric shunt damping is a promising technology for noise reduction over a broadband of frequencies.

Acoustics↗

A speech enhancement scheme incorporating spectral expansion evaluated with simulated loss of frequency selectivity.

Hearing-impaired listeners often suffer from supra-threshold speech perception deficits. One such deficit is reduced frequency selectivity. We applied a speech enhancement scheme that incorporated spectral expansion in an attempt to reduce the effects of this deficit. The speech processing could contain up to three stages, a first in which the peak-valley ratio of the speech spectrum was enlarged to counteract the broadening of the auditory filtering, and a second in which the overall speech spectrum was modified to counteract the effects of upward-spread-of-masking, using a linear filter. The third stage was a noise suppression stage, applied before the spectral enhancement. The effectiveness of the speech processing with and without noise suppression was evaluated for various parameter settings by measuring the speech reception threshold (SRT) in noise, i.e., the signal-to-noise ratio at which listeners repeat 50% of presented sentences correctly. We used normal-hearing subjects. To simulate the loss of frequency selectivity we applied spectral smearing to the stimuli presented to the subjects. The speech material of the SRT tests was mixed with the noise before processing, and, when present, the smearing was applied last. The results indicated that for one specific parameter setting the SRT values decreased (i.e., improved) by approximately 1 dB when incorporating the spectral expansion together with the linear filtering. Employing either of these two stages separately did not improve the SRT. The application of the noise suppression stage did not further improve the SRT. A pilot study using hearing-impaired listeners showed more promising results for a female than for a male speaker.

Adult↗

Enhancing sensitivity to interaural delays at high frequencies by using "transposed stimuli".

It is well-known that thresholds for ongoing interaural temporal disparities (ITDs) at high frequencies are larger than threshold ITDs obtained at low frequencies. These differences could reflect true differences in the binaural mechanisms that mediate performance. Alternatively, as suggested by Colburn and Esquissaud [J. Acoust. Soc. Am. Suppl. 1 59, S23 (1976)], they could reflect differences in the peripheral processing of the stimuli. In order to investigate this issue, threshold ITDs were measured using three types of stimuli: (1) low-frequency pure tones; (2) 100% sinusoidally amplitude-modulated (SAM) high-frequency tones, and (3) special, "transposed" high-frequency stimuli whose envelopes were designed to provide the high-frequency channels with information similar to that available in low-frequency channels. The data and their interpretation can be characterized by two general statements. First, threshold ITDs obtained with the transposed stimuli were generally smaller than those obtained with SAM tones and, at modulation frequencies of 128 and 64 Hz, were equal to or smaller than threshold ITDs obtained with their low-frequency pure-tone counterparts. Second, quantitative analyses revealed that the data could be well accounted for via a model based on normalized interaural correlations computed subsequent to known stages of peripheral auditory processing augmented by low-pass filtering of the envelopes within the high-frequency channels of each ear. The data and the results of the quantitative analyses appear to be consistent with the general ideas comprising Colburn and Esquissaud's hypothesis.

Adult↗

A quasiarticulatory approach to controlling acoustic source parameters in a Klatt-type formant synthesizer using HLsyn.

The HLsyn speech synthesizer uses models of the vocal tract to map higher-level quasiarticulatory parameters to the acoustic parameters of a Klatt-type formant synthesizer. The benefits of this system are several. In addition to requiring a relatively small number of parameters, the HLsyn model includes constraints on source-filter relations that occur naturally during speech production. Such constraints help to prevent combinations of sources and filter that are impossible to achieve with the human vocal tract. Thus, HLsyn could lead to reductions in the complexity of formant synthesis and result in better quality synthesis. HLsyn can also be a useful tool for speech-science education and speech research. This paper focuses on the generation of acoustic sources in HLsyn. Described in detail are the equations and methods used to estimate Klatt-type source parameters from HLsyn parameters. Several examples illustrating the generation of source parameters for obstruents (voiced and voiceless) and sonorants are provided. Future papers will describe the filtering components of HLsyn.

Communication Devices for People with Disabilities↗