Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sound Spectrography”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,675 records · Page 93Linked to original sources

Measurement and reproduction accuracy of computer-controlled grand pianos.

The recording and reproducing capabilities of a Yamaha Disklavier grand piano and a Bösendorfer SE290 computer-controlled grand piano were tested, with the goal of examining their reliability for performance research. An experimental setup consisting of accelerometers and a calibrated microphone was used to capture key and hammer movements, as well as the acoustic signal. Five selected keys were played by pianists with two types of touch ("staccato" and "legato"). Timing and dynamic differences between the original performance, the corresponding MIDI file recorded by the computer-controlled pianos, and its reproduction were analyzed. The two devices performed quite differently with respect to timing and dynamic accuracy. The Disklavier's onset capturing was slightly more precise (+/- 10 ms) than its reproduction (-20 to +30 ms); the Bösendorfer performed generally better, but its timing accuracy was slightly less precise for recording (-10 to 3 ms) than for reproduction (+/- 2 ms). Both devices exhibited a systematic (linear) error in recording over time. In the dynamic dimension, the Bösendorfer showed higher consistency over the whole dynamic range, while the Disklavier performed well only in a wide middle range. Neither device was able to capture or reproduce different types of touch.

Equipment Design↗

An approximate transfer function for the dual-resonance nonlinear filter model of auditory frequency selectivity.

The dual-resonance nonlinear filter [Meddis et al., J. Acoust. Soc. Am. 109, 2852-2861 (2001)] was presented as a digital time-domain algorithm to model nonlinear auditory frequency selectivity. This report extends previous work by presenting an approximate analytic transfer function that allows calculating and analyzing its level-dependent frequency-domain response. The transfer function is derived on the assumption that the filter behaves linearly for any given input amplitude. It matches accurately the response (gain and phase) of the digital filter for tones. Practical uses for the transfer function are suggested.

Algorithms↗

Spectral pattern, harmonic relations, and the perceptual grouping of low-numbered components.

Mistuning a harmonic increases its salience and produces an exaggerated change in its pitch. The effects on component grouping of spectral pattern, global pitch, and local harmonicity were explored using these phenomena. Stimuli were either harmonic (F0 = 200 Hz) or frequency shifted by 25% of F0. Component 1 or 2 was replaced by one of a set of sinusoidal probes in the same spectral region. Listeners either matched the probe pitch by adjusting the frequency of a pure tone (experiments 1 and 3) or matched the probe loudness by adjusting the level of a tone of identical frequency (experiment 2). Probe positions corresponding to greatest perceptual fusion were estimated from the variations in pitch shift and loudness across frequency. Both measures gave similar estimates. For harmonic stimuli, fusion was greatest at harmonic values. For shifted stimuli, fusion was greatest close to the suboctave (225 Hz) and the frequency (450 Hz) of component 2. The latter value moved downwards to near 433 Hz (2:3 ratio with component 3) when component 1 was removed. Together, these results indicate that the lowest component of a shifted complex is grouped by local harmonicity, whereas the higher components are grouped by common spectral spacing. Global pitch did not influence component grouping.

Cochlear Nerve↗

Scale-model study of the effectiveness of highway noise barriers.

A scale-model facility was developed to test the insertion loss (IL) of highway noise barriers. Three model materials were utilized to simulate packed-earth berms and ground (expanded polystyrene), vertical walls (dense polystyrene), and roadways (varnished particleboard). Thirty-eight noise-barrier configurations were tested and used to compare how IL varied with changes to the barrier profile for walls, berms, and combinations of walls and berms for receivers at a representative, highway-adjacent location. The atmospheric conditions were assumed to be homogeneous and nonrefracting. Changes of barrier surface impedance were also assessed. A highway line source was simulated by positioning both an air-jet point source and a receiver microphone at a series of equally spaced points, in order to form an array of source-receiver measurement pairs making differing angles of propagation to the noise-barrier crest line. The IL measurement results are presented in unweighted third-octave bands. In addition, total A-weighted insertion losses (ILA) were obtained by applying an A-weighted, traffic-noise spectrum. When a berm was modeled with surface impedance closely matching that of packed earth, it was found that walls outperformed berms by 1 to 2 dBA. When the surface impedance of a berm was modeled to be acoustically soft, the ILA increased sufficiently to favor berms by about 2 dBA. The result for an acoustically soft berm does not support the long-standing practice of assuming that earth berms outperform walls by 3 dBA, but is consistent with the performance predicted by newer prediction algorithms. When the slopes of berms were made shallower, the IL generally decreased for a berm alone, but generally increased in cases with a wall atop the berm.

Humans↗

Objective measures of breathy voice quality obtained using an auditory model.

While several acoustic measures have been proposed to quantify listener ratings of breathy voice quality, most have failed to give a consistent and high correlation with perceptual ratings of breathiness. One reason for these limitations is that most acoustic measures do not address the nonlinear processes that occur in the peripheral auditory system during the auditory perceptual process. It was hypothesized that modeling such nonlinear events during signal processing may provide objective parameters that better correspond to perceptual ratings of breathy voice quality. Ten listeners rated 27 voice stimuli using a five-point rating scale. Acoustic measures were determined from these stimuli and were selected based on their history of having a moderate to strong correlation to perceptual ratings of breathiness. The stimuli were also analyzed using an auditory model proposed by Moore, Glasberg, and Baer [J. Audio Eng. Soc. 45(4), 224-239 (1997)], and new measures were calculated from the output of this model. These measures included the partial loudness of the signal and the loudness of the aspiration noise. Measures obtained from the output of the auditory model were found to account for a high amount of variance in the perceptual ratings of breathiness.

Adult↗

Hearing protection: surpassing the limits to attenuation imposed by the bone-conduction pathways.

With louder and louder weapon systems being developed and military personnel being exposed to steady noise levels approaching and sometimes exceeding 150 dB, a growing interest in greater amounts of hearing protection is evident. When the need for communications is included in the equation, the situation is even more extreme. New initiatives are underway to design improved hearing protection, including active noise reduction (ANR) earplugs and perhaps even active cancellation of head-borne vibration. With that in mind it may be useful to explore the limits to attenuation, and whether they can be approached with existing technology. Data on the noise reduction achievable with high-attenuation foam earplugs, as a function of insertion depth, will be reported. Previous studies will be reviewed that provide indications of the bone-conduction (BC) limits to attenuation that, in terms of mean values, range from 40 to 60 dB across the frequencies from 125 Hz to 8 kHz. Additionally, new research on the effects of a flight helmet on the BC limits, as well as the potential attenuation from deeply inserted passive foam earplugs, worn with passive earmuffs, or with active-noise reduction (ANR) earmuffs, will be examined. The data demonstrate that gains in attenuation exceeding 10 dB above the head-not-covered limits can be achieved if the head is effectively shielded from acoustical stimulation.

Auditory Threshold↗

Effects of contrast between onsets of speech and other complex spectra.

Previous studies using speech and nonspeech analogs have shown that auditory mechanisms which serve to enhance spectral contrast contribute to perception of coarticulated speech for which spectral properties assimilate over time. In order to better understand the nature of contrastive auditory processes, a series of CV syllables varying acoustically in F2-onset frequency and perceptually from /ba/ to /da/ was identified following a variety of spectra including three-peak renditions of [e] and [o], one-peak simulations of only F2, and spectral complements of these spectra for which peaks are replaced with troughs. Results for three-versus one-peak (or trough) precursor spectra were practically indistinguishable, suggesting that effects were spectrally local and not dependent upon perception of precursors as speech. Effects of complementary (trough) spectra had complementary effects on perception of following stops; however, effects for spectral complements were particularly dependent upon the interval between precursor and CV onsets. Results from these studies cannot be explained by simple masking or adaptation of suppression. Instead, they provide evidence for the existence of processes that selectively enhance contrast between onset spectra of neighboring sounds, and these processes are relevant for perception of connected speech.

Adult↗

Phase effects in masking: within- versus across-channel processes.

The effects of bandwidth and component phase on masking were investigated using 200-ms narrowband (1-ERB(N)) and broadband (5-ERB(N)) cosine-phase (CP) and random-phase (RP) harmonic complex maskers, centered at 1 or 6 kHz. A continuous notched-noise was used to restrict off-frequency listening. The masker fundamental frequency (F0) was 25 Hz. In experiment 1, thresholds were measured for sinusoidal signals at 1 and 6 kHz, gated with the maskers. Thresholds were lower in the CP than in the RP masker, for both bandwidths, but the effect was markedly greater for the wider bandwidth. For the CP maskers, thresholds were markedly lower for the 5-ERB(N) than for the 1-ERB(N) bandwidth; for the RP maskers, there was a small effect in the opposite direction. Experiment 2 used 1- and 6-kHz CP maskers. The masker components in the ERB(N) around the signal frequency were presented to one ear, and the remaining components were presented contralaterally. Thresholds were much higher than when all components were presented to the same ear, and were higher than for the 1-ERB(N) masker alone, suggesting that the low thresholds for broadband monaural presentation do not depend on "high level" across-channel comparisons. Simultaneous masked thresholds could be predicted well using a model based on a simulated auditory filter, a level-dependent compressive nonlinearity, and a sliding temporal integrator; it was not necessary to assume the involvement of across-channel processes or of selective listening in the masker dips.

Adult↗

A phenomenological model for the responses of auditory-nerve fibers. II. Nonlinear tuning with a frequency glide.

A computational model was developed to simulate the responses of auditory-nerve (AN) fibers in cat. The model's signal path consisted of a time-varying bandpass filter; the bandwidth and gain of the signal path were controlled by a nonlinear feed-forward control path. This model produced realistic response features to several stimuli, including pure tones, two-tone combinations, wideband noise, and clicks. Instantaneous frequency glides in the reverse-correlation (revcor) function of the model's response to broadband noise were achieved by carefully restricting the locations of the poles and zeros of the bandpass filter. The pole locations were continuously varied as a function of time by the control signal to change the gain and bandwidth of the signal path, but the instantaneous frequency profile in the revcor function was independent of sound pressure level, consistent with physiological data. In addition, this model has other important properties, such as nonlinear compression, two-tone suppression, and reasonable Q10 values for tuning curves. The incorporation of both the level-independent frequency glide and the level-dependent compressive nonlinearity into a phenomenological model for the AN was the primary focus of this work. The ability of this model to process arbitrary sound inputs makes it a useful tool for studying peripheral auditory processing.

Auditory Pathways↗

A quasi-glottogram signal.

A novel, noninvasive experiment is proposed that reliably shows the strength of glottal oscillations. The quasi-glottogram (QGG) signal is generated from a microphone array that is trained to approximate the electroglottogram signal. The QGG may be useful to improve estimates of whether speech is voiced, to quantify partial voicing, and to reduce the phoneme effect when measuring the amplitude of speech signals. The technique is well adapted to the generation of text-to-speech systems, as it allows an estimate of the glottal flow during undisturbed, natural speech. For prosody studies, it can be used to provide an estimate of amplitude which is relatively unaffected by changes in phonemes, and is at least as reliable as standard estimators of amplitude.

Algorithms↗

Perceptual segregation of competing speech sounds: the role of spatial location.

Culling and Summerfield [J. Acoust Soc. Am. 92, 785-797 (1995)] showed that listeners could not use ongoing interaural time differences (ITDs) to achieve source segregation. The present experiments tested a free-field analog of their experiment. The stimuli consisted of narrow bands of noise, pairs of which represented the first and second formants of the whispered vowels "ar," "ee," "er," and "oo." A target noise-band pair (vowel) was presented at various angles on the listeners' left while a complementary distracter was presented on the listeners' right. Listeners correctly identified the target vowel in the free-field well above chance. Performance remained well above chance in headphone experiments that retained spatial cues but eliminated reverberations and head movements. The full range of cues that normally determine perceived spatial location provided sufficient information for segregation. Further experiments, which systematically evaluated the contribution of these cues in isolation and in combination, showed that some listeners, following training, exhibited the ability to segregate based on ongoing ITDs alone. Substantial individual differences were observed. The results show that listeners can use spatial cues to segregate simultaneous sound sources.

Adolescent↗

A measure of internal noise based on sample discrimination.

Internal noise is often inferred from the difference between observed performance and optimum performance in detection and discrimination tasks. It can be measured directly in some cases by observing the extent to which a change in external variability impacts performance. In the studies reported here, external variability was added to an intensity discrimination task by adding a Gaussian random variable with zero mean to the overall level presented in each interval of a two-interval forced-choice task. The standard deviation of the random variable was set to half the mean difference between the levels in the two intervals, resulting in d'(ideal) = 2. As the mean difference and the corresponding standard deviation of the random variable decreased in size, performance was increasingly limited by internal noise, permitting a reliable estimate of internal noise to be obtained. This can be viewed as a sample discrimination task, with one component per sample. In the first study, performance was measured using 2-kHz tones presented at an average level of 70 dB SPL, with mean differences between distributions ranging from 0.1 to 2.2 dB in steps of 0.3 dB. The distributions were either Gaussian in level or in power. Conditions with no external variability were used to obtain a psychometric function. In the second study, performance was measured using 2-kHz tones presented at average levels of 50 and 90 dB SPL, with mean differences ranging from 0.4 to 2.2 dB in steps of 0.6 dB. In both studies, the measure of internal noise was highly reliable and in good agreement with the intensity difference limen (DL) estimated from the psychometric function. Analyses suggest that this measure could be used to estimate the mean difference between the decision distributions as well as the amount of internal noise in cases where the mean difference between the distributions is unknown.

Adolescent↗

Nonlinear dynamics of phonations in excised larynx experiments.

Nonlinear dynamic methods including correlation dimension and Lyapunov exponents are applied to quantitatively analyze phonations in excised larynx experiments. Irregular phonations are typically characterized by aperiodic waveforms and broadband spectra. Finite correlation dimensions and positive Lyapunov exponents of irregular phonations demonstrate the existence of chaos in excised larynx phonations. Furthermore, the correlation dimension, maximal Lyapunov exponent, jitter, shimmer, and peak prominence ratio are used to statistically distinguish irregular phonations from normal phonations. The correlation dimension and maximal Lyapunov exponent indicate a significant difference between irregular and normal phonations; however, jitter, shimmer, and peak prominence ratio do not reveal such a significant difference and thus are unsuitable to differentiate between irregular phonations and normal phonations. These findings might potentially assist investigators in understanding rough phonations and developing clinically valuable methodologies for the diagnosis of voice disorders.

Animals↗

Speech segregation based on sound localization.

At a cocktail party, one can selectively attend to a single voice and filter out all the other acoustical interferences. How to simulate this perceptual ability remains a great challenge. This paper describes a novel, supervised learning approach to speech segregation, in which a target speech signal is separated from interfering sounds using spatial localization cues: interaural time differences (ITD) and interaural intensity differences (IID). Motivated by the auditory masking effect, the notion of an "ideal" time-frequency binary mask is suggested, which selects the target if it is stronger than the interference in a local time-frequency (T-F) unit. It is observed that within a narrow frequency band, modifications to the relative strength of the target source with respect to the interference trigger systematic changes for estimated ITD and IID. For a given spatial configuration, this interaction produces characteristic clustering in the binaural feature space. Consequently, pattern classification is performed in order to estimate ideal binary masks. A systematic evaluation in terms of signal-to-noise ratio as well as automatic speech recognition performance shows that the resulting system produces masks very close to ideal binary ones. A quantitative comparison shows that the model yields significant improvement in performance over an existing approach. Furthermore, under certain conditions the model produces large speech intelligibility improvements with normal listeners.

Adult↗

Threshold differences for interaural time delays carried by double vowels.

Experimental measurements were made of threshold interaural time differences (ITDs) for a "target" vowel presented simultaneously with a fixed-ITD "distracter" vowel. Three double-vowel pairs were used, comprising an "er" (/e/) together with either an "ai," "ar," or "oo" (respectively, /e/, /c/, and /u/). Threshold ITDs were found to be larger for the target vowel when it was part of a double-vowel pair than in control conditions in which it was presented alone. The effect size depended upon the choice of target vowel and distracter vowel, the level of the target relative to the distracter, and whether the two vowels had the same or different fundamental frequencies. The experiment was analyzed using a multichannel modification of Heller and Trahiotis' [J. Acoust. Soc. Am. 99, 3632-3637 (1996)] model, which used a weighted combination of the detectabilities of the ITD of the target and the distracter. It gave predictions consistent with the observed effects of level and with some of the effects of the choice of target vowel, but it could not describe the effect of the target-distracter differences in fundamental frequency. It was found that a single-channel version of the model, in which the chosen channel was allowed to depend upon fundamental frequency (which could be derived using a monaural autocorrelation model) did give a set of predictions in qualitative accord with the data.

Adult↗

Cutoff frequencies and cross fingerings in baroque, classical, and modern flutes.

Baroque, classical, and modern flutes have successively more and larger tone holes. This paper reports measurements of the standing waves in the bores of instruments representing these three classes. It presents the frequency dependence of propagation of standing waves in lattices of open tone holes and compares these measurements with the cutoff frequency: the frequency at which, in an idealized system, the standing waves propagate without loss in such a lattice. It also reports the dependence of the sound field in the bore of the instrument as a function of both frequency and position along the bore for both simple and "cross fingerings" (configurations in which one or more tone holes are closed below an open hole). These measurements show how "cross fingerings" produce a longer standing wave, a technique used to produce the nondiatonic notes on instruments with a small number of tone holes closed only by the unaided fingers. They also show why the changes from baroque to classical to modern gave the instruments a louder, brighter sound and a greater range.

Humans↗

Application of loudness models to sound processing for cochlear implants.

A new paradigm for processing sound signals for multiple-electrode cochlear implants is introduced, and results are presented from an initial psychophysical evaluation of its effect on the perceived loudness of complex sounds. A real-time processing scheme based on this paradigm, called SpeL, has been developed primarily to improve control of loudness for implant users. SpeL differs from previous schemes in several ways. Most importantly, it incorporates a published numerical model which predicts the loudness perceived by implant users for complex patterns of pulsatile electric stimulation as a function of the pulses' physical parameters. This model is controlled by the output of a corresponding model that estimates the loudness perceived by normally hearing listeners for complex sounds. The latter model produces an estimate of the specific loudness arising from an acoustic signal. In SpeL, the specific loudness function, which describes the contribution to total loudness of each of a number of frequency bands (or cochlear positions), is converted to a pattern of electric stimulation on an appropriate set of electrodes. By application of the loudness model for electric stimulation, this pattern is designed to produce a specific loudness function for the implant user which approximates that produced by the normal-hearing model for the same input signal. The results of loudness magnitude estimation experiments with five users of the SpeL scheme confirmed that the psychophysical functions relating overall loudness perceived to input sound level for five complex acoustic signals were, on average, very similar to those for normal hearing.

Aged↗