Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sound Spectrography”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,459 records · Page 81Linked to original sources

Instantaneous frequency and short term fourier transforms: application to piano sounds.

For more than 30 years and until nowadays, development of a system reproducing the functioning of human hearing has remained an aim difficult to reach. Recent methods for identification of the fundamental frequency of musical sounds obtain good results using information about the temporal evolution of the amplitude and frequency of individual sound partials. Piano sounds and polyphonic sounds, which can have several partials that are closely spaced in the frequency domain, have not been extensively tested by these procedures. In this paper, the Instantaneous Frequency (IF), as defined by the Hilbert transform, is used to obtain the frequency variations of piano sounds partials. The result implies that, for these sounds, the IF may contain modulations resulting in the separation of an apparent single sinusoid signal into two or more sinusoidal components at various times in the analysis process, which makes it impossible to use the temporal evolution of the frequency of partials for the procedure of note identification. The separation phenomenon also appears when the short term Fourier transform is used and can induce the detection of short-lived parasitic spectral peaks that must be taken into account by any note identification procedure based on the use of spectral information.

Fourier Analysis↗

Measuring vocal quality with speech synthesis.

Much previous research has demonstrated that listeners do not agree well when using traditional rating scales to measure pathological voice quality. Although these findings may indicate that listeners are inherently unable to agree in their perception of such complex auditory stimuli, another explanation implicates the particular measurement method-rating scale judgments-as the culprit. An alternative method of assessing quality-listener-mediated analysis-synthesis-was devised to assess this possibility. In this new approach, listeners explicitly compare synthetic and natural voice samples, and adjust speech synthesizer parameters to create auditory matches to voice stimuli. This method is designed to replace unstable internal standards for qualities like breathiness and roughness with externally presented stimuli, to overcome major hypothetical sources of disagreement in rating scale judgments. In a preliminary test of the reliability of this method, listeners were asked to adjust the signal-to-noise ratio for 12 synthetic pathological voices so that the resulting stimuli matched the natural target voices as well as possible For comparison to the synthesis judgments, listeners also judged the noisiness of the natural stimuli in a separate task using a traditional visual-analog rating scale. For 9 of the 12 voices, agreement among listeners was significantly (and substantially) greater for the synthesis task than for the rating scale task. Response variances for the two tasks did not differ for the remaining three voices. However, a second experiment showed that the synthesis settings that listeners selected for these three voices were within a difference limen, and therefore observed differences were perceptually insignificant. These results indicate that listeners can in fact agree in their perceptual assessments of voice quality, and that analysis-synthesis can measure perception reliably.

Adult↗

Perceptual fusion and fragmentation of complex tones made inharmonic by applying different degrees of frequency shift and spectral stretch.

Global pitch depends on harmonic relations between components, but the perceptual coherence of a complex tone cannot be explained in the same way. Instead, it has been proposed that the auditory system responds to a common pattern of equal spacing between components, but is only sensitive to deviations from this pattern over a limited range [Roberts and Brunstrom, J. Acoust. Soc. Am. 104, 2326-2338 (1998)]. This hypothesis predicts that spectral fusion will be largely unaffected either by frequency shifting a harmonic stimulus (because equal spacing is preserved), or by small degrees of spectral stretch (because significant deviations from equal spacing only cumulate over large spectral distances). Complex tones were either shifted by 0%-50% of F0 (200 Hz+/-10%) or stretched by 0%-12% of F0 (100 Hz+/-10%). Subjects heard a complex followed by a pure tone in a continuous loop. One of the components 2-11 was mistuned by +/- 4%, and subjects adjusted the pure tone to match its pitch. Broadly consistent with our hypothesis, frequency shifts had relatively little effect on hit rates and only large degrees of stretch reduced them substantially. The implications for simultaneous grouping are explored with reference to an autocorrelation model of auditory processing.

Adult↗

Quantitative assessment of vocal development in the zebra finch using self-organizing neural networks.

To understand the mechanisms of song learning by songbirds it is necessary to have in hand tools for extracting, describing, and quantifying features of the developing vocalizations. The extremely large number of vocalizations produced by juvenile zebra finches and the variability in these vocalizations during the sensorimotor learning period preclude manual scoring methods. Here we describe an approach for classification of vocalizations produced during sensorimotor learning based on self-organizing neural networks. This approach allowed us to construct probability distributions of spectrotemporal features recorded on each day. By training the network with samples obtained across the course of vocal development in individual birds, we observed developmental trajectories of these features. The emergence of stereotypy in sequences of song elements was captured by computing the entropy in the matrices of first- and second-order transition probabilities. Self-organizing maps may assist in classifying large libraries of zebra finch vocalizations and shedding light on mechanisms of vocal development.

Aging↗

Surrogate analysis for detecting nonlinear dynamics in normal vowels.

Normal vowels are known to have irregularities in the pitch-to-pitch variation which is quite important for speech signals to be perceived as natural human sound. Such pitch-to-pitch variation of vowels is studied in the light of nonlinear dynamics. For the analysis, five normal vowels recorded from three male and two female subjects are exploited, where the vowel signals are shown to have normal levels of the pitch-to-pitch variation. First, by the false nearest-neighbor analysis, nonlinear dynamics of the vowels are shown to be well analyzed by using a relatively low-dimensional reconstructing dimension of 4 < or = d < or = 7. Then, we further studied nonlinear dynamics of the vowels by spike-and-wave surrogate analysis. The results imply that there exists nonlinear dynamical correlation between one pitch-waveform pattern to another in the vowel signals. On the basis of the analysis results, applicability of the nonlinear prediction technique to vowel synthesis is discussed.

Humans↗

A two-microphone dual delay-line approach for extraction of a speech sound in the presence of multiple interferers.

This paper describes algorithms for signal extraction for use as a front-end of telecommunication devices, speech recognition systems, as well as hearing aids that operate in noisy environments. The development was based on some independent, hypothesized theories of the computational mechanics of biological systems in which directional hearing is enabled mainly by binaural processing of interaural directional cues. Our system uses two microphones as input devices and a signal processing method based on the two input channels. The signal processing procedure comprises two major stages: (i) source localization, and (ii) cancellation of noise sources based on knowledge of the locations of all sound sources. The source localization, detailed in our previous paper [Liu et al., J. Acoust. Soc. Am. 108, 1888 (2000)], was based on a well-recognized biological architecture comprising a dual delay-line and a coincidence detection mechanism. This paper focuses on description of the noise cancellation stage. We designed a simple subtraction method which, when strategically employed over the dual delay-line structure in the broadband manner, can effectively cancel multiple interfering sound sources and consequently enhance the desired signal. We obtained an 8-10 dB enhancement for the desired speech in the situations of four talkers in the anechoic acoustic test (or 7-10 dB enhancement in the situations of six talkers in the computer simulation) when all the sounds were equally intense and temporally aligned.

Algorithms↗

An overlapping-feature-based phonological model incorporating linguistic constraints: applications to speech recognition.

Modeling phonological units of speech is a critical issue in speech recognition. In this paper, our recent development of an overlapping-feature-based phonological model that represents long-span contextual dependency in speech acoustics is reported. In this model, high-level linguistic constraints are incorporated in automatic construction of the patterns of feature-overlapping and of the hidden Markov model (HMM) states induced by such patterns. The main linguistic information explored includes word and phrase boundaries, morpheme, syllable, syllable constituent categories, and word stress. A consistent computational framework developed for the construction of the feature-based model and the major components of the model are described. Experimental results on the use of the overlapping-feature model in an HMM-based system for speech recognition show improvements over the conventional triphone-based phonological model.

Humans↗

Detecting stop consonants in continuous speech.

The problem of implementing a detector for stop consonants in continuously spoken speech is considered. The problem is posed as one of finding an optimal filter (linear or nonlinear) that operates on a particular appropriately chosen representation, and ideally outputs a 1 when a stop occurs and 0 otherwise. The performance of several variants of a canonical stop detector is discussed and its implications for human and machine speech recognition is considered.

Humans↗

Effects of prosodic factors on spectral dynamics. II. Synthesis.

In Paper I [J. Wouters and M. Macon, J. Acoust. Soc. Am. 111, 417-427 (2002)], the effects of prosodic factors on the spectral rate of change of phoneme transitions were analyzed for a balanced speech corpus. The results showed that the spectral rate of change, defined as the root-mean-square of the first three formant slopes, increased with linguistic prominence, i.e., in stressed syllables, in accented words, in sentence-medial words, and in clearly articulated speech. In the present paper, an initial approach is described to integrate the results of Paper I in a concatenative synthesis framework. The target spectral rate of change of acoustic units is predicted based on the prosodic structure of utterances to be synthesized. Then, the spectral shape of the acoustic units is modified according to the predicted spectral rate of change. Experiments show that the proposed approach provides control over the degree of articulation of acoustic units, and improves the naturalness and intelligibility of concatenated speech in comparison to standard concatenation methods.

Humans↗

Frequency specificity of chirp-evoked auditory brainstem responses.

This study examines the usefulness of the upward chirp stimulus developed by Dau et al. [J. Acoust. Soc. Am. 107, 1530-1540 (2000)] for retrieving frequency-specific information. The chirp was designed to produce simultaneous displacement maxima along the cochlear partition by compensating for frequency-dependent traveling-time differences. In the first experiment, auditory brainstem responses (ABR) elicited by the click and the broadband chirp were obtained in the presence of high-pass masking noise, with cutoff frequencies of 0.5, 1, 2, 4, and 8 kHz. Results revealed a larger wave-V amplitude for chirp than for click stimulation in all masking conditions. Wave-V amplitude for the chirp increased continuously with increasing high-pass cutoff frequency while it remains nearly constant for the click for cutoff frequencies greater than 1 kHz. The same two stimuli were tested in the presence of a notched-noise masker with one-octave wide spectral notches corresponding to the cutoff frequencies used in the first experiment. The recordings were compared with derived responses, calculated offline, from the high-pass masking conditions. No significant difference in response amplitude between click and chirp stimulation was found for the notched-noise responses as well as for the derived responses. In the second experiment, responses were obtained using narrow-band stimuli. A low-frequency chirp and a 250-Hz tone pulse with comparable duration and magnitude spectrum were used as stimuli. The narrow-band chirp elicited a larger response amplitude than the tone pulse at low and medium stimulation levels. Overall, the results of the present study further demonstrate the importance of considering peripheral processing for the formation of ABR. The chirp might be of particular interest for assessing low-frequency information.

Acoustic Stimulation↗

Acoustic features of male baboon loud calls: influences of context, age, and individuality.

The acoustic structure of loud calls ("wahoos") recorded from free-ranging male baboons (Papio cynocephalus ursinus) in the Moremi Game Reserve, Botswana, was examined for differences between and within contexts, using calls given in response to predators (alarm wahoos), during male contests (contest wahoos), and when a male had become separated from the group (contact wahoos). Calls were recorded from adolescent, subadult, and adult males. In addition, male alarm calls were compared with those recorded from females. Despite their superficial acoustic similarity, the analysis revealed a number of significant differences between alarm, contest, and contact wahoos. Contest wahoos are given at a much higher rate, exhibit lower frequency characteristics, have a longer "hoo" duration, and a relatively louder "hoo" portion than alarm wahoos. Contact wahoos are acoustically similar to contest wahoos, but are given at a much lower rate. Both alarm and contest wahoos also exhibit significant differences among individuals. Some of the acoustic features that vary in relation to age and sex presumably reflect differences in body size, whereas others are possibly related to male stamina and endurance. The finding that calls serving markedly different functions constitute variants of the same general call type suggests that the vocal production in nonhuman primates is evolutionarily constrained.

Age Factors↗

Intensity-importance functions for bandlimited monosyllabic words.

A study was carried out to determine the relative importance to speech intelligibility of different intensities within the speech dynamic range. The functions that were derived are analogous to previous descriptions of the relative importance of different frequencies and are referred to here as intensity-importance functions (IIFs). They were obtained as follows. Sharply filtered bands of speech (NU6 monosyllabic words) were mixed with filtered noise and presented alone or in pairs at 19 signal-to-noise ratios (-25 to 41 dB). When paired bands were tested, the level and signal-to-noise ratio (SNR) of one band were held constant while the level and SNR of the other band were varied. The listeners were 100 normal hearers, organized into five 20-person groups. Each group provided speech recognition data for one of five frequency regions (141-562, 562-1122, 1122-1778, 1778-2818, and 2818-8913 Hz). Comparisons of the results for each group indicated that IIFs vary with frequency and SNR. Current methods for predicting intelligibility from physical measurements of speech audibility would need to be revised in order to take such findings into consideration.

Adult↗

Maximum speed of pitch change and how it may relate to speech.

How fast speakers can change pitch voluntarily is potentially an important articulatory constraint for speech production. Previous attempts at assessing the maximum speed of pitch change have helped improve understanding of certain aspects of pitch production in speech. However, since only "response time"--time needed to complete the middle 75% of a pitch shift--was measured in previous studies, direct comparisons with speech data have been difficult. In the present study, a new experimental paradigm was adopted in which subjects produced rapid successions of pitch shifts by imitating synthesized model pitch undulation patterns. This permitted the measurement of the duration of entire pitch shifts. Native speakers of English and Mandarin participated as subjects. The speed of pitch change was measured both in terms of response time and excursion time-time needed to complete the entire pitch shift. Results show that excursion time is nearly twice as long as response time. This suggests that physiological limitation on the speed of pitch movement is greater than has been recognized. Also, it is found that the maximum speed of pitch change varies quite linearly with excursion size, and that it is different for pitch rises and falls. Comparisons of present data with data on speed of pitch change from studies of real speech found them to be largely comparable. This suggests that the maximum speed of pitch change is often approached in speech, and that the role of physiological constraints in determining the shape and alignment of F0 contours in speech is probably greater than has been appreciated.

Adolescent↗

Learning to perceive pitch differences.

This paper reports two experiments concerning the stimulus specificity of pitch discrimination learning. In experiment 1, listeners were initially trained, during ten sessions (about 11,000 trials), to discriminate a monaural pure tone of 3000 Hz from ipsilateral pure tones with slightly different frequencies. The resulting perceptual learning (improvement in discrimination thresholds) appeared to be frequency-specific since, in subsequent sessions, new learning was observed when the 3000-Hz standard tone was replaced by a standard tone of 1200 Hz, or 6500 Hz. By contrast, a subsequent presentation of the initial tones to the contralateral ear showed that the initial learning was not, or was only weakly, ear-specific. In experiment 2, training in pitch discrimination was initially provided using complex tones that consisted of harmonics 3-7 of a missing fundamental (near 100 Hz for some listeners, 500 Hz for others). Subsequently, the standard complex was replaced by a standard pure tone with a frequency which could be either equal to the standard complex's missing fundamental or remote from it. In the former case, the two standard stimuli were matched in pitch. However, this perceptual relationship did not appear to favor the transfer of learning. Therefore, the results indicated that pitch discrimination learning is, at least to some extent, timbre-specific, and cannot be viewed as a reduction of an internal noise which would affect directly the output of a neural device extracting pitch from both pure tones and complex tones including low-rank harmonics.

Adult↗

Interaction between adenosine triphosphate and mechanically induced modulation of electrically evoked otoacoustic emissions.

It was shown previously that electrically evoked otoacoustic emissions (EEOAEs) can be amplitude modulated by low-frequency bias tones and enhanced by application of adenosine triphosphate (ATP) to scala media. These effects were attributed, respectively, to the mechano-electrical transduction (MET) channels and ATP-gated ion channels on outer hair cell (OHC) stereocilia, two conductance pathways that appear to be functionally independent and additive in their effects on ionic current through the OHC. In the experiments described here, the separate influences of ATP and MET channel bias on EEOAEs did not combine linearly. Modulated EEOAEs increased in amplitude, but lost modulation at the phase and frequency of the bias tone (except at very high sound levels) after application of ATP to scala media, even though spectral components at the modulation sideband frequencies were still present. Some sidebands underwent phase shifts after ATP. In EEOAEs modulated by tones at lower sound levels, substitution of the original phase values restored modulation to the waveform, which then resembled a linear summation of the separate effects of ATP and low-frequency bias. While the physiological meaning of this procedure is not clear, the result raises the possibility that a secondary effect of ATP on one or more nonlinear stages in the transduction process, which may have caused the phase shifts, obscured linear summation at lower sound levels. In addition, "acoustic enhancement" of the EEOAE may have introduced nonlinear interaction at higher levels of the bias tones.

Adenosine Triphosphate↗

Relationship between low-frequency aircraft noise and annoyance due to rattle and vibration.

A near-replication of a study of the annoyance of rattle and vibration attributable to aircraft noise [Fidell et al., J. Acoust. Soc. Am. 106, 1408-1415 (1999)] was conducted in the vicinity of Minneapolis-St. Paul International Airport (MSP). The findings of the current study were similar to those reported earlier with respect to the types of objects cited as sources of rattle in homes, frequencies of notice of rattle, and the prevalence of annoyance due to aircraft noise-induced rattle. A reliably lower prevalence rate of annoyance (but not of complaints) with rattle and vibration was noted among respondents living in homes that had been treated to achieve a 5-dB improvement in A-weighted noise reduction than among respondents living in untreated homes. This difference is not due to any substantive increase in low-frequency noise reduction of acoustically treated homes, but may be associated with installation of nonrattling windows. Common interpretations of the prevalence of a consequential degree of annoyance attributable to low-frequency aircraft noise may be developed from the combined results of the present and prior studies.

Aircraft↗

Similarity, uncertainty, and masking in the identification of nonspeech auditory patterns.

This study examined whether increasing the similarity between informational maskers and signals would increase the amount of masking obtained in a nonspeech pattern identification task. The signals were contiguous sequences of pure-tone bursts arranged in six narrow-band spectro-temporal patterns. The informational maskers were sequences of multitone bursts played synchronously with the signal tones. The listener's task was to identify the patterns in a 1-interval 6-alternative forced-choice procedure. Three types of multitone maskers were generated according to different randomization rules. For the least signal-like informational masker, the components in each multitone burst were chosen at random within the frequency range of 200-6500 Hz, excluding a "protected region" around the signal frequencies. For the intermediate masker, the frequency components in the first burst were chosen quasirandomly, but the components in successive bursts were constrained to fall in narrow frequency bands around the frequencies of the components in the initial burst. Within the narrow bands the frequencies were randomized. This masker was considered to be more similar to the signal patterns because it consisted of a set of narrow-band sequences any one of which might be mistaken for a signal pattern. The most signal-like masker was similar to the intermediate masker in that it consisted of a set of synchronously played narrow-band sequences, but the variation in frequency within each sequence was sinusoidal, completing roughly one period in a sequence. This masker consisted of discernible patterns but not patterns that were part of the set of signals. In addition, masking produced by Gaussian noise bursts--thought to produce primarily peripherally based "energetic masking"--was measured and compared to the informational masking results. For the three informational maskers, more masking was produced by the maskers comprised of narrow-band sequences than for the masker in which the frequencies were not constrained to narrow bands. Also, the slopes of the performance-level functions for the three informational maskers were much shallower than for the Gaussian noise masker or for no masker. The findings provided qualified support for the hypothesis that increasing the similarity between signals and maskers, or parts of the maskers, causes greater informational masking. However, it is also possible that the greater masking was a consequence of increasing the number of perceptual "streams" that had to be evaluated by the listener.

Adult↗