Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sound Spectrography”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

Audio-visual speech perception is special.

In face-to-face conversation speech is perceived by ear and eye. We studied the prerequisites of audio-visual speech perception by using perceptually ambiguous sine wave replicas of natural speech as auditory stimuli. When the subjects were not aware that the auditory stimuli were speech, they showed only negligible integration of auditory and visual stimuli. When the same subjects learned to perceive the same auditory stimuli as speech, they integrated the auditory and visual stimuli in a similar manner as natural speech. These results demonstrate the existence of a multisensory speech-specific mode of perception.

Association Learning↗

An interaction between prosody and statistics in the segmentation of fluent speech.

Sensitivity to prosodic cues might be used to constrain lexical search. Indeed, the prosodic organization of speech is such that words are invariably aligned with phrasal prosodic edges, providing a cue to segmentation. In this paper we devise an experimental paradigm that allows us to investigate the interaction between statistical and prosodic cues to extract words from a speech stream. We provide evidence that statistics over the syllables are computed independently of prosody. However, we also show that trisyllabic sequences with high transition probabilities that straddle two prosodic constituents appear not to be recognized. Taken together, our findings suggest that prosody acts as a filter, suppressing possible word-like sequences that span prosodic constituents.

Adult↗

Nonlinear analysis of wheezes using wavelet bicoherence.

Wheezes, as being abnormal breath sounds, are observed in patients with obstructive pulmonary diseases, such as asthma. The aim of this study was to capture and analyze the nonlinear characteristics of asthmatic wheezes, reflected in the quadrature phase coupling of their harmonics, as they evolve over time within the breathing cycle. To achieve this, the continuous wavelet transform (CWT) was combined with third-order statistics/spectra. Wheezes from patients with diagnosed asthma were drawn from a lung sound database and analyzed in the time-bi-frequency domain. The analysis results justified the efficient performance of this combinatory approach to reveal and quantify the evolution of the nonlinearities of wheezes with time.

Algorithms↗

Wavelet time-frequency analysis and least squares support vector machines for the identification of voice disorders.

This work describes a novel algorithm to identify laryngeal pathologies, by the digital analysis of the voice. It is based on Daubechies' discrete wavelet transform (DWT-db), linear prediction coefficients (LPC), and least squares support vector machines (LS-SVM). Wavelets with different support-sizes and three LS-SVM kernels are compared. Particularly, the proposed approach, implemented with modest computer requirements, leads to an adequate larynx pathology classifier to identify nodules in vocal folds. It presents over 90% of classification accuracy and has a low order of computational complexity in relation to the speech signal's length.

Adolescent↗

Long-term signal detection, segmentation and summarization using wavelets and fractal dimension: a bioacoustics application in gastrointestinal-motility monitoring.

The current paper describes a wavelet-based method for long-term processing and analysis of gastrointestinal sounds (GIS). Windowing techniques are used to select sequential blocks of the prolonged multi-channel recordings and proceed to various wavelet-domain processing stages. De-noising, significant-activity detection, automated segmentation and extraction of summary curves are applied in an integrated mode, allowing for enhanced content manipulation and analysis. The proposed analysis scheme combines flexible long-term graphical representation tools, while maintaining the ability of quick browsing via visualization and auralization of the detected short-term events. This work is part of a project aiming to implement non-invasive diagnosis over gastrointestinal-motility (GIM) physiology. However, the proposed techniques might be applied to any study of long-term bioacoustics time series.

Algorithms↗

Optimal selection of wavelet-packet-based features using genetic algorithm in pathological assessment of patients' speech signal with unilateral vocal fold paralysis.

Unilateral vocal fold paralysis (UVFP) is one of the most severe types of neurogenic laryngeal disorder in which the patients, due to their vocal cords malfunction, are confronted by some serious problems. As the effect of such pathologies would be significantly evident in the reduced quality and feature variation of dysphonic voices, this study is designed to scrutinize the piecewise variation of some specific types of these features, known as energy and entropy, all over the frequency range of pathological speech signals. In order to do so, the wavelet-packet coefficients, in five consecutive levels of decomposition, are used to extract the energy and entropy measures at different spectral sub-bands. As the decomposition procedure leads to a set of high-dimensional feature vectors, genetic algorithm is invoked to search for a group of optimal sub-band indexes for which the extracted features result in the highest recognition rate for pathological and normal subjects' classification. The results of our simulations, using support vector machine classifier, show that the highest recognition rate, for both optimized energy and entropy measures, is achieved at the fifth level of wavelet-packet decomposition. It is also found that entropy feature, with the highest recognition rate of 100% vs. 93.62% for energy, is more prominent in discriminating patients with UVFP from normal subjects. Therefore, entropy feature, in comparison with energy, demonstrates a more efficient description of such pathological voices and provides us a valuable tool for clinical diagnosis of unilateral laryngeal paralysis.

Adolescent↗

Vocal-tract filtering by lingual articulation in a parrot.

Human speech and bird vocalization are complex communicative behaviors with notable similarities in development and underlying mechanisms. However, there is an important difference between humans and birds in the way vocal complexity is generally produced. Human speech originates from independent modulatory actions of a sound source, e.g., the vibrating vocal folds, and an acoustic filter, formed by the resonances of the vocal tract (formants). Modulation in bird vocalization, in contrast, is thought to originate predominantly from the sound source, whereas the role of the resonance filter is only subsidiary in emphasizing the complex time-frequency patterns of the source (e.g., but see ). However, it has been suggested that, analogous to human speech production, tongue movements observed in parrot vocalizations modulate formant characteristics independently from the vocal source. As yet, direct evidence of such a causal relationship is lacking. In five Monk parakeets, Myiopsitta monachus, we replaced the vocal source, the syrinx, with a small speaker that generated a broad-band sound, and we measured the effects of tongue placement on the sound emitted from the beak. The results show that tongue movements cause significant frequency changes in two formants and cause amplitude changes in all four formants present between 0.5 and 10 kHz. We suggest that lingual articulation may thus in part explain the well-known ability of parrots to mimic human speech, and, even more intriguingly, may also underlie a speech-like formant system in natural parrot vocalizations.

Acoustics↗

Monkeys match the number of voices they hear to the number of faces they see.

Convergent evidence demonstrates that adult humans possess numerical representations that are independent of language [1, 2, 3, 4, 5 and 6]. Human infants and nonhuman animals can also make purely numerical discriminations, implicating both developmental and evolutionary bases for adult humans' language-independent representations of number [7 and 8]. Recent evidence suggests that the nonverbal representations of number held by human adults are not constrained by the sensory modality in which they were perceived [9]. Previous studies, however, have yielded conflicting results concerning whether the number representations held by nonhuman animals and human infants are tied to the modality in which they were established [10, 11, 12, 13, 14 and 15]. Here, we report that untrained monkeys preferentially looked at a dynamic video display depicting the number of conspecifics that matched the number of vocalizations they heard. These findings suggest that number representations held by monkeys, like those held by adult humans, are unfettered by stimulus modality.

Acoustic Stimulation↗

Functionally referential communication in a chimpanzee.

The evolutionary origins of the use of speech signals to refer to events or objects in the world have remained obscure. Although functionally referential calls have been described in some monkey species, studies with our closest living relatives, the great apes, have not generated comparable findings. These negative results have been taken to suggest that ape vocalizations are not the product of their otherwise sophisticated mentality and that ape gestural communication is more informative for theories of language evolution. We tested whether chimpanzee rough grunts, which are produced during feeding contexts, functioned as referential signals. Individuals produced acoustically distinct types of "rough grunts" when encountering different foods. In a naturalistic playback experiment, a focal subject was able to use the information conveyed by these calls produced by several group mates to guide his search for food, demonstrating that the different grunt types were meaningful to him. This study provides experimental evidence that our closest living relatives can produce and understand functionally referential calls as part of their natural communication. We suggest that these findings give support to the vocal rather than gestural theories of language evolution.

Acoustic Stimulation↗

Tuning to natural stimulus dynamics in primary auditory cortex.

The amplitude and pitch fluctuations of natural soundscapes often exhibit "1/f spectra", which means that large, abrupt changes in pitch or loudness occur proportionally less frequently in nature than gentle, gradual fluctuations. Furthermore, human listeners reportedly prefer 1/f distributed random melodies to melodies with faster (1/f0) or slower (1/f2) dynamics. One might therefore suspect that neurons in the central auditory system may be tuned to 1/f dynamics, particularly given that recent reports provide evidence for tuning to 1/f dynamics in primary visual cortex. To test whether neurons in primary auditory cortex (A1) are tuned to 1/f dynamics, we recorded responses to random tone complexes in which the fundamental frequency and the envelope were determined by statistically independent "1/f(gamma) random walks," with gamma set to values between 0.5 and 4. Many A1 neurons showed clear evidence of tuning and responded with higher firing rates to stimuli with gamma between 1 and 1.5. Response patterns elicited by 1/f(gamma) stimuli were more reproducible for values of gamma close to 1. These findings indicate that auditory cortex is indeed tuned to the 1/f dynamics commonly found in the statistical distributions of natural soundscapes.

Acoustic Stimulation↗

In vitro fracture resistance of fiber reinforced cusp-replacing composite restorations.

OBJECTIVES: To assess the fracture resistance and failure mode of fiber reinforced composite (FRC) cusp-replacing restorations in premolars. METHODS: Forty-five extracted sound upper premolars were randomly divided into three groups. Identical MOD cavities with simulated buccal cusp fracture and height reduction of the palatal cusp were prepared. In Group A two layers of resin impregnated woven continuous FRC (EverStick Net) were applied. In Group B one layer of unidirectional continuous FRC (EverStick) was used. In Group C no fibers were applied (control). Subsequently, all teeth were restored with resin composite (Clearfil Photo Posterior), subjected to thermocycling (6000 x 5-55 degrees C) and static load tests. Load until fracture was registered for each tooth. Simultaneously, fracture propagation was monitored using acoustic emission analysis (AE). Failure modes were visually assessed. RESULTS: Weibull analysis revealed a characteristic strength and Weibull modulus (m) at 2364.8 N for Group A (m=8.9), 2437.9 N for Group B (m=5.9) and 2160.3N for Group C (m=13.6). Fracture loads were not significantly different (ANOVA, p>0.05). Teeth with FRC showed less fractures below the cemento-enamel junction (CEJ) (38% and 23% for Groups A and B, respectively) than teeth without FRC (93%) (chi-square, p<0.05). The control group showed the least AE energy signals. SIGNIFICANCE: The results suggest that glass FRC does not increase fracture load of premolars with cusp-replacing restorations. However, FRC has a beneficial effect on the failure mode. Woven fibers give more consistent results than unidirectional fibers.

Analysis of Variance↗

Pitch is determined by naturally occurring periodic sounds.

The phenomenology of pitch has been difficult to rationalize and remains the subject of much debate. Here we test the hypothesis that audition generates pitch percepts by relating inherently ambiguous sound stimuli to their probable sources in the human auditory environment. A database of speech sounds, the principal source of periodic sound energy for human listeners, was compiled and the dominant periodicity of each speech sound determined. A set of synthetic test stimuli were used to assess whether the major pitch phenomena described in the literature could be explained by the probabilistic relationship between the stimuli and their probable sources (i.e., speech sounds). The phenomena tested included the perception of the missing fundamental, the pitch-shift of the residue, spectral dominance and the perception of pitch strength. In each case, the conditional probability distribution of speech sound periodicities accurately predicted the pitches normally heard in response to the test stimuli. We conclude from these findings that pitch entails an auditory process that relates inevitably ambiguous sound stimuli to their probable natural sources.

Acoustic Stimulation↗

MMN to natural Arabic CV syllables: 1-normative data.

Mismatch negativity response parameters; latency, amplitude, and duration to natural Arabic CV syllables differing in durational change (Baa-Waa) and in spectrotemporal change (Gaa-Daa) were obtained from normal hearing young adult Egyptians. The aim was to get normative data for MMN response parameters and to find any differences between both primary and non-primary auditory pathways in encoding and processing speech signals. Statistically significant differences between durational and spectrotemporal contrasts for latency and duration were found. This was attributed to acoustic differences and to physiological differences between primary and non-primary auditory pathways.

Acoustic Stimulation↗

Representation of the purr call in the guinea pig primary auditory cortex.

Guinea pigs produce the low-frequency purr or rumble call as an alerting signal. A digitised example of the call was presented to anaesthetised guinea pigs via a closed sound system while recording from the primary auditory cortex. The exemplar used in this study had 9 regular phrases each spaced with their centres about 80 ms apart. Low-frequency (1.1 kHz) units responded best to the call but within this population there were four separate groups: (1) cells that responded vigorously to many or all of the 9 phrases; (2) cells that gave an onset response; (3) cells that only responded to a click embedded in the call; (4) cells that did not respond. Particular response types were often grouped together. Thus when orthogonal electrode tracks were used most units gave a similar response. There was no correlation between the type of response and the cortical depth. A similar range of response types was also found in the thalamus and there was no evidence of a distinct response in the cortex that was due to intracortical processing. Cells in the cortex were able to represent the temporal structure of the purr with the same fidelity as cells in the thalamus.

Acoustic Stimulation↗

Responses to species-specific vocalizations in the auditory cortex of awake and anesthetized guinea pigs.

Species-specific vocalizations represent an important acoustical signal that must be decoded in the auditory system of the listener. We were interested in examining to what extent anesthesia may change the process of signal decoding in neurons of the auditory cortex in the guinea pig. With this aim, the multiple-unit activity, either spontaneous or acoustically evoked, was recorded in the auditory cortex of guinea pigs, at first in the awake state and then after the injection of anesthetics (33 mg/kg ketamine with 6.6 mg/kg xylazine). Acoustical stimuli, presented in free-field conditions, consisted of four typical guinea pig calls (purr, chutter, chirp and whistle), a time-reversed version of the whistle and a broad-band noise burst. The administration of anesthesia typically resulted in a decrease in the level of spontaneous activity and in changes in the strength of the neuronal response to acoustical stimuli. The effect of anesthesia was mostly, but not exclusively, suppressive. Diversity in the effects of anesthesia led in some recordings to an enhanced response to one call accompanied by a suppressed response to another call. The temporal pattern of the response to vocalizations was changed in some cases under anesthesia, which may indicate a change in the synaptic input of the recorded neurons. In summary, our results suggest that anesthesia must be considered as an important factor when investigating the processing of complex sounds such as species-specific vocalizations in the auditory cortex.

Acoustic Stimulation↗

Auditory nerve representation of naturally-produced vowels with variable acoustics.

This investigation compared the encoding of naturally-produced, whispered and normally-voiced vowels by auditory nerve fibers. Speech syllables containing the vowels /open o/ and /ae/ were produced by two female speakers and presented at three intensities to ketamine-anesthetized chinchillas. Six different representations of the spectral components in the vowels in the responses of the auditory nerve fibers were evaluated. For both normal and whispered vowels over a 30 dB range, the formant peaks in the vowel were best displayed using rate-place representations. The spectral detail in the vowel was revealed by average localized synchronized rates (ALSR) and autocorrelations of individual peristimulus time histograms. The average localized interval rates (ALIR), autocorrelations of ensemble responses, and autocorrelations of individual spike trains demonstrated poor representations of vowel spectra, although the frequency components of normally-voiced vowels had better representations than those of whispered vowels. These analyses suggest that rate-based and synchronization-based measures yields two very different pieces of information, but only a normalized rate-based measure consistently identified the formants of both the whispered and normally-voiced vowels.

Acoustic Stimulation↗

The representation of noise vocoded speech in the auditory nerve of the chinchilla: physiological correlates of the perception of spectrally reduced speech.

This study investigated the neural representation of naturally produced and noise vocoded speech signals in the auditory nerve of the chinchilla. The syllables [see text] produced by male speakers were used to synthesize noise vocoded speech stimuli containing one, two, three and four bands of envelope modulated noise. The ensemble response of the auditory nerve, computed by pooling the PST histograms across many auditory nerve fibers, revealed temporal patterns in the responses to the natural tokens that uniquely identified the stop consonants. The responses to the 3- and 4-band noise vocoded tokens contained temporal patterns that were nearly identical to those observed for the natural tokens, while the responses to the 1- and 2-band tokens were significantly different (p<0.0001). The ALSR, ALIR and autocorrelation of the pooled PST histograms represented the detail of the frequency spectrum for a naturally produced vowel, while the driven rate was unreliable. Each of these spectral analyses failed to reveal significant information about the noise vocoded vowels. These results suggest that temporal patterns in the responses of the auditory nerve can provide the cues necessary for the recognition of noise vocoded stop consonants.

Acoustic Stimulation↗

The temporal representation of the delay of iterated rippled noise with positive or negative gain by chopper units in the cochlear nucleus.

The role of chopper units in representing the pitch of complex sounds is unresolved. Traditionally chopper units have been regarded as primarily responding to the stimulus envelope of complex stimuli. This has been supported by the response of chopper units to iterated rippled noise (IRN) as they can provide a robust representation of the delay of IRN with positive gain (+) in their first-order interspike intervals and for some chopper units this representation is relatively level independent. The envelope modulation of IRN(+), and pitch, is at the reciprocal of the delay, the pitch of IRN with negative gain (IRN(-)) is often at twice the delay. This distinction between IRN(+) and IRN(-) can be used to help determine whether a unit is simply responding to modulation or to stimulus fine structure. Chopper units with relatively high best frequencies (BF) are unable to represent the distinction between IRN(+) and IRN(-). However, in this study it is shown that at least some chopper units, with low BFs (<1.25 kHz), can represent the pitch of the IRN(-) as perceived perceptually.

Animals↗