Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sound Spectrography”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,297 records · Page 72Linked to original sources

Underdetermined blind source separation of temporomandibular joint sounds.

The underdetermined blind source separation problem using a filtering approach is addressed. An extension of the FastICA algorithm is devised which exploits the disparity in the kurtoses of the underlying sources to estimate the mixing matrix and thereafter achieves source recovery by employing the ll-norm algorithm. Besides, we demonstrate how promising FastICA can be to extract the sources. Furthermore, we illustrate how this scenario is particularly appropriate for the separation of temporomandibular joint (TMJ) sounds.

Algorithms↗

Information extraction from sound for medical telemonitoring.

Today, the growth of the aging population in Europe needs an increasing number of health care professionals and facilities for aged persons. Medical telemonitoring at home (and, more generally, telemedicine) improves the patient's comfort and reduces hospitalization costs. Using sound surveillance as an alternative solution to video telemonitoring, this paper deals with the detection and classification of alarming sounds in a noisy environment. The proposed sound analysis system can detect distress or everyday sounds everywhere in the monitored apartment, and is connected to classical medical telemonitoring sensors through a data fusion process. The sound analysis system is divided in two stages: sound detection and classification. The first analysis stage (sound detection) must extract significant sounds from a continuous signal flow. A new detection algorithm based on discrete wavelet transform is proposed in this paper, which leads to accurate results when applied to nonstationary signals (such as impulsive sounds). The algorithm presented in this paper was evaluated in a noisy environment and is favorably compared to the state of the art algorithms in the field. The second stage of the system is sound classification, which uses a statistical approach to identify unknown sounds. A statistical study was done to find out the most discriminant acoustical parameters in the input of the classification module. New wavelet based parameters, better adapted to noise, are proposed in this paper. The telemonitoring system validation is presented through various real and simulated test sets. The global sound based system leads to a 3% missed alarm rate and could be fused with other medical sensors to improve performance.

Activities of Daily Living↗

Identification of the defective transmission devices using the wavelet transform.

In this paper, a system is described that uses the wavelet transform to automatically identify the particular failure mode of a known defective transmission device. The problem of identifying a particular failure mode within a costly failed assembly is of benefit in practical applications. In this system, external acoustic sensors, instead of intrusive vibrometers, are used to record the acoustic data of the operating transmission device. A skilled factory worker, who is unfamiliar with statistical classification, helps to determine the feature vector of the particular failure mode in the feature extraction process. In the automatic identification part, an improved learning vector quantization (LVQ) method with normalizing the inputting feature vectors is proposed to compensate for variations in practical data. Some acoustic data, which are collected from the manufacturing site, are utilized to test the effectiveness of the described identification system. The experimental results show that this system can identify the particular failure mode of a defective transmission device and find out the causes of failure successfully.

Acoustics↗

Activity recognition of assembly tasks using body-worn microphones and accelerometers.

In order to provide relevant information to mobile users, such as workers engaging in the manual tasks of maintenance and assembly, a wearable computer requires information about the user's specific activities. This work focuses on the recognition of activities that are characterized by a hand motion and an accompanying sound. Suitable activities can be found in assembly and maintenance work. Here, we provide an initial exploration into the problem domain of continuous activity recognition using on-body sensing. We use a mock "wood workshop" assembly task to ground our investigation. We describe a method for the continuous recognition of activities (sawing, hammering, filing, drilling, grinding, sanding, opening a drawer, tightening a vise, and turning a screwdriver) using microphones and three-axis accelerometers mounted at two positions on the user's arms. Potentially "interesting" activities are segmented from continuous streams of data using an analysis of the sound intensity detected at the two different locations. Activity classification is then performed on these detected segments using linear discriminant analysis (LDA) on the sound channel and hidden Markov models (HMMs) on the acceleration data. Four different methods at classifier fusion are compared for improving these classifications. Using user-dependent training, we obtain continuous average recall and precision rates (for positive activities) of 78 percent and 74 percent, respectively. Using user-independent training (leave-one-out across five users), we obtain recall rates of 66 percent and precision rates of 63 percent. In isolation, these activities were recognized with accuracies of 98 percent, 87 percent, and 95 percent for the user-dependent, user-independent, and user-adapted cases, respectively.

Acceleration↗

Enhanced sound localization.

A new approach to sound localization, known as enhanced sound localization, is introduced, offering two major benefits over state-of-the-art algorithms. First, higher localization accuracy can be achieved compared to existing methods. Second, an estimate of the source orientation is obtained jointly, as a consequence of the proposed sound localization technique. The orientation estimates and improved localizations are a result of explicitly modeling the various factors that affect a microphone's level of access to different spatial positions and orientations in an acoustic environment. Three primary factors are accounted for, namely the source directivity, microphone directivity, and source-microphone distances. Using this model of the acoustic environment, several different enhanced sound localization algorithms are derived. Experiments are carried out in a real environment whose reverberation time is 0.1 seconds, with the average microphone SNR ranging between 10-20 dB. Using a 24-element microphone array, a weighted version of the SRP-PHAT algorithm is found to give an average localization error of 13.7 cm with 3.7% anomalies, compared to 14.7 cm and 7.8% anomalies with the standard SRP-PHAT technique.

Algorithms↗

Phase-based dual-microphone robust speech enhancement.

A dual-microphone speech-signal enhancement algorithm, utilizing phase-error based filters that depend only on the phase of the signals, is proposed. This algorithm involves obtaining time-varying, or alternatively, time-frequency (TF), phase-error filters based on prior knowledge regarding the time difference of arrival (TDOA) of the speech source of interest and the phases of the signals recorded by the microphones. It is shown that by masking the TF representation of the speech signals, the noise components are distorted beyond recognition while the speech source of interest maintains its perceptual quality. This is supported by digit recognition experiments which show a substantial recognition accuracy rate improvement over prior multimicrophone speech enhancement algorithms. For example, for a case with two speakers with a 0.1 s reverberation time, the phase-error based technique results in a 28.9% recognition rate gain over the single channel noisy signal, a gain of 22.0% over superdirective beamforming, and a gain of 8.5% over postfiltering.

Algorithms↗

A brain-like neural network for periodicity analysis.

This paper introduces a brain-like neural model for sound processing. The periodicity analyzing network (PAN) is a bio-inspired neural network of spiking neurons. The PAN consists of complex models of neurons, which can be used for understanding the dynamics of individual neurons and neuronal networks. On a technical level, the PAN is able to compute the ratio of modulation and carrier frequency of harmonic sound signals. The PAN model may, therefore, be used in audio signal processing applications, such as sound source separation, periodicity analysis, and the cocktail party problem.

Algorithms↗

A stable learning algorithm for block-diagonal recurrent neural networks: application to the analysis of lung sounds.

A novel learning algorithm, the Recurrent Neural Network Constrained Optimization Method (RENNCOM) is suggested in this paper, for training block-diagonal recurrent neural networks. The training task is formulated as a constrained optimization problem, whose objective is twofold: (1) minimization of an error measure, leading to successful approximation of the input/output mapping and (2) optimization of an additional functional, the payoff function, which aims at ensuring network stability throughout the learning process. Having assured the network and training stability conditions, the payoff function is switched to an alternative form with the scope to accelerate learning. Simulation results on a benchmark identification problem demonstrate that, compared to other learning schemes with stabilizing attributes, the RENNCOM algorithm has enhanced qualities, including, improved speed of convergence, accuracy and robustness. The proposed algorithm is also applied to the problem of the analysis of lung sounds. Particularly, a filter based on block-diagonal recurrent neural networks is developed, trained with the RENNCOM method. Extensive experimental results are given and performance comparisons with a series of other models are conducted, underlining the effectiveness of the proposed filter.

Algorithms↗

Robust speaker's location detection in a vehicle environment using GMM models.

Abstract-Human-computer interaction (HCI) using speech communication is becoming increasingly important, especially in driving where safety is the primary concern. Knowing the speaker's location (i.e., speaker localization) not only improves the enhancement results of a corrupted signal, but also provides assistance to speaker identification. Since conventional speech localization algorithms suffer from the uncertainties of environmental complexity and noise, as well as from the microphone mismatch problem, they are frequently not robust in practice. Without a high reliability, the acceptance of speech-based HCI would never be realized. This work presents a novel speaker's location detection method and demonstrates high accuracy within a vehicle cabinet using a single linear microphone array. The proposed approach utilize Gaussian mixture models (GMM) to model the distributions of the phase differences among the microphones caused by the complex characteristic of room acoustic and microphone mismatch. The model can be applied both in near-field and far-field situations in a noisy environment. The individual Gaussian component of a GMM represents some general location-dependent but content and speaker-independent phase difference distributions. Moreover, the scheme performs well not only in nonline-of-sight cases, but also when the speakers are aligned toward the microphone array but at difference distances from it. This strong performance can be achieved by exploiting the fact that the phase difference distributions at different locations are distinguishable in the environment of a car. The experimental results also show that the proposed method outperforms the conventional multiple signal classification method (MUSIC) technique at various SNRs.

Acoustics↗

A probabilistic model for binaural sound localization.

This paper proposes a biologically inspired and technically implemented sound localization system to robustly estimate the position of a sound source in the frontal azimuthal half-plane. For localization, binaural cues are extracted using cochleagrams generated by a cochlear model that serve as input to the system. The basic idea of the model is to separately measure interaural time differences and interaural level differences for a number of frequencies and process these measurements as a whole. This leads to two-dimensional frequency versus time-delay representations of binaural cues, so-called activity maps. A probabilistic evaluation is presented to estimate the position of a sound source over time based on these activity maps. Learned reference maps for different azimuthal positions are integrated into the computation to gain time-dependent discrete conditional probabilities. At every timestep these probabilities are combined over frequencies and binaural cues to estimate the sound source position. In addition, they are propagated over time to improve position estimation. This leads to a system that is able to localize audible signals, for example human speech signals, even in reverberating environments.

Artificial Intelligence↗

Is infant-directed speech prosody a result of the vocal expression of emotion?

Many studies have found that infant-directed (ID) speech has higher pitch, has more exaggerated pitch contours, has a larger pitch range, has a slower tempo, and is more rhythmic than typical adult-directed (AD) speech. We show that the ID speech style reflects free vocal expression of emotion to infants, in comparison with more inhibited expression of emotion in typical AD speech. When AD speech does express emotion, the same acoustic features are used as in ID speech. We recorded ID and AD samples of speech expressing love-comfort, fear, and surprise. The emotions were equally discriminable in the ID and AD samples. Acoustic analyses showed few differences between the ID and AD samples, but robust differences across the emotions. We conclude that ID prosody itself is not special. What is special is the widespread expression of emotion to infants in comparison with the more inhibited expression of emotion in typical adult interactions.

Adult↗

Good pitch memory is widespread.

Here we show that good pitch memory is widespread among adults with no musical training. We tested unselected college students on their memory for the pitch level of instrumental soundtracks from familiar television programs. Participants heard 5-s excerpts either at the original pitch level or shifted upward or downward by 1 or 2 semitones. They successfully identified the original pitch levels. Other participants who heard comparable excerpts from unfamiliar recordings could not do so. These findings reveal that ordinary listeners retain fine-grained information about pitch level over extended periods. Adults' reportedly poor memory for pitch is likely to be a by-product of their inability to name isolated pitches.

Adult↗

The "ticktock" of our internal clock: direct brain evidence of subjective accents in isochronous sequences.

The phenomenon commonly known as subjective accenting refers to the fact that identical sound events within purely isochronous sequences are perceived as unequal. Although subjective accenting has been extensively explored using behavioral methods, no physiological evidence has ever been provided for it. In the present study, we tested the notion that these perceived irregularities are related to the dynamic deployment of attention. We disrupted listeners' expectancies in different positions of auditory equitone sequences and measured their responses through brain event-related potentials (ERPs). Significant differences in a late parietal (P3-like) ERP component were found between the responses elicited on odd-numbered versus even-numbered positions, suggesting that a default binary metric structure was perceived. Our findings indicate that this phenomenon has a rather cognitive, attention-dependent origin, partly affected by musical expertise.

Adult↗

Sexual selection in female perceptual space: how female túngara frogs perceive and respond to complex population variation in acoustic mating signals.

Female preferences for male mating signals are often evaluated on single parameters in isolation or small suites of characters. Most signals, however, are composites of many individual parameters. In this study we quantified multivariate traits in the advertisement call of the túngara frog, Physalaemus pustulosus. We represented the calls in multidimensional scaling space and chose nine test calls to represent the range of population variation. We then tested females for phonotactic preference between calls in each pair of the nine test calls. We used statistics developed for paired comparisons in such "round robin" competitions to evaluate the null hypothesis of equal attractiveness, and to examine the degree to which females responded to calls as being different from or similar to one another in attractiveness. We then examined the attractiveness of each test call relative to all other test calls as a function of their location in multivariate acoustic space (the acoustic landscape) to visualize sexual selection on calls. Finally, we used methods from cognitive psychology to illustrate the females' perception of call attractiveness in multivariate space, and compared this perceptual landscape to the acoustic landscape of quantitative call variation. We show that correlations between individual call characters are not strong and thus there are few biomechanical constraints on their independent evolution. Most call variables differed among males, and there was high repeatability of call characters within males. Females often discriminated between pairs of calls from the population, and there were significant differences among calls in their attractiveness. Female preferences for calls were not stabilizing. The region of the acoustic landscape that was most attractive to females included the mean call but was not centered around it. The females' perceptual or preference landscape did not correlate with the call's acoustic landscape, and female perception of calls decreased rather than enhanced call differences.

Animal Communication↗

Temporally nonadjacent nonlinguistic sounds affect speech categorization.

Speech perception is an ecologically important example of the highly context-dependent nature of perception; adjacent speech, and even nonspeech, sounds influence how listeners categorize speech. Some theories emphasize linguistic or articulation-based processes in speech-elicited context effects and peripheral (cochlear) auditory perceptual interactions in non-speech-elicited context effects. The present studies challenge this division. Results of three experiments indicate that acoustic histories composed of sine-wave tones drawn from spectral distributions with different mean frequencies robustly affect speech categorization. These context effects were observed even when the acoustic context temporally adjacent to the speech stimulus was held constant and when more than a second of silence or multiple intervening sounds separated the nonlinguistic acoustic context and speech targets. These experiments indicate that speech categorization is sensitive to statistical distributions of spectral information, even if the distributions are composed of nonlinguistic elements. Acoustic context need be neither linguistic nor local to influence speech perception.

Adult↗

Visual prosody and speech intelligibility: head movement improves auditory speech perception.

People naturally move their heads when they speak, and our study shows that this rhythmic head motion conveys linguistic information. Three-dimensional head and face motion and the acoustics of a talker producing Japanese sentences were recorded and analyzed. The head movement correlated strongly with the pitch (fundamental frequency) and amplitude of the talker's voice. In a perception study, Japanese subjects viewed realistic talking-head animations based on these movement recordings in a speech-in-noise task. The animations allowed the head motion to be manipulated without changing other characteristics of the visual or acoustic speech. Subjects correctly identified more syllables when natural head motion was present in the animation than when it was eliminated or distorted. These results suggest that nonverbal gestures such as head movements play a more direct role in the perception of speech than previously known.

Adult↗

An objective method of assessing nasality: a possible aid in the selection of patients for adenoidectomy.

We present an appraisal of an objective technique for assessing nasality, or the nasal component of speech. Evidence suggests that a subjective impression of hyponasal speech is related to the adenoid volume and the radiographic palatal airway, although clinical assessments may have poor inter- and intra-observer agreement. Determination of the oral and nasal acoustic ratio or 'Nasalance' is quick, painless, and non-invasive. There was good agreement and reproducibility within normal subjects when test phrases were used. Words such as 'bananas' which contain nasal consonants showed large reductions in the Nasalance score when the nostrils were occluded and are of use in clinical assessment. This method may be of use in refining the selection of children for adenoidectomy.

Adenoidectomy↗

The use of sound recording and oxygen saturation in screening snorers for obstructive sleep apnoea.

It is desirable to screen snoring patients for obstructive sleep apnoea (OSA) prior to surgical treatment. We postulated that the addition of a sound profile would increase the value of overnight oxygen saturation (SaO2) as a screening method. Thirty-nine polysomnographic studies including sound level measured by calibrated meter were performed on snorers being considered for uvulopalato-pharyngoplasty (UPPP). Polysomnography showed an apnoea/hypopnoea index (AHI) > or = 15 per hour of sleep in seven subjects. Two experienced observers independently, without knowledge of other data, classified paper records of SaO2 alone and SaO2 plus sound level obtained during polysomnography as OSA 'unlikely', 'equivocal' or 'definite'. The addition of sound to SaO2 reduced the number of equivocal results from 14 to six and increased the number classified as 'definite' or 'unlikely'. The sensitivity of oximetry +/- sound increased as the threshold AHI used in the definition of OSA increased; addition of sound improved recognition of mild OSA without impairing specificity.

Adult↗