Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sound Spectrography”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,279 records · Page 71Linked to original sources

Wavelet-based enhancement of lung and bowel sounds using fractal dimension thresholding--Part I: methodology.

An efficient method for the enhancement of lung sounds (LS) and bowel sounds (BS), based on wavelet transform (WT), and fractal dimension (FD) analysis is presented in this paper. The proposed method combines multiresolution analysis with FD-based thresholding to compose a WT-FD filter, for enhanced separation of explosive LS (ELS) and BS (EBS) from the background noise. In particular, the WT-FD filter incorporates the WT-based multiresolution decomposition to initially decompose the recorded bioacoustic signal into approximation and detail space in the WT domain. Next, the FD of the derived WT coefficients is estimated within a sliding window and used to infer where the thresholding of the WT coefficients has to happen. This is achieved through a self-adjusted procedure that iteratively "peels" the estimated FD signal and isolates its peaks produced by the WT coefficients corresponding to ELS or EBS. In this way, two new signals are constructed containing the useful and the undesired WT coefficients, respectively. By applying WT-based multiresolution reconstruction to these two signals, a first version of the desired signal and the background noise is provided, accordingly. This procedure is repeated until a stopping criterion is met, finally resulting in efficient separation of the ELS or EBS from the background noise. The proposed WT-FD filter introduces an alternative way to the enhancement of bioacoustic signals, applicable to any separation problem involving nonstationary transient signals mixed with uncorrelated stationary background noise. The results from the application of the WT-FD filter to real bioacoustic data are presented and discussed in an accompanying paper.

Algorithms↗

Wavelet-based enhancement of lung and bowel sounds using fractal dimension thresholding--Part II: application results.

The application of the wavelet transform-fractal dimension-based (WT-FD) filter of Part I of this paper to real bioacoustic data, which include explosive lung sounds (ELS) and explosive bowel sounds (EBS) recorded from patients with pulmonary or gastrointestinal dysfunction, respectively, is presented in this paper. The objective of the latter is the evaluation of the performance of the WT-FD filter on different types of bioacoustic signals, varying not only in their structural morphology but also in the degree of their noise contamination. As it is thoroughly described in Part I of this paper, the WT-FD filter uses the fractal dimension to form an efficient way of thresholding the WT coefficients at different resolution scales, keeping, thus, only those that can contribute to the accurate reconstruction of the ELS and EBS signals. Quantitative and qualitative analysis of the experimental results show an efficient performance of the WT-FD filter to circumvent the noise presence (100% detectability rate, 100% sensitivity, 100% specificity) by faithfully extracting the authentic structure of ELS and EBS from the background noise. The WT-FD filter does not require any noise reference signal or noise reference templates. The results from a noise stress test (mean cross-correlation index of the original and the estimated signal converging to 100%; mean normalized maximum amplitude error converging to 0.7%) prove its robustness to various noise levels (0-20 dB), enabling its potential use in similar noise cases met in everyday clinical medicine. Furthermore, the efficient performance of the WT-FD filter facilitates the physician to better interpret the auscultation findings. Due to its simplicity and low computational cost, the WT-FD filter can possibly be implemented in a real-time context to serve as a tool for the continuous ELS and EBS screening.

Adult↗

Speaker normalization for chinese vowel recognition in cochlear implants.

Because of the limited spectra-temporal resolution associated with cochlear implants, implant patients often have greater difficulty with multitalker speech recognition. The present study investigated whether multitalker speech recognition can be improved by applying speaker normalization techniques to cochlear implant speech processing. Multitalker Chinese vowel recognition was tested with normal-hearing Chinese-speaking subjects listening to a 4-channel cochlear implant simulation, with and without speaker normalization. For each subject, speaker normalization was referenced to the speaker that produced the best recognition performance under conditions without speaker normalization. To match the remaining speakers to this "optimal" output pattern, the overall frequency range of the analysis filter bank was adjusted for each speaker according to the ratio of the mean third formant frequency values between the specific speaker and the reference speaker. Results showed that speaker normalization provided a small but significant improvement in subjects' overall recognition performance. After speaker normalization, subjects' patterns of recognition performance across speakers changed, demonstrating the potential for speaker-dependent effects with the proposed normalization technique.

Artificial Intelligence↗

SVD-based optimal filtering for noise reduction in dual microphone hearing aids: a real time implementation and perceptual evaluation.

In this paper, the first real-time implementation and perceptual evaluation of a singular value decomposition (SVD)-based optimal filtering technique for noise reduction in a dual microphone behind-the-ear (BTE) hearing aid is presented. This evaluation was carried out for a speech weighted noise and multitalker babble, for single and multiple jammer sound source scenarios. Two basic microphone configurations in the hearing aid were used. The SVD-based optimal filtering technique was compared against an adaptive beamformer, which is known to give significant improvements in speech intelligibility in noisy environment. The optimal filtering technique works without assumptions about a speaker position, unlike the two-stage adaptive beamformer. However this strategy needs a robust voice activity detector (VAD). A method to improve the performance of the VAD was presented and evaluated physically. By connecting the VAD to the output of the noise reduction algorithms, a good discrimination between the speech-and-noise periods and the noise-only periods of the signals was obtained. The perceptual experiments demonstrated that the SVD-based optimal filtering technique could perform as well as the adaptive beamformer in a single noise source scenario, i.e., the ideal scenario for the latter technique, and could outperform the adaptive beamformer in multiple noise source scenarios.

Adult↗

A mathematical model for source separation of MMG signals recorded with a coupled microphone-accelerometer sensor pair.

Recent advances in sensor technology for muscle activity monitoring have resulted in the development of a coupled microphone-accelerometer sensor pair for physiological acousti signal recording. This sensor can be used to eliminate interfering sources in practical settings where the contamination of an acoustic signal by ambient noise confounds detection but cannot be easily removed [e.g., mechanomyography (MMG), swallowing sounds, respiration, and heart sounds]. This paper presents a mathematical model for the coupled microphone-accelerometer vibration sensor pair, specifically applied to muscle activity monitoring (i.e., MMG) and noise discrimination in externally powered prostheses for below-elbow amputees. While the model provides a simple and reliable source separation technique for MMG signals, it can also be easily adapted to other aplications where the recording of low-frequency (< 1 kHz) physiological vibration signals is required.

Acceleration↗

Acoustical signal properties for cardiac/respiratory activity and apneas.

Traditionally, auscultation is applied to the diagnosis of either respiratory disturbances by respiratory sounds or cardiac disturbances by cardiac sounds. In addition, for sleep apnea syndrome diagnosis, snoring sounds are also monitored. The present study was aimed at synchronous detection of all three sound components (cardiac, respiratory, and snoring) from a single spot. The sounds were analyzed with respect to the cardiorespiratory activity, and to the detection and classification of apneas. Sound signals from 30 subjects including 10 apnea patients were detected by means of a microphone connected to a chestpiece which was applied to the heart region. The complex nature of the signal was investigated using time, spectral, and statistical approaches, in connection with self-defined time-based and event-based characteristics. The results show that the obstruction is accompanied by an increase of statistically relevant spectral components in the range of 300 to 2000 Hz, however, not within the range up to 300 Hz. Signal properties are discussed with respect to different breathing types, as well as to the presence and the type of apneas. Principal component analysis of the event-based characteristics shows significant properties of the sound signal with respect to different types of apneas and different patient groups, respectively. The analysis reflects apneas with an obstructive segment and those with a central segment. In addition, aiming for an optimum detection of all three sound components, alternative regions on the thorax and on the neck were investigated on two subjects. The results suggest that the right thorax region in the seventh intercostal space and the neck are optimal regions. It is concluded that for patient assessment, extensive acoustic analysis offers a reduction in the number of required sensor components, especially with respect to compact home monitoring of apneas.

Algorithms↗

A new insight into postsurgical objective voice quality evaluation: application to thyroplastic medialization.

This paper aims at providing new objective parameters and plots, easily understandable and usable by clinicians and logopaedicians, in order to assess voice quality recovering after vocal fold surgery. The proposed software tool performs presurgical and postsurgical comparison of main voice characteristics (fundamental frequency, noise, formants) by means of robust analysis tools, specifically devoted to deal with highly degraded speech signals as those under study. Specifically, we address the problem of quantifying voice quality, before and after medialization thyroplasty, for patients affected by glottis incompetence. Functional evaluation after thyroplastic medialization is commonly based on several approaches: videolaryngostroboscopy (VLS), for morphological aspects evaluation, GRBAS scale and Voice Handicap Index (VHI), relative to perceptive and subjective voice analysis respectively, and Multi-Dimensional Voice Program (MDVP), that provides objective acoustic parameters. While GRBAS has the drawback to entirely rely on perceptive evaluation of trained professionals, MDVP often fails in performing analysis of highly degraded signals, thus preventing from presurgical/postsurgical comparison in such cases. On the contrary, the new tool, being capable to deal with severely corrupted signals, always allows a complete objective analysis. The new parameters are compared to scores obtained with the GRBAS scale and to some MDVP parameters, suitably modified, showing good correlation with them. Hence, the new tool could successfully replace or integrate existing ones. With the proposed approach, deeper insight into voice recovering and its possible changes after surgery can thus be obtained and easily evaluated by the clinician.

Diagnosis, Computer-Assisted↗

Telephony-based voice pathology assessment using automated speech analysis.

A system for remotely detecting vocal fold pathologies using telephone-quality speech is presented. The system uses a linear classifier, processing measurements of pitch perturbation, amplitude perturbation and harmonic-to-noise ratio derived from digitized speech recordings. Voice recordings from the Disordered Voice Database Model 4337 system were used to develop and validate the system. Results show that while a sustained phonation, recorded in a controlled environment, can be classified as normal or pathologic with accuracy of 89.1%, telephone-quality speech can be classified as normal or pathologic with an accuracy of 74.2%, using the same scheme. Amplitude perturbation features prove most robust for telephone-quality speech. The pathologic recordings were then subcategorized into four groups, comprising normal, neuromuscular pathologic, physical pathologic and mixed (neuromuscular with physical) pathologic. A separate classifier was developed for classifying the normal group from each pathologic subcategory. Results show that neuromuscular disorders could be detected remotely with an accuracy of 87%, physical abnormalities with an accuracy of 78% and mixed pathology voice with an accuracy of 61%. This study highlights the real possibility for remote detection and diagnosis of voice pathology.

Algorithms↗

Modified local discriminant bases algorithm and its application in analysis of human knee joint vibration signals.

Knee joint disorders are common in the elderly population, athletes, and outdoor sports enthusiasts. These disorders are often painful and incapacitating. Vibration signals [vibroarthrographic (VAG)] are emitted at the knee joint during the swinging movement of the knee. These VAG signals contain information that can be used to characterize certain pathological aspects of the knee joint. In this paper, we present a noninvasive method for screening knee joint disorders using the VAG signals. The proposed approach uses wavelet packet decompositions and a modified local discriminant bases algorithm to analyze the VAG signals and to identify the highly discriminatory basis functions. We demonstrate the effectiveness of using a combination of multiple dissimilarity measures to arrive at the optimal set of discriminatory basis functions, thereby maximizing the classification accuracy. A database of 89 VAG signals containing 51 normal and 38 abnormal samples were used in this study. The features extracted from the coefficients of the selected basis functions were analyzed and classified using a linear-discriminant-analysis-based classifier. A classification accuracy as high as 80% was achieved using this true nonstationary signal analysis approach.

Algorithms↗

A robust method for heart sounds localization using lung sounds entropy.

Heart sounds are the main unavoidable interference in lung sound recording and analysis. Hence, several techniques have been developed to reduce or cancel heart sounds (HS) from lung sound records. The first step in most HS cancellation techniques is to detect the segments including HS. This paper proposes a novel method for HS localization using entropy of the lung sounds. We investigated both Shannon and Renyi entropies and the results of the method using Shannon entropy were superior. Another HS localization method based on multiresolution product of lung sounds wavelet coefficients adopted from was also implemented for comparison. The methods were tested on data from 6 healthy subjects recorded at low (7.5 ml/s/kg) and medium 115 ml/s/kg) flow rates. The error of entropy-based method using Shannon entropy was found to be 0.1 +/- 0.4% and 1.0 +/- 0.7% at low and medium flow rates, respectively, which is significantly lower than that of multiresolution product method and those of other methods reported in previous studies. The proposed method is fully automated and detects HS included segments in a completely unsupervised manner.

Adult↗

Multiexpert automatic speech recognition using acoustic and myoelectric signals.

Classification accuracy of conventional automatic speech recognition (ASR) systems can decrease dramatically under acoustically noisy conditions. To improve classification accuracy and increase system robustness a multiexpert ASR system is implemented. In this system, acoustic speech information is supplemented with information from facial myoelectric signals (MES). A new method of combining experts, known as the plausibility method, is employed to combine an acoustic ASR expert and a MES ASR expert. The plausibility method of combining multiple experts, which is based on the mathematical framework of evidence theory, is compared to the Borda count and score-based methods of combination. Acoustic and facial MES data were collected from 5 subjects, using a 10-word vocabulary across an 18-dB range of acoustic noise. As expected the performance of an acoustic expert decreases with increasing acoustic noise; classification accuracies of the acoustic ASR expert are as low as 11.5%. The effect of noise is significantly reduced with the addition of the MES ASR expert. Classification accuracies remain above 78.8% across the 18-dB range of acoustic noise, when the plausibility method is used to combine the opinions of multiple experts. In addition, the plausibility method produced classification accuracies higher than any individual expert at all noise levels, as well as the highest classification accuracies, except at the 9-dB noise level. Using the Borda count and score-based multiexpert systems, classification accuracies are improved relative to the acoustic ASR expert but are as low as 51.5% and 59.5%, respectively.

Algorithms↗

A robust method for estimating respiratory flow using tracheal sounds entropy.

The relationship between respiratory sounds and flow is of great interest for researchers and physicians due to its diagnostic potentials. Due to difficulties and inaccuracy of most of the flow measurement techniques, several researchers have attempted to estimate flow from respiratory sounds. However, all of the proposed methods heavily depend on the availability of different rates of flow for calibrating the model, which makes their use limited by a large degree. In this paper, a robust and novel method for estimating flow using entropy of the band pass filtered tracheal sounds is proposed. The proposed method is novel in terms of being independent of the flow rate chosen for calibration; it requires only one breath for calibration and can estimate any flow rate even out of the range of calibration flow. After removing the effects of heart sounds (which distort the low-frequency components of tracheal sounds) on the calculated entropy of the tracheal sounds, the performance of the method at different frequency ranges were investigated. Also, the performance of the proposed method was tested using 6 different segment sizes for entropy calculation and the best segment sizes during inspiration and expiration were found. The method was tested on data of 10 healthy subjects at five different flow rates. The overall estimation error was found to be 8.3 +/- 2.8% and 9.6 +/- 2.8% for inspiration and expiration phases, respectively.

Adult↗

Dimensionality reduction of a pathological voice quality assessment system based on Gaussian mixture models and short-term cepstral parameters.

Voice diseases have been increasing dramatically in recent times due mainly to unhealthy social habits and voice abuse. These diseases must be diagnosed and treated at an early stage, especially in the case of larynx cancer. It is widely recognized that vocal and voice diseases do not necessarily cause changes in voice quality as perceived by a listener. Acoustic analysis could be a useful tool to diagnose this type of disease. Preliminary research has shown that the detection of voice alterations can be carried out by means of Gaussian mixture models and short-term mel cepstral parameters complemented by frame energy together with first and second derivatives. This paper, using the F-Ratio and Fisher's discriminant ratio, will demonstrate that the detection of voice impairments can be performed using both mel cesptral vectors and their first derivative, ignoring the second derivative.

Computer Simulation↗

Detection of cough signals in continuous audio recordings using hidden Markov models.

Cough is a common symptom of many respiratory diseases. The evaluation of its intensity and frequency of occurrence could provide valuable clinical information in the assessment of patients with chronic cough. In this paper we propose the use of hidden Markov models (HMMs) to automatically detect cough sounds from continuous ambulatory recordings. The recording system consists of a digital sound recorder and a microphone attached to the patient's chest. The recognition algorithm follows a keyword-spotting approach, with cough sounds representing the keywords. It was trained on 821 min selected from 10 ambulatory recordings, including 2473 manually labeled cough events, and tested on a database of nine recordings from separate patients with a total recording time of 3060 min and comprising 2155 cough events. The average detection rate was 82% at a false alarm rate of seven events/h, when considering only events above an energy threshold relative to each recording's average energy. These results suggest that HMMs can be applied to the detection of cough sounds from ambulatory patients. A postprocessing stage to perform a more detailed analysis on the detected events is under development, and could allow the rejection of some of the incorrectly detected events.

Algorithms↗

Voice low tone to high tone ratio: a potential quantitative index for vowel [a:] and its nasalization.

Hypernasality is associated with various diseases and interferes with speech intelligibility. A recently developed quantitative index called voice low tone to high tone ratio (VLHR) was used to estimate nasalization. The voice spectrum is divided into low-frequency power (LFP) and high-frequency power (HFP) by a specific cutoff frequency (600 Hz). VLHR is defined as the division of LFP into HFP and is expressed in decibels. Voice signals of the sustained vowel [a :] and its nasalization in eight subjects with hypernasality were collected for analysis of nasalance and VLHR. The correlation of VLHR with nasalance scores was significant (r = 0.76, p < 0.01), and so was the correlation between VLHR and perceptual hypernasality scores (r = 0.80, p < 0.01). Simultaneous recordings of nasal airflow temperature with a thermistor and voice signals in another 8 healthy subjects showed a significant correlation between temperature rate of nasal airflow and VLHR (r = 0.76, p < 0.01), as well. We conclude that VLHR may become a potential quantitative index of hypernasal speech and can be applied in either basic or clinical studies.

Adult↗

Effects of diameter, length, and circuit pressure on sound conductance through endotracheal tubes.

We evaluated the acoustic frequency response of endotracheal tubes (ETs) to assess their effect on respiratory system sound transmission studies. White noise 150-3300 Hz was introduced into 4.0-, 6.0-, and 8.0-mm ETs and recorded at their proximal and distal ends. Four tubes of each size were studied at their original and normalized lengths, in straight and bent configurations, and at circuit pressures from 0 to 20 cmH2O. The characteristics of the sound transmission were compared using an analysis of variance for repeated measures. The average transmission amplitude varied directly with tube diameter. The position of peaks and troughs on the amplitude frequency distribution depended on tube length but not on tube diameter. The angle of the phase-frequency plot correlated well with the length of the tube and was independent of its diameter. A 90 degrees bend in the tube had no effect on its sound transmission. Increasing the circuit pressure above ambient modified the frequency response only if volume changes occurred in the test lung. When used to conduct sound into the respiratory system an ET affects the incident signal predictably depending on its length and diameter but not on its curvature or circuit pressure.

Animals↗

Quantitative indices for the assessment of the repeatability of distortion product otoacoustic emissions in laboratory animals.

Distortion product otoacoustic emissions (DPOAE) can be used to study cochlear function in an objective and non-invasive manner. One practical and essential aspect of any investigating measure is the consistency of its results upon repeated testing of the same individual/animal (i.e., its test/retest repeatability). The goal of the present work is to propose two indices to quantitatively assess the repeatability of DPOAE in laboratory animals. The methodology is here illustrated using two data sets which consist of DPOAE subsequently collected from Sprague-Dawley rats. The results of these experiments showed that the proposed indices are capable of estimating both the repeatability of the true emission level and the inconsistencies associated with measurement error. These indices could be a significantly useful tool to identify real and even small changes in the cochlear function exerted by potential ototoxic agents.

Algorithms↗