Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sound Spectrography”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Influence of estuarine hypoxia on feeding and sound production by two sympatric pipefish species (Syngnathidae).

This research utilizes the acoustic behavior of two sympatric pipefish species to assess the impact of hypoxia on feeding. We collected northern, Syngnathus fuscus, and dusky pipefishes, Syngnathus floridae, from the relatively pristine Chincoteague Bay, Virginia, USA and audiovisually recorded behavior in the laboratory of fish held in normoxic (>5 mg/L O(2)) and hypoxic (2 and 1 mg/L O(2)) conditions. Both species produced high frequency ( approximately 0.9-1.4 kHz), short duration (3-22 msec) clicks. Feeding strikes were significantly correlated with both wet weight of ingested food and click production. Thus, sound production serves as an accurate measure of feeding activity. In hypoxic conditions, reduced food intake corresponded with decreased sound production. Significant declines in both behaviors were evident after 1 day and continued as long as hypoxic conditions were maintained. Interspecific differences in sensitivity were detected. Specifically, S. floridae showed a tendency to perform head snaps at the surface. S. fuscus exhibited a breakdown in the coupling of sound production with food intake in 2 mg/L O(2) with clicks produced in other contexts, particularly choking and food expulsion. Reductions in feeding will ultimately impact growth, health, and eventually reproduction as resources are devoted to survival instead of gamete production and courtship. This work suggests acoustic monitoring of field sites with adverse environmental conditions may reflect changes in feeding behavior in addition to population dispersal.

Analysis of Variance↗

Spectral pattern complexity analysis and the quantification of voice normality in healthy and radiotherapy patient groups.

Vocal fold functionality may alter in response to direct radiotherapy or indirectly by perturbation of the hypothalamic-pituitary axis. Perceptual assessment of voice quality is difficult to summarise in a single, reliable figure of normality and normality itself is undefined. In this study spectral analysis of vocal fold vibration, based on impedance variations measured across the larynx using an electro-glottogram, is used to build a single parameter description of standard vowel phonation in the normal male population. Patient data and perceptual assessment are then compared to this standard. The spectral pattern of the vowel/i/ electro-glottogram time series is analysed using approximate entropy after dynamic fundamental-harmonic frequency normalisation. The approximate entropy provides a single estimate of the spectral pattern complexity. A cohort of 89 normal males formed two statistically distinct groups, G1, with strong spectral pattern and high complexity 0.338 (+/-0.036), and G2 with a weak spectral pattern and low complexity 0.175 (+/-0.049). Membership ratio G1:G2 was 2:1. A cohort of 30 male larynx cancer cases were analysed approximately 3-6 months after irradiation, and three male prophylactic cranial irradiation cases some years after treatment. Two-thirds of patients had G2 or lower levels of complexity. The lower G2 complexity level appears to be the subjective, as well as the objective, threshold for voice normality.

Algorithms↗

An amplitude-bandwidth expansion method for hearing-aid adjustment.

The basic concept for a technique that facilitates the adjustment of a hearing aid by a person of normal hearing is proposed. The technique involves processing using an amplitude-bandwidth expansion method to expand the hearing-aid output to fit to the user's hearing requirements and loudness recruitment. The expansion is a reversal of the cochlear compression model. As the expansion method introduces distortion of the waveform and a reduction of the expansion rate, the support technique presented here includes solutions to these problems. The difference between the original speech signal and the expanded output generated from the compressed output of the original signal is virtually inaudible. This technique, which effectively simulates a patient's hearing characteristics in order to allow an audiologist to set up a patient's hearing aid, is worth further investigation.

Acoustic Stimulation↗

An integrated tool for the diagnosis of voice disorders.

A PC-based integrated aid tool has been developed for the analysis and screening of pathological voices. With it the user can simultaneously record speech, electroglottographic (EGG), and videoendoscopic signals, and synchronously edit them to select the most significant segments. These multimedia data are stored on a relational database, together with a patient's personal information, anamnesis, diagnosis, visits, explorations and any other comment the specialist may wish to include. The speech and EGG waveforms are analysed by means of temporal representations and the quantitative measurements of parameters such as spectrograms, frequency and amplitude perturbation measurements, harmonic energy, noise, etc. are calculated using digital signal processing techniques, giving an idea of the degree of hoarseness and quality of the voice register. Within this framework, the system uses a standard protocol to evaluate and build complete databases of voice disorders. The target users of this system are speech and language therapists and ear nose and throat (ENT) clinicians. The application can be easily configured to cover the needs of both groups of professionals. The software has a user-friendly Windows style interface. The PC should be equipped with standard sound and video capture cards. Signals are captured using common transducers: a microphone, an electroglottograph and a fiberscope or telelaryngoscope. The clinical usefulness of the system is addressed in a comprehensive evaluation section.

Computer Graphics↗

Acoustic-structural coupled finite element analysis for sound transmission in human ear--pressure distributions.

A three-dimensional (3D) finite element (FE) model of human ear with accurate structural geometry of the external ear canal, tympanic membrane (TM), ossicles, middle ear suspensory ligaments, and middle ear cavity has been recently reported by our group. In present study, this 3D FE model was modified to include acoustic-structural interfaces for coupled analysis from the ear canal through the TM to middle ear cavity. Pressure distributions in the canal and middle ear cavity at different frequencies were computed under input sound pressure applied at different locations in the canal. The spectral distributions of middle ear pressure at the oval window, round window, and medial site of the umbo were calculated and the results demonstrated that there was no significant difference of pressures between those locations at frequency below 3.5 kHz. Finally, the influence of TM perforation on pressure distributions in the canal and middle ear cavity was investigated for perforations in the inferior-posterior and inferior sites of the TM in the FE model and human temporal bones. The results show that variation of middle ear pressure is related to the perforation type and location, and is sensitive to frequency.

Acoustic Stimulation↗

Investigation of an HMM/ANN hybrid structure in pattern recognition application using cepstral analysis of dysarthric (distorted) speech signals.

Computer speech recognition of individuals with dysarthria, such as cerebral palsy patients requires a robust technique that can handle conditions of very high variability and limited training data. In this study, application of a 10 state ergodic hidden Markov model (HMM)/artificial neural network (ANN) hybrid structure for a dysarthric speech (isolated word) recognition system, intended to act as an assistive tool, was investigated. A small size vocabulary spoken by three cerebral palsy subjects was chosen. The effect of such a structure on the recognition rate of the system was investigated by comparing it with an ergodic hidden Markov model as a control tool. This was done in order to determine if this modified technique contributed to enhanced recognition of dysarthric speech. The speech was sampled at 11 kHz. Mel frequency cepstral coefficients were extracted from them using 15 ms frames and served as training input to the hybrid model setup. The subsequent results demonstrated that the hybrid model structure was quite robust in its ability to handle the large variability and non-conformity of dysarthric speech. The level of variability in input dysarthric speech patterns sometimes limits the reliability of the system. However, its application as a rehabilitation/control tool to assist dysarthric motor impaired individuals holds sufficient promise.

Algorithms↗

A speech-controlled environmental control system for people with severe dysarthria.

Automatic speech recognition (ASR) can provide a rapid means of controlling electronic assistive technology. Off-the-shelf ASR systems function poorly for users with severe dysarthria because of the increased variability of their articulations. We have developed a limited vocabulary speaker dependent speech recognition application which has greater tolerance to variability of speech, coupled with a computerised training package which assists dysarthric speakers to improve the consistency of their vocalisations and provides more data for recogniser training. These applications, and their implementation as the interface for a speech-controlled environmental control system (ECS), are described. The results of field trials to evaluate the training program and the speech-controlled ECS are presented. The user-training phase increased the recognition rate from 88.5% to 95.4% (p<0.001). Recognition rates were good for people with even the most severe dysarthria in everyday usage in the home (mean word recognition rate 86.9%). Speech-controlled ECS were less accurate (mean task completion accuracy 78.6% versus 94.8%) but were faster to use than switch-scanning systems, even taking into account the need to repeat unsuccessful operations (mean task completion time 7.7s versus 16.9s, p<0.001). It is concluded that a speech-controlled ECS is a viable alternative to switch-scanning systems for some people with severe dysarthria and would lead, in many cases, to more efficient control of the home.

Algorithms↗

Acoustic characteristics of air puff-induced 22-kHz alarm calls in direct recordings.

Alarm calls were induced in adult Wistar rats by an air puff. Emitted calls were digitized and directly recorded on a computer hard drive. The long-duration 22-kHz calls were emitted almost exclusively in series. Initial calls in the series tended to have the longest durations, higher frequency range, and the highest degree of frequency modulation, as compared to other calls. The frequency modulation always appeared as a downward sweep and seemed to represent a tuning of individual calls to a 3 kHz communicatory band. Regardless of the maximum frequency, rats always reached approximately the same minimum frequency, common to all calls. Thus, the broader was the frequency range of a given call, the longer the call duration. It is postulated, therefore, that rats emit 22-kHz calls at the minimum possible ultrasonic frequency they are able to produce, which is synonymous with peak frequency. It is further postulated that production of alarm calls in series, with long call duration and the invariably low ultrasonic frequency, maximizes successful communication in dangerous situations. Exceptions to this rule were observed immediately following air puffs, suggesting that acoustic parameters of the initial calls may differ from the alarming properties of the remaining 22-kHz calls.

Acoustics↗

ARTSTREAM: a neural network model of auditory scene analysis and source segregation.

Multiple sound sources often contain harmonics that overlap and may be degraded by environmental noise. The auditory system is capable of teasing apart these sources into distinct mental objects, or streams. Such an 'auditory scene analysis' enables the brain to solve the cocktail party problem. A neural network model of auditory scene analysis, called the ARTSTREAM model, is presented to propose how the brain accomplishes this feat. The model clarifies how the frequency components that correspond to a given acoustic source may be coherently grouped together into a distinct stream based on pitch and spatial location cues. The model also clarifies how multiple streams may be distinguished and separated by the brain. Streams are formed as spectral-pitch resonances that emerge through feedback interactions between frequency-specific spectral representations of a sound source and its pitch. First, the model transforms a sound into a spatial pattern of frequency-specific activation across a spectral stream layer. The sound has multiple parallel representations at this layer. A sound's spectral representation activates a bottom-up filter that is sensitive to the harmonics of the sound's pitch. This filter activates a pitch category which, in turn, activates a top-down expectation that is also sensitive to the harmonics of the pitch. Resonance develops when the spectral and pitch representations mutually reinforce one another. Resonance provides the coherence that allows one voice or instrument to be tracked through a noisy multiple source environment. Spectral components are suppressed if they do not match harmonics of the top-down expectation that is read-out by the selected pitch, thereby allowing another stream to capture these components, as in the 'old-plus-new heuristic' of Bregman. Multiple simultaneously occurring spectral-pitch resonances can hereby emerge. These resonance and matching mechanisms are specialized versions of Adaptive Resonance Theory, or ART, which clarifies how pitch representations can self-organize during learning of harmonic bottom-up filters and top-down expectations. The model also clarifies how spatial location cues can help to disambiguate two sources with similar spectral cues. Data are simulated from psychophysical grouping experiments, such as how a tone sweeping upwards in frequency creates a bounce percept by grouping with a downward sweeping tone due to proximity in frequency, even if noise replaces the tones at their intersection point. Illusory auditory percepts are also simulated, such as the auditory continuity illusion of a tone continuing through a noise burst even if the tone is not present during the noise, and the scale illusion of Deutsch whereby downward and upward scales presented alternately to the two ears are regrouped based on frequency proximity, leading to a bounce percept. Since related sorts of resonances have been used to quantitatively simulate psychophysical data about speech perception, the model strengthens the hypothesis that ART-like mechanisms are used at multiple levels of the auditory system. Proposals for developing the model to explain more complex streaming data are also provided.

Acoustic Stimulation↗

Learning new sounds of speech: reallocation of neural substrates.

Functional magnetic resonance imaging (fMRI) was used to investigate changes in brain activity related to phonetic learning. Ten monolingual English-speaking subjects were scanned while performing an identification task both before and after five sessions of training with a Hindi dental-retroflex nonnative contrast. Behaviorally, training resulted in an improvement in the ability to identify the nonnative contrast. Imaging results suggest that the successful learning of a nonnative phonetic contrast results in the recruitment of the same areas that are involved during the processing of native contrasts, including the left superior temporal gyrus, insula-frontal operculum, and inferior frontal gyrus. Additionally, results of correlational analyses between behavioral improvement and the blood-oxygenation-level-dependent (BOLD) signal obtained during the posttraining Hindi task suggest that the degree of success in learning is accompanied by more efficient neural processing in classical frontal speech regions, and by a reduction of deactivation relative to a noise baseline condition in left parietotemporal speech regions.

Adult↗

Dynamics of brain activity in motor and frontal cortical areas during music listening: a magnetoencephalographic study.

There are formidable problems in studying how 'real' music engages the brain over wide ranges of temporal scales extending from milliseconds to a lifetime. In this work, we recorded the magnetoencephalographic signal while subjects listened to music as it unfolded over long periods of time (seconds), and we developed and applied methods to correlate the time course of the regional brain activations with the dynamic aspects of the musical sound. We showed that frontal areas generally respond with slow time constants to the music, reflecting their more integrative mode; motor-related areas showed transient-mode responses to fine temporal scale structures of the sound. The study combined novel analysis techniques designed to capture and quantify fine temporal sequencing from the authentic musical piece (characterized by a clearly defined rhythm and melodic structure) with the extraction of relevant features from the dynamics of the regional brain activations. The results demonstrated that activity in motor-related structures, specifically in lateral premotor areas, supplementary motor areas, and somatomotor areas, correlated with measures of rhythmicity derived from the music. These correlations showed distinct laterality depending on how the musical performance deviated from the strict tempo of the music score, that is, depending on the musical expression.

Adult↗

Relating neuronal dynamics for auditory object processing to neuroimaging activity: a computational modeling and an fMRI study.

We investigated the neural basis of auditory object processing in the cerebral cortex by combining neural modeling and functional neuroimaging. We developed a large-scale, neurobiologically realistic network model of auditory pattern recognition that relates the neuronal dynamics of cortical auditory processing of frequency modulated (FM) sweeps to functional neuroimaging data of the type obtained using PET and fMRI. Areas included in the model extend from primary auditory to prefrontal cortex. The electrical activities of the neuronal units of the model were constrained to agree with data from the neurophysiological literature regarding the perception of FM sweeps. We also conducted an fMRI experiment using stimuli and tasks similar to those used in our simulations. The integrated synaptic activity of the neuronal units in each region of the model, convolved with a hemodynamic response function, was used as a correlate of the simulated fMRI activity, and generally agreed with the experimentally observed fMRI data in the brain areas corresponding to the regions of the model. Our results demonstrate that the model is capable of exhibiting the salient features of both electrophysiological neuronal activities and fMRI values that are in agreement with empirically observed data. These findings provide support for our hypotheses concerning how auditory objects are processed by primate neocortex.

Adult↗

Systematic latency variation of the auditory evoked M100: from average to single-trial data.

Standard analyses of neurophysiologically evoked response data rely on signal averaging across many epochs associated with specific events. The amplitudes and latencies of these averaged events are subsequently interpreted in the context of the given perceptual, motor, or cognitive tasks. Can such critical timing properties of event-related responses be recovered from single-trial data? Here, we make use of the M100 latency paradigm used in previous magnetoencephalography (MEG) research to evaluate a novel single-trial analysis approach. Specifically, the latency of the auditory evoked M100 varies systematically with stimulus frequency over a well-defined time range (lower frequencies, e.g., 125 Hz, yield up to 25 ms longer latencies than higher frequencies, e.g., 1000 Hz). Here, we show that the complex filtering approach to single-trial analysis recovers this key characteristic of the M100 response, as well as some other important response properties relating to lateralization. The results illustrate (i) the utility of the complex filtering method and (ii) the potential of the M100 latency to be used for stimulus encoding, since the relevant variation can be observed in single trials.

Acoustic Stimulation↗

Hemispheric roles in the perception of speech prosody.

Speech prosody is processed in neither a single region nor a specific hemisphere, but engages multiple areas comprising a large-scale spatially distributed network in both hemispheres. It remains to be elucidated whether hemispheric lateralization is based on higher-level prosodic representations or lower-level encoding of acoustic cues, or both. A cross-language (Chinese; English) fMRI study was conducted to examine brain activity elicited by selective attention to Chinese intonation (I) and tone (T) presented in three-syllable (I3, T3) and one-syllable (I1, T1) utterance pairs in a speeded response, discrimination paradigm. The Chinese group exhibited greater activity than the English in a left inferior parietal region across tasks (I1, I3, T1, T3). Only the Chinese group exhibited a leftward asymmetry in inferior parietal and posterior superior temporal (I1, I3, T1, T3), anterior temporal (I1, I3, T1, T3), and frontopolar (I1, I3) regions. Both language groups shared a rightward asymmetry in the mid portions of the superior temporal sulcus and middle frontal gyrus irrespective of prosodic unit or temporal interval. Hemispheric laterality effects enable us to distinguish brain activity associated with higher-order prosodic representations in the Chinese group from that associated with lower-level acoustic/auditory processes that are shared among listeners regardless of language experience. Lateralization is influenced by language experience that shapes the internal prosodic representation of an external auditory signal. We propose that speech prosody perception is mediated primarily by the RH, but is left-lateralized to task-dependent regions when language processing is required beyond the auditory analysis of the complex sound.

Adult↗

Sound frequency representation in cat auditory cortex.

Using the intrinsic signal optical recording technique, we reconstructed the two-dimensional pattern of stimulus-evoked neuronal activities in the auditory cortex of anesthetized and paralyzed cats. The average magnitude of intrinsic signal in response to a pure tone stimulus increased steadily as the sound pressure level increased. A detailed analysis demonstrated that the evoked signals at early frames were scaled by the sound pressure level, which in turn indicated the presence of a minimum level of sound pressure beyond which stimulus-related intrinsic signal can be generated. Intrinsic signals evoked significantly by pure tone stimuli of different frequencies were localized and arranged in an orderly manner in the middle ectosylvian gyrus, which indicates that the primary auditory field (AI) is tonotopically organized. The arrangement of optimal frequencies obtained from optical recordings of the same auditory cortex, which were conducted on different days, was highly reproducible. Furthermore, other auditory fields surrounding AI in the recorded area were allocated based on the observed tonotopicity. We also conducted unit recordings on the cats used for optical recording with the same set of acoustic stimuli. The gross feature of the arrangement of optimal frequencies determined by unit recordings agreed with the tonotopic arrangement determined by the optical recording, although the precise agreement was not obtained.

Acoustic Stimulation↗

Gamma-band activity dissociates between matching and nonmatching stimulus pairs in an auditory delayed matching-to-sample task.

Electro- and magnetoencephalography studies have suggested that increased gamma-band activity (GBA) is a correlate of activated neural stimulus representations. In this study, a delayed matching-to-sample paradigm for auditory spatial information was employed to investigate the role of magnetoencephalographic gamma-band activity in the differentiation between matching and nonmatching stimulus pairs. Twelve subjects made same-different judgments about the lateralization angle of pairs of filtered noise stimuli (S1 and S2) presented with 0.8-s delays. One half of the subjects had to respond to matching stimulus pairs, the other half to nonmatching stimulus pairs. Cortical oscillatory activity in the memory task was compared to a control task requiring the detection of background noise intensity changes. Memory-related GBA increases were revealed over midline parietal areas in the middle of the delay phase and during the presentation of S2 and over frontocentral areas at the end of the delay phase. This replicated previous findings. In addition, nonmatching trials were associated with increased GBA over right parietal areas in response to S2. The midline parietal GBA increase during S2 in the memory condition may have reflected the representation of S1 needed for a comparison between S1 and S2. When S1 and S2 were identical, no further representation was required. In contrast, for nonmatching pairs, a second representation was activated over right parietal areas.

Adult↗

Phonetic processing areas revealed by sinewave speech and acoustically similar non-speech.

The neural substrates underlying speech perception are still not well understood. Previously, we found dissociation of speech and nonspeech processing at the earliest cortical level (AI), using speech and nonspeech complexity dimensions. Acoustic differences between speech and nonspeech stimuli in imaging studies, however, confound the search for linguistic-phonetic regions. Presently, we used sinewave speech (SWsp) and nonspeech (SWnon), which replace speech formants with sinewave tones, in order to match acoustic spectral and temporal complexity while contrasting phonetics. Chord progressions (CP) were used to remove the effects of auditory coherence and object processing. Twelve normal RH volunteers were scanned with fMRI while listening to SWsp, SWnon, CP, and a baseline condition arranged in blocks. Only two brain regions, in bilateral superior temporal sulcus, extending more posteriorly on the left, were found to prefer the SWsp condition after accounting for acoustic modulation and coherence effects. Two regions responded preferentially to the more frequency-modulated stimuli, including one that overlapped the right temporal phonetic area and another in the left angular gyrus far from the phonetic area. These findings are proposed to form the basis for the two subtypes of auditory word deafness. Several brain regions, including auditory and non-auditory areas, preferred the coherent auditory stimuli and are likely involved in auditory object recognition. The design of the current study allowed for separation of acoustic spectrotemporal, object recognition, and phonetic effects resulting in distinct and overlapping components.

Adolescent↗

Locating the initial stages of speech-sound processing in human temporal cortex.

It is commonly assumed that, in the cochlea and the brainstem, the auditory system processes speech sounds without differentiating them from any other sounds. At some stage, however, it must treat speech sounds and nonspeech sounds differently, since we perceive them as different. The purpose of this study was to delimit the first location in the auditory pathway that makes this distinction using functional MRI, by identifying regions that are differentially sensitive to the internal structure of speech sounds as opposed to closely matched control sounds. We analyzed data from nine right-handed volunteers who were scanned while listening to natural and synthetic vowels, or to nonspeech stimuli matched to the vowel sounds in terms of their long-term energy and both their spectral and temporal profiles. The vowels produced more activation than nonspeech sounds in a bilateral region of the superior temporal sulcus, lateral and inferior to regions of auditory cortex that were activated by both vowels and nonspeech stimuli. The results suggest that the perception of vowel sounds is compatible with a hierarchical model of primate auditory processing in which early cortical stages of processing respond indiscriminately to speech and nonspeech sounds, and only higher regions, beyond anatomically defined auditory cortex, show selectivity for speech sounds.

Adult↗