Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sound Spectrography”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,477 records · Page 82Linked to original sources

Auditory filter nonlinearity in mild/moderate hearing impairment.

Sensorineural hearing loss has frequently been shown to result in a loss of frequency selectivity. Less is known about its effects on the level dependence of selectivity that is so prominent a feature of normal hearing. The aim of the present study is to characterize such changes in nonlinearity as manifested in the auditory filter shapes of listeners with mild/moderate hearing impairment. Notched-noise masked thresholds at 2 kHz were measured over a range of stimulus levels in hearing-impaired listeners with losses of 20-50 dB. Growth-of-masking functions for different notch widths are more parallel for hearing-impaired than for normal-hearing listeners, indicating a more linear filter. Level-dependent filter shapes estimated from the data show relatively little change in shape across level. The loss of nonlinearity is also evident in the input/output functions derived from the fitted filter shapes. Reductions in nonlinearity are clearly evident even in a listener with only 20-dB hearing loss.

Acoustic Stimulation↗

Spectral loudness summation as a function of duration.

Loudness was measured as a function of signal bandwidth for 10-, 100-, and 1000-ms-long signals. The test and reference signals were bandpass-filtered noise spectrally centered at 2 kHz. The bandwidth of the test signal was varied from 200 to 6400 Hz. The reference signal had a bandwidth of 3200 Hz. The reference levels were 45, 55, and 65 dB SPL. The level to produce equal loudness was measured with an adaptive, two-interval, two-alternative forced-choice procedure. A loudness matching procedure was used, where the tracks for all signal pairs to be compared were interleaved. Mean results for nine normal-hearing subjects showed that the magnitude of spectral loudness summation depends on signal duration. For all reference levels, a 6- to 8-dB larger level difference between equally loud signals with the smallest (delta f = 200 Hz) and largest (delta f = 6400 Hz) bandwidth is found for 10-ms-long signals than for the 1000-ms-long signals. The duration effect slightly decreases with increasing reference loudness. As a consequence, loudness models should include a duration-dependent compression stage. Alternatively, if a fixed loudness ratio between signals of different duration is assumed, this loudness ratio should depend on the signal spectrum.

Adult↗

Decision strategies of hearing-impaired listeners in spectral shape discrimination.

The ability to discriminate between sounds with different spectral shapes was evaluated for normal-hearing and hearing-impaired listeners. Listeners detected a 920-Hz tone added in phase to a single component of a standard consisting of the sum of five tones spaced equally on a logarithmic frequency scale ranging from 200 to 4200 Hz. An overall level randomization of 10 dB was either present or absent. In one subset of conditions, the no-perturbation conditions, the standard stimulus was the sum of equal-amplitude tones. In the perturbation conditions, the amplitudes of the components within a stimulus were randomly altered on every presentation. For both perturbation and no-perturbation conditions, thresholds for the detection of the 920-Hz tone were measured to compare sensitivity to changes in spectral shape between normal-hearing and hearing-impaired listeners. To assess whether hearing-impaired listeners relied on different regions of the spectrum to discriminate between sounds, spectral weights were estimated from the perturbed standards by correlating the listener's responses with the level differences per component across two intervals of a two-alternative forced-choice task. Results showed that hearing-impaired and normal-hearing listeners had similar sensitivity to changes in spectral shape. On average, across-frequency correlation functions also were similar for both groups of listeners, suggesting that as long as all components are audible and well separated in frequency, hearing-impaired listeners can use information across frequency as well as normal-hearing listeners. Analysis of the individual data revealed, however, that normal-hearing listeners may be better able to adopt optimal weighting schemes. This conclusion is only tentative, as differences in internal noise may need to be considered to interpret the results obtained from weighting studies between normal-hearing and hearing-impaired listeners.

Adult↗

Two-tone suppression in the cricket, Eunemobius carolinus (Gryllidae, Nemobiinae).

Sounds with frequencies >15 kHz elicit an acoustic startle response (ASR) in flying crickets (Eunemobius carolinus). Although frequencies <15 kHz do not elicit the ASR when presented alone, when presented with ultrasound (40 kHz), low-frequency stimuli suppress the ultrasound-induced startle. Thus, using methods similar to those in masking experiments, we used two-tone suppression to assay sensitivity to frequencies in the audio band. Startle suppression was tuned to frequencies near 5 kHz, the frequency range of male calling songs. Similar to equal loudness contours measured in humans, however, equal suppression contours were not parallel, as the equivalent rectangular bandwidth of suppression tuning changed with increases in ultrasound intensity. Temporal integration of suppressor stimuli was measured using nonsimultaneous presentations of 5-ms pulses of 6 and 40 kHz. We found that no suppression occurs when the suppressing tone is >2 ms after and >5 ms before the ultrasound stimulus, suggesting that stimulus overlap is a requirement for suppression. When considered together with our finding that the intensity of low-frequency stimuli required for suppression is greater than that produced by singing males, the overlap requirement suggests that two-tone suppression functions to limit the ASR to sounds containing only ultrasound and not to broadband sounds that span the audio and ultrasound range.

Acoustic Stimulation↗

Auditory stream segregation on the basis of amplitude-modulation rate.

In this study, auditory stream segregation based on differences in the rate of envelope fluctuations--in the absence of spectral and temporal fine structure cues--was tested. The temporal sequences to segregate were composed of fully amplitude-modulated (AM) bursts of broadband noises A and B. All sequences were built by the reiteration of a ABA triplet where A modulation rate was fixed at 100 Hz and B modulation rate was variable. The first experiment was devoted to measuring the threshold difference in AM rate leading subjects to perceive the sequence as two streams as opposed to just one. The results of this first experiment revealed that subjects generally perceived the sequences as a single perceptual stream when the difference in AM rate between the A and B noises was smaller than 0.75 oct, and as two streams when the difference was larger than about 1.00 oct. These streaming thresholds were found to be substantially larger than, and not related to, the subjects' modulation-rate discrimination thresholds. The results of a second experiment demonstrated that AM-rate-based streaming was adversely affected by decreases in AM depth, but that segregation remained possible as long as the AM of either the A or B noises was above the subject's AM-detection threshold. The results of a third experiment indicated that AM-rate-based streaming effects were still observed when the modulations applied to the A and B noises were set individually, either at a constant level in dB above AM-detection threshold, or at levels at which they were of the same perceived strength. This finding suggests that AM-rate-based streaming is not necessarily mediated by perceived differences in AM depth. Altogether, the results of this study indicate that sequential sounds can be segregated on the sole basis of differences in the rate of their temporal fluctuations in the absence of other temporal or spectral cues.

Adult↗

Modeling sound transmission through the pulmonary system and chest with application to diagnosis of a collapsed lung.

A theoretical and experimental study was undertaken to examine the feasibility of using audible-frequency vibro-acoustic waves for diagnosis of pneumothorax, a collapsed lung. The hypothesis was that the acoustic response of the chest to external excitation would change with this condition. In experimental canine studies, external acoustic energy was introduced into the trachea via an endotracheal tube. For the control (nonpneumothorax) state, it is hypothesized that sound waves primarily travel through the airways, couple to the lung parenchyma, and then are transmitted directly to the chest wall. In contradistinction, when a pneumothorax is present the intervening air presents an added barrier to efficient acoustic energy transfer. Theoretical models of sound transmission through the pulmonary system and chest region to the chest wall surface are developed to more clearly understand the mechanisms of intensity loss when a pneumothorax is present, relative to a baseline case. These models predict significant decreases in acoustic transmission strength when a pneumothorax is present, in qualitative agreement with experimental measurements. Development of the models, their extension via finite element analysis, and comparisons with experimental canine studies are reviewed.

Acoustic Stimulation↗

Rhythmic masking release: contribution of cues for perceptual organization to the cross-spectral fusion of concurrent narrow-band noises.

The contribution of temporal asynchrony, spatial separation, and frequency separation to the cross-spectral fusion of temporally contiguous brief narrow-band noise bursts was studied using the Rhythmic Masking Release paradigm (RMR). RMR involves the discrimination of one of two possible rhythms, despite perceptual masking of the rhythm by an irregular sequence of sounds identical to the rhythmic bursts, interleaved among them. The release of the rhythm from masking can be induced by causing the fusion of the irregular interfering sounds with concurrent "flanking" sounds situated in different frequency regions. The accuracy and the rated clarity of the identified rhythm in a 2-AFC procedure were employed to estimate the degree of fusion of the interferring sounds with flanking sounds. The results suggest that while synchrony fully fuses short-duration noise bursts across frequency and across space (i.e., across ears and loudspeakers), an asynchrony of 20-40 ms produces no fusion. Intermediate asynchronies of 10-20 ms produce partial fusion, where the presence of other cues is critical for unambiguous grouping. Though frequency and spatial separation reduced fusion, neither of these manipulations was sufficient to abolish it. For the parameters varied in this study, stimulus onset asynchrony was the dominant cue determining fusion, but there were additive effects of the other cues. Temporal synchrony appears to be critical in determining whether brief sounds with abrupt onsets and offsets are heard as one event or more than one.

Acoustic Stimulation↗

Sources of DPOAEs revealed by suppression experiments, inverse fast Fourier transforms, and SFOAEs in impaired ears.

DPOAE sources are modeled by intermodulation distortion generated near the f2 place and a reflection of this distortion near the DP place. In a previous paper, inverse fast Fourier transforms (IFFTs) of DPOAE filter functions in normal ears were consistent with this model [Konrad-Martin et al., J. Acoust. Soc. Am. 109, 2862-2879 (2001)]. In the present article, similar measurements were made in ears with specific hearing-loss configurations. It was hypothesized that hearing loss at f2 or DP frequencies would influence the relative contributions to the DPOAE from the corresponding basilar membrane places, and would affect the relative magnitudes of SFOAEs at frequencies equal to f2 and fDP. DPOAEs were measured with f2 = 4 kHz, f1 varied, and a suppressor near fDP. L2 was 25-55 dB SPL (L1 = L2 + 10 dB). SFOAEs were measured at f2 and at 2.7 kHz (the average fDP produced by the f1 sweep) for stimulus levels of 20-60 dB SPL. SFOAE results supported predictions of the pattern of amplitude differences between SFOAEs at 4 and 2.7 kHz for sloping losses, but did not support predictions for the rising- and flat-loss categories. Unsuppressed IFFTs for rising losses typically had one peak. IFFTs for flat or sloping losses typically have two or more peaks; later peaks were more prominent in ears with sloping losses compared to normal ears. Specific predictions were unambiguously supported by the results for only four of ten cases, and were generally supported in two additional cases. Therefore, the relative contributions of the two DPOAE sources often were abnormal in impaired ears, but not always in the predicted manner.

Acoustic Stimulation↗

YIN, a fundamental frequency estimator for speech and music.

An algorithm is presented for the estimation of the fundamental frequency (F0) of speech or musical sounds. It is based on the well-known autocorrelation method with a number of modifications that combine to prevent errors. The algorithm has several desirable features. Error rates are about three times lower than the best competing methods, as evaluated over a database of speech recorded together with a laryngograph signal. There is no upper limit on the frequency search range, so the algorithm is suited for high-pitched voices and music. The algorithm is relatively simple and may be implemented efficiently and with low latency, and it involves few parameters that must be tuned. It is based on a signal model (periodic signal) that may be extended in several ways to handle various forms of aperiodicity that occur in particular applications. Finally, interesting parallels may be drawn with models of auditory processing.

Algorithms↗

Toward a model for lexical access based on acoustic landmarks and distinctive features.

This article describes a model in which the acoustic speech signal is processed to yield a discrete representation of the speech stream in terms of a sequence of segments, each of which is described by a set (or bundle) of binary distinctive features. These distinctive features specify the phonemic contrasts that are used in the language, such that a change in the value of a feature can potentially generate a new word. This model is a part of a more general model that derives a word sequence from this feature representation, the words being represented in a lexicon by sequences of feature bundles. The processing of the signal proceeds in three steps: (1) Detection of peaks, valleys, and discontinuities in particular frequency ranges of the signal leads to identification of acoustic landmarks. The type of landmark provides evidence for a subset of distinctive features called articulator-free features (e.g., [vowel], [consonant], [continuant]). (2) Acoustic parameters are derived from the signal near the landmarks to provide evidence for the actions of particular articulators, and acoustic cues are extracted by sampling selected attributes of these parameters in these regions. The selection of cues that are extracted depends on the type of landmark and on the environment in which it occurs. (3) The cues obtained in step (2) are combined, taking context into account, to provide estimates of "articulator-bound" features associated with each landmark (e.g., [lips], [high], [nasal]). These articulator-bound features, combined with the articulator-free features in (1), constitute the sequence of feature bundles that forms the output of the model. Examples of cues that are used, and justification for this selection, are given, as well as examples of the process of inferring the underlying features for a segment when there is variability in the signal due to enhancement gestures (recruited by a speaker to make a contrast more salient) or due to overlap of gestures from neighboring segments.

Attention↗

Assessing auditory distance perception using virtual acoustics.

In most naturally occurring situations, multiple acoustic properties of the sound reaching a listener's ears change as sound source distance changes. Because many of these acoustic properties, or cues, can be confounded with variation in the acoustic properties of the source and the environment, the perceptual processes subserving distance localization likely combine and weight multiple cues in order to produce stable estimates of sound source distance. Here, this cue-weighting process is examined psychophysically, using a method of virtual acoustics that allows precise measurement and control of the acoustic cues thought to be salient for distance perception in a representative large-room environment. Though listeners' judgments of sound source distance are found to consistently and exponentially underestimate true distance, the perceptual weight assigned to two primary distance cues (intensity and direct-to-reverberant energy ratio) varies substantially as a function of both sound source type (noise and speech) and angular position (0 degrees and 90 degrees relative to the median plane). These results suggest that the cue-weighting process is flexible, and able to adapt to individual distance cues that vary as a result of source properties and environmental conditions.

Adult↗

Auditory normalization of French vowels synthesized by an articulatory model simulating growth from birth to adulthood.

The present article aims at exploring the invariant parameters involved in the perceptual normalization of French vowels. A set of 490 stimuli, including the ten French vowels /i y u e ø o E oe (inverted c) a/ produced by an articulatory model, simulating seven growth stages and seven fundamental frequency values, has been submitted as a perceptual identification test to 43 subjects. The results confirm the important effect of the tonality distance between F1 and f0 in perceived height. It does not seem, however, that height perception involves a binary organization determined by the 3-3.5-Bark critical distance. Regarding place of articulation, the tonotopic distance between F1 and F2 appears to be the best predictor of the perceived front-back dimension. Nevertheless, the role of the difference between F2 and F3 remains important. Roundedness is also examined and correlated to the effective second formant, involving spectral integration of higher formants within the 3.5-Bark critical distance. The results shed light on the issue of perceptual invariance, and can be interpreted as perceptual constraints imposed on speech production.

Adolescent↗

American and Swedish children's acquisition of vowel duration: effects of vowel identity and final stop voicing.

Vowel durations typically vary according to both intrinsic (segment-specific) and extrinsic (contextual) specifications. It can be argued that such variations are due to both predisposition and cognitive learning. The present report utilizes acoustic phonetic measurements from Swedish and American children aged 24 and 30 months to investigate the hypothesis that default behaviors may precede language-specific learning effects. The predicted pattern is the presence of final consonant voicing effects in both languages as a default, and subsequent learning of intrinsic effects most notably in the Swedish children. The data, from 443 monosyllabic tokens containing high-front vowels and final stop consonants, are analyzed in statistical frameworks at group and individual levels. The results confirm that Swedish children show an early tendency to vary vowel durations according to final consonant voicing, followed only six months later by a stage at which the intrinsic influence of vowel identity grows relatively more robust. Measures of vowel formant structure from selected 30-month-old children also revealed a tendency for children of this age to focus on particular acoustic contrasts. In conclusion, the results indicate that early acquisition of vowel specifications involves an interaction between language-specific features and articulatory predispositions associated with phonetic context.

Child Language↗

Acoustic competition in the gulf toadfish Opsanus beta: acoustic tagging.

Nesting male gulf toadfish Opsanus beta produce a boatwhistle advertisement call used in male-male competition and to attract females and an agonistic grunt call. The grunt is a short-duration pulsatile call, and the boatwhistle is a complex call typically consisting of zero to three introductory grunts, a long tonal boop note, and zero to three shorter boops. The beginning of the boop note is also gruntlike. Anomalous boatwhistles contain a short-duration grunt embedded in the tonal portion of the boop or between an introductory grunt and the boop. Embedded grunts have sound-pressure levels and frequency spectra that correspond with those of recognized neighbors, suggesting that one fish is grunting during another's call, a phenomenon here termed acoustic tagging. Snaps of nearby pistol shrimp may also be tagged, and chains of tags involving more than two fish occur. The stimulus to tag is a relatively intense sound with a rapid rise time, and tags are generally produced within 100 ms of a trigger stimulus. Time between the trigger and the tag decreases with increased trigger amplitude. Tagging is distinct from increased calling in response to natural calls or stimulatory playbacks since calls rarely overlap other calls or playbacks. Tagging is not generally reciprocal between fish, suggesting parallels to dominance displays.

Acoustics↗

Transformation of external-ear spectral cues into perceived delays by the big brown bat, Eptesicus fuscus.

The external-ear transfer function for big brown bats (Eptesicus fuscus) contains two prominent notches that vary from 30 to 55 kHz and from 70 to 100 kHz, respectively, as sound-source elevation moves from -40 to +10 degrees. These notches resemble a higher-frequency version of external-ear cues for vertical localization in humans and other mammals. However, they also resemble interference notches created in echoes when reflected sounds overlap at short time separations of 30-50 micros. Psychophysical experiments have shown that bats actually perceive small time separations from interference notches, and here we used the same technique to test whether external-ear notches are recognized as a corresponding time separation, too. The bats' performance reveals the elevation dependence of a time-separation estimate at 25-45 micros in perceived delay. Convergence of target-shape and external-ear cues onto echo spectra creates ambiguity about whether a particular notch relates to the object or to its location, which the bat could resolve by ignoring the presence of notches at external-ear frequencies. Instead, the bat registers the frequencies of notches caused by the external ear along with notches caused by the target's structure and employs spectrogram correlation and transformation (SCAT) to convert them all into a family of delay estimates that includes elevation.

Animals↗

Representations of sound that are insensitive to spectral filtering and parametrization procedures.

This paper describes representations of time-dependent signals that are invariant under any invertible signal distortion. Such a representation can be created by rescaling the signal in a nonlinear dynamic manner that is determined by recently encountered signal levels. Information that is encoded in such representations will be faithfully communicated in the presence of severe signal distortions, which may originate in the transmitter, receiver, or the channel between them. As in speech communication, the receiver is "blind" and need not characterize the form of the signal distortion, which remains unknown. The method is applied to analytical examples, acoustic waveforms of human speech, and the short-term Fourier spectra of a bird song. The results suggest that the rescaled representation of a sound is insensitive to the way its spectra have been filtered and parametrized, as long as those processes do not obliterate the differences between the various spectra in the sound. Finally, the possible "speaker" independence of these representations is explored in the context of a simple linear prediction model of vocal tracts with a single degree of freedom.

Acoustics↗

Quantitative assessment of second language learners' fluency: comparisons between read and spontaneous speech.

This paper describes two experiments aimed at exploring the relationship between objective properties of speech and perceived fluency in read and spontaneous speech. The aim is to determine whether such quantitative measures can be used to develop objective fluency tests. Fragments of read speech (Experiment 1) of 60 non-native speakers of Dutch and of spontaneous speech (Experiment 2) of another group of 57 non-native speakers of Dutch were scored for fluency by human raters and were analyzed by means of a continuous speech recognizer to calculate a number of objective measures of speech quality known to be related to perceived fluency. The results show that the objective measures investigated in this study can be employed to predict fluency ratings, but the predictive power of such measures is stronger for read speech than for spontaneous speech. Moreover, the adequacy of the variables to be employed appears to be dependent on the specific type of speech material investigated and the specific task performed by the speaker.

Humans↗

Nonlinear analysis of irregular animal vocalizations.

Animal vocalizations range from almost periodic vocal-fold vibration to completely atonal turbulent noise. Between these two extremes, a variety of nonlinear dynamics such as limit cycles, subharmonics, biphonation, and chaotic episodes have been recently observed. These observations imply possible functional roles of nonlinear dynamics in animal acoustic communication. Nonlinear dynamics may also provide insight into the degree to which detailed features of vocalizations are under close neural control, as opposed to more directly reflecting biomechanical properties of the vibrating vocal folds themselves. So far, nonlinear dynamical structures of animal voices have been mainly studied with spectrograms. In this study, the deterministic versus stochastic (DVS) prediction technique was used to quantify the amount of nonlinearity in three animal vocalizations: macaque screams, piglet screams, and dog barks. Results showed that in vocalizations with pronounced harmonic components (adult macaque screams, certain piglet screams, and dog barks), deterministic nonlinear prediction was clearly more powerful than stochastic linear prediction. The difference, termed low-dimensional nonlinearity measure (LNM), indicates the presence of a low-dimensional attractor. In highly irregular signals such as juvenile macaque screams, piglet screams, and some dog barks, the detectable amount of nonlinearity was comparatively small. Analyzing 120 samples of dog barks, it was further shown that the harmonic-to-noise ratio (HNR) was positively correlated with LNM. It is concluded that nonlinear analysis is primarily useful in animal vocalizations with strong harmonic components (including subharmonics and biphonation) or low-dimensional chaos.

Animals↗