Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sound Spectrography”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,441 records · Page 80Linked to original sources

Application of multidimensional scaling to subjective evaluation of coded speech.

We present results from a pilot study directed at developing an anchorable subjective speech quality test. The test uses multidimensional scaling techniques to obtain quantitative information about the perceptual attributes of speech. In the first phase of the study, subjects ranked perceptual distances between samples of speech produced by two different talkers, one male and one female, processed by a variety of codecs. The resulting distance matrices were processed to obtain, for each talker, a stimulus space for the various speech samples. This stimulus space has the properties that distances between stimuli in this space correspond to perceptual distances between stimuli and that the dimensions of this space correspond to attributes used by the subjects in determining perceptual distances. Mean opinion scores (MOS) scores obtained in an earlier study were found to be highly correlated with position in the stimulus space, and the three dimensions of the stimulus space were found to have identifiable physical and perceptual correlates. In the second phase of the study, we developed techniques for fitting speech generated by a new codec under investigation into a previously established stimulus space. The user is provided with a collection of speech samples and with the stimulus space for these speech samples as determined by a large-scale listening test. The user then carries out a much smaller listening test to determine the position of the new stimulus in the previously established stimulus space. This system is anchorable, so that different versions of a codec under development can be compared directly, and it provides more detailed information than the single number provided by MOS testing. We suggest that this information could be used to advantage in algorithm development and in development of objective measures of speech quality.

Adult↗

Representation of harmonic complex stimuli in the ventral cochlear nucleus of the chinchilla.

The representation of Schroeder-phase harmonic complex sounds in the ventral cochlear nucleus (VCN) of the anesthetized chinchilla was studied. Stimuli consisted of a series of harmonically related sinusoids, multiples of a fundamental frequency (f0), summed in either negative (-SCHR) or positive (+SCHR) Schroeder phase. Psychoacoustic experiments performed in humans by other investigators have revealed that masking effects of -SCHR stimuli are larger than those found using +SCHR stimuli as maskers. In our laboratory, basilar membrane measurements at the base of the chinchilla cochlea show that responses to -SCHR stimuli are less "peaked," or modulated, than responses to +SCHR stimuli. We also found that suppression of a characteristic-frequency (CF) tone by -SCHR stimuli is larger than that evoked by +SCHR stimuli. Rate-intensity functions display higher firing rates in responses to -SCHR stimuli than in those produced by +SCHR stimuli. Firing rates evoked by either -SCHR or +SCHR stimuli saturate at lower values than those obtained in responses to CF tones. Rate and synchrony suppressions by -SCHR stimuli were larger than those evoked by +SCHR stimuli. Auditory nerve fiber responses to Schroeder complex stimuli share most of the properties of VCN responses, indicating little additional processing by the VCN.

Animals↗

Inner hair cell response patterns: implications for low-frequency hearing.

Inner hair cell (IHC) responses to tone-burst stimuli were measured from three locations in the apical half of the guinea pig cochlea. In addition to the measurement of ac receptor potentials, average intracellular voltages, reflecting both ac and dc components of the receptor potential, were computed and compared to determine how bandwidth changes with level. Companion phase measures were also obtained and evaluated. Data collected from turn 2, where best frequency (BF) is approximately 4000 Hz, indicate that frequency response functions are asymmetrical with steeper slopes above the best frequency of the cell. However, in turn 4, where BF is around 250 Hz, the opposite behavior is observed and the steepest slopes are measured below BF. The data imply that cochlear filters are generally asymmetrical with steeper slopes above BF. High-pass filtering by the middle ear serves to reduce this asymmetry in turn 3 and to reverse it in turn 4. Apical response patterns are used to assess the degree to which the middle ear transfer function, the IHC's velocity dependence and the shunting effect of the helicotrema influence low-frequency hearing in guinea pigs. Implications for low-frequency hearing in man are also discussed.

Animals↗

Auditory scene analysis by echolocation in bats.

Echolocating bats transmit ultrasonic vocalizations and use information contained in the reflected sounds to analyze the auditory scene. Auditory scene analysis, a phenomenon that applies broadly to all hearing vertebrates, involves the grouping and segregation of sounds to perceptually organize information about auditory objects. The perceptual organization of sound is influenced by the spectral and temporal characteristics of acoustic signals. In the case of the echolocating bat, its active control over the timing, duration, intensity, and bandwidth of sonar transmissions directly impacts its perception of the auditory objects that comprise the scene. Here, data are presented from perceptual experiments, laboratory insect capture studies, and field recordings of sonar behavior of different bat species, to illustrate principles of importance to auditory scene analysis by echolocation in bats. In the perceptual experiments, FM bats (Eptesicus fuscus) learned to discriminate between systematic and random delay sequences in echo playback sets. The results of these experiments demonstrate that the FM bat can assemble information about echo delay changes over time, a requirement for the analysis of a dynamic auditory scene. Laboratory insect capture experiments examined the vocal production patterns of flying E. fuscus taking tethered insects in a large room. In each trial, the bats consistently produced echolocation signal groups with a relatively stable repetition rate (within 5%). Similar temporal patterning of sonar vocalizations was also observed in the field recordings from E. fuscus, thus suggesting the importance of temporal control of vocal production for perceptually guided behavior. It is hypothesized that a stable sonar signal production rate facilitates the perceptual organization of echoes arriving from objects at different directions and distances as the bat flies through a dynamic auditory scene. Field recordings of E. fuscus, Noctilio albiventris, N. leporinus, Pippistrellus pippistrellus, and Cormura brevirostris revealed that spectral adjustments in sonar signals may also be important to permit tracking of echoes in a complex auditory scene.

Animals↗

The harmonic-to-noise ratio applied to dog barks.

Dog barks are typically a mixture of regular components and irregular (noisy) components. The regular part of the signal is given by a series of harmonics and is most probably due to regular vibrations of the vocal folds, whereas noise refers to any nonharmonic (irregular) energy in the spectrum of the bark signal. The noise components might be due to chaotic vibrations of the vocal-fold tissue or due to turbulence of the air. The ratio of harmonic to nonharmonic energy in dog barks is quantified by applying the harmonics-to-noise ratio (HNR). Barks of a single dog breed were recorded in the same behavioral context. Two groups of dogs were considered: a group of ten healthy dogs (the normal sample), and a group of ten unhealthy dogs, i.e., dogs treated in a veterinary clinic (the clinic sample). Although the unhealthy dogs had no voice disease, differences in emotion or pain or impacts of surgery might have influenced their barks. The barks of the dogs were recorded for a period of 6 months. The HNR computation is based on the Fourier spectrum of a 50-ms section from the middle of the bark. A 10-point moving average curve of the spectrum on a logarithmic scale is considered as estimator of the noise level in the bark, and the maximum difference of the original spectrum and the moving average is defined as the HNR measure. It is shown that a reasonable ranking of the voices is achievable based on the measurement of the HNR. The HNR-based classification is found to be consistent with perceptual evaluation of the barks. In addition, a multiparametric approach confirms the classification based on the HNR. Hence, it may be concluded that the HNR might be useful as a novel parameter in bioacoustics for quantifying the noise within a signal.

Animal Communication↗

Aero-acoustics of silicone rubber lip reeds for alternative voice production in laryngectomees.

To improve voice quality after laryngectomy, a small pneumatic sound source to be incorporated in a regular tracheoesophageal shunt valve was designed. This artificial voice source consists of a single floppy lip reed, which performs self-sustaining flutter-type oscillations driven by the expired pulmonary air that flows through the tracheoesophageal shunt valve along the outward-striking lip reed. In this in vitro study, aero-acoustic data and detailed high-speed digital image sequences of lip reed behavior are obtained for 10 lip configurations. The high-speed visualizations provide a more explicit understanding and reveal details of lip reed behavior, such as the onset of vibration, beating of the lip against the walls of its housing, and chaotic behavior at high volume flow. We discuss several aspects of lip reed behavior in general and implications for its application as an artificial voice source. For pressures above the sounding threshold, volume flow, fundamental frequency and sound pressure level generated by the floppy lip reed are almost linear functions of the driving force, static pressure difference across the lip. Observed irregularities in these relations are mainly caused by transitions from one type of beating behavior of the lip against the walls of its housing to another. This beating explains the wide range and the driving force dependence of fundamental frequency, and seems to have a strong effect on the spectral content. The thickness of the lip base is linearly related to the fundamental frequency of lip reed oscillation.

Humans↗

Noisy speech recognition using de-noised multiresolution analysis acoustic features.

This paper describes a novel application of multiresolution analysis (MRA) in extracting acoustic features that possess de-noising capability for robust speech recognition. The MRA algorithm is used to construct a mel-scaled wavelet packet filter-bank, from which subband powers are computed as the feature parameters for speech recognition. Wiener filtering is applied to a few selected subbands at some intermediate stages of decomposition. For high-frequency bands, Wiener filters are designed based on a reduced fraction of the estimated noise power, making the consonant features much more prominent and contrastive. The proposed method is evaluated in phone recognition experiments with the TIMIT database. In the presence of stationary white noise at 10-dB SNR, the de-noised MRA features attain a phone recognition rate of 32%. There is a noticeable improvement compared with the accuracy of 29% and 20% attained by the commonly used mel-frequency cepstral coefficients (MFCC) with and without cepstral mean normalization (CMN), respectively. The effectiveness of the MRA features is also verified by the fact that they exhibit smaller distortion from clean speech.

Attention↗

Vowel formant discrimination II: Effects of stimulus uncertainty, consonantal context, and training.

This study is one in a series that has examined factors contributing to vowel perception in everyday listening. Four experimental variables have been manipulated to examine systematical differences between optimal laboratory testing conditions and those characterizing everyday listening. These include length of phonetic context, level of stimulus uncertainty, linguistic meaning, and amount of subject training. The present study investigated the effects of stimulus uncertainty from minimal to high uncertainty in two phonetic contexts, /V/ or /bVd/, when listeners had either little or extensive training. Thresholds for discriminating a small change in a formant for synthetic female vowels /I,E,ae,a,inverted v,o/ were obtained using adaptive tracking procedures. Experiment I optimized extensive training for five listeners by beginning under minimal uncertainty (only one formant tested per block) and then increasing uncertainty from 8-to-16-to-22 formants per block. Effects of higher uncertainty were less than expected; performance only decreased by about 30%. Thresholds for CVCs were 25% poorer than for isolated vowels. A previous study using similar stimuli [Kewley-Port and Zheng. J. Acoust. Soc. Am. 106, 2945-2958 (1999)] determined that the ability to discriminate formants was degraded by longer phonetic context. A comparison of those results with the present ones indicates that longer phonetic context degrades formant frequency discrimination more than higher levels of stimulus uncertainty. In experiment 2, performance in the 22-formant condition was tracked over 1 h for 37 typical listeners without formal laboratory training. Performance for typical listeners was initially about 230% worse than for trained listeners. Individual listeners' performance ranged widely with some listeners occasionally achieving performance similar to that of the trained listeners in just one hour.

Adult↗

Effect of stimulus bandwidth on the perception of /s/ in normal- and hearing-impaired children and adults.

Recent studies with adults have suggested that amplification at 4 kHz and above fails to improve speech recognition and may even degrade performance when high-frequency thresholds exceed 50-60 dB HL. This study examined the extent to which high frequencies can provide useful information for fricative perception for normal-hearing and hearing-impaired children and adults. Eighty subjects (20 per group) participated. Nonsense syllables containing the phonemes /s/, /f/, and /O/, produced by a male, female, and child talker, were low-pass filtered at 2, 3, 4, 5, 6, and 9 kHz. Frequency shaping was provided for the hearing-impaired subjects only. Results revealed significant differences in recognition between the four groups of subjects. Specifically, both groups of children performed more poorly than their adult counterparts at similar bandwidths. Likewise, both hearing-impaired groups performed more poorly than their normal-hearing counterparts. In addition, significant talker effects for /s/ were observed. For the male talker, optimum performance was reached at a bandwidth of approximately 4-5 kHz, whereas optimum performance for the female and child talkers did not occur until a bandwidth of 9 kHz.

Adult↗

Transforming echoes into pseudo-action potentials for classifying plants.

Animals perceive their environment by converting sensory stimuli into action potentials, or temporal point processes, that are interpreted by the brain. This paper investigates the information content of point processes extracted from echoes from in situ plants in an effort to understand how bats recognize landmarks in the field. A mobile sonar converts echoes into biologically similar temporal point processes. termed pseudo-action potentials (PAPs), whose inter-PAP interval relates to echo amplitude. The sonar forms a sector scan of an object to produce a spatial-temporal PAP field. Classifier neurons apply delays and coincidence detection to the PAP field to identify three distinct echo types, glints, blobs, and fuzz, which characterize plant features. Glints are large amplitude echoes exhibiting coherence over successive echoes in the sector scan, typically produced by favorably oriented isolated specular reflectors. Blobs are large echoes lacking coherence, typically bordering glints or formed by collections of interfering reflectors. Fuzz represents weak echoes, typically produced by collection of weak scatterers or by reflectors on the beam periphery. A small mirror reflector models a flat leaf surface and motivates the glint criteria. Classifiers are applied to experimental data from two types of tree trunks, a glint-producing sycamore (Platanus occidenatalis) and a glint-absent Norway maple (Acer platanoides) and two plants, a glint-producing rhododendron (Rhododendron maximus) and a glint-absent yew (Taxus media). We speculate that our narrow-band sonar models the activity of a single frequency bin in the frequency-modulated (FM) sweep emitted by bats, and that one function of the frequency bins in the FM sweep is to form a sector scan of the environment.

Animals↗

Evaluation of loudness-level weightings for assessing the annoyance of environmental noise.

Assessment of the annoyance of combined noise environments has been the subject of much research and debate. Currently, most countries use some form of the A-weighted equivalent level (ALEQ) to assess the annoyance of most noises. It provides a constant filter that is independent of sound level. Schomer [Acust. Acta Acust. 86(1), 49-61 (2000)] suggested the use of the equal loudness-level contours (ISO 226, 1987) as a dynamic filter that changes with both sound level and frequency. He showed that loudness-level-weighted sound-exposure level (LLSEL) and loudness-level-weighted equivalent level (LL-LEQ) can be used to assess the annoyance of environmental noise. Compared with A-weighting, loudness-level weighting better orders and assesses transportation noise sources, sounds with strong low-frequency content and, with the addition of a 12-dB adjustment, it better orders and assesses highly impulsive sounds vis-a-vis transportation sounds. This paper compares the LLSEL method with two methods based on loudness calculations using ISO 532b (1975). It shows that in terms of correlation with subjective judgments of annoyance-not loudness-the LLSEL formulation performs much better than do the loudness calculations. This result is true across a range of sources that includes aircraft, helicopters, motor vehicles, trains, and impulsive sources. It also is true within several of the sources separately.

Humans↗

Cross-spectral methods for processing speech.

We present time-frequency methods which are well suited to the analysis of nonstationary multicomponent FM signals, such as speech. These methods are based on group delay, instantaneous frequency, and higher-order phase derivative surfaces computed from the short time Fourier transform (STFT). Unlike more conventional approaches, these methods do not assume a locally stationary approximation of the signal model. We describe the computation of the phase derivatives, the physical interpretation of these derivatives, and a re-mapping algorithm based on these phase derivatives. We show analytically, and by example, the convergence of the re-mapping to the FM representation of the signal. The methods are applied to speech to estimate signal parameters, such as the group delay of a transmission channel and speech formant frequencies. Our goal is to develop a unified method which can accurately estimate speech components in both time and frequency and to apply these methods to the estimation of instantaneous formant frequencies, effective excitation time, vocal tract group delay, and channel group delay. The proposed method has several interesting properties, the most important of which is the ability to simultaneously resolve all FM components of a multicomponent signal, as long as the STFT of the composite signal satisfies a simple separability condition. The method can provide super-resolution in both time and frequency in the sense that it can simultaneously provide time and frequency estimates of FM components, which have much better accuracy than the Heisenberg uncertainty of the STFT. Super-resolution provides the capability to accurately "re-map" each component of the STFT surface to the time and frequency of the FM signal component it represents. To attain high resolution and accuracy, the signal must be jointly estimated simultaneously in time and frequency. This is accomplished by estimating two surfaces, which are essentially the derivatives of the STFT phase with respect to time and frequency. To avoid phase ambiguities, the differentiation is performed as a cross-spectral product.

Fourier Analysis↗

A new procedure for measuring peripheral compression in normal-hearing and hearing-impaired listeners.

Forward-masking growth functions for on-frequency (6-kHz) and off-frequency (3-kHz) sinusoidal maskers were measured in quiet and in a high-pass noise just above the 6-kHz probe frequency. The data show that estimates of response-growth rates obtained from those functions in quiet, which have been used to infer cochlear compression, are strongly dependent on the spread of probe excitation toward higher frequency regions. Therefore, an alternative procedure for measuring response-growth rates was proposed, one that employs a fixed low-level probe and avoids level-dependent spread of probe excitation. Fixed-probe-level temporal masking curves (TMCs) were obtained from normal-hearing listeners at a test frequency of 1 kHz, where the short 1-kHz probe was fixed in level at about 10 dB SL. The level of the preceding forward masker was adjusted to obtain masked threshold as a function of the time delay between masker and probe. The TMCs were obtained for an on-frequency masker (1 kHz) and for other maskers with frequencies both below and above the probe frequency. From these measurements, input/output response-growth curves were derived for individual ears. Response-growth slopes varied from >1.0 at low masker levels to <0.2 at mid masker levels. In three subjects, response growth increased again at high masker levels (>80 dB SPL). For the fixed-level probe, the TMC slopes changed very little in the presence of a high-pass noise masking upward spread of probe excitation. A greater effect on the TMCs was observed when a high-frequency cueing tone was used with the masking tone. In both cases, however, the net effects on the estimated rate of response growth were minimal.

Audiometry, Pure-Tone↗

Second-order modulation detection thresholds for pure-tone and narrow-band noise carriers.

Modulation perception has typically been characterized by measuring detection thresholds for sinusoidally amplitude-modulated (SAM) signals. This study uses multicomponent modulations. "Second-order" temporal modulation transfer functions (TMTFs) measure detection thresholds for a sinusoidal modulation of the modulation waveform of a SAM signal [Lorenzi et al., J. Acoust. Soc. Am. 110, 1030-2038 (2001)]. The SAM signal therefore acts as a "carrier" stimulus of frequency fm, and sinusoidal modulation of the SAM signal's modulation depth (at rate f'm) generates two additional components in the modulation spectrum at fm - f'm and fm + f'm. There is no spectral energy at the envelope beat frequency f'm in the modulation spectrum of the "physical" stimulus. In the present study, second-order TMTFs were measured for three listeners when fm was 16, 64, and 256 Hz. The carrier was either a 5-kHz pure tone or a narrow-band noise with center frequency and bandwidth of 5 kHz and 2 Hz, respectively. The narrow-band noise carrier was used to prevent listeners from detecting spectral energy at the beat frequency f'm in the "internal" stimuli's modulation spectrum. The results show that, for the 5-kHz pure-tone carrier, second-order TMTFs are nearly low pass in shape; the overall sensitivity and cutoff frequency measured on these second-order TMTFs increase when fm increases from 16 to 256 Hz. For the 2-Hz-wide narrow-band noise carrier, second-order TMTFs are nearly flat in shape for fm = 16 and 64 Hz, and they show a high-pass segment for fm = 256 Hz. These results suggest that detection of spectral energy at the envelope beat frequency contributes in part to the detection of second-order modulation. This is consistent with the idea that nonlinear mechanisms in the auditory pathway produce an audible distortion component at the envelope beat frequency in the internal modulation spectrum of the sounds.

Adult↗

Sources of variation in profile analysis. II. Component spacing, dynamic changes, and roving level.

Profile-analysis experiments have typically employed static profiles with constant frequency components spaced at equal intervals along a logarithmic frequency axis. Most periodic, naturally occurring stimuli, however, have components that are harmonically related and vary dynamically in time. One goal of these studies was to determine whether amplitude-increment detection thresholds are different in dynamic, harmonically spaced profiles compared to those for static-log profiles, and why such differences might exist. A second goal was to determine the impact of roving levels (within-trial variation of level). Thresholds for static-log profiles were, on average, 8.7 dB lower than for static-harmonic profiles. A traditional filter-bank model could not account for this result. No consistent effect of dynamic contour (an exponential rising frequency glide) was observed. Thresholds were consistently poorer by 4 to 7 dB when the level was roved, but the differences in thresholds among the different profiles varied little. It is proposed that the higher thresholds observed in static-harmonic profiles may be accounted for by the more intense pitch strength associated with the harmonic profiles.

Adolescent↗

Informational and energetic masking effects in the perception of multiple simultaneous talkers.

Although many researchers have examined the role that binaural cues play in the perception of spatially separated speech signals, relatively little is known about the cues that listeners use to segregate competing speech messages in a monaural or diotic stimulus. This series of experiments examined how variations in the relative levels and voice characteristics of the target and masking talkers influence a listener's ability to extract information from a target phrase in a 3-talker or 4-talker diotic stimulus. Performance in this speech perception task decreased systematically when the level of the target talker was reduced relative to the masking talkers. Performance also generally decreased when the target and masking talkers had similar voice characteristics: the target phrase was most intelligible when the target and masking phrases were spoken by different-sex talkers, and least intelligible when the target and masking phrases were spoken by the same talker. However, when the target-to-masker ratio was less than 3 dB, overall performance was usually lower with one different-sex masker than with all same-sex maskers. In most of the conditions tested, the listeners performed better when they were exposed to the characteristics of the target voice prior to the presentation of the stimulus. The results of these experiments demonstrate how monaural factors may play an important role in the segregation of speech signals in multitalker environments.

Adult↗

The intensity-difference limen for Gaussian-enveloped stimuli as a function of level: tones and broadband noise.

Van Schijndel et al. [J. Acoust. Soc. Am. 105, 3425-3435 (1999)] have proposed that the internal excitation evoked by an auditory stimulus is segmented into "windows" according to the stimulus spectrum and stimulus length. This "multiple looks" model accounts for the mid-duration hump they observed in plots of intensity-difference limens (DLs) versus pip duration for Gaussian-shaped 1- and 4-kHz tones, an effect replicated by Baer et al. [J. Acoust. Soc. Am. 106, 1907-1916 (1999)]. However, van Schijndel et al. and Baer et al. used few levels. A greater number of levels were used by Nizami (1999) for Gaussian-shaped 2-kHz tone-pips whose equivalent rectangular duration (D) was 1.25 ms. The DLs show the mid-level hump known for clicks [Raab and Taub, J. Acoust. Soc. Am. 46, 965-968 (1969)]. At some duration this pattern must become the "near-miss to Weber's law." To determine this duration, as well as the level-dependence of the mid-duration hump, DLs were established for Gaussian-shaped 2-kHz tone-pips of D = 1.25, 2.51, and 10.03 ms at levels of 30-90 dB SPL. The across-subject average DLs for the tone-pips rise up at mid-levels for D= 1.25 and D = 2.51 ms. The DLs for D=2.51 ms are larger, creating the mid-duration hump. At all durations, the new DLs are smaller at high levels than at low levels, consistent with the near-miss to Weber's law. DLs were also obtained here for Gaussian-shaped broadband-noise pips of D=0.63, 1.25, 2.51, 5.02, and 10.03 ms. The DLs for the noise-pip show a mid-level hump for all pip durations. The noise-pip DLs decrease as the pip lengthens, such that the plot of DL versus log duration shows a linear decline, with no mid-duration hump. Analysis of variance reveals that the mid-level hump coexists with the classical patterns of level-dependence, perhaps reflecting the existence of two level-encoding mechanisms, one that depends on firing-rates counted over single neurons and which is responsible for the classical patterns, and one that depends on the initial coordinated burst of neuronal spikes caused by rapid ramping, and which presumably causes the mid-level hump.

Adult↗