Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sound Spectrography”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,405 records · Page 78Linked to original sources

Phonetically trained models for speaker recognition.

In this paper, a speaker recognition system that introduces acoustic information into a Gaussian mixture model (GMM)-based recognizer is presented. This is achieved by using a phonetic classifier during the training phase. The experimental results show that, while maintaining the recognition rate, the decrease in the computational load is between 65% and 80% depending on the number of mixtures of the models.

Adult↗

Binaural coherence edge pitch.

The binaural coherence edge pitch (BICEP) is a dichotic broadband noise pitch effect similar to the binaural edge pitch (BEP). The BICEP stimulus is made by summing spectrally dense sine wave components with random phases. The interaural phase angle is a constant (0 or pi) for components with frequencies below (or above) a chosen edge frequency, and it is a random variable for the remaining components. The chosen edge frequency is a coherence edge because the noises to the two ears are mutually coherent within any band of frequencies on one side of the edge and they are mutually incoherent in any band on the other side. Pitch-matching experiments show that the BICEP exists for coherence edge frequencies between about 300 and 1000 Hz. It is matched by a pure-tone frequency that differs from the edge frequency by 5% to 10%. The matching frequency lies on the incoherent side of the edge, an important result that is consistent with the way that the equalization-cancellation model has been applied to binaural pitch effects, especially the BEP. The results of BICEP experiments depend upon whether the coherent components are presented in 0 or pi interaural phase for some listeners but not for all. The BICEP persists if the noise to one of the ears is delayed, but it becomes weaker and less well matched as the delay increases beyond 2 ms. The BICEP does not depend on whether the component amplitudes are all created equal or are given a Rayleigh distribution. Some reliable pitch sensation exists even when the component amplitudes are entirely independent in the two ears, so long as the phase coherence conditions of the BICEP stimulus are maintained. The existence of the BICEP is a challenge for current models of dichotic pitch because none of them predicts all its features.

Adult↗

Enhancing maximum measurable sound reduction index using sound intensity method and strong receiving room absorption.

The sound intensity method is usually recommended instead of the pressure method in the presence of strong flanking transmission. Especially when small and/or heavy specimens are tested, the flanking often causes problems in laboratories practicing only the pressure method. The purpose of this study was to determine experimentally the difference between the maximum sound reduction indices obtained by the intensity method, RI,max, and by the pressure method, Rmax. In addition, the influence of adding room absorption to the receiving room was studied. The experiments were carried out in an ordinary two-room test laboratory. The exact value of RI,max was estimated by applying a fitting equation to the measured data points. The fitting equation involved the dependence of the pressure-intensity indicator on measured acoustical parameters. In an empty receiving room, the difference between RI,max and Rmax was 4-15 dB, depending on frequency. When the average reverberation time was reduced from 3.5 to 0.6 s, the values of RI,max increased by 2-10 dB compared to the results in the empty room. Thus, it is possible to measure wall structures having 9-22 dB better sound reduction index using the intensity method than with the pressure method. This facilitates the measurements of small and/or heavy specimens in the presence of flanking. Moreover, when new laboratories are designed, the intensity method is an alternative to the pressure method which presupposes expensive isolation structures between the rooms.

Building Codes↗

A linear relation between the compressibility and density of blood.

By considering the blood as a mixture of ultrafiltrate and protein concentrate, the additive nature of compressibility and density from the components is utilized to deduce a linear relation between the compressibility and density for blood. This deduction also indicates that the intercept and slope of the linear relation are independent of the hematocrit, plasma protein concentration, and hemoglobin concentration of red blood cells. To verify experimentally this linear relation, saline and plasma dilutions on porcine or canine blood flowing in an extracorporeal circuit were carried out. The hematocrit of the experiments ranges from 0% to 55% and the plasma protein concentration ranges from 10 to 90 g/l. A resonance device in the circuit measured the density rhob of blood at 37 degrees C and an ultrasound system measured the sound velocity cb. The range of density is from 1,010 to 1,060 g/l and that of sound velocity is from 1,530 to 1,580 m/s. The linear relation that best fits the data of compressibility [computed as (rhob cb(2))-1] and density has a correlation coefficient of 0.9978. The linear relation is found to fit well the dependence of compressibility on density derived from the sound velocity data of human, horse, and porcine blood in the literature.

Animals↗

Acoustic and linguistic factors in the perception of bandpass-filtered speech.

Speech can remain intelligible for listeners with normal hearing when processed by narrow bandpass filters that transmit only a small fraction of the audible spectrum. Two experiments investigated the basis for the high intelligibility of narrowband speech. Experiment 1 confirmed reports that everyday English sentences can be recognized accurately (82%-98% words correct) when filtered at center frequencies of 1500, 2100, and 3000 Hz. However, narrowband low predictability (LP) sentences were less accurately recognized than high predictability (HP) sentences (20% lower scores), and excised narrowband words were even less intelligible than LP sentences (a further 23% drop). While experiment 1 revealed similar levels of performance for narrowband and broadband sentences at conversational speech levels, experiment 2 showed that speech reception thresholds were substantially (>30 dB) poorer for narrowband sentences. One explanation for this increased disparity between narrowband and broadband speech at threshold (compared to conversational speech levels) is that spectral components in the sloping transition bands of the filters provide important cues for the recognition of narrowband speech, but these components become inaudible as the signal level is reduced. Experiment 2 also showed that performance was degraded by the introduction of a speech masker (a single competing talker). The elevation in threshold was similar for narrowband and broadband speech (11 dB, on average), but because the narrowband sentences required considerably higher sound levels to reach their thresholds in quiet compared to broadband sentences, their target-to-masker ratios were very different (+23 dB for narrowband sentences and -12 dB for broadband sentences). As in experiment 1, performance was better for HP than LP sentences. The LP-HP difference was larger for narrowband than broadband sentences, suggesting that context provides greater benefits when speech is distorted by narrow bandpass filtering.

Humans↗

Recognition of spectrally asynchronous speech by normal-hearing listeners and Nucleus-22 cochlear implant users.

This experiment examined the effects of spectral resolution and fine spectral structure on recognition of spectrally asynchronous sentences by normal-hearing and cochlear implant listeners. Sentence recognition was measured in six normal-hearing subjects listening to either full-spectrum or noise-band processors and five Nucleus-22 cochlear implant listeners fitted with 4-channel continuous interleaved sampling (CIS) processors. For the full-spectrum processor, the speech signals were divided into either 4 or 16 channels. For the noise-band processor, after band-pass filtering into 4 or 16 channels, the envelope of each channel was extracted and used to modulate noise of the same bandwidth as the analysis band, thus eliminating the fine spectral structure available in the full-spectrum processor. For the 4-channel CIS processor, the amplitude envelopes extracted from four bands were transformed to electric currents by a power function and the resulting electric currents were used to modulate pulse trains delivered to four electrode pairs. For all processors, the output of each channel was time-shifted relative to other channels, varying the channel delay across channels from 0 to 240 ms (in 40-ms steps). Within each delay condition, all channels were desynchronized such that the cross-channel delays between adjacent channels were maximized, thereby avoiding local pockets of channel synchrony. Results show no significant difference between the 4- and 16-channel full-spectrum speech processor for normal-hearing listeners. Recognition scores dropped significantly only when the maximum delay reached 200 ms for the 4-channel processor and 240 ms for the 16-channel processor. When fine spectral structures were removed in the noise-band processor, sentence recognition dropped significantly when the maximum delay was 160 ms for the 16-channel noise-band processor and 40 ms for the 4-channel noise-band processor. There was no significant difference between implant listeners using the 4-channel CIS processor and normal-hearing listeners using the 4-channel noise-band processor. The results imply that when fine spectral structures are not available, as in the implant listener's case, increased spectral resolution is important for overcoming cross-channel asynchrony in speech signals.

Adult↗

Effects of stimulus frequency and complexity on the mismatch negativity and other components of the cortical auditory-evoked potential.

This study investigated, first, the effect of stimulus frequency on mismatch negativity (MMN), N1, and P2 components of the cortical auditory event-related potential (ERP) evoked during passive listening to an oddball sequence. The hypothesis was that these components would show frequency-related changes, reflected in their latency and magnitude. Second, the effect of stimulus complexity on those same ERPs was investigated using words and consonant-vowel tokens (CVs) discriminated on the basis of formant change. Twelve normally hearing listeners were tested with tone bursts in the speech frequency range (400/440, 1,500/1,650, and 3,000/3,300 Hz), words (/baed/ vs /daed/) and CVs (/bae/ vs /dae/). N1 amplitude and latency decreased as frequency increased. P2 amplitude, but not latency, decreased as frequency increased. Frequency-related changes in MMN were similar to those for N1, resulting in a larger MMN area to low frequency contrasts. N1 amplitude and latency for speech sounds were similar to those found for low tones but MMN had a smaller area. Overall, MMN was present in 46%-71% of tests for tone contrasts but for only 25%-32% of speech contrasts. The magnitude of N1 and MMN for tones appear to be closely related, and both reflect the tonotopicity of the auditory cortex.

Adult↗

Characteristics of whistles from the acoustic repertoire of resident killer whales (Orcinus orca) off Vancouver Island, British Columbia.

The acoustic repertoire of killer whales (Orcinus orca) consists of pulsed calls and tonal sounds, called whistles. Although previous studies gave information on whistle parameters, no study has presented a detailed quantitative characterization of whistles from wild killer whales. Thus an interpretation of possible functions of whistles in killer whale underwater communication has been impossible so far. In this study acoustic parameters of whistles from groups of individually known killer whales were measured. Observations in the field indicate that whistles are close-range signals. The majority of whistles (90%) were tones with several harmonics with the main energy concentrated in the fundamental. The remainder were tones with enhanced second or higher harmonics and tones without harmonics. Whistles had an average bandwidth of 4.5 kHz, an average dominant frequency of 8.3 kHz, and an average duration of 1.8 s. The number of frequency modulations per whistle ranged between 0 and 71. The study indicates that whistles in wild killer whales serve a different function than whistles of other delphinids. Their structure makes whistles of killer whales suitable to function as close-range motivational sounds.

Acoustics↗

Acoustic-phonetic features for the automatic classification of fricatives.

In this article, the acoustic-phonetic characteristics of the American English fricative consonants are investigated from the automatic classification standpoint. The features studied in the literature are evaluated and new features are proposed. To test the value of the extracted features, a statistically guided, knowledge-based, acoustic-phonetic system for the automatic classification of fricatives in speaker-independent continuous speech is proposed. The system uses an auditory-based front-end processing system and incorporates new algorithms for the extraction and manipulation of the acoustic-phonetic features that proved to be rich in their information content. Classification experiments are performed using hard-decision algorithms on fricatives extracted from the TIMIT database continuous speech of 60 speakers (not used in the design/training process) from seven different dialects of American English. An accuracy of 93% is obtained for voicing detection, 91% for place of articulation detection, and 87% for the overall classification of fricatives.

Algorithms↗

Emphasis of short-duration acoustic speech cues for cochlear implant users.

A new speech-coding strategy for cochlear implant users, called the transient emphasis spectral maxima (TESM), was developed to aid perception of short-duration transient cues in speech. Speech-perception scores using the TESM strategy were compared to scores using the spectral maxima sound processor (SMSP) strategy in a group of eight adult users of the Nucleus 22 cochlear implant system. Significant improvements in mean speech-perception scores for the group were obtained on CNC open-set monosyllabic word tests in quiet (SMSP: 53.6% TESM: 61.3%, p<0.001), and on MUSL open-set sentence tests in multitalker noise (SMSP: 64.9% TESM: 70.6%, p<0.001). Significant increases were also shown for consonant scores in the word test (SMSP: 75.1% TESM: 80.6%, p<0.001) and for vowel scores in the word test (SMSP: 83.1% TESM: 85.7%, p<0.05). Analysis of consonant perception results from the CNC word tests showed that perception of nasal, stop, and fricative consonant discrimination was most improved. Information transmission analysis indicated that place of articulation was most improved, although improvements were also evident for manner of articulation. The increases in discrimination were shown to be related to improved coding of short-duration acoustic cues, particularly those of low intensity.

Adult↗

The relationship between spectral characteristics and perceived hypernasality in children.

The purpose of this study was to quantify perceived hypernasality in children. One-third octave spectra of the isolated vowel [i] were obtained from 32 children with cleft palate and 5 children without cleft palate. Four experienced listeners rated the severity of hypernasality of the 37 speech samples using a 6-point equal-appearing interval scale. When the average 1/3-octave spectra from the hypernasal group and the normal resonance group were compared, spectral characteristics of hypernasality were identified as increased amplitudes between F1 and F2 and decreased amplitudes in the region of F2. Based on the findings of the children's speech, 36 speech samples with manipulated spectral characteristics were used to minimize the influences of voice source characteristics on perceived hypernasality. Multiple regression analysis revealed a high correlation (R = 0.84) between the amplitudes of 1/3-octave bands (1 k, 1.6 k, and 2.5 kHz) and the perceptual ratings. Increased amplitudes of bands between F1 and F2 (1 k, 1.6 kHz) and decreased amplitude of the band of F2 (2.5 kHz) was associated with an increasing perceived hypernasality. These results suggest that the amplitudes of the three 1/3-octave bands are appropriate acoustic parameters to quantify hypernasality in the isolated vowel [i].

Adolescent↗

Modulation detection interference: effects of concurrent and sequential streaming.

The presence of amplitude fluctuations in one frequency region can interfere with our ability to detect similar fluctuations in another (remote) frequency region. This effect is known as modulation detection interference (MDI). Gating the interfering and target sounds asynchronously is known to lead to a reduction in MDI, presumably because the two sounds become perceptually segregated. The first experiment examined the relative effects of carrier and modulator gating asynchrony in producing a release from MDI. The target carrier was a 900-ms, 4.3-kHz sinusoid, modulated in amplitude by a 500-ms, 16-Hz sinusoid, with 200-ms unmodulated fringes preceding and following the modulation. The interferer (masker) was a 1-kHz sinusoid, modulated by a narrowband noise with a 16-Hz bandwidth, centered around 16 Hz. Extending the masker carrier for 200 ms before and after the signal carrier reduced MDI, regardless of whether the target and masker modulators were gated synchronously or were gated with onset and offset asynchronies of 200 ms. Similarly, when the carriers were gated synchronously, asynchronous gating of the modulators did not produce a release from MDI. The second experiment measured MDI with a synchronous target and masker and investigated the effect of adding a series of precursor tones, which were designed to promote the forming of a perceptual stream with the masker, thereby leaving the target perceptually isolated. Four modulated or unmodulated precursor tones presented at the masker frequency were sufficient to completely eliminate MDI. The results support the idea that MDI is due to a perceptual grouping of the masker and target, and show that conditions promoting sufficient perceptual segregation of the masker and target can lead to a total elimination of MDI.

Adult↗

Melody lead in piano performance: expressive device or artifact?

As reported in the recent literature on piano performance, an emphasized voice (the melody) tends to be played not only louder than the other voices, but also about 30 ms earlier (melody lead). It remains unclear whether pianists deliberately apply melody lead to separate different voices, or whether it occurs because the melody is played louder (velocity artifact). The velocity artifact explanation implies that pianists initially strike the keys simultaneously; it is only different velocities that make the hammers arrive at different points in time. The measured note onsets in these studies, mostly derived from computer-monitored pianos, represent the hammer-string impact times. In the present study, the finger-key contact times are calculated and analyzed as well. If the velocity artifact hypothesis is correct, the melody lead phenomenon should disappear at the finger-key level. Chopin's Ballade op. 38 (45 measures) and Etude op. 10/3 (21 measures) were performed on a Bösendorfer computer-monitored grand piano by 22 skilled pianists. The hammer-string asynchronies among voices closely resemble the results reported in the literature. However, the melody lead decreases almost to zero at the finger-key level, which supports the velocity artifact hypothesis. In addition to this, expected onset asynchronies are predicted from differences in hammer velocity, if finger-key asynchronies are assumed to be zero. They correlate highly with the observed melody lead.

Artifacts↗

Basilar-membrane response to multicomponent stimuli in chinchilla.

The response of chinchilla basilar membrane in the basal region of the cochlea to multicomponent (1, 3, 5, 6, or 7) stimuli was studied using a laser interferometer. Three-component stimuli were amplitude-modulated signals with modulation depths that varied from 25% to 200% and the modulation frequency varied from 100 to 2000 Hz while the carrier frequency was set to the characteristic frequency of the region under study (approximately 6.3 to 9 kHz). Results indicate that, for certain modulation frequencies and depths, there is enhancement of the response. Responses to five equal-amplitude sine wave stimuli indicated the occurrence of nonlinear phenomena such as spectral edge enhancement, present when the frequency spacing was less than 200 Hz, and mutual suppression. For five-component stimuli, the first, third, or fifth component was placed at the characteristic frequency and the component frequency separation was varied over a 2-kHz range. Responses to seven component stimuli were similar to those of five-component stimuli. Six-component stimuli were generated by leaving out the center component of the seven-component stimuli. In the latter case, the center component was restored in the basilar-membrane response as a result of distortion-product generation in the nonlinear cochlea.

Animals↗

Category restructuring during second-language speech acquisition.

This study examined the production of English /b/ and the perception of short-lag English /b d g/ tokens by four groups of bilinguals who differed according to their age of arrival (AOA) in Canada from Italy and amount of self-reported native language (L1) use. A clear difference emerged between early bilinguals (mean AOA= 8 years) and late bilinguals (mean AOA= 20 years). The late bilinguals showed a stronger L1 influence than the early bilinguals did on both the production and perception of English stops. In experiment 2, the late bilinguals produced a larger percentage of prevoiced English /b/ tokens than early bilinguals and native English (NE) speakers did. In experiment 3, the late bilinguals misidentified short-lag English /b d g/ tokens as /p t k/ more often than the early bilinguals and NE speakers did. Experiment 4 revealed that the frequencies with which the bilinguals prevoiced /b d g/ in Italian and English were correlated. The observed differences between the early and late bilinguals were attributed to differences in the quantity and quality of English phonetic input they had received, not to a greater likelihood by the early than late bilinguals to establish new phonetic categories for English /b d g/.

Adult↗

A model of echolocation of multiple targets in 3D space from a single emission.

Bats, using frequency-modulated echolocation sounds, can capture a moving target in real 3D space. The process by which they are able to accomplish this, however, is not completely understood. This work offers and analyzes a model for description of one mechanism that may play a role in the echolocation process of real bats. This mechanism allows for the localization of targets in 3D space from the echoes produced by a single emission. It is impossible to locate multiple targets in 3D space by using only the delay time between an emission and the resulting echoes received at two points (i.e., two ears). To locate multiple targets in 3D space requires directional information for each target. The frequency of the spectral notch, which is the frequency corresponding to the minimum of the external ear's transfer function, provides a crucial cue for directional localization. The spectrum of the echoes from nearly equidistant targets includes spectral components of both the interference between the echoes and the interference resulting from the physical process of reception at the external ear. Thus, in order to extract the spectral component associated with the external ear, this component must first be distinguished from the spectral components associated with the interference of echoes from nearly equidistant targets. In the model presented, a computation that consists of the deconvolution of the spectrum is used to extract the external-ear-dependent component in the time domain. This model describes one mechanism that can be used to locate multiple targets in 3D space.

Animals↗

Effects of degradation of intensity, time, or frequency content on speech intelligibility for normal-hearing and hearing-impaired listeners.

Many hearing-impaired listeners suffer from distorted auditory processing capabilities. This study examines which aspects of auditory coding (i.e., intensity, time, or frequency) are distorted and how this affects speech perception. The distortion-sensitivity model is used: The effect of distorted auditory coding of a speech signal is simulated by an artificial distortion, and the sensitivity of speech intelligibility to this artificial distortion is compared for normal-hearing and hearing-impaired listeners. Stimuli (speech plus noise) are wavelet coded using a complex sinusoidal carrier with a Gaussian envelope (1/4 octave bandwidth). Intensity information is distorted by multiplying the modulus of each wavelet coefficient by a random factor. Temporal and spectral information are distorted by randomly shifting the wavelet positions along the temporal or spectral axis, respectively. Measured were (1) detection thresholds for each type of distortion, and (2) speech-reception thresholds for various degrees of distortion. For spectral distortion, hearing-impaired listeners showed increased detection thresholds and were also less sensitive to the distortion with respect to speech perception. For intensity and temporal distortion, this was not observed. Results indicate that a distorted coding of spectral information may be an important factor underlying reduced speech intelligibility for the hearing impaired.

Adult↗