Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sound Spectrography”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,423 records · Page 79Linked to original sources

SIM--simultaneous inverse filtering and matching of a glottal flow model for acoustic speech signals.

A new method "simultaneous inverse filtering and model matching" (SIM) is proposed that allows one to calculate voice source measures without any user interaction. It is based on the discrete all-pole modeling (DAP) technique for inverse filtering (IF), which is modified to include a model of the glottal flow as integral part [LF model, Fant et al., STL-QPSR (Stockholm) 4/1985, 1-13 (1986)]. As the correct LF parameters are initially unknown, they are estimated in an iterative procedure using multi-dimensional optimization techniques that are initialized according to the results of an exhaustive search. The error criteria applied reflect how well the IF is performed after the spectral contribution of the glottal flow has been removed. The resulting optimal LF parameter constellation serves as the basis to calculate 11 voice source measures. The performance was evaluated using synthesized signals and recordings of natural utterances. For the synthesized signals, the accuracy to reproduce the original parameters was high (correlations exceeding 0.88) for measures where the starting point of the glottal cycle did not enter explicitly. Errors were smaller compared to conventional estimation methods where the measures were estimated from the IF signal. The analysis of natural utterances indicates that problems still exist with regard to robustness, but that under advantageous conditions the open quotient, the speed quotient, the closing quotient, the parabolic spectral parameter, and the negative peak amplitude of the glottal flow derivative can indeed be determined automatically by the SIM method.

Fourier Analysis↗

Formant-frequency matching between sounds with different bandwidths and on different fundamental frequencies.

The two experiments described here use a formant-matching task to investigate what abstract representations of sound are available to listeners. The first experiment examines how veridically and reliably listeners can adjust the formant frequency of a single-formant sound to match the timbre of a target single-formant sound that has a different bandwidth and either the same or a different fundamental frequency (F0). Comparison with previous results [Dissard and Darwin, J. Acoust. Soc. Am. 106, 960-969 (2000)] shows that (i) for sounds on the same F0, introducing a difference in bandwidth increases the variability of matches regardless of whether the harmonics close to the formant are resolved or unresolved; (ii) for sounds on different F0's, introducing a difference in bandwidth only increases variability for sounds that have unresolved harmonics close to the formant. The second experiment shows that match variability for sounds differing in F0, but with the same bandwidth and with resolved harmonics near the formant peak, is not influenced by the harmonic spacing or by the alignment of harmonics with the formant peak. Overall, these results indicate that match variability increases when the match cannot be made on the basis of the excitation pattern, but match variability does not appear to depend on whether ideal matching performance requires simply interpolation of a spectral envelope or also the extraction of the envelope's peak frequency.

Adult↗

Sex-specific fundamental and formant frequency patterns in a cross-sectional study.

An extensive developmental acoustic study of the speech patterns of children and adults was reported by Lee and colleagues [Lee et al., J. Acoust. Soc. Am. 105, 1455-1468 (1999)]. This paper presents a reexamination of selected fundamental frequency and formant frequency data presented in their report for ten monophthongs by investigating sex-specific and developmental patterns using two different approaches. The first of these includes the investigation of age- and sex-specific formant frequency patterns in the monophthongs. The second, the investigation of fundamental frequency and formant frequency data using the critical band rate (bark) scale and a number of acoustic-phonetic dimensions of the monophthongs from an age- and sex-specific perspective. These acoustic-phonetic dimensions include: vowel spaces and distances from speaker centroids; frequency differences between the formant frequencies of males and females; vowel openness/closeness and frontness/backness; the degree of vocal effort; and formant frequency ranges. Both approaches reveal both age- and sex-specific development patterns which also appear to be dependent on whether vowels are peripheral or nonperipheral. The developmental emergence of these sex-specific differences are discussed with reference to anatomical, physiological, sociophonetic, and culturally determined factors. Some directions for further investigation into the age-linked sex differences in speech across the lifespan are also proposed.

Adolescent↗

Target spectral, dynamic spectral, and duration cues in infant perception of German vowels.

Previous studies of vowel perception have shown that adult speakers of American English and of North German identify native vowels by exploiting at least three types of acoustic information contained in consonant-vowel-consonant (CVC) syllables: target spectral information reflecting the articulatory target of the vowel, dynamic spectral information reflecting CV- and -VC coarticulation, and duration information. The present study examined the contribution of each of these three types of information to vowel perception in prelingual infants and adults using a discrimination task. Experiment 1 examined German adults' discrimination of four German vowel contrasts (see text), originally produced in /dVt/ syllables, in eight experimental conditions in which the type of vowel information was manipulated. Experiment 2 examined German-learning infants' discrimination of the same vowel contrasts using a comparable procedure. The results show that German adults and German-learning infants appear able to use either dynamic spectral information or target spectral information to discriminate contrasting vowels. With respect to duration information, the removal of this cue selectively affected the discriminability of two of the vowel contrasts for adults. However, for infants, removal of contrastive duration information had a larger effect on the discrimination of all contrasts tested.

Adult↗

Forward- and simultaneous-masked thresholds in bandlimited maskers in subjects with normal hearing and cochlear hearing loss.

Forward- and simultaneous-masked thresholds were measured at 0.5 and 2.0 kHz in bandpass maskers as a function of masker bandwidth and in a broadband masker with the goal of estimating psychophysical suppression. Suppression was operationally defined in two ways: (1) as a change in forward-masked threshold as a function of masker bandwidth, and (2) as a change in effective masker level with increased masker bandwidth, taking into account the nonlinear growth of forward masking. Subjects were younger adults with normal hearing and older adults with cochlear hearing loss. Thresholds decreased as a function of masker bandwidth in forward masking, which was attributed to effects of suppression; thresholds remained constant or increased slightly with increasing masker bandwidth in simultaneous masking. For subjects with normal hearing, slightly larger estimates of suppression were obtained at 2.0 kHz rather than at 0.5 kHz. For hearing-impaired subjects, suppression was reduced in regions of hearing loss. The magnitude of suppression was strongly correlated with the absolute threshold at the signal frequency, but did not vary with thresholds at frequencies remote from the signal. The results suggest that measuring forward-masked thresholds in bandlimited and broadband maskers may be an efficient psychophysical method for estimating suppression.

Adult↗

Psychophysical suppression measured with bandlimited noise extended below and/or above the signal: effects of age and hearing loss.

The objectives of this study were to measure suppression with bandlimited noise extended below and above the signal, at lower and higher signal frequencies, between younger and older subjects, and between subjects with normal hearing and cochlear hearing loss. Psychophysical suppression was assessed by measuring forward-masked thresholds at 0.8 and 2.0 kHz in bandlimited maskers as a function of masker bandwidth. Bandpass-masker bandwidth was increased by introducing noise components below and above the signal frequency while keeping the noise centered on the signal frequency, and also by adding noise below the signal only, and above the signal only. Subjects were younger and older adults with normal hearing and older adults with cochlear hearing loss. For all subjects, suppression was larger when noise was added below the signal than when noise was added above the signal, consistent with some physiological evidence of stronger suppression below a fiber's characteristic frequency than above. For subjects with normal hearing, suppression was greater at higher than at lower frequencies. For older subjects with hearing loss, suppression was reduced to a greater extent above the signal than below and where thresholds were elevated. Suppression for older subjects with normal hearing was poorer than would be predicted from their absolute thresholds, suggesting that age may have contributed to reduced suppression or that suppression was sensitive to changes in cochlear function that did not result in significant threshold elevation.

Adult↗

Exploring the temporal mechanism involved in the pitch of unresolved harmonics.

This paper continues a line of research initiated by Kaernbach and Demany [J. Acoust. Soc. Am. 104, 2298-2306 (1998)], who employed filtered click sequences to explore the temporal mechanism involved in the pitch of unresolved harmonics. In a first experiment, the just noticeable difference (jnd) for the fundamental frequency (F0) of high-pass filtered and low-pass masked click trains was measured, with F0 (100 to 250 Hz) and the cut frequency (0.5 to 6 kHz) being varied orthogonally. The data confirm the result of Houtsma and Smurzynski [J. Acoust. Soc. Am. 87, 304-310 (1990)] that a pitch mechanism working on the temporal structure of the signal is responsible for analyzing frequencies higher than ten times the fundamental. Using high-pass filtered click trains, however, the jnd for the temporal analysis is at 1.2% as compared to 2%-3% found in studies using band-pass filtered stimuli. Two further experiments provide evidence that the pitch of this stimulus can convey musical information. A fourth experiment replicates the finding of Kaernbach and Demany on first- and second-order regularities with a cut frequency of 2 kHz and extends the paradigm to binaural aperiodic click sequences. The result suggests that listeners can detect first-order temporal regularities in monaural click streams as well as in binaurally fused click streams.

Adult↗

Speech recognition in noise as a function of the number of spectral channels: comparison of acoustic hearing and cochlear implants.

Speech recognition was measured as a function of spectral resolution (number of spectral channels) and speech-to-noise ratio in normal-hearing (NH) and cochlear-implant (CI) listeners. Vowel, consonant, word, and sentence recognition were measured in five normal-hearing listeners, ten listeners with the Nucleus-22 cochlear implant, and nine listeners with the Advanced Bionics Clarion cochlear implant. Recognition was measured as a function of the number of spectral channels (noise bands or electrodes) at signal-to-noise ratios of + 15, + 10, +5, 0 dB, and in quiet. Performance with three different speech processing strategies (SPEAK, CIS, and SAS) was similar across all conditions, and improved as the number of electrodes increased (up to seven or eight) for all conditions. For all noise levels, vowel and consonant recognition with the SPEAK speech processor did not improve with more than seven electrodes, while for normal-hearing listeners, performance continued to increase up to at least 20 channels. Speech recognition on more difficult speech materials (word and sentence recognition) showed a marginally significant increase in Nucleus-22 listeners from seven to ten electrodes. The average implant score on all processing strategies was poorer than scores of NH listeners with similar processing. However, the best CI scores were similar to the normal-hearing scores for that condition (up to seven channels). CI listeners with the highest performance level increased in performance as the number of electrodes increased up to seven, while CI listeners with low levels of speech recognition did not increase in performance as the number of electrodes was increased beyond four. These results quantify the effect of number of spectral channels on speech recognition in noise and demonstrate that most CI subjects are not able to fully utilize the spectral information provided by the number of electrodes used in their implant.

Adult↗

Second-order temporal modulation transfer functions.

Detection thresholds were measured for a sinusoidal modulation applied to the modulation depth of a sinusoidally amplitude-modulated (SAM) white noise carrier as a function of the frequency of the modulation applied to the modulation depth (referred to as f'm). The SAM noise acted therefore as a "carrier" stimulus of frequency fm, and sinusoidal modulation of the SAM-noise modulation depth generated two additional components in the modulation spectrum: fm-f'm and fm+f'm. The tracking variable was the modulation depth of the sinusoidal variation applied to the "carrier" modulation depth. The resulting "second-order" temporal modulation transfer functions (TMTFs) measured on four listeners for "carrier" modulation frequencies fm of 16, 64, and 256 Hz display a low-pass segment followed by a plateau. This indicates that sensitivity to fluctuations in the strength of amplitude modulation is best for fluctuation rates f'm below about 2-4 Hz when using broadband noise carriers. Measurements of masked modulation detection thresholds for the lower and upper modulation sideband suggest that this capacity is possibly related to the detection of a beat in the sound's temporal envelope. The results appear qualitatively consistent with the predictions of an envelope detector model consisting of a low-pass filtering stage followed by a decision stage. Unlike listeners' performance, a modulation filterbank model using Q values > or = 2 should predict that second-order modulation detection thresholds should decrease at high values of f'm due to the spectral resolution of the modulation sidebands (in the modulation domain). This suggests that, if such modulation filters do exist, their selectivity is poor. In the latter case, the Q value of modulation filters would have to be less than 2. This estimate of modulation filter selectivity is consistent with the results of a previous study using a modulation-masking paradigm [S. D. Ewert and T. Dau, J. Acoust. Soc. Am. 108, 1181-1196 (2000)].

Adult↗

Binaural processing model based on contralateral inhibition. II. Dependence on spectral parameters.

This and two accompanying articles [Breebaart et al., J. Acoust. Soc. Am. 110, 1074-1088 (2001); 110, 1105-1117 (2001)] describe a computational model for the signal processing in the binaural auditory system. The model consists of several stages of monaural and binaural preprocessing combined with an optimal detector. In the present article the model is tested and validated by comparing its predictions with experimental data for binaural discrimination and masking conditions as a function of the spectral parameters of both masker and signal. For this purpose, the model is used as an artificial observer in a three-interval, forced-choice adaptive procedure. All model parameters were kept constant for all simulations described in this and the subsequent article. The effects of the following experimental parameters were investigated: center frequency of both masker and target, bandwidth of masker and target, the interaural phase relations of masker and target, and the level of the masker. Several phenomena that occur in binaural listening conditions can be accounted for. These include the wider effective binaural critical bandwidth observed in band-widening NoS(pi) conditions, the different masker-level dependence of binaural detection thresholds for narrow- and for wide-band maskers, the unification of IID and ITD sensitivity with binaural detection data, and the dependence of binaural thresholds on frequency.

Auditory Threshold↗

Binaural processing model based on contralateral inhibition. III. Dependence on temporal parameters.

This paper and two accompanying papers [Breebaart et al., J. Acoust. Soc. Am. 110, 1074-1088 (2001); 110, 1089-1104 (2001)] describe a computational model for the signal processing of the binaural auditory system. The model consists of several stages of monaural and binaural preprocessing combined with an optimal detector. Simulations of binaural masking experiments were performed as a function of temporal stimulus parameters and compared to psychophysical data adapted from literature. For this purpose, the model was used as an artificial observer in a three-interval, forced-choice procedure. All model parameters were kept constant for all simulations. Model predictions were obtained as a function of the interaural correlation of a masking noise and as a function of both masker and signal duration. Furthermore, maskers with a time-varying interaural correlation were used. Predictions were also obtained for stimuli with time-varying interaural time or intensity differences. Finally, binaural forward-masking conditions were simulated. The results show that the combination of a temporal integrator followed by an optimal detector in the time domain can account for all conditions that were tested, except for those using periodically varying interaural time differences (ITDs) and those measuring interaural correlation just-noticeable differences (jnd's) as a function of bandwidth.

Auditory Threshold↗

On the effectiveness of whole spectral shape for vowel perception.

The formant hypothesis of vowel perception, where the lowest two or three formant frequencies are essential cues for vowel quality perception, is widely accepted. There has, however, been some controversy suggesting that formant frequencies are not sufficient and that the whole spectral shape is necessary for perception. Three psychophysical experiments were performed to study this question. In the first experiment, the first or second formant peak of stimuli was suppressed as much as possible while still maintaining the original spectral shape. The responses to these stimuli were not radically different from the ones for the unsuppressed control. In the second experiment, F2-suppressed stimuli, whose amplitude ratios of high- to low-frequency components were systemically changed, were used. The results indicate that the ratio changes can affect perceived vowel quality, especially its place of articulation. In the third experiment, the full-formant stimuli, whose amplitude ratios were changed from the original and whose F2's were kept constant, were used. The results suggest that the amplitude ratio is equal to or more effective than F2 as a cue for place of articulation. We conclude that formant frequencies are not exclusive cues and that the whole spectral shape can be crucial for vowel perception.

Humans↗

Consonant identification under maskers with sinusoidal modulation: masking release or modulation interference?

The present study investigated the effect of envelope modulations in a background masker on consonant recognition by normal hearing listeners. It is well known that listeners understand speech better under a temporally modulated masker than under a steady masker at the same level, due to masking release. The possibility of an opposite phenomenon, modulation interference, whereby speech recognition could be degraded by a modulated masker due to interference with auditory processing of the speech envelope, was hypothesized and tested under various speech and masker conditions. It was of interest whether modulation interference for speech perception, if it were observed, could be predicted by modulation masking, as found in psychoacoustic studies using nonspeech stimuli. Results revealed that masking release measurably occurred under a variety of conditions, especially when the speech signal maintained a high degree of redundancy across several frequency bands. Modulation interference was also clearly observed under several circumstances when the speech signal did not contain a high redundancy. However, the effect of modulation interference did not follow the expected pattern from psychoacoustic modulation masking results. In conclusion, (1) both factors, modulation interference and masking release, should be accounted for whenever a background masker contains temporal fluctuations, and (2) caution needs to be taken when psychoacoustic theory on modulation masking is applied to speech recognition.

Adult↗

Temporal modulation transfer functions obtained using sinusoidal carriers with normally hearing and hearing-impaired listeners.

Temporal modulation transfer functions were obtained using sinusoidal carriers for four normally hearing subjects and three subjects with mild to moderate cochlear hearing loss. Carrier frequencies were 1000, 2000 and 5000 Hz, and modulation frequencies ranged from 10 to 640 Hz in one-octave steps. The normally hearing subjects were tested using levels of 30 and 80 dB SPL. For the higher level, modulation detection thresholds varied only slightly with modulation frequency for frequencies up to 80 Hz, but decreased for high modulation frequencies. The decrease can be attributed to the detection of spectral sidebands. For the lower level, thresholds varied little with modulation frequency for all three carrier frequencies. The absence of a decrease in the threshold for large modulation frequencies can be explained by the low sensation level of the spectral sidebands. The hearing-impaired subjects were tested at 80 dB SPL, except for two cases where the absolute threshold at the carrier frequency was greater than 70 dB SPL; in these cases a level of 90 dB was used. The results were consistent with the idea that spectral sidebands were less detectable for the hearing-impaired than for the normally hearing subjects. For the two lower carrier frequencies, there were no large decreases in threshold with increasing modulation frequency, and where decreases did occur, this happened only between 320 and 640 Hz. For the 5000-Hz carrier, thresholds were roughly constant for modulation frequencies from 10 to 80 or 160 Hz, and then increased monotonically, becoming unmeasurable at 640 Hz. The results for this carrier may reflect "pure" effects of temporal resolution, without any influence from the detection of spectral sidebands. The results suggest that temporal resolution for deterministic stimuli is similar for normally hearing and hearing-impaired listeners.

Adult↗

Auditory detection of hollowness.

The airborne sounds produced by freely vibrating hollow and solid bars were synthesized according to the equations of bar motion from theoretical acoustics, and were presented to listeners over headphones. In a two-interval, forced-choice task, listeners were asked to distinguish between the hollow and solid bar sounds as bar length was varied at random from one presentation to the next. All other physical properties of the bar were held constant across trials. Listener decision strategies for detecting hollowness in iron, aluminum, and wood bars were determined from regression weights describing the relation between the listener's response and the frequency, intensity, and decay modulus of the individual partials comprising these sounds. The obtained weights were compared to those of a hypothetical listener that bases judgments on the acoustic relations intrinsic to hollowness, as determined from the equations for motion. Results indicate that listeners adopt roughly one of two decision strategies, either basing judgments on the appropriate acoustic relations, or basing judgments predominantly on frequency alone. The decision strategy of some listeners also changed from one type to the other with a change in bar material or upon replication of the same condition. The results are interpreted in terms of the vulnerability of the intrinsic acoustic relations to small perturbations in acoustic parameters, as would be associated with listener internal noise. They demonstrate that basic limits of human sensitivity can have a profound effect on the identification of rudimentary source attributes from sound, even in conditions where acoustic variation is largely dictated by physical variation in the source.

Auditory Perception↗

Spatial unmasking of nearby speech sources in a simulated anechoic environment.

Spatial unmasking of speech has traditionally been studied with target and masker at the same, relatively large distance. The present study investigated spatial unmasking for configurations in which the simulated sources varied in azimuth and could be either near or far from the head. Target sentences and speech-shaped noise maskers were simulated over headphones using head-related transfer functions derived from a spherical-head model. Speech reception thresholds were measured adaptively, varying target level while keeping the masker level constant at the "better" ear. Results demonstrate that small positional changes can result in very large changes in speech intelligibility when sources are near the listener as a result of large changes in the overall level of the stimuli reaching the ears. In addition, the difference in the target-to-masker ratios at the two ears can be substantially larger for nearby sources than for relatively distant sources. Predictions from an existing model of binaural speech intelligibility are in good agreement with results from all conditions comparable to those that have been tested previously. However, small but important deviations between the measured and predicted results are observed for other spatial configurations, suggesting that current theories do not accurately account for speech intelligibility for some of the novel spatial configurations tested.

Adult↗

Asymmetry of masking between noise and iterated rippled noise: evidence for time-interval processing in the auditory system.

This study describes the masking asymmetry between noise and iterated rippled noise (IRN) as a function of spectral region and the IRN delay. Masking asymmetry refers to the fact that noise masks IRN much more effectively than IRN masks noise, even when the stimuli occupy the same spectral region. Detection thresholds for IRN masked by noise and for noise masked by IRN were measured with an adaptive two-alternative, forced choice (2AFC) procedure with signal level as the adaptive parameter. Masker level was randomly varied within a 10-dB range in order to reduce the salience of loudness as a cue for detection. The stimuli were filtered into frequency bands, 2.2-kHz wide, with lower cutoff frequencies ranging from 0.8 to 6.4 kHz. IRN was generated with 16 iterations and with varying delays. The reciprocal of the delay was 16, 32, 64, or 128 Hz. When the reciprocal of the IRN delay was within the pitch range, i.e., above 30 Hz, there was a substantial masking asymmetry between IRN and noise for all filter cutoff frequencies; threshold for IRN masked by noise was about 10 dB larger than threshold for noise masked by IRN. For the 16-Hz IRN, the masking asymmetry decreased progressively with increasing filter cutoff frequency, from about 9 dB for the lowest cutoff frequency to less than 1 dB for the highest cutoff frequency. This suggests that masking asymmetry may be determined by different cues for delays within and below the pitch range. The fact that masking asymmetry exists for conditions that combine very long IRN delays with very high filter cutoff frequencies means that it is unlikely that models based on the excitation patterns of the stimuli would be successful in explaining the threshold data. A range of time-domain models of auditory processing that focus on the time intervals in phase-locked neural activity patterns is reviewed. Most of these models were successful in accounting for the basic masking asymmetry between IRN and noise for conditions within the pitch range, and one of the models produced an exceptionally good fit to the data.

Adult↗