Search PubMed⌕ Search

Biomedical subjects

Q Summerfield

Publications and source records attributed to Q Summerfield.

32 records · Page 2Linked to original sources

Auditory enhancement of changes in spectral amplitude.

An auditory enhancement effect occurs when one component of a harmonic series is omitted for a few hundred milliseconds and then reintroduced: The reintroduced harmonic stands out perceptually. Three experiments are reported that studied a version of this effect in which several components of a harmonic series are enhanced to define the formants of a vowel. Using the accuracy of vowel identification to measure the prominence of the formant peaks in the effective auditory representation, forms of the effect were identified that are qualitatively similar to the incremental and decremental responses seen in primary auditory-nerve fibers. These results are compatible with an origin for the enhancement effect in peripheral auditory adaptation. However, an additional mechanism is required to account for the demonstration [Viemeister and Bacon, J. Acoust. Soc. Am. 71, 1502-1507 (1982)] that enhancement can involve a true gain in the frequency region of the reintroduced component. These effects demonstrate one way in which the auditory system may attenuate the prominence of background noises while preserving the ability to represent changes in spectral amplitude produced by newly arriving signals.

Adult↗

Minimum spectral contrast for vowel identification by normal-hearing and hearing-impaired listeners.

To determine the minimum difference in amplitude between spectral peaks and troughs sufficient for vowel identification by normal-hearing and hearing-impaired listeners, four vowel-like complex sounds were created by summing the first 30 harmonics of a 100-Hz tone. The amplitudes of all harmonics were equal, except for two consecutive harmonics located at each of three "formant" locations. The amplitudes of these harmonics were equal and ranged from 1-8 dB more than the remaining components. Normal-hearing listeners achieved greater than 75% accuracy when peak-to-trough differences were 1-2 dB. Normal-hearing listeners who were tested in a noise background sufficient to raise their thresholds to the level of a flat, moderate hearing loss needed a 4-dB difference for identification. Listeners with a moderate, flat hearing loss required a 6- to 7-dB difference for identification. The results suggest, for normal-hearing listeners, that the peak-to-trough amplitude difference required for identification of this set of vowels is very near the threshold for detection of a change in the amplitude spectrum of a complex signal. Hearing-impaired listeners may have difficulty using closely spaced formants for vowel identification due to abnormal smoothing of the internal representation of the spectrum by broadened auditory filters.

Adult↗

Quantifying the contribution of vision to speech perception in noise.

The intelligibility of sentences presented in noise improves when the listener can view the talker's face. Our aims were to quantify this benefit, and to relate it to individual differences among subjects in lipreading ability and among sentences in lipreading difficulty. Auditory and audiovisual speech-reception thresholds (SRTs) were measured in 20 listeners with normal hearing. Sixty sentences, selected to range in the difficulty with which they could be lipread (with vision alone) from easy to hard, were presented for identification in white noise. Using the ascending method of limits, the SRT was defined as the lowest signal-to-noise ratio at which all three 'key words' in each sentence could be identified correctly. Measured as the difference in dB between auditory-alone and audiovisual SRTs, 'audiovisual benefit' averaged 11 dB, ranging from 6 to 15 dB among subjects, and from 3 to 22 dB among sentences. As predicted, audiovisual benefit is a measure of lipreading ability. It was highly correlated with visual-alone performance (n = 20, r = 0.86, P less than 0.01). Likewise, those sentences which were easiest to lipread gave a higher measure of benefit from vision in audiovisual conditions than did sentences that were hard to lipread (n = 60, r = 0.92, P less than 0.01). The results establish the basis of an efficient test of speech-reception disability in which measures are freed from the floor and ceiling effects encountered when percentage correct is used as the dependent variable.

Adult↗

Intermodal timing relations and audio-visual speech recognition by normal-hearing adults.

Audio-visual identification of sentences was measured as a function of audio delay in untrained observers with normal hearing; the soundtrack was replaced by rectangular pulses originally synchronized to the closing of the talker's vocal folds and then subjected to delay. When the soundtrack was delayed by 160 ms, identification scores were no better than when no acoustical information at all was provided. Delays of up to 80 ms had little effect on group-mean performance, but a separate analysis of a subgroup of better lipreaders showed a significant trend of reduced scores with increased delay in the range from 0-80 ms. A second experiment tested the interpretation that, although the main disruptive effect of the delay occurred on a syllabic time scale, better lipreaders might be attempting to use intermodal timing cues at a phonemic level. Normal-hearing observers determined whether a 120-Hz complex tone started before or after the opening of a pair of liplike Lissajou figures. Group-mean difference limens (70.7% correct DLs) were - 79 ms (sound leading) and + 138 ms (sound lagging), with no significant correlation between DLs and sentence lipreading scores. It was concluded that most observers, whether good lipreaders or not, possess insufficient sensitivity to intermodal timing cues in audio-visual speech for them to be used analogously to voice onset time in auditory speech perception. The results of both experiments imply that delays of up to about 40 ms introduced by signal-processing algorithms in aids to lipreading should not materially affect audio-visual speech understanding.

Adult↗

Differences between spectral dependencies in auditory and phonetic temporal processing: Relevance to the perception of voicing in initial stops.

Untrained listeners can reliably judge the temporal order of the onset of (a) pairs of coterminous tones [forming tone-onset-time (TOT) continua], and (b) higher-frequency bandlimited noises and lower-frequency bandlimited pulse trains [forming noise-onset-time (NOT) continua], but only if the onset of the second sound lags the first by at least 15-20 ms. It has been argued that the limitation of auditory temporal-order resolution that gives rise to this threshold also underlies the distinction between voiced [b, d, g] and voiceless aspirated [ph, th, kh] syllable-initial stop constants [which can be expressed in differences of voice-onset-time. (VOT)]. The positions of boundaries between phonetic categories on VOT continua depend on the values of a variety of spectral parameters, including the onset frequency of the first formant; lowering this results in boundaries shifting to longer values of VOT. The present experiment demonstrated that analogous spectral manipulations applied to the members of TOT and NOT continua do not result in systematic shifts in the location of the simultaneity-successivity threshold. The result suggest that the role of F1 in the perception of voicing does not have a purely auditory basis, a conclusion compatible with certain development and cross-language studies that have demonstrated that sensitivity to F1 is acquired and language dependent. The threshold may determine ranges of VOT between which auditory contrast is heightened, and so have helped to shape the preferred phonetic forms of phonological distinctions in the world's languages. However, other factors, such a production constraints or arbitrary processes of cultural development, appear to be required to account for the positions of voicing boundaries in particular languages.

Auditory Perception↗

Articulatory rate and perceptual constancy in phonetic perception.

The perception of syllable-initial stop consonants as voiced (/b/, /d/, /g/) or voiceless (/p/, /t/, /k/) was shown to depend on the prevailing rate of articulation. Reducing the articulatory rate of a precursor phrase causes a greater proportion of test consonants to be identified as voiced. Subsequent experiments demonstrated that this effect depends almost entirely on variation in the duration of the syllable immediately preceding the test syllable; this, the duration of the intervening silent stop closure, and the duration of the test syllable itself all influenced the identification of the stop as voiced or voiceless. Variation in the tempo of a nonspeech melody produced no effect on the perception of embedded test syllables. Those manipulations which produce the major part of the influence of rate do so not by changing the context in which the stop is perceived, but rather by changing temporal concomitants of the constriction, occlusion, and release phases of the articulation of the stop itself. For this reason, an explanation for such effects based on extrinsic timing in perception is found to be wanting. Timing should, in the main, be regarded as intrinsic to the acoustical specifications of phonetic events, a view that is compatible with recent reformulations of the problem of timing control in speech production.

Adolescent↗

Information in speech: observations on the perception of [s]-stop clusters.

A series of experiments is reported that investigated the pattern of acoustic information specifying place and manner of stop consonants in medial position after [s]. In both production and perception, information for stop place includes the spectrum of the fricative at offset, the duration of the silent closure interval, the spectral relationship between the frequency of the stop release burst and the following periodically excited formants, and the spectral and temporal characteristics of the first formant transition. Similarly, the information for stop manner includes the duration of silent closure, the frequency of the first formant at the release, the magnitude of the first formant transition, and the proximity of the second and third formants at release. A relationship was shown to exist in perception between the spectral characteristics of the first formant and the duration of the silent closure required to hear a stop. This appears to reciprocate the covariation of these parameters in production across different places of articulation and different vocalic contexts. The existence of perceptual sensitivity to a wide range of the acoustic consequences of production questions the efficacy of accounts of speech perception in terms of the fractionation of the signal into elemental acoustic cues, which are then integrated to yield a phonetic percept. It is argued that it is inappropriate to ascribe a psychological status to cues whose only reality is their operational role as physical parameters whose manipulation can change the phenotic interpretation of a signal. It is suggested that the metric of the information for phonetic perception cannot be that of the cues; rather, a metric should be sought in which acoustic and articulatory dynamics are isomorphic.

Humans↗

Identification of synthetic /bdg/ by hearing-impaired listeners under monotic and dichotic formant presentation.

Individuals with sensorineural hearing losses of both flat and sloping configuration evidence difficulty in identifying stop consonant place of articulation. To assess whether upward spread of masking is responsible for this difficulty, we presented hearing-impaired listeners with stimuli from a /ba da ga/ continuum in both monotic and dichotic (F1 to one ear; F2/F3 to the other ear) listening conditions. In the monotic conditions, listeners with similar audiograms evidence great variability in identification performance. In the dichotic conditions performance did not generally improve. For a few listeners, however, the improvement was striking. At moderate levels of signal presentation, upward spread of masking does not appear to be responsible for the poor identification of place by the majority of listeners with moderate hearing losses.

Adolescent↗