Search PubMedSearch

Biomedical subjects

K R Kluender

Publications and source records attributed to K R Kluender.

18 recordsLinked to original sources

Depolarizing the perceptual magnet effect.

In recent years there has been a great deal of interest in demonstrations of the so-called "Perceptual-Magnet Effect" (PME). In these studies, AX-discrimination tasks purportedly reveal that discriminability of speech sounds from a single category varies with judged phonetic "goodness" of the sounds. However, one possible confound is that category membership is determined by identification of sounds in isolation, whereas, discrimination tasks include pairs of stimuli. In the first experiment of the current study, identification and goodness judgments were obtained for vowels (/i/-/e/) presented in pairs. A substantial shift in phonetic identity was evidenced with changes in the context vowel. In a second experiment, listeners participated in an AX-discrimination task with the vowel pairs from the first experiment. Using the contextual identification functions from the first experiment, predictions of discriminability were calculated using the classic tenets of Categorical Perception. Obtained discriminability functions were well accounted for by predictions from identification. There was no additional unexplained variance that required the proposal of "perceptual magnets." These results suggest that PME may be nothing more than further demonstration that general discriminability is greater for cross-category stimulus pairs than for within-category pairs.

Humans

Role of experience for language-specific functional mappings of vowel sounds.

Studies involving human infants and monkeys suggest that experience plays a critical role in modifying how subjects respond to vowel sounds between and within phonemic classes. Experiments with human listeners were conducted to establish appropriate stimulus materials. Then, eight European starlings (Sturnus vulgaris) were trained to respond differentially to vowel tokens drawn from stylized distributions for the English vowels /i/ and /I/, or from two distributions of vowel sounds that were orthogonal in the F1-F2 plane. Following training, starlings' responses generalized with facility to novel stimuli drawn from these distributions. Responses could be predicted well on the bases of frequencies of the first two formants and distributional characteristics of experienced vowel sounds with a graded structure about the central "prototypical" vowel of the training distributions. Starling responses corresponded closely to adult human judgments of "goodness" for English vowel sounds. Finally, a simple linear association network model trained with vowels drawn from the avian training set provided a good account for the data. Findings suggest that little more than sensitivity to statistical regularities of language input (probability-density distributions) together with organizational processes that serve to enhance distinctiveness may accommodate much of what is known about the functional equivalence of vowel sounds.

Adult

General contrast effects in speech perception: effect of preceding liquid on stop consonant identification.

When members of a series of synthesized stop consonants varying acoustically in F3 characteristics and varying perceptually from /da/ to /ga/ are preceded by /al/, subjects report hearing more /ga/ syllables relative to when each member is preceded by /ar/ (Mann, 1980). It has been suggested that this result demonstrates the existence of a mechanism that compensates for coarticulation via tacit knowledge of articulatory dynamics and constraints, or through perceptual recovery of vocal-tract dynamics. The present study was designed to assess the degree to which these perceptual effects are specific to qualities of human articulatory sources. In three experiments, series of consonant-vowel (CV) stimuli varying in F3-onset frequency (/da/-/ga/) were preceded by speech versions or nonspeech analogues of /al/ and /ar/. The effect of liquid identity on stop consonant labeling remained when the preceding VC was produced by a female speaker and the CV syllable was modeled after a male speaker's productions. Labeling boundaries also shifted when the CV was preceded by a sine wave glide modeled after F3 characteristics of /al/ and /ar/. Identifications shifted even when the preceding sine wave was of constant frequency equal to the offset frequency of F3 from a natural production. These results suggest an explanation in terms of general auditory processes as opposed to recovery of or knowledge of specific articulatory dynamics.

Adult

Perceptual compensation for coarticulation by Japanese quail (Coturnix coturnix japonica).

When members of a series of synthesized stop consonants varying in third-formant (F3) characteristics and varying perceptually from /da/ to /ga/ are preceded by /al/, human listeners report hearing more /ga/ syllables than when the members of the series are preceded by /ar/. It has been suggested that this shift in identification is the result of specialized processes that compensate for acoustic consequences of coarticulation. To test the species-specificity of this perceptual phenomenon, data were collected from nonhuman animals in a syllable "labeling" task. Four Japanese quail (Coturnix coturnix japonica) were trained to peck a key differentially to identify clear /da/ and /ga/ exemplars. After training, ambiguous members of a /da/-/ga/ series were presented in the context of /al/ and /ar/ syllables. Pecking performance demonstrated a shift which coincided with data from humans. These results suggest that processes underlying "perceptual compensation for coarticulation" are species-general. In addition, the pattern of response behavior expressed is rather common across perceptual systems.

Animals

Effect of voice quality on perceived height of English vowels.

Across a variety of languages, phonation type and vocal-tract shape systematically covary in vowel production. Breathy phonation tends to accompany vowels produced with a raised tongue body and/or advanced tongue root. A potential explanation for this regularity, based on a hypothesized interaction between the acoustic effects of vocal-tract shape and phonation type, is evaluated. It is suggested that increased spectral tilt and first-harmonic amplitude resulting from breathy phonation interact with the lower-frequency first formant resulting from a raised tongue body to produce a perceptually 'higher' vowel. To test this hypothesis, breathy and modal versions of vowel series modelled after male and female productions of English vowel pairs /i/ and /i/, /u/ and /[symbol: see text]/, and /lamda/ and /a/ were synthesized. Results indicate that for most cases, breathy voice quality led to more tokens being identified as the higher vowel (i.e. /i/, /u/, /lamda/). In addition, the effect of voice quality is greater for vowels modelled after female productions. These results are consistent with a hypothesized perceptual explanation for the covariation of phonation type and tongue-root advancement in West African languages. The findings may also be relevant to gender differences in phonation type.

Female

Spectral discontinuities and the vowel length effect.

Perception of voicing for stop consonants in consonant-vowel syllables can be affected by the duration of the following vowel so that longer vowels lead to more "voiced" responses. On the basis of several experiments, Green, Stevens, and Kuhl (1994) concluded that continuity of fundamental frequency (f0), but not continuity of formant structure, determined the effective length of the following vowel. In an extension of those efforts, we found here that both effects were critically dependent on particular f0s and formant values. First, discontinuity in f0 does not necessarily preclude the vowel length effect because the effect maintains when f0 changes from 200 to 100 Hz, and 200-Hz partials extend continuously through test syllables. Second, spectral discontinuity does preclude the vowel length effect when formant changes result in a spectral peak shifting to another harmonic. The results indicate that the effectiveness of stimulus changes for sustaining or diminishing the vowel length effect depends critically on particulars of spectral composition.

Adult

Perception of voicing for syllable-initial stops at different intensities: does synchrony capture signal voiceless stop consonants?

In response to stop consonants with longer F1-cutback duration, the dominant synchronization of mid- and high-CF chinchilla auditory-nerve fibers changes from frequencies near F2 to frequencies near F1 at onset of voicing [D. G. Sinex and L. P. McDonald, J. Acoust. Soc. Am. 85, 1995-2004 (1989)]. If this change in neural synchronization is perceptually relevant for human listeners, then it may be predicted that changes in stimulus intensity and changes in the frequency difference between lower (F1) and higher (F2/F3) stimulus components should both affect perception of voicing. In a series of experiments, multiple continua of synthesized CVs varying in F1 cutback of the consonantal portion were played to listeners at levels ranging from 40 to 80 dB SPL. Across experiments, the frequency difference between F1 and F2 was manipulated by changing the onset frequency of F1 or F2. Subjects labeled more initial stops as voiceless as a function of increasing stimulus level and of decreasing frequency difference between F1 and F2. There was also an interaction between stimulus intensity and the frequency difference between F1 and F2 such that the effect of intensity was greater for smaller differences. These effects were reliable across a number of synthetic F1-cutback series, and the effect of intensity extended to a digitally edited series of hybrid CVs in which F1-cutback was varied by cross splicing naturally produced /da/ and /ta/. The effect of overall stimulus intensity was not affected by amplitude of prevocalic aspiration energy or by the presence or absence of release bursts. The results provide evidence for the perceptual significance of synchrony encoding of voicing for stop consonants.

Humans

Effects of first formant onset frequency on [-voice] judgments result from auditory processes not specific to humans.

When F1-onset frequency is lower, longer F1 cut-back (VOT) is required for human listeners to perceive synthesized stop consonants as voiceless. K. R. Kluender [J. Acoust. Soc. Am. 90, 83-96 (1991)] found comparable effects of F1-onset frequency on the "labeling" of stop consonants by Japanese quail (coturnix coturnix japonica) trained to distinguish stop consonants varying in F1 cut-back. In that study, CVs were synthesized with natural-like rising F1 transitions, and endpoint training stimuli differed in the onset frequency of F1 because a longer cut-back resulted in a higher F1 onset. In order to assess whether earlier results were due to auditory predispositions or due to animals having learned the natural covariance between F1 cut-back and F1-onset frequency, the present experiment was conducted with synthetic continua having either a relatively low (375 Hz) or high (750 Hz) constant-frequency F1. Six birds were trained to respond differentially to endpoint stimuli from three series of synthesized /CV/s varying in duration of F1 cut-back. Second and third formant transitions were appropriate for labial, alveolar, or velar stops. Despite the fact that there was no opportunity for animal subjects to use experienced covariation of F1-onset frequency and F1 cut-back, quail typically exhibited shorter labeling boundaries (more voiceless stops) for intermediate stimuli of the continua when F1 frequency was higher. Responses by human subjects listening to the same stimuli were also collected. Results lend support to the earlier conclusion that part or all of the effect of F1 onset frequency on perception of voicing may be adequately explained by general auditory processes.(ABSTRACT TRUNCATED AT 250 WORDS)

Animals

Amplitude rise time and the perception of the voiceless affricate/fricative distinction.

Variation of amplitude envelope at stimulus onset has been considered to be of primary importance for distinguishing voiceless affricates from fricatives (e.g., [symbol: see text]). In earlier perceptual experiments, however, variation in amplitude rise time was confounded with variation in frication duration. In two experiments, these variables were independently manipulated, and their individual and combined effects for perception of magnitude of [symbol: see text] were examined. Variation in amplitude rise time alone was not sufficient to signal the voiceless affricate/fricative contrast in these experiments, but variation in frication duration alone was sufficient.

Humans

Effects of glide slope, noise intensity, and noise duration on the extrapolation of FM glides through noise.

Listeners are quite adept at maintaining integrated perceptual events in environments that are frequently noisy. Three experiments were conducted to assess the mechanisms by which listeners maintain continuity for upward sinusoidal glides that are interrupted by a period of broadband noise. The first two experiments used stimulus complexes consisting of three parts: prenoise glide, broadband noise interval, and postnoise glide. For a given prenoise glide and noise interval, the subject's task was to adjust the onset frequency of a same-slope postnoise glide so that, together with the prenoise glide and noise, the complex sounded as "smooth and continuous as possible." The slope of the glides (1.67, 3.33, 5, and 6.67 Bark/sec) as well as the duration (50, 200, and 350 msec) and relative level of the interrupting noise (0, -6, and -12 dB S/N) were varied. For all but the shallowest glides, subjects consistently adjusted the offset portion of the glide to frequencies lower than predicted by accurate interpolation of the prenoise portion. Curiously, for the shallowest glides, subjects consistently selected postnoise glide onset-frequency values higher than predicted by accurate extrapolation of the prenoise glide. There was no effect of noise level on subjects' adjustments in the first two experiments. The third experiment used a signal detection task to measure the phenomenal experience of continuity through the noise. Frequency glides were either present or absent during the noise for stimuli like those use in the first two experiments as well as for stimuli that had no prenoise or postnoise glides. Subjects were more likely to report the presence of glides in the noise when none occurred (false positives) when noise was shorter or of greater relative level and when glides were present adjacent to the noise.

Adult

On the interpretability of speech/nonspeech comparisons: a reply to Fowler.

Fowler [J. Acoust. Soc. Am. 88, 1236-1249 (1990)] makes a set of claims on the basis of which she denies the general interpretability of experiments that compare the perception of speech sounds to the perception of acoustically analogous nonspeech sound. She also challenges a specific auditory hypothesis offered by Diehl and Walsh [J. Acoust. Soc. Am. 85, 2154-2164 (1989)] to explain the stimulus-length effect in the perception of stops and glides. It will be argued that her conclusions are unwarranted.

Auditory Perception

A composite model of the auditory periphery for the processing of speech based on the filter response functions of single auditory-nerve fibers.

A composite model of the auditory periphery, based upon a unique analysis technique for deriving filter response characteristics from cat auditory-nerve fibers, is presented. The model is distinctive in its ability to capture a significant broadening of auditory-nerve fiber frequency selectivity as a function of increasing sound-pressure level within a computationally tractable time-invariant structure. The output of the model shows the tonotopic distribution of synchrony activity of single fibers in response to the steady-state vowel [e] presented over a 40-dB range of sound-pressure levels and is compared with the population-response data of Young and Sachs (1979). The model, while limited by its time invariance, accurately captures most of the place-synchrony response patterns reported by the Johns Hopkins group. In both the physiology and in the model, auditory-nerve fibers spanning a broad tonotopic range synchronize to the first formant (F1), with the proportion of units phase-locked to F1 increasing appreciably at moderate to high sound-pressure levels. A smaller proportion of fibers maintain phase locking to the second and third formants across the same intensity range. At sound-pressure levels of 60 dB and above, the vast majority of fibers with characteristic frequencies greater than 3 kHz synchronize to F1 (512 Hz), rather than to frequencies in the most sensitive portion of their response range. On the basis of these response patterns it is suggested that neural synchrony is the dominant auditory-nerve representation of formant information under "normal" listening conditions in which speech signals occur across a wide range of intensities and against a background of unpredictable and frequently intense acoustic interference.

Animals

Effects of first formant onset properties on voicing judgments result from processes not specific to humans.

Both first formant (F1) transition duration and F1 onset frequency have been proposed to be perceptually significant in categorization of voiced and voiceless syllable-initial stops. Transition duration per se may not, however, explain the fact that, for longer transitions, longer F1 cutback is required in order to perceive a stop as voiceless. Longer transitions result in lower F1 onsets at any duration of cutback greater than zero, and it is possible that the major effect of F1 is determined by its frequency at onset. In this study, F1-transition duration, onset frequency, and slope were varied across four types of F1 transition in which one of the three variables (onset frequency, duration, slope) was held constant while the other two were allowed to vary. Each of the four F1 types was used in syllables with higher formants appropriate for labial, alveolar, and velar places of articulation. By far, the best predictor of identification of these stimuli by human listeners was F1 onset frequency. F1 duration, F1 slope, and place of articulation had little or no effect on labeling boundaries. In a second experiment using Japanese quail (Coturnix coturnix japonica), birds were trained to respond differentially to voiced versus voiceless stops. The differential effects of F1 onset frequency on the "labeling" behavior of these birds was strikingly similar to that of humans listening to the same stimuli. These results are taken to provide strong evidence that F1 onset frequency is the primary determinant of shifts in voicing boundaries across place of articulation, and that general mechanisms not unique to humans appear adequate to account for the effects of F1 onset frequency on perception of voicing for syllable-initial stops.

Adult

Japanese quail can learn phonetic categories.

Japanese quail (Coturnix coturnix) learned a category for syllable-initial [d] followed by a dozen different vowels. After learning to categorize syllables consisting of [d], [b], or [g] followed by four different vowels, quail correctly categorized syllables in which the same consonants preceded eight novel vowels. Acoustic analysis of the categorized syllables revealed no single feature or pattern of features that could support generalization, suggesting that the quail adopted a more complex mapping of stimuli into categories. These results challenge theories of speech sound classification that posit uniquely human capacities.

Animals

Are selective adaptation and contrast effects really distinct?

Although there is evidence that selective adaptation and contrast effects in speech perception are produced by the same mechanisms, Sawusch and Jusczyk (1981) reported a dissociation between the effects and concluded that adaptation and contrast occur at separate processing levels. They found that an ambiguous test stimulus was more likely to be labeled b following adaptation with [pha] and more likely to be labeled p following adaptation with [ba] or [spa] (the latter consisting of [ba] preceded by [s] noise). In the contrast session, where a single context stimulus occurred with a single test item, the [ba] and [pha] contexts had contrastive effects similar to those of the [ba] and [pha] adaptors, but the [spa] context produced an increase in b responses to the test stimulus, an effect opposite to that of the [spa] adaptor. One interpretation of this difference is that the rapid presentation of the [spa] adaptor gave rise to "streaming," whereby the [s] was perceptually segregated from the [ba]. In our experiment, we essentially replicated the results of Sawusch and Jusczyk (1981), using procedures similar to theirs. Next, we increased the interadaptor interval to remove the likelihood of stream segregation and found that the adaptation and contrast effects converged.

Adaptation, Physiological

A possible auditory basis for internal structure of phonetic categories.

We used a selective adaptation procedure to investigate the possibility that differences in the degree to which stimuli within a phonetic category are considered to be good exemplars of the category--that is, differences in perceived category goodness--have a basis at a prephonetic, auditory level of processing. For three different phonetic contrasts (/b-p/, /d-g/, /b-w/), we assessed the relative magnitude of adaptation along a stimulus continuum produced by a variety of stimuli from the continuum belonging to a given phonetic category. For all three phonetic contrasts, nonmonotonic adaptation functions were obtained: As the adaptor moved away from the category boundary, there was an initial increase in adaptation, followed by a subsequent decrease. On the assumption that selective adaptation taps a prephonetic, auditory level of processing, these findings permit the following conclusions. First, at an auditory level there is a limit on the range of stimuli along a continuum that is treated as relevant to a given contrast; that is, the stimuli along a continuum are effectively grouped into auditory categories. Second, stimuli within an auditory category vary in their effectiveness as category members, providing an internal structure to the categories. Finally, this internal category structure at the auditory level, revealed by the adaptation procedure, may provide a basis for differences in perceived category goodness at the phonetic level.

Adaptation, Physiological