Search PubMedSearch

Biomedical subjects

T M Nearey

Publications and source records attributed to T M Nearey.

6 recordsLinked to original sources

Speech perception as pattern recognition.

This work provides theoretical and empirical arguments in favor of an approach to phonetics that is called double-weak. It is so called because it assumes relatively weak constraints both on the articulatory gestures and on the auditory patterns that map phonological elements. This approach views speech production and perception as distinct but cooperative systems. Like the motor theory of speech perception, double-weak theory accepts that phonological units are modified by context in ways that are important to perception. It further agrees that many aspects of such context dependency have their origin in natural articulatory processes. However, double-weak theory sides with proponents of auditory theories of phonetics by accepting that the real-time objects of perception are well-defined auditory patterns. Because speakers find ways to "orderly" output conditions" (Sussman et al., 1995), listeners are able to successfully decode speech using relatively simple pattern-recognition mechanisms. It is suggested that this situation has arisen through a stylization of gestural patterns to accommodate real-time limits of the perceptual system. Results from a new perceptual experiment, involving a four-dimensional stimulus continuum and a 10-category/hVC/response set, are shown to be largely compatible with this framework.

Attention

On the sufficiency of compound target specification of isolated vowels and vowels in /bVb/ syllables.

It has been suggested [e.g., Strange et al., J. Acoust. Soc. Am. 74, 695-705 (1983); Verbrugge and Rakerd, Language Speech 29, 39-57 (1986)] that the temporal margins of vowels in consonantal contexts, consisting mainly of the rapid CV and VC transitions of CVC's, contain dynamic cues to vowel identity that are not available in isolated vowels and that may be perceptually superior in some circumstances to cues which are inherent to the vowels proper. However, this study shows that vowel-inherent formant targets and cues to vowel-inherent spectral change (measured from nucleus to offglide sections of the vowel itself) persist in the margins of /bVb/ syllables, confirming a hypothesis of Nearey and Assmann [J. Acoust. Soc. Am. 80, 1297-1308 (1986)]. Experiments were conducted to test whether listeners might be using such vowel-inherent, rather than coarticulatory information to identify the vowels. In the first experiment, perceptual tests using "hybrid silent center" syllables (i.e., syllables which contain only brief initial and final portions of the original syllable, and in which speaker identity changes from the initial to the final portion) show that listeners' error rates and confusion matrices for vowels in /bVb/ syllables are very similar to those for isolated vowels. These results suggest that listeners are using essentially the same type of information in essentially the same way to identify both kinds of stimuli. Statistical pattern recognition models confirm the relative robustness of nucleus and vocalic offglide cues and can predict reasonably well listeners' error patterns in all experimental conditions, though performance for /bVb/ syllables is somewhat worse than for isolated vowels. The second experiment involves the use of simplified synthetic stimuli, lacking consonantal transitions, which are shown to provide information that is nearly equivalent phonetically to that of the natural silent center /bVb/ syllables (from which the target measurements were extracted). Although no conclusions are drawn about other contexts, for speakers of Western Canadian English coarticulatory cues appear to play at best a minor role in the perception of vowels in /bVb/ context, while vowel-inherent factors dominate listeners' perception.

Female

Static, dynamic, and relational properties in vowel perception.

The present work reviews theories and empirical findings, including results from two new experiments, that bear on the perception of English vowels, with an emphasis on the comparison of data analytic "machine recognition" approaches with results from speech perception experiments. Two major sources of variability (viz., speaker differences and consonantal context effects) are addressed from the classical perspective of overlap between vowel categories in F1 x F2 space. Various approaches to the reduction of this overlap are evaluated. Two types of speaker normalization are considered. "Intrinsic" methods based on relationships among the steady-state properties (F0, F1, F2, and F3) within individual vowel tokens are contrasted with "extrinsic" methods, involving the relationships among the formant frequencies of the entire vowel system of a single speaker. Evidence from a new experiment supports Ainsworth's (1975) conclusion [W. Ainsworth, Auditory Analysis and Perception of Speech (Academic, London, 1975)] that both types of information have a role to play in perception. The effects of consonantal context on formant overlap are also considered. A new experiment is presented that extends Lindblom and Studdert-Kennedy's finding [B. Lindblom and M. Studdert-Kennedy, J. Acoust. Soc. Am. 43, 840-843 (1967)] of perceptual effects of consonantal context on vowel perception to /dVd/ and /bVb/ contexts. Finally, the role of vowel-inherent dynamic properties, including duration and diphthongization, is briefly reviewed. All of the above factors are shown to have reliable influences on vowel perception, although the relative weight of such effects and the circumstances that alter these weights remain far from clear. It is suggested that the design of more complex perceptual experiments, together with the development of quantitative pattern recognition models of human vowel perception, will be necessary to resolve these issues.

Humans

Perception of front vowels: the role of harmonics in the first formant region.

Vowel matching and identification experiments were carried out to investigate the perceptual contribution of harmonics in the first formant region of synthetic front vowels. In the first experiment, listeners selected the best phonetic match from an F1 continuum, for reference stimuli in which a band of two to five adjacent harmonics of equal intensity replaced the F1 peak; F1 values of best matches were near the frequency of the highest frequency harmonic in the band. Attenuation of the highest harmonic in the band resulted in lower F1 matches. Attenuation of the lowest harmonic had no significant effects, except in the case of a 2-harmonic band, where higher F1 matches were selected. A second experiment investigated the shifts in matched F1 resulting from an intensity increment to either one of a pair of harmonics in the F1 region. These shifts were relatively invariant over different harmonic frequencies and proportional to the fundamental frequency. A third experiment used a vowel identification task to determine phoneme boundaries on an F1 continuum. These boundaries were not substantially altered when the stimuli comprised only the two most prominent harmonics in the F1 region, or these plus either the higher or lower frequency subset of the remaining F1 harmonics. The results are consistent with an estimation procedure for the F1 peak which assigns greatest weight to the two most prominent harmonics in the first formant region.

Female

Vowel identification: orthographic, perceptual, and acoustic aspects.

This study investigates conditions under which vowels are well recognized and relates perceptual identification of individual tokens to acoustic characteristics. Results support recent finding that isolated vowels may be readily identified by listeners. Two experiments provided evidence that certain response tasks result in inflated error rates. Subsequent experiments showed improved identification in a fixed speaker context, compared with randomized speakers, for isolated vowels and gated centers. Performance was worse for gated vowels, suggesting that dynamic properties (such as duration and diphthongization) supplement steady-state cues. However, even-speaker-randomized gated vowels were well identified (14% errors). Measures of "steady-state information" (formant frequencies and f0), "dynamic information" (formant slopes and duration), and "speaker information" (normalization) were adopted. Discriminant analyses of acoustic measurements indicated relatively little overlap between vowel categories. Using a new technique for relating acoustic measurements of individual tokens with identification by listeners, it is shown that (a) identification performance is clearly related to acoustic characteristics; (b) improvement in the fixed speaker context is correlated with improved statistical separation resulting from formant normalization, for the gated vowels; and (c) "dynamic information" is related to identification differences between full and gated isolated vowels.

Adolescent

Context effects in a double-weak theory of speech perception.

The present study provides an elaboration of the "double-weak" theory of speech perception proposed by Nearey (1990, 1991). In this framework, the objects of speech perception (and production) are viewed as neither primarily auditory nor primarily gestural; rather, they are abstract, symbolic elements lawfully constrained to map onto relatively simple (but not entirely transparent) patterns in both domains. Speech cannot be understood unless both articulation and acoustics are considered: Many production strategies appear to be directed at achieving acoustically-oriented goals, yet most context effects in speech perception seem to be motivated by the consequences of gestural overlap. The double-weak framework suggests that speech perception and speech production are less-than-perfect inverses of each other. Despite long-term accommodation of each for the demands of the other, real-time production and perception may operate as autonomous subsystems. A family of perceptual models is discussed that provides varying degrees of approximation to "ideal solutions" (in the sense of minimizing error rate) to classifying production data exhibiting contextual interactions. Members of this family that provide substantial, yet incomplete, compensation for the consequences of gestural overlap appear to be adequate to account for the results of many speech perception experiments. Such partial perceptual compensation allows, in principle, for the kind of imperfect "error correction" discussed by Ohala (1981, 1990) in conjunction with hypo- and hypercorrection phenomena.

Female