Search PubMed⌕ Search

Biomedical subjects

E Vatikiotis-Bateson

Publications and source records attributed to E Vatikiotis-Bateson.

7 recordsLinked to original sources

Spatial frequency requirements for audiovisual speech perception.

Spatial frequency band-pass and low-pass filtered images of a talker were used in an audiovisual speech-in-noise task. Three experiments tested subjects' use of information contained in the different filter bands with center frequencies ranging from 2.7 to 44.1 cycles/face (c/face). Experiment 1 demonstrated that information from a broad range of spatial frequencies enhanced auditory intelligibility. The frequency bands differed in the degree of enhancement, with a peak being observed in a mid-range band (11-c/face center frequency). Experiment 2 showed that this pattern was not influenced by viewing distance and, thus, that the results are best interpreted in object spatial frequency, rather than in retinal coordinates. Experiment 3 showed that low-pass filtered images could produce a performance equivalent to that produced by unfiltered images. These experiments are consistent with the hypothesis that high spatial resolution information is not necessary for audiovisual speech perception and that a limited range of spatial frequency spectrum is sufficient.

Face↗

Multimodal contribution to speech perception revealed by independent component analysis: a single-sweep EEG case study.

In this single-sweep electroencephalographic case study, independent component analysis (ICA) was used to investigate multimodal processes underlying the enhancement of speech intelligibility in noise (for monosyllabic English words) by visualizing facial motion concordant with the audio speech signal. Wavelet analysis of the single-sweep IC activation waveforms revealed increased high-frequency energy for two ICs underlying the visual enhancement effect. For one IC, current source density analysis localized activity mainly to the superior temporal gyrus, consistent with principles of multimodal integration. For the other IC, activity was distributed across multiple cortical areas perhaps reflecting global mappings underlying the visual enhancement effect.

Adult↗

Eye movement of perceivers during audiovisual speech perception.

Perceiver eye movements were recorded during audiovisual presentations of extended monologues. Monologues were presented at different image sizes and with different levels of acoustic masking noise. Two clear targets of gaze fixation were identified, the eyes and the mouth. Regardless of image size, perceivers of both Japanese and English gazed more at the mouth as masking noise levels increased. However, even at the highest noise levels and largest image sizes, subjects gazed at the mouth only about half the time. For the eye target, perceivers typically gazed at one eye more than the other, and the tendency became stronger at higher noise levels. English perceivers displayed more variety of gaze-sequence patterns (e.g., left eye to mouth to left eye to right eye) and persisted in using them at higher noise levels than did Japanese perceivers. No segment-level correlations were found between perceiver eye motions and phoneme identity of the stimuli.

Eye Movements↗

An examination of the degrees of freedom of human jaw motion in speech and mastication.

The kinematics of human jaw movements were assessed in terms of the three orientation angles and three positions that characterize the motion of the jaw as a rigid body. The analysis focused on the identification of the jaw's independent movement dimensions, and was based on an examination of jaw motion paths that were plotted in various combinations of linear and angular coordinate frames. Overall, both behaviors were characterized by independent motion in four degrees of freedom. In general, when jaw movements were plotted to show orientation in the sagittal plane as a function of horizontal position, relatively straight paths were observed. In speech, the slopes and intercepts of these paths varied depending on the phonetic material. The vertical position of the jaw was observed to shift up or down so as to displace the overall form of the sagittal plane motion path of the jaw. Yaw movements were small but independent of pitch, and vertical and horizontal position. In mastication, the slope and intercept of the relationship between pitch and horizontal position were affected by the type of food and its size. However, the range of variation was less than that observed in speech. When vertical jaw position was plotted as a function of horizontal position, the basic form of the path of the jaw was maintained but could be shifted vertically. In general, larger bolus diameters were associated with lower jaw positions throughout the movement. The timing of pitch and yaw motion differed. The most common pattern involved changes in pitch angle during jaw opening followed by a phase predominated by lateral motion (yaw). Thus, in both behaviors there was evidence of independent motion in pitch, yaw, horizontal position, and vertical position. This is consistent with the idea that motions in these degrees of freedom are independently controlled.

Humans↗

A computational theory for movement pattern recognition based on optimal movement pattern generation.

We have previously proposed an optimal trajectory and control theory for continuous movements, such as reaching or cursive handwriting. According to Marr's three-level description of brain function, our theory can be summarized as follows: (1) The computational theory is the minimum torque-change model; (2) the intermediate representation of a pattern is given as a set of via-points extracted from an example pattern; and (3) algorithm and hardware are provided by FIRM, a neural network that can generate and control minimum torque-change trajectories. In this paper, we propose a computational theory for movement pattern recognition that is based on our theory for optimal movement pattern generation. The three levels of the description of brain function in the recognition theory are tightly coupled with those for pattern generation. In recognition, the generation process and the recognition process are actually two flows of information in opposite directions within a single functional unit. In our theory, if the input movement trajectory data are identical to the optimal movement pattern reconstructed from an intermediate representation of some symbol, the input data are recognized as that symbol. If an error exists between the movement trajectory data and the generated trajectory, the putative symbol is corrected, and the generation is repeated. In particular, we present concrete computational procedures for the recognition of connected cursive handwritten characters, as well as for the estimation of phonemic timing in natural speech. Our most important contribution is to demonstrate the computational realizability for the 'motor theory of movement pattern perception': the movement-pattern recognition process can be realized by actively recruiting the movement-pattern formation process. The way in which the formation process is utilized in pattern recognition in our theory suggests a duality between movement pattern formation and movement pattern perception.

Algorithms↗

A qualitative dynamic analysis of reiterant speech production: phase portraits, kinematics, and dynamic modeling.

The departure point of the present paper is our effort to characterize and understand the spatiotemporal structure of articulatory patterns in speech. To do so, we removed segmental variation as much as possible while retaining the spoken act's stress and prosodic structure. Subjects produced two sentences from the "rainbow passage" using reiterant speech in which normal syllables were replaced by /ba/ or /ma/. This task was performed at two self-selected rates, conversational and fast. Infrared LEDs were placed on the jaw and lips and monitored using a modified SELSPOT optical tracking system. As expected, when pauses marking major syntactic boundaries were removed, a high degree of rhythmicity within rate was observed, characterized by well-defined periodicities and small coefficients of variation. When articulatory gestures were examined geometrically on the phase plane, the trajectories revealed a scaling relation between a gesture's peak velocity and displacement. Further quantitative analysis of articulator movement as a function of stress and speaking rate was indicative of a language-modulated dynamical system with linear stiffness and equilibrium (or rest) position as key control parameters. Preliminary modeling was consonant with this dynamical perspective which, importantly, does not require that time per se be a controlled variable.

Female↗

Functionally specific articulatory cooperation following jaw perturbations during speech: evidence for coordinative structures.

In three experiments we show that articulatory patterns in response to jaw perturbations are specific to the utterance produced. In Experiments 1 and 2, an unexpected constant force load (5.88 N) applied during upward jaw motion for final /b/ closure in the utterance /baeb/ revealed nearly immediate compensation in upper and lower lips, but not the tongue, on the first perturbation trial. The same perturbation applied during the utterance /baez/ evoked rapid and increased tongue-muscle activity for /z/ frication, but no active lip compensation. Although jaw perturbation represented a threat to both utterances, no perceptible distortion of speech occurred. In Experiment 3, the phase of the jaw perturbation was varied during the production of bilabial consonants. Remote reactions in the upper lip were observed only when the jaw was perturbed during the closing phase of motion. These findings provide evidence for flexibly assembled coordinative structures in speech production.

Adult↗