Search PubMed⌕ Search

PubMed · 9304693

Model-based interpretation of complex and variable images.

Abstract

The ultimate goal of machine vision is image understanding-the ability not only to recover image structure but also to know what it represents. By definition, this involves the use of models which describe and label the expected structure of the world. Over the past decade, model-based vision has been applied successfully to images of man-made objects. It has proved much more difficult to develop model-based approaches to the interpretation of images of complex and variable structures such as faces or the internal organs of the human body (as visualized in medical images). In such cases it has been problematic even to recover image structure reliably, without a model to organize the often noisy and incomplete image evidence. The key problem is that of variability. To be useful, a model needs to be specific-that is, to be capable of representing only 'legal' examples of the modelled object(s). It has proved difficult to achieve this whilst allowing for natural variability. Recent developments have overcome this problem; it has been shown that specific patterns of variability in shape and grey-level appearance can be captured by statistical models that can be used directly in image interpretation. The details of the approach are outlined and practical examples from medical image interpretation and face recognition are used to illustrate how previously intractable problems can now be tackled successfully. It is also interesting to ask whether these results provide any possible insights into natural vision; for example, we show that the apparent changes in shape which result from viewing three-dimensional objects from different viewpoints can be modelled quite well in two dimensions; this may lend some support to the 'characteristic views' model of natural vision.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

C J Taylor, T F Cootes, A Lanitis, G Edwards, P Smyth, A C Kotcheff. 1997-08-29. Model-based interpretation of complex and variable images.. https://doi.org/10.1098/rstb.1997.0109

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Absence of flash-lag when judging global shape from local positions.

When a flash is presented aligned with a moving stimulus, the former is perceived to lag behind the latter (the flash-lag effect). We study whether this mislocalization occurs when a positional judgment is not required, but a veridical spatial relationship between moving and flashed stimuli is needed to perceive a global shape. To do this, we used Glass patterns that are formed by pairs of correlated dots. One dot of each pair was presented moving and, at a given moment, the other dot of each pair was flashed in order to build the Glass pattern. If a flash-lag effect occurs between each pair of dots, we expect the best perception of the global shape to occur when the flashed dots are presented before the moving dots arrive at the position that physically builds the Glass pattern. Contrary to this, we found that the best detection of Glass patterns occurred for the situation of physical alignment. This result is not consistent with a low-level contribution to the flash-lag effect.

Form Perception↗

Object recognition and segmentation by a fragment-based hierarchy.

How do we learn to recognize visual categories, such as dogs and cats? Somehow, the brain uses limited variable examples to extract the essential characteristics of new visual categories. Here, I describe an approach to category learning and recognition that is based on recent computational advances. In this approach, objects are represented by a hierarchy of fragments that are extracted during learning from observed examples. The fragments are class-specific features and are selected to deliver a high amount of information for categorization. The same fragments hierarchy is then used for general categorization, individual object recognition and object-parts identification. Recognition is also combined with object segmentation, using stored fragments, to provide a top-down process that delineates object boundaries in complex cluttered scenes. The approach is computationally effective and provides a possible framework for categorization, recognition and segmentation in human vision.

Form Perception↗

Dynamics of shape interaction in human vision.

Spatial context can alter perceived shape, and temporal context can influence the perception of a stimulus. We sought to determine the time course of shape interactions by using a paradigm in which closed shape contours are laterally displaced over space and time. Target and masks are separated by various stimulus onset asynchrony (SOA) values, yielding forward, backward, and simultaneous masking conditions. Results indicate that spatial lateral interactions of shape are amplified by temporal asynchrony, reaching a peak at SOAs of 80-110 ms. Mask amplitude scales all effects and masking is shape specific. When a single mask follows the target, both spatial configuration and mask onset transient are critical in determining depth of masking. When the target is followed by two sequential masks, the possibility of apparent motion determines whether one or both masks drive masking. These findings suggest that temporal interactions of shape are dependent on an interactive combination of shape specificity and transients, that apparent motion plays a modulatory role, and that target shape is determined after a temporal window, not at its onset.

Form Perception↗