Search PubMedSearch

Biomedical subjects

B S Tjan

Publications and source records attributed to B S Tjan.

4 recordsLinked to original sources

The viewpoint complexity of an object-recognition task.

There is an ongoing debate about the nature of perceptual representation in human object recognition. Resolution of this debate has been hampered by the lack of a metric for assessing the representational requirements of a recognition task. To recognize a member of a given set of 3-D objects, how much detail must the objects' representations contain in order to achieve a specific accuracy criterion? From the performance of an ideal observer, we derived a quantity called the view complexity (VX) to measure the required granularity of representation. VX is an intrinsic property of the object-recognition task, taking into account both the object ensemble and the type of decision required of an observer. It does not depend on the visual representation or processing used by the observer. VX can be interpreted as the number of randomly selected 2-D images needed to represent the decision boundaries in the image space of a 3-D object-recognition task. A low VX means the task is inherently more viewpoint invariant and a high VX means it is inherently more viewpoint dependent. By measuring the VX of recognition tasks with different object sets, we show that the current confusion about the nature of human perceptual representation is partly due to a failure in distinguishing between human visual processing and the properties of a task and its stimuli. We find general correspondence between the VX of a recognition task and the published human data on viewpoint dependence. Exceptions in this relationship motivated us to propose the view-rate hypothesis: human visual performance is limited by the equivalent number of 2-D image views that can be processed per unit time.

Algorithms

Mr. Chips: an ideal-observer model of reading.

The integration of visual, lexical, and oculomotor information is a critical part of reading. Mr. Chips is an ideal-observer model that combines these sources of information optimally to read simple texts in the minimum number of saccades. In the model, the concept of the visual span (the number of letters that can be identified in a single fixation) plays a key, unifying role. The behavior of the model provides a computational framework for reexamining the literature on human reading saccades. Emergent properties of the model, such as regressive saccades and an optimal-viewing position, suggest new interpretations of human behavior. Because Mr. Chip's "retina" can have any (one-dimensional) arrangement of high-resolution regions and scotomas, the model can simulate common visual disorders. Surprising saccade strategies are linked to the pattern of scotomas. For example, Mr. Chips sometimes plans a saccade that places a decisive letter in a scotoma. This article provides the first quantitative model of the effects of scotomas on reading.

Algorithms

Human efficiency for recognizing 3-D objects in luminance noise.

The purpose of this study was to establish how efficiently humans use visual information to recognize simple 3-D objects. The stimuli were computer-rendered images of four simple 3-D objects--wedge, cone, cylinder, and pyramid--each rendered from 8 randomly chosen viewing positions as shaded objects, line drawings, or silhouettes. The objects were presented in static, 2-D Gaussian luminance noise. The observer's task was to indicate which of the four objects had been presented. We obtained human contrast thresholds for recognition, and compared these to an ideal observer's thresholds to obtain efficiencies. In two auxiliary experiments, we measured efficiencies for object detection and letter recognition. Our results showed that human object-recognition efficiency is low (3-8%) when compared to efficiencies reported for some other visual-information processing tasks. The low efficiency means that human recognition performance is limited primarily by factors intrinsic to the observer rather than the information content of the stimuli. We found three factors that play a large role in accounting for low object-recognition efficiency: stimulus size, spatial uncertainty, and detection efficiency. Four other factors play a smaller role in limiting object-recognition efficiency: observers' internal noise, stimulus rendering condition, stimulus familiarity, and categorization across views.

Contrast Sensitivity

Human efficiency for recognizing and detecting low-pass filtered objects.

Recently, Tjan, Braje, Legge and Kersten [(1995) Vision Research, 35, 3053-3069] found that human efficiency for object recognition was less than 10%, indicating that humans fail to use much of the information available to an ideal observer. We examine two explanations for these low efficiencies: (1) humans are inefficient in using high spatial-frequency information; and (2) humans are inefficient in detecting image samples. We tested the first possibility by measuring human efficiency for recognizing low-pass filtered objects, rendered as line drawings and silhouettes, in luminance noise. Efficiency did not improve when high frequencies were removed, and the first explanation was rejected. We tested the second explanation by comparing efficiencies for object detection and recognition. Recognition efficiency was higher than detection efficiency for silhouettes but not line drawings, showing that detection efficiency does not place a ceiling on recognition efficiency. The results indicate that human vision is designed to extract image features, such as contours, that enhance recognition. A computer simulation suggests that this can occur if the observer views the world through a band-pass spatial-frequency channel.

Algorithms