The data problem for color objectivism.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to D D Hoffman.
Explore the source record for details and available documents.
Color-from-motion displays consist of a sparse array of dots which never move but change color according to various algorithms. Yet such displays can trigger human vision to construct apparent motion of a subjective surface which is uniformly colored and bounded by a subjective contour. We show that the perceptual strength of this construction depends on the density and regularity of dot placement. We studied three objective measures of density and regularity: nearest-neighbor distance, mean of maximal disks, and variance of maximal disks. We found that nearest-neighbor mechanisms alone are inadequate to account for the perceptual strength of the subjective surfaces and contours. Mechanisms sensitive to areal gaps provide a more adequate account.
Visual images are ambiguous. Any image, or collection of images, is consistent with an infinite number of possible scenes in the world. Yet we are generally unaware of this ambiguity. During ordinary perception we are generally aware of only one, or perhaps a few of these possibilities. Human vision evidently exploits certain constraints--assumptions about the world and images formed of it--in order to generate its perceptions. One constraint that has been widely studied by researchers in human and machine vision is the generic-viewpoint assumption. We show that this assumption can help to explain the widely discussed fact that outlines of blobs are ineffective inducers of illusory contours. We also present a number of novel effects and report an experiment suggesting that the generic-viewpoint assumption strongly influences illusory-contour perception.
Many researchers have proposed that, for the purpose of recognition, human vision parses shapes into component parts. Precisely how is not yet known. The minima rule for silhouettes (Hoffman & Richards, 1984) defines boundary points at which to parse but does not tell how to use these points to cut silhouettes and, therefore, does not tell what the parts are. In this paper, we propose the short-cut rule, which states that, other things being equal, human vision prefers to use the shortest possible cuts to parse silhouettes. We motivate this rule, and the well-known Petter's rule for modal completion, by the principle of transversality. We present five psychophysical experiments that test the short-cut rule, show that it successfully predicts part cuts that connect boundary points given by the minima rule, and show that it can also create new boundary points.
Visual completion is a ubiquitous phenomenon: Human vision often constructs contours and surfaces in regions that have no sharp gradients in any image property. When does human vision interpolate a contour between a given pair of luminance-defined edges? Two different answers have been proposed: relatability and minimizing inflections. We state and prove a proposition that links these two proposals by showing that, under appropriate conditions, relatability is mathematically equivalent to the existence of a smooth curve with no inflection points that interpolates between the two edges. The proposition thus provides a set of necessary and sufficient conditions for two edges to be relatable. On the basis of these conditions, we suggest a way to extend the definition of relatability (1) to include the role of genericity, and (2) to extend the current all-or-none character of relatability to a graded measure that can track the gradedness in psychophysical data.
Many objects have component parts, and these parts often differ in their visual salience. In this paper we present a theory of part salience. The theory builds on the minima rule for defining part boundaries. According to this rule, human vision defines part boundaries at negative minima of curvature on silhouettes, and along negative minima of the principal curvatures on surfaces. We propose that the salience of a part depends on (at least) three factors: its size relative to the whole object, the degree to which it protrudes, and the strength of its boundaries. We present evidence that these factors influence visual processes which determine the choice of figure and ground. We give quantitative definitions for the factors, visual demonstrations of their effects, and results of psychophysical experiments.
'Color from motion' describes the perception of a spread of subjective color over achromatic regions seen as moving. The effect can be produced in a display of multiple frames shown in quick succession, each frame consisting of a fixed, random placement of colored dots on a high-luminance white background with color assignments of some dots, but not dot locations, changing from frame to frame. Evidence is presented that the perception of apparent motion and the spread of subjective color can be activated by binocular combination of disjoint signals to each eye. The dichoptic presentation of every odd-numbered frame of the full stimulus sequence presented to one eye and, out of phase, every even-numbered frame to the other eye produces a compelling perception of color from motion equal to that seen with the full sequence presented to each eye alone. This is consistent with the idea that color from motion is regulated in sites at or beyond the convergence of monocular pathways. When the background field in the stimulus display is of low luminance, an amodally complete object, fully colored and matching the dots defining the moving region in hue and saturation, is seen to move behind a partially occluding screen. Observers do not perceive such an object in still view. Hence, color from motion can be used by the visual system to produce amodal completion, which suggests that it may play a role in enhancing the visibility of camouflaged objects.
We introduce and explore a color phenomenon which requires the prior perception of motion to produce a spread of color over a region defined by motion. We call this motion-induced spread of color dynamic color spreading. The perception of dynamic color spreading is yoked to the perception of apparent motion: As the ratings of perceived motion increase, the ratings of color spreading increase. The effect is most pronounced if the region defined by motion is near 1 degree of visual angle. As the luminance contrast between the region defined by motion and the surround changes, perceived saturation of color spreading changes while perceived hue remains roughly constant. Dynamic color spreading is sometimes, but not always, bounded by a subjective contour. We discuss these findings in terms of interactions between color and motion pathways.
The ability of subjects to detect whether a structure-from-motion display depicts one or two rigid objects was examined in the presence or the absence of noise points. Each object was composed of a set of points chosen randomly within the volume of a sphere. The objects rotated rigidly about different axes passing through the center of the sphere. For displays without noise points, detection increased with larger angles between the rotation axes and with more points in each object. For displays in which noise points were present, detection was above chance but, in general, worse than that for displays without noise points. The implications of these results for image segmentation in complex motion patterns is discussed.
Interpolation across orientation discontinuities in simulated three-dimensional (3-D) surfaces was studied in three experiments with the use of structure-from-motion (SFM) displays. The displays depicted dots on two slanted planes with a region devoid of dots (a gap) between them. If extended through the gap at constant slope, the planes would meet at a dihedral edge. Subjects were required to place an SFM probe dot, located within the gap, on the perceived surface. Probe dot placements indicated that subjects perceived a smooth surface connecting the planes rather than a surface with a discontinuity. Probe dot placements varied with slope of the planes, density of the dots, and gap size, but not with orientation (horizontal or vertical) of the dihedral edge or of the axis of rotation. Smoothing was consistent with models of 2-D interpolation proposed by Ullman (1976) and Kellman and Shipley (1991) and with a model of 3-D interpolation proposed by Grimson (1981).
Five experiments were conducted to examine constraints used to interpret structure-from-motion displays. Theoretically, two orthographic views of four or more points in rigid motion yield a one-parameter family of rigid three-dimensional (3-D) interpretations. Additional views yield a unique rigid interpretation. Subjects viewed two-view and thirty-view displays of five-point objects in apparent motion. The subjects selected the best 3-D interpretation from a set of 89 compatible alternatives (experiments 1-3) or judged depth directly (experiment 4). In both cases the judged depth increased when relative image motion increased, even when the increased motion was due to increased simulation rotation. Subjects also judged rotation to be greater when either simulated depth or simulated rotation increased (experiment 4). The results are consistent with a heuristic analysis in which perceived depth is determined by relative motion.
We investigated surface interpolation in displays of structure from motion (SFM). To do so, we introduced a new method for measuring surface perception in dynamic displays--the SFM probe. An SFM probe is a dot that moves rigidly with the dots on a simulated surface, and whose distance from that surface can be adjusted with a joystick or similar control. The displays we studied were random-dot cylinders containing a vertical strip devoid of feature points (the gap). Subjects adjusted an SFM probe, presented in the gap, until the probe dot appeared to be on the surface. Variability in probe-dot placement decreased with increasing texture density on the cylinder and increased with increasing gap width. Subjects showed a consistent bias to place the probe dot outside the cylinder. This bias increased with increasing texture density for the SFM displays. (The opposite bias was found in a static two-dimensional interpolation task with an arc whose curvature matched that of the cylinder: Subjects placed the probe dot inside the arc.) This outside bias is inconsistent with several theoretical approaches to surface interpolation.
Perceptual scientists have recently enjoyed success in constructing mathematical theories for specific perceptual capacities, capacities such as stereovision, auditory localization, and color perception. Analysis of these theories suggests that they all share a common mathematical structure. If this is true, the elucidation of this structure, the study of its properties, the derivation of its consequences, and the empirical testing of its predictions are promising directions for perceptual research. We consider a candidate for the common structure, a candidate called an "observer". Observers, in essence, perform inferences; each observer has a characteristic class of perceptual premises, a characteristic class of perceptual conclusions, and its own functional relationship between these premises and conclusions. If observers indeed capture the structure common to perceptual capacities, then each capacity, regardless of its modality or manner of instantiation, can be described as some observer. In this paper we develop the definition of an observer. We first consider two examples of perceptual capacities: the measurement of visual motion, and the perception of depth from visual motion. In each case, we review a formal theory of the capacity and abstract its structural essence. From this essence we construct the definition of observer. We then exercise the definition in discussions of transduction, perceptual illusions, perceptual uncertainty, regularization theory, the cognitive penetrability of perception, and the theory neutrality of observation.
Theoretical investigations of structure from motion have demonstrated that an ideal observer can discriminate rigid from nonrigid motion from two views of as few as four points. We report three experiments that demonstrate similar abilities in human observers: In one experiment, 4 of 6 subjects made this discrimination from two views of four points; the remaining subjects required five points. Accuracy in discriminating rigid from nonrigid motion depended on the amount of nonrigidity (variance of the interpoint distances over views) in the nonrigid structure. The ability to detect a rigid group dropped sharply as noise points (points not part of the rigid group) were added to the display. We conclude that human observers do extremely well in discriminating between nonrigid and fully rigid motion, but that they do quite poorly at segregating points in a display on the basis of rigidity.
Three experiments were conducted to test Hoffman and Richards's (1984) hypothesis that, for purposes of visual recognition, the human visual system divides three-dimensional shapes into parts at negative minima of curvature. In the first two experiments, subjects observed a simulated object (surface of revolution) rotating about a vertical axis, followed by a display of four alternative parts. They were asked to select a part that was from the object. Two of the four parts were divided at negative minima of curvature and two at positive maxima. When both a minima part and a maxima part from the object were presented on each trial (experiment 1), most of the correct responses were minima parts (101 versus 55). When only one part from the object--either a minima part or a maxima part--was shown on each trial (experiment 2), accuracy on trials with correct minima parts and correct maxima parts did not differ significantly. However, some subjects indicated that they reversed figure and ground, thereby changing maxima parts into minima parts. In experiment 3, subjects marked apparent part boundaries. 81% of these marks indicated minima parts, 10% of the marks indicated maxima parts, and 9% of the marks were at other positions. These results provide converging evidence, from two different methods, which supports Hoffman and Richard's minima rule.
We study the inference of rigid three-dimensional interpretations for the structure and motion of four or more moving points from but two orthographic views of the points. We develop an algorithm to determine whether image data are compatible with a rigid interpretation. As a corollary of this result we find that the measure of false targets (roughly, nonrigid objects that appear rigid) is zero. We find that if the two views have at least one rigid interpretation, then in fact there is a canonical one-parameter family of rigid interpretations; we show how to compute this family, and we describe precisely how the rigid interpretations vary within it. Since only two views are used, this analysis is relevant also to stereo vision.
Mathematical analyses of motion perception have established minimum combinations of points and distinct views that are sufficient to recover three-dimensional (3D) structure from two-dimensional (2D) images, using such regularities as rigid motion, fixed axis of rotation, and constant angular velocity. To determine whether human subjects could recover 3D information at these theoretical levels, we presented subjects with pairs of displays and asked them to determine whether they represented the same or different 3D structures. Number of points was varied between two and five; number of views was varied between two and six; and the motion was fixed axis with constant angular velocity, fixed axis with variable velocity, or variable axis with variable velocity. Accuracy increased with views, decreased with points, and was greater with fixed-axis motion. Subjects performed above chance levels even when motion was eliminated, indicating that they exploited regularities in addition to those in the theoretical analyses.
We show that four orthographic projections of two rigidly linked points are compatible with at most four interpretations of the relative three-dimensional positions of the points if the points rotate about a fixed axis--even when the points as a system undergo arbitrary rigid translations. A fifth view (projection) yields a unique interpretation and makes zero the probability that randomly chosen image points will receive a three-dimensional interpretation. Assuming that the points rotate at a constant angular velocity, instead of adding a fifth view, also yields a unique interpretation and makes zero the probability that randomly chosen image points will receive a three-dimensional interpretation.