Search PubMed⌕ Search

Biomedical subjects

M S Landy

Publications and source records attributed to M S Landy.

At least 19 recordsLinked to original sources

Long range interactions between oriented texture elements.

Long range interactions between texture elements (short, oriented line segments) were examined. Specifically, we studied the influence of a background array of texture elements on the detectability of a target element (separated from the background by an intermediate textured region) using textures like those of Caputo (Vis. Res. 1996, 36, 2815-2826). We found that, in general, when the background elements were oriented orthogonally to the target element, detection of the target element was better than when the background elements had the same orientation as the target element. We discuss these interactions in terms of inhibitory and excitatory connections between orientation and spatial frequency selective linear filters (e.g. filters which mimic V1 simple cells) which would respond to the individual texture elements.

Choice Behavior↗

Interaction between the perceived shape of two objects.

The difference between the way in which binocular disparity scales with viewing distance and the way in which motion parallax scales with viewing distance introduces a potential indirect cue for viewing distance: the viewing distance is the only distance at which disparity and motion specify the same depth. The present study examines whether this information is used. Two simulated ellipsoids were presented on a computer screen in complete darkness. The two ellipsoids were 6 degrees to the left and right of straight ahead. Subjects set the width and depth of each ellipsoid to match a tennis ball, and set the distance of the one on the right to half that of the one on the left. The distance of the left ellipsoid varied between trials. On half of the trials it was static. On the other half it was rotating up and down around its frontal horizontal axis. Rotating the left ellipsoid influenced its set depth: rotating ellipsoids were set to be much more spherical. There was no influence on the set depth of the other ellipsoid, or on the set width of either. The set distance of the right ellipsoid was also unaffected. We conclude that subjects do not combine binocular disparity and motion parallax to obtain more veridical information about viewing distance.

Distance Perception↗

Examining edge- and region-based texture analysis mechanisms.

Instantaneous texture discrimination performance was examined for different texture stimuli to uncover the use of edge-based and region-based texture analysis mechanisms. Textures were composed of randomly placed, short, oriented line segments. Line segment orientation was chosen randomly using a Gaussian distribution (described by a mean and a standard deviation). One such distribution determined the orientations on the left side of the image, and a second distribution was used for the right side. The two textures either abutted to form an edge or were separated by a blank region. A texture difference in mean orientation led to superior discrimination performance when the textures abutted. On the other hand, when the textures differed in the standard deviation of the orientation distribution, performance was similar in the two conditions. These results suggest that edge-based texture analysis mechanisms were used (i.e. were the most sensitive) in the abutting difference-in-mean case, but region-based texture analysis mechanisms were used in the other three cases.

Differential Threshold↗

Observer biases in the 3D interpretation of line drawings.

Line drawings produced by contours traced on a surface can produce a vivid impression of the surface shape. The stability of this perception is notable considering that the information provided by the surface contours is quite ambiguous. We have studied the stability of line drawing perception from psychophysical and computational standpoints. For a given family of simple line drawings, human observers could perceive the drawings as depicting either an elliptic (egg-shaped) or hyperbolic (saddle-shaped) smooth surface patch. Rotation of the image along the line of sight and change in aspect ratio of the line drawing could bias the observer toward either interpretation. The results were modeled by a simple Bayesian observer that computes the probability to choose either interpretation given the information in the image and prior preferences. The model's decision rule is noncommitting: for a given input image its responses are still probabilistic, reflecting variability in the modeled observers' judgements. A good fit to the data was obtained when three observer assumptions were introduced: a preference for convex surfaces, a preference for surface contours aligned with the principal lines of curvature, and a preference for a surface orientation consistent with an object viewed from above. We discuss how these assumptions might reflect regularities of the visual world.

Adult↗

Measurement and modeling of depth cue combination: in defense of weak fusion.

Various visual cues provide information about depth and shape in a scene. When several of these cues are simultaneously available in a single location in the scene, the visual system attempts to combine them. In this paper, we discuss three key issues relevant to the experimental analysis of depth cue combination in human vision: cue promotion, dynamic weighting of cues, and robustness of cue combination. We review recent psychophysical studies of human depth cue combination in light of these issues. We organize the discussion and review as the development of a model of the depth cue combination process termed modified weak fusion (MWF). We relate the MWF framework to Bayesian theories of cue combination. We argue that the MWF model is consistent with previous experimental results and is a parsimonious summary of these results. While the MWF model is motivated by normative considerations, it is primarily intended to guide experimental analysis of depth cue combination in human vision. We describe experimental methods, analogous to perturbation analysis, that permit us to analyze depth cue combination in novel ways. In particular these methods allow us to investigate the key issues we have raised. We summarize recent experimental tests of the MWF framework that use these methods.

Bayes Theorem↗

Discrimination of orientation-defined texture edges.

Preattentive texture segregation was examined using textures composed of randomly placed, oriented line segments. A difference in texture element orientation produced an illusory, or orientation-defined, texture edge. Subjects discriminated between two textures, one with a straight texture edge and one with a "wavy" texture edge. Across conditions the orientation of the texture elements and the orientation of the texture edge varied. Although the orientation difference across the texture edge (the "texture gradient") is an important determinant of texture segregation performance, it is not the only one. Evidence from several experiments suggests that configural effects are also important. That is, orientation-defined texture edges are strongest when the texture elements (on one side of the edge) are parallel to the edge. This result is not consistent with a number of texture segregation models including feature- and filter-based models. One possible explanation is that the second-order channel used to detect a texture edge of a particular orientation gives greater weight to first-order input channels of that same orientation.

Discrimination, Psychological↗

Integration of stereopsis and motion shape cues.

A global shape judgement task was used to investigate the combination of stereopsis and kinetic depth. With both cues present, there were no distortions of shape perception, even under conditions where either cue alone did show such distortions. We suggest that the addition of motion information overcomes the stereo distance scaling problem. However, when incongruent combinations of disparity and motion were used, the results did not match predictions of a number of combination theories. These data could be described by a model which used weighted linear combination after correctly scaling disparities for viewing distance. When the motion cue was weakened by presenting only two frames of each motion sequence, stereo was weighted more heavily.

Cues↗

Histogram contrast analysis and the visual segregation of IID textures.

A new psychophysical methodology is introduced, histogram contrast analysis, that allows one to measure stimulus transformations, f, used by the visual system to draw distinctions between different image regions. The method involves the discrimination of images constructed by selecting texture micropatterns randomly and independently (across locations) on the basis of a given micropattern histogram. Different components of f are measured by use of different component functions to modulate the micropattern histogram until the resulting textures are discriminable. When no discrimination threshold can be obtained for a given modulating component function, a second titration technique may be used to measure the contribution of that component to f. The method includes several strong tests of its own assumptions. An example is given of the method applied to visual textures composed of small, uniform squares with randomly chosen gray levels. In particular, for a fixed mean gray level mu and a fixed gray-level variance sigma 2, histogram contrast analysis is used to establish that the class S of all textures composed of small squares with jointly independent, identically distributed gray levels with mean mu and variance sigma 2 is perceptually elementary in the following sense: there exists a single, real-valued function f S of gray level, such that two textures I and J in S are discriminable only if the average value of f S applied to the gray levels in I is significantly different from the average value of f S applied to the gray levels in J. Finally, histogram contrast analysis is used to obtain a seventh-order polynomial approximation of f S.

Contrast Sensitivity↗

A perturbation analysis of depth perception from combinations of texture and motion cues.

We examined how depth information from two different cue types (object motion and texture gradient) is integrated into a single estimate in human vision. Two critical assumptions of a recent model of depth cue combination (termed modified weak fusion) were tested. The first assumption is that the overall depth estimate is a weighted linear combination of the estimates derived from the individual cues, after initial processing needed to bring them to a common format. The second assumption is that the weight assigned to a cue reflects the apparent reliability of that cue in a particular scene. By this account, the depth combination rule is linear and dynamic, changing in a predictable fashion in response to the particular scene and viewing conditions. A novel procedure was used to measure the weights assigned to the texture and motion cues across experimental conditions. This procedure uses a type of perturbation analysis. The results are consistent with the weighted linear combination rule. In addition, when either cue is corrupted by added noise, the weighted linear combination rule shifts in favor of the uncontaminated cue.

Cues↗

Role of chromatic and luminance contrast in inferring structure from motion.

We measured the ability to infer structure from motion (SFM) in several directions in three-dimensional color space. Only motion cues are useful to subjects in performing this three-dimensional shape-identification task. We report the following results: (1) SFM performance is at chance for equiluminant stimuli that isolate short-wavelength-sensitive cones. Hence the short-wavelength-sensitive-cone input to SFM is negligible. (2) SFM performance increases with the magnitude of delta L - delta M signal when delta L + delta M = 0 (i.e., only chromatic and no luminance contrast is available). We reject the hypothesis that SFM obtains input from a single chromatic mechanism combining the long- and medium-wavelength-sensitive cones linearly. Our data are compatible with SFM that uses the output of two mechanisms, one taking the difference between the long- and medium-wavelength-sensitive-cone signals and the other taking the respective sum. We reject the particular hypothesis that SFM utilizes only the magnitude and not the sign of the long- and medium-wavelength-sensitive-cone signal. (3) We compare SFM performance with threshold performance for velocity and motion discrimination. Stimuli with luminance contrast yield SFM performance that is superior to stimuli without luminance contrast when they are expressed as multiples of velocity discrimination threshold. This superiority is even greater when SFM performance is compared with motion-direction discrimination thresholds.

Color Perception↗

Texture segregation and orientation gradient.

Rapid texture segregation is examined using filtered noise textures. The stimuli consist of a foreground region of filtered noise with one dominant texture orientation against a background region with a different dominant orientation. Shape discrimination of the foreground region is measured as a function of the difference in orientation between the two regions (delta theta), the distance over which the dominant orientation rotates from the background to the foreground value (delta chi), and the dominant spatial frequency of the textures (f). Performance declines with smaller delta theta, larger delta chi, and lower f. These effects are partially independent of viewing distance, which implies that it is the relative or object spatial frequency, not retinal spatial frequency, which determines performance in this task. We present a model consisting of channels tuned for orientation and spatial frequency which compute local oriented energy, followed by (texture) edge detection and a cross-correlator which performs the shape discrimination. Monte Carlo simulations of this model are in accord with the degradation in performance with increased delta chi and decreased delta theta.

Humans↗

The kinetic depth effect and optic flow--II. First- and second-order motion.

We use a difficult shape identification task to analyze how humans extract 3D surface structure from dynamic 2D stimuli--the kinetic depth effect (KDE). Stimuli composed of luminous tokens moving on a less luminous background yield accurate 3D shape identification regardless of the particular token used (either dots, lines, or disks). These displays stimulate both the 1st-order (Fourier-energy) motion detectors and 2nd-order (nonFourier) motion detectors. To determine which system supports KDE, we employ stimulus manipulations that weaken or distort 1st-order motion energy (e.g. frame-to-frame alternation of the contrast polarity of tokens) and manipulations that create microbalanced stimuli which have no useful 1st-order motion energy. All manipulations that impair 1st-order motion energy correspondingly impair 3D shape identification. In certain cases, 2nd-order motion could support limited KDE, but it was not robust and was of low spatial resolution. We conclude that 1st-order motion detectors are the primary input to the kinetic depth system. To determine minimal conditions for KDE, we use a two frame display. Under optimal conditions, KDE supports shape identification performance at 63-94% of full-rotation displays (where baseline is 5%). Increasing the amount of 3D rotation portrayed or introducing a blank inter-stimulus interval impairs performance. Together, our results confirm that the human KDE computation of surface shape uses a global optic flow computed primarily by 1st-order motion detectors with minor 2nd-order inputs. Accurate 3D shape identification requires only two views and therefore does not require knowledge of acceleration.

Depth Perception↗

Nonadditivity of masking by narrow-band noises.

Characterization of the visual system as a linear system has many consequences. One property implied by such a characterization is that signal-to-noise ratio at threshold is constant, so that the contrast energy of a sinusoidal grating at threshold depends on the effective noise passed by the filter used to detect the target. When the contrast energy in an external noise is sufficiently high, the contribution of internal noise may be conveniently ignored. In such circumstances, the effectiveness of a given noise is measured by its ability to mask the target. One consequence of the linearity assumption is that if the energy of the effective noise is increased by a factor k, then threshold energy of the signal is increased by the same amount. The contrast energy at threshold in the presence of a masker created by adding two maskers should be the sum of the threshold energies in the individual maskers. We have tested this hypothesis for spectrally nonoverlapping maskers. We find that the contrast energy required to detect the target in the presence of the combined maskers is much greater than the sum of the threshold energies for the two maskers. This "excess masking" violates the linearity assumption. On the other hand, when maskers do have spectral overlap, no excess masking is found.

Contrast Sensitivity↗

Applications of the EVE software for visual modeling.

EVE, the Early Vision Emulation software, is a system for the stimulation of early visual processing. EVE has the ability to carry out the operations of a wide variety of models of spatial vision, motion detection and processing, and spatial sampling. We introduce the EVE software and illustrate some of its applications for models of pattern detection, pattern discrimination, and motion detection.

Humans↗

Intelligent temporal subsampling of American Sign Language using event boundaries.

How well can a sequence of frames be represented by a subset of the frames? Video sequences of American Sign Language (ASL) were investigated in two modes: dynamic (ordinary video) and static (frames printed side by side on the display). An activity index was used to choose critical frames at event boundaries, times when the difference between successive frames is at a local minimum. Sign intelligibility was measured for 32 experienced ASL signers who viewed individual signs. For full gray-scale dynamic signs activity-index subsampling yielded sequences that were significantly more intelligible than when every mth frame was chosen. This result was even more pronounced for static images. For binary images, the relative advantage of activity subsampling was smaller. We conclude that event boundaries can be defined computationally and that subsampling from event boundaries is better than choosing at regular intervals.

Adolescent↗

How to study the kinetic depth effect experimentally.

Sperling, Landy, Dosher, and Perkins (1989) proposed an objective 3D shape identification task with 2D artifactual cues removed and with full feedback (FB) to the subjects to measure KDE and to circumvent algorithmically equivalent KDE-alternative computations and artifactual non-KDE processing. (1) The 2D velocity flow-field was necessary and sufficient for true KDE. (2) Only the first-order (Fourier-based) perceptual motion system could solve our task because the second-order (rectifying) system could not simultaneously process more than two locations. (3) To ensure first-order motion processing, KDE tasks must require simultaneous processing at more than two locations. (4) Practice with FB is essential to measure ultimate capacity (aptitude) and, thereby, to enable comparisons with ideal observers. Experiments without FB measure ecological achievement--the ability of subjects to extrapolate their past experience to the current stimuli.

Algorithms↗

Kinetic depth effect and optic flow--I. 3D shape from Fourier motion.

Fifty-three different 3D shapes were defined by sequences of 2D views (frames) of dots on a rotating 3D surface. (1) Subjects' accuracy of shape identifications dropped from over 90% to less than 10% when either the polarity of the stimulus dots was alternated from light-on-gray to dark-on-gray on successive frames or when neutral gray interframe intervals were interposed. Both manipulations interfere with motion extraction by spatio-temporal (Fourier) and gradient first-order detectors. Second-order (non-Fourier) detectors that use full-wave rectification are unaffected by alternating-polarity but disrupted by interposed gray frames. (2) To equate the accuracy of two-alternative forced-choice (2AFC) planar direction-of-motion discrimination in standard and polarity-alternated stimuli, standard contrast was reduced. 3D shape discrimination survived contrast reduction in standard stimuli whereas it failed completely with polarity-alternation even at full contrast. (3) When individual dots were permitted to remain in the image sequence for only two frames, performance showed little loss compared to standard displays where individual dots had an expected lifetime of 20 frames, showing that 3D shape identification does not require continuity of stimulus tokens. (4) Performance in all discrimination tasks is predicted (up to a monotone transformation) by considering the quality of first-order information (as given by a simple computation on Fourier power) and the number of locations at which motion information is required. Perceptual first-order analysis of optic flow is the primary substrate for structure-from-motion computations in random dot displays because only it offers sufficient quality of perceptual motion at a sufficient number of locations.

Depth Perception↗

Ratings of kinetic depth in multidot displays.

Subjects saw kinetic depth displays whose shape (sphere or cylinder) was defined by luminous dots distributed randomly on the surface or in the volume of the object. Subjects rated perceived 3-D depth, rigidity, and coherence. Despite individual differences, all 3 ratings increased with the number of dots. Dots in the volume yielded ratings equal to or greater than surface dots. Each rating varied with 3 of 4 factors (shape, distribution, numerosity, and perspective), but the ratings either between trials or between conditions were often uncorrelated. Object shape affected rigidity but not depth ratings. Veridically perceived polar displays had slightly lower rigidity but higher depth ratings than parallel projection displays. (Reversed polar displays were always grossly nonrigid.) The interaction of ratings and stimulus parameters requires theories and experiments in which different KDE ratings are not treated interchangeably.

Attention↗