Search PubMedSearch

Biomedical subjects

G Sperling

Publications and source records attributed to G Sperling.

At least 19 recordsLinked to original sources

Attention-generated apparent motion.

Motion perception mechanisms have recently been divided into three categories. First-order mechanisms primarily extract motion from moving objects or features that differ from the background in luminance. Second-order mechanism extract motion from moving properties, such as a moving area of flicker in which there is no difference in mean luminance between target and background. These first- and second-order motion mechanisms are primarily monocular. The existence of purely binocular, interocular and various other unusual kinds of apparent motion has promoted conjectures of a third-order mechanism, but there has been no clear suggestion as to the actual computations that such a mechanism might perform. Here we demonstrate 'alternating feature' stimuli that produce apparent motion only when the observer selectively attends to one of the embedded features in the display. The latent motion in the alternating feature stimuli is invisible to first- or second-order motion mechanisms, and the direction of apparent motion depends on the particular feature attended. These findings suggest the mechanism of third-order motion: the locations of the most significant features are registered in a salience map, and motion is computed directly from this map.

Attention

Measuring the spatial frequency selectivity of second-order texture mechanisms.

Recent investigations of texture and motion perception suggest two early filtering stages: an initial stage of selective linear filtering followed by rectification and a second stage of linear filtering. Here we demonstrate that there are differently scaled second-stage filters, and we measure their contrast modulation sensitivity as a function of spatial frequency. Our stimuli are Gabor modulations of a suprathreshold, bandlimited, isotropic carrier noise. The subjects' task is to discriminate between two possible orientations of the Gabor. Carrier noises are filtered into four octave-wide bands, centered at m = 2, 4, 8, and 16 c/deg. The Gabor test signals are w = 0.5, 1, 2, 4 and 8 c/deg. The threshold modulation of the test signal is measured for all 20 combinations of m and w. For each carrier frequency m, the Gabor test frequency w to which subjects are maximally sensitive appears to be approximately 3-4 octaves below m. The consistent m x w interaction suggests that each second-stage spatial filter may be differentially tuned to a particular first-stage spatial frequency. The most sensitive combination is a second-stage filter of 1 c/deg with first-stage inputs of 8-16 c/deg. We conclude that second-order texture perception appears to utilize multiple channels tuned to spatial frequency and orientation, with channels tuned to low modulation frequencies appearing to be best served by carrier frequencies 8 to 16 times higher than the modulations they are tuned to detect.

Contrast Sensitivity

1st- and 2nd-order motion and texture resolution in central and peripheral vision.

STIMULI. The 1st-order stimuli are moving sine gratings. The 2nd-order stimuli are fields of static visual texture, whose contrasts are modulated by moving sine gratings. Neither the spatial slant (orientation) nor the direction of motion of these 2nd-order (microbalanced) stimuli can be detected by a Fourier analysis; they are invisible to Reichardt and motion-energy detectors. METHOD. For these dynamic stimuli, when presented both centrally and in an annular window extending from 8 to 10 deg in eccentricity, we measured the highest spatial frequency for which discrimination between +/- 45 deg texture slants and discrimination between opposite directions of motion were each possible. RESULTS. For sufficiently low spatial frequencies, slant and direction can be discriminated in both central and peripheral vision, for both 1st- and for 2nd-order stimuli. For both 1st- and 2nd-order stimuli, at both retinal locations, slant discrimination is possible at higher spatial frequencies than direction discrimination. For both 1st- and 2nd-order stimuli, motion resolution decreases 2-3 times more rapidly with eccentricity than does texture resolution. CONCLUSIONS. (1) 1st- and 2nd-order motion scale similarly with eccentricity. (2) 1st- and 2nd-order texture scale similarly with eccentricity. (3) The central/peripheral resolution fall-off is 2-3 times greater for motion than for texture.

Contrast Sensitivity

The functional architecture of human visual motion perception.

UNLABELLED: A powerful paradigm (the pedestal-plus-test display) is combined with several subsidiary paradigms (interocular presentation, stimulus superpositions with varying phases, and attentional manipulations) to determine the functional architecture of visual motion perception: i.e. the nature of the various mechanisms of motion perception and their relations to each other. Three systems are isolated: a first-order system that uses a primitive motion energy computation to extract motion from moving luminance modulations; a second-order system that uses motion energy to extract motion from moving texture-contrast modulations; and a third-order system that tracks features. Pedestal displays exclude feature-tracking and thereby yield pure measures of the first- and second-order systems which are found to be exclusively monocular. Interocular displays exclude the first- and second-order systems and thereby to yield pure measures of feature-tracking. RESULTS: both first- and second-order systems are fast (with temporal frequency cutoff at 12 Hz) and sensitive. Feature tracking operates interocularly almost as well as monocularly. It is slower (cutoff frequency is 3 Hz) and it requires much more stimulus contrast than the first- and second-order systems. Feature tracking is both bottom-up (it computes motion from luminance modulation, texture-contrast modulation, depth modulation, motion modulation, flicker modulation, and from other types of stimuli) and top-down--e.g. attentional instructions can determine the direction of perceived motion.

Attention

Non-Fourier motion analysis.

It has been realized for some time that the visual system performs at least two general sorts of motion processing. First-order motion processing applies some variant of standard motion analysis (i.e. spatiotemporal Fourier energy analysis) directly to stimulus luminance, whereas second-order motion processing applies standard motion analysis to one or another grossly non-linear transformation of stimulus luminance. We have developed a method for disentangling the different sorts of mechanisms that may operate in human vision to detect second-order motion. This method hinges on an empirical condition called transition invariance that may or may not be satisfied by a family psi of textures. Any failure of this condition indicates that more than one mechanism is involved in detecting the motion of stimuli composed of the textures in psi. We have shown that the family of sinusoidal gratings oriented orthogonally to the direction of motion and varying in contrast and spatial frequency is transition invariant. We modelled the results in terms of a single-channel motion computation. We have new results indicating that a specific class of textures differing in texture element density and texture element contrast decisively fails the test of transition invariance. These findings suggest that in addition to the single second-order motion channel required by our earlier results there exists at least one other second-order motion channel. We argue that the preprocessing transformation used by this channel is a pointwise non-linearity that maps stimulus contrasts of absolute value less than some relatively high threshold tau onto 0, but increases with magnitude of c-tau for contrasts. c of absolute value greater than tau.

Animals

Full-wave and half-wave processes in second-order motion and texture.

A theory of human second-order motion perception is proposed and further applied to the discrimination of texture slant. The computational algorithms for deriving the direction of left-right motion from a sequence of images are equivalent to the algorithms for deriving the direction of slant (e.g. from top left to bottom right or top right to bottom left) in a single 2D image. There is a broad range of phenomena for which Fourier analysis of the image plus a few simple rules gives a good account of human perception. The problem with this first-order analysis is that there exists a broad class of 'microbalanced' stimuli in which the motion or slant is completely obvious to human subjects but is invisible to first-order analysis. Microbalanced stimuli require second-order analysis which consists of non-linear preprocessing (spatiotemporal filtering followed by rectification of the input signal) before standard motion or slant analysis. To determine whether the second-order rectification is half-wave or full-wave, we construct two special microbalanced stimulus types: 'half-wave stimuli' whose motion (or texture slant) is interpretable by a half-wave rectifying system but not by full-wave or a first-order (Fourier) analysis and 'full-wave stimuli' which are interpretable only after full-wave rectification. Such experiments show that second-order texture-slant perception utilizes both half-wave and full-wave processes, second-order motion-direction discrimination depends predominantly on full-wave rectification and second-order spatial interactions such as lateral contrast-contrast inhibition and second-order Mach bands are exclusively full-wave.

Algorithms

Full-wave and half-wave rectification in second-order motion perception.

UNLABELLED: Microbalanced stimuli are dynamic displays which do not stimulate motion mechanisms that apply standard (Fourier-energy or autocorrelational) motion analysis directly to the visual signal. In order to extract motion information from microbalanced stimuli, Chubb and Sperling [(1988) Journal of the Optical Society of America, 5, 1986-2006] proposed that the human visual system performs a rectifying transformation on the visual signal prior to standard motion analysis. The current research employs two novel types of microbalanced stimuli: half-wave stimuli preserve motion information following half-wave rectification (with a threshold) but lose motion information following full-wave rectification; full-wave stimuli preserve motion information following full-wave rectification but lose motion information following half-wave rectification. Additionally, Fourier stimuli, ordinary square-wave gratings, were used to stimulate standard motion mechanisms. Psychometric functions (direction discrimination vs stimulus contrast) were obtained for each type of stimulus when presented alone, and when masked by each of the other stimuli (presented as moving masks and also as nonmoving, counterphase-flickering masks). RESULTS: given sufficient contrast, all three types of stimulus convey motion. However, only one-third of the population can perceive the motion of the half-wave stimulus. Observers are able to process the motion information contained in the Fourier stimulus slightly more efficiently than the information in the full-wave stimulus but are much less efficient in processing half-wave motion information. Moving masks are more effective than counterphase masks at hampering direction discrimination, indicating that some of the masking effect is interference between motion mechanisms, and some occurs at earlier stages. When either full-wave and Fourier or half-wave and Fourier gratings are presented simultaneously, there is a wide range of relative contrasts within which the motion directions of both gratings are easily determinable. Conversely, when half-wave and full-wave gratings are combined, the direction of only one of these gratings can be determined with high accuracy. CONCLUSIONS: the results indicate that three motion computations are carried out, any two in parallel: one standard ("first order") and two non-Fourier ("second-order") computations that employ full-wave and half-wave rectification.

Contrast Sensitivity

Perception of apparent motion between dissimilar gratings: spatiotemporal properties.

What determines the strength of texture-defined apparent motion perception when the stimulus has no net directional energy in the Fourier domain? In a previous paper [Werkhoven, Sperling & Chubb (1993) Vision Research, 33, 463-485] we demonstrated the counterintuitive finding that the correspondence in spatial frequency and in modulation amplitude between neighboring patches of texture in a spatiotemporal motion path are irrelevant to motion strength. Instead, we found strong support for what we call a single channel or one-dimensional motion computation: a simple nonlinear transformation of the image, followed by standard motion analysis. Here, we further studied the dimensionality of the motion computation in a parameter space that includes texture orientation and stimulus display rate in addition to texture spatial frequency and modulation amplitude. We used ambiguous motion displays in which one motion path, consisting of patches of nonsimilar texture, competes with another motion path comprised entirely of similar texture patches. The data show that motion between dissimilar patches of texture that are orthogonally oriented, have a two octave difference in spatial frequency and differ 50% in modulation amplitude can easily dominate motion between similar patches of texture. A single channel accounts for more than 70% of texture-from-motion strength for the parameter space examined and this channel is invariant for stimulus display rates varying over a four-fold range.

Humans

Object spatial frequencies, retinal spatial frequencies, noise, and the efficiency of letter discrimination.

To determine which spatial frequencies are most effective for letter identification, and whether this is because letters are objectively more discriminable in these frequency bands or because can utilize the information more efficiently, we studied the 26 upper-case letters of English. Six two-octave wide filters were used to produce spatially filtered letters with 2D-mean frequencies ranging from 0.4 to 20 cycles per letter height. Subjects attempted to identify filtered letters in the presence of identically filtered, added Gaussian noise. The percent of correct letter identifications vs s/n (the root-mean-square ratio of signal to noise power) was determined for each band at four viewing distances ranging over 32:1. Object spatial frequency band and s/n determine presence of information in the stimulus; viewing distance determines retinal spatial frequency, and affects only ability to utilize. Viewing distance had no effect upon letter discriminability: object spatial frequency, not retinal spatial frequency, determined discriminability. To determine discrimination efficiency, we compared human discrimination to an ideal discriminator. For our two-octave wide bands, s/n performance of humans and of the ideal detector improved with frequency mainly because linear bandwidth increased as a function of frequency. Relative to the ideal detector, human efficiency was 0 in the lowest frequency bands, reached a maximum of 0.42 at 1.5 cycles per object and dropped to about 0.104 in the highest band. Thus, our subjects best extract upper-case letter information from spatial frequencies of 1.5 cycles per object height, and they can extract it with equal efficiency over a 32:1 range of retinal frequencies, from 0.074 to more than 2.3 cycles per degree of visual angle.

Adult

The kinetic depth effect and optic flow--II. First- and second-order motion.

We use a difficult shape identification task to analyze how humans extract 3D surface structure from dynamic 2D stimuli--the kinetic depth effect (KDE). Stimuli composed of luminous tokens moving on a less luminous background yield accurate 3D shape identification regardless of the particular token used (either dots, lines, or disks). These displays stimulate both the 1st-order (Fourier-energy) motion detectors and 2nd-order (nonFourier) motion detectors. To determine which system supports KDE, we employ stimulus manipulations that weaken or distort 1st-order motion energy (e.g. frame-to-frame alternation of the contrast polarity of tokens) and manipulations that create microbalanced stimuli which have no useful 1st-order motion energy. All manipulations that impair 1st-order motion energy correspondingly impair 3D shape identification. In certain cases, 2nd-order motion could support limited KDE, but it was not robust and was of low spatial resolution. We conclude that 1st-order motion detectors are the primary input to the kinetic depth system. To determine minimal conditions for KDE, we use a two frame display. Under optimal conditions, KDE supports shape identification performance at 63-94% of full-rotation displays (where baseline is 5%). Increasing the amount of 3D rotation portrayed or introducing a blank inter-stimulus interval impairs performance. Together, our results confirm that the human KDE computation of surface shape uses a global optic flow computed primarily by 1st-order motion detectors with minor 2nd-order inputs. Accurate 3D shape identification requires only two views and therefore does not require knowledge of acceleration.

Depth Perception

The visible persistence of stimuli in stroboscopic motion.

This paper reports an improved paradigm to measure visible persistence. The stimulus is a pair of lines stroboscopically displayed in successive positions moving in opposite directions. The subjects' judgement of simultaneous appearance of all the presented lines is used to estimate visible persistence. This paradigm permitted independent manipulation of spatial and temporal stimulus separations in linear motion. The resulting estimates of visible persistence increase with spatial separation up to 0.24 deg of visual angle and approaches a maximum value at larger spatial separations. The results are consistent with the existence of a hypothetical visual gain mechanism that operates over small retinal distances to effectively decrease persistence duration with decreasing spatial separation.

Afterimage

Intelligent temporal subsampling of American Sign Language using event boundaries.

How well can a sequence of frames be represented by a subset of the frames? Video sequences of American Sign Language (ASL) were investigated in two modes: dynamic (ordinary video) and static (frames printed side by side on the display). An activity index was used to choose critical frames at event boundaries, times when the difference between successive frames is at a local minimum. Sign intelligibility was measured for 32 experienced ASL signers who viewed individual signs. For full gray-scale dynamic signs activity-index subsampling yielded sequences that were significantly more intelligible than when every mth frame was chosen. This result was even more pronounced for static images. For binary images, the relative advantage of activity subsampling was smaller. We conclude that event boundaries can be defined computationally and that subsampling from event boundaries is better than choosing at regular intervals.

Adolescent

[Stenosing ureteritis and factor XIII deficiency in anaphylactoid purpura].

A case report of a 6-year-old boy is presented. The patient suffered from a severe Schönlein-Henoch purpura. It was demonstrated that ureteral stenosis develops during the clinical course of this disease. This complication has to be considered in anaphylactoid purpura, since it is usually self-limiting and does not require surgical intervention. This confirms once again the necessity to look for a decrease of factor XIII activity in these patients and of the value of substituting this compound if there are severe abdominal complaints. The theoretical background of this therapeutical intervention is discussed.

Child

How to study the kinetic depth effect experimentally.

Sperling, Landy, Dosher, and Perkins (1989) proposed an objective 3D shape identification task with 2D artifactual cues removed and with full feedback (FB) to the subjects to measure KDE and to circumvent algorithmically equivalent KDE-alternative computations and artifactual non-KDE processing. (1) The 2D velocity flow-field was necessary and sufficient for true KDE. (2) Only the first-order (Fourier-based) perceptual motion system could solve our task because the second-order (rectifying) system could not simultaneously process more than two locations. (3) To ensure first-order motion processing, KDE tasks must require simultaneous processing at more than two locations. (4) Practice with FB is essential to measure ultimate capacity (aptitude) and, thereby, to enable comparisons with ideal observers. Experiments without FB measure ecological achievement--the ability of subjects to extrapolate their past experience to the current stimuli.

Algorithms

Kinetic depth effect and optic flow--I. 3D shape from Fourier motion.

Fifty-three different 3D shapes were defined by sequences of 2D views (frames) of dots on a rotating 3D surface. (1) Subjects' accuracy of shape identifications dropped from over 90% to less than 10% when either the polarity of the stimulus dots was alternated from light-on-gray to dark-on-gray on successive frames or when neutral gray interframe intervals were interposed. Both manipulations interfere with motion extraction by spatio-temporal (Fourier) and gradient first-order detectors. Second-order (non-Fourier) detectors that use full-wave rectification are unaffected by alternating-polarity but disrupted by interposed gray frames. (2) To equate the accuracy of two-alternative forced-choice (2AFC) planar direction-of-motion discrimination in standard and polarity-alternated stimuli, standard contrast was reduced. 3D shape discrimination survived contrast reduction in standard stimuli whereas it failed completely with polarity-alternation even at full contrast. (3) When individual dots were permitted to remain in the image sequence for only two frames, performance showed little loss compared to standard displays where individual dots had an expected lifetime of 20 frames, showing that 3D shape identification does not require continuity of stimulus tokens. (4) Performance in all discrimination tasks is predicted (up to a monotone transformation) by considering the quality of first-order information (as given by a simple computation on Fourier power) and the number of locations at which motion information is required. Perceptual first-order analysis of optic flow is the primary substrate for structure-from-motion computations in random dot displays because only it offers sufficient quality of perceptual motion at a sufficient number of locations.

Depth Perception

Ratings of kinetic depth in multidot displays.

Subjects saw kinetic depth displays whose shape (sphere or cylinder) was defined by luminous dots distributed randomly on the surface or in the volume of the object. Subjects rated perceived 3-D depth, rigidity, and coherence. Despite individual differences, all 3 ratings increased with the number of dots. Dots in the volume yielded ratings equal to or greater than surface dots. Each rating varied with 3 of 4 factors (shape, distribution, numerosity, and perspective), but the ratings either between trials or between conditions were often uncorrelated. Object shape affected rigidity but not depth ratings. Veridically perceived polar displays had slightly lower rigidity but higher depth ratings than parallel projection displays. (Reversed polar displays were always grossly nonrigid.) The interaction of ratings and stimulus parameters requires theories and experiments in which different KDE ratings are not treated interchangeably.

Attention

Kinetic depth effect and identification of shape.

We introduce an objective shape-identification task for measuring the kinetic depth effect (KDE). A rigidly rotating surface consisting of hills and valleys on an otherwise flat ground was defined by 300 randomly positioned dots. On each trial, 1 of 53 shapes was presented; the observer's task was to identify the shape and its overall direction of rotation. Identification accuracy was an objective measure, with a low guessing base rate of the observer's perceptual ability to extract 3D structure from 2D motion via KDE. (1) Objective accuracy data were consistent with previously obtained subjective rating judgments of depth and coherence. (2) Along with motion cues, rotating real 3D dot-defined shapes inevitably produced a cue of changing dot density. By shortening dot lifetimes to control dot density, we showed that changing density was neither necessary nor sufficient to account for accuracy; motion alone sufficed. (3) Our shape task was solvable with motion cues from the 6 most relevant locations. We extracted the dots from these locations and used them in a simplified 2D direction-labeling motion task with 6 perceptually flat flow fields. Subjects' performance in the 2D and 3D tasks was equivalent, indicating that the information processing capacity of KDE is not unique. (4) Our proposed structure-from-motion algorithm for the shape task first finds relative minima and maxima of local velocity and then assigns 3D depths proportional to velocity.

Adult