Search PubMed⌕ Search

Biomedical subjects

Laurent Itti

Publications and source records attributed to Laurent Itti.

11 recordsLinked to original sources

Rapid biologically-inspired scene classification using features shared with visual attention.

We describe and validate a simple context-based scene recognition algorithm for mobile robotics applications. The system can differentiate outdoor scenes from various sites on a college campus using a multiscale set of early-visual features, which capture the "gist" of the scene into a low-dimensional signature vector. Distinct from previous approaches, the algorithm presents the advantage of being biologically plausible and of having low-computational complexity, sharing its low-level features with a model for visual attention that may operate concurrently on a robot. We compare classification accuracy using scenes filmed at three outdoor sites on campus (13,965 to 34,711 frames per site). Dividing each site into nine segments, we obtain segment classification rates between 84.21 percent and 88.62 percent. Combining scenes from all sites (75,073 frames in total) yields 86.45 percent correct classification, demonstrating the generalization and scalability of the approach.

Algorithms↗

Visual causes versus correlates of attentional selection in dynamic scenes.

What are the visual causes, rather than mere correlates, of attentional selection and how do they compare to each other during natural vision? To address these questions, we first strung together semantically unrelated dynamic scenes into MTV-style video clips, and performed eye tracking experiments with human observers. We then quantified predictions of saccade target selection based on seven bottom-up models, including intensity variance, orientation contrast, intensity contrast, color contrast, flicker contrast, motion contrast, and integrated saliency. On average, all tested models predicted saccade target selection well above chance. Dynamic models were particularly predictive of saccades that were most likely bottom-up driven-initiated shortly after scene onsets, leading to maximal inter-observer similarity. Static models showed mixed results in these circumstances, with intensity variance and orientation contrast featuring particularly weak prediction accuracy (lower than their own average, and approximately 4 times lower than dynamic models). These results indicate that dynamic visual cues play a dominant causal role in attracting attention. In comparison, some static visual cues play a weaker causal role, while other static cues are not causal at all, and may instead reflect top-down causes.

Adult↗

Top-down attention selection is fine grained.

Although much is known about the sources and modulatory effects of top-down attentional signals, the information capacity of these signals is less known. Here, we investigate the granularity of top-down attentional signals. Previous theories in psychophysics have provided conflicting evidence on whether top-down guidance is coarse grained (i.e., one gain control term per feature dimension) or fine grained (i.e., multiple gain control terms per dimension). We resolve the conflict by designing new experiments that disentangle top-down from bottom-up contributions, thereby avoiding confounds existing in previous studies. The results of our eye-tracking experiments show that subjects can selectively saccade to items belonging to the relevant feature interval compared with irrelevant intervals within a dimension. This suggests that top-down signals can specify not only the relevant feature dimension but also the relevant feature interval within a dimension. We conclude that top-down signals are fine grained and can specify multiple gain control terms per dimension.

Attention↗

The role of memory in guiding attention during natural vision.

What is the time frame in which perceptual memory guides attention? Current estimates range from a few hundred milliseconds to several seconds, minutes, or even days. Here, we answer this question by establishing the time course of attentional selection in realistic viewing conditions. First, we transformed continuous video clips into MTV-style video clips by stringing together continuous clip segments using abrupt transitions (jump cuts). We then asked participants to visually explore either continuous or MTV-style clips and recorded their saccades as objective behavioral indicators of attentional selections. The utilization of perceptual memory was estimated across viewing conditions and over time by quantifying the agreement between human attentional selections and predictions made by a neurally grounded computational model. In the critical condition, jump cuts led to sharp declines in the impact of perceptual memory on attentional selection, followed by monotonic increases in memory utilization across seven consecutive saccades and 2.5 s. These results demonstrate that perceptual memory traces play an important role in guiding attention across several saccades during natural vision. We propose novel hypotheses and experiments using hybrid natural-artificial stimuli to further elucidate neurocomputational mechanisms of attentional selection.

Adult↗

Computational modeling and exploration of contour integration for visual saliency.

We propose a computational model of contour integration for visual saliency. The model uses biologically plausible devices to simulate how the representations of elements aligned collinearly along a contour in an image are enhanced. Our model adds such devices as a dopamine-like fast plasticity, local GABAergic inhibition and multi-scale processing of images. The fast plasticity addresses the problem of how neurons in visual cortex seem to be able to influence neurons they are not directly connected to, for instance, as observed in contour closure effect. Local GABAergic inhibition is used to control gain in the system without using global mechanisms which may be non-plausible given the limited reach of axonal arbors in visual cortex. The model is then used to explore not only its validity in real and artificial images, but to discover some of the mechanisms involved in processing of complex visual features such as junctions and end-stops as well as contours. We present evidence for the validity of our model in several phases, starting with local enhancement of only a few collinear elements. We then test our model on more complex contour integration images with a large number of Gabor elements. Sections of the model are also extracted and used to discover how the model might relate contour integration neurons to neurons that process end-stops and junctions. Finally, we present results from real world images. Results from the model suggest that it is a good current approximation of contour integration in human vision. As well, it suggests that contour integration mechanisms may be strongly related to mechanisms for detecting end-stops and junction points. Additionally, a contour integration mechanism may be involved in finding features for objects such as faces. This suggests that visual cortex may be more information efficient and that neural regions may have multiple roles.

Form Perception↗

Perceptual consequences of feature-based attention.

Attention modulates visual processing along at least two dimensions: a spatial dimension, which enhances the representation of stimuli within the focus of attention, and a feature dimension, which is thought to enhance attended visual features (e.g., upward motion) throughout the visual field. We investigate the consequences of feature-based attention onto visual perception, using dual-task human psychophysics and two distant drifting Gabor stimuli to systematically explore 64 combinations of visual features (orientations and drift speeds) and tasks (discriminating orientation or drift speed). The resulting single, consistent data set suggests a functional model, which predicts a maximum rule by which only the dominant product of feature enhancement and feature benefit by feature relevance may benefit perception.

Attention↗

Modeling the influence of task on attention.

We propose a computational model for the task-specific guidance of visual attention in real-world scenes. Our model emphasizes four aspects that are important in biological vision: determining task-relevance of an entity, biasing attention for the low-level visual features of desired targets, recognizing these targets using the same low-level features, and incrementally building a visual map of task-relevance at every scene location. Given a task definition in the form of keywords, the model first determines and stores the task-relevant entities in working memory, using prior knowledge stored in long-term memory. It attempts to detect the most relevant entity by biasing its visual attention system with the entity's learned low-level features. It attends to the most salient location in the scene, and attempts to recognize the attended object through hierarchical matching against object representations stored in long-term memory. It updates its working memory with the task-relevance of the recognized entity and updates a topographic task-relevance map with the location and relevance of the recognized entity. The model is tested on three types of tasks: single-target detection in 343 natural and synthetic images, where biasing for the target accelerates target detection over twofold on average; sequential multiple-target detection in 28 natural images, where biasing, recognition, working memory and long term memory contribute to rapidly finding all targets; and learning a map of likely locations of cars from a video clip filmed while driving on a highway. The model's performance on search for single features and feature conjunctions is consistent with existing psychophysical data. These results of our biologically-motivated architecture suggest that the model may provide a reasonable approximation to many brain processes involved in complex task-driven visual behaviors.

Attention↗

Components of bottom-up gaze allocation in natural images.

Recent research [Parkhurst, D., Law, K., & Niebur, E., 2002. Modeling the role of salience in the allocation of overt visual attention. Vision Research 42 (1) (2002) 107-123] showed that a model of bottom-up visual attention can account in part for the spatial locations fixated by humans while free-viewing complex natural and artificial scenes. That study used a definition of salience based on local detectors with coarse global surround inhibition. Here, we use a similar framework to investigate the roles of several types of non-linear interactions known to exist in visual cortex, and of eccentricity-dependent processing. For each of these, we added a component to the salience model, including richer interactions among orientation-tuned units, both at spatial short range (for clutter reduction) and long range (for contour facilitation), and a detailed model of eccentricity-dependent changes in visual processing. Subjects free-viewed naturalistic and artificial images while their eye movements were recorded, and the resulting fixation locations were compared with the models' predicted salience maps. We found that the proposed interactions indeed play a significant role in the spatiotemporal deployment of attention in natural scenes; about half of the observed inter-subject variance can be explained by these different models. This suggests that attentional guidance does not depend solely on local visual features, but must also include the effects of interactions among features. As models of these interactions become more accurate in predicting behaviorally-relevant salient locations, they become useful to a range of applications in computer vision and human-machine interface design.

Adolescent↗

Automatic foveation for video compression using a neurobiological model of visual attention.

We evaluate the applicability of a biologically-motivated algorithm to select visually-salient regions of interest in video streams for multiply-foveated video compression. Regions are selected based on a nonlinear integration of low-level visual cues, mimicking processing in primate occipital, and posterior parietal cortex. A dynamic foveation filter then blurs every frame, increasingly with distance from salient locations. Sixty-three variants of the algorithm (varying number and shape of virtual foveas, maximum blur, and saliency competition) are evaluated against an outdoor video scene, using MPEG-1 and constant-quality MPEG-4 (DivX) encoding. Additional compression radios of 1.1 to 8.5 are achieved by foveation. Two variants of the algorithm are validated against eye fixations recorded from four to six human observers on a heterogeneous collection of 50 video clips (over 45 000 frames in total). Significantly higher overlap than expected by chance is found between human and algorithmic foveations. With both variants, foveated clips are, on average, approximately half the size of unfoveated clips, for both MPEG-1 and MPEG-4. These results suggest a general-purpose usefulness of the algorithm in improving compression ratios of unconstrained video.

Adult↗

Functional neuroimaging provides evidence of anomalous cerebral laterality in adults with Klinefelter's syndrome.

This study aimed to characterize cerebral perfusion in men with Klinefelter's syndrome, known to present specific deficits in language, using (99m)Tc- hexamethylpropylene-amine-oxime scintigraphy and Talairach normalization. While a perfusion asymmetry toward the left hemisphere was found in controls, perfusion was mostly symmetrical in Klinefelter patients in the upper temporal and lower parietal areas. Scores on verbal tests were inversely correlated with perfusion changes, providing neurobiological substrate of anomalous cerebral laterality.

Adolescent↗