Search PubMed⌕ Search

Biomedical subjects

Rama Chellappa

Publications and source records attributed to Rama Chellappa.

17 recordsLinked to original sources

Appearance characterization of linear Lambertian objects, generalized photometric stereo, and illumination-invariant face recognition.

Traditional photometric stereo algorithms employ a Lambertian reflectance model with a varying albedo field and involve the appearance of only one object. In this paper, we generalize photometric stereo algorithms to handle all appearances of all objects in a class, in particular the human face class, by making use of the linear Lambertian property. A linear Lambertian object is one which is linearly spanned by a set of basis objects and has a Lambertian surface. The linear property leads to a rank constraint and, consequently, a factorization of an observation matrix that consists of exemplar images of different objects (e.g., faces of different subjects) under different, unknown illuminations. Integrability and symmetry constraints are used to fully recover the subspace bases using a novel linearized algorithm that takes the varying albedo field into account. The effectiveness of the linear Lambertian property is further investigated by using it for the problem of illumination-invariant face recognition using just one image. Attached shadows are incorporated in the model by a careful treatment of the inherent nonlinearity in Lambert's law. This enables us to extend our algorithm to perform face recognition in the presence of multiple illumination sources. Experimental results using standard data sets are presented.

Algorithms↗

Robust ego-motion estimation and 3-D model refinement using surface parallax.

We present an iterative algorithm for robustly estimating the ego-motion and refining and updating a coarse depth map using parametric surface parallax models and brightness derivatives extracted from an image pair. Given a coarse depth map acquired by a range-finder or extracted from a digital elevation map (DEM), ego-motion is estimated by combining a global ego-motion constraint and a local brightness constancy constraint. Using the estimated camera motion and the available depth estimate, motion of the three-dimensional (3-D) points is compensated. We utilize the fact that the resulting surface parallax field is an epipolar field, and knowing its direction from the previous motion estimates, estimate its magnitude and use it to refine the depth map estimate. The parallax magnitude is estimated using a constant parallax model (CPM) which assumes a smooth parallax field and a depth based parallax model (DBPM), which models the parallax magnitude using the given depth map. We obtain confidence measures for determining the accuracy of the estimated depth values which are used to remove regions with potentially incorrect depth estimates for robustly estimating ego-motion in subsequent iterations. Experimental results using both synthetic and real data (both indoor and outdoor sequences) illustrate the effectiveness of the proposed algorithm.

Algorithms↗

Principal components null space analysis for image and video classification.

We present a new classification algorithm, principal component null space analysis (PCNSA), which is designed for classification problems like object recognition where different classes have unequal and nonwhite noise covariance matrices. PCNSA first obtains a principal components subspace (PCA space) for the entire data. In this PCA space, it finds for each class "i," an Mi-dimensional subspace along which the class' intraclass variance is the smallest. We call this subspace an approximate null space (ANS) since the lowest variance is usually "much smaller" than the highest. A query is classified into class "i" if its distance from the class' mean in the class' ANS is a minimum. We derive upper bounds on classification error probability of PCNSA and use these expressions to compare classification performance of PCNSA with that of subspace linear discriminant analysis (SLDA). We propose a practical modification of PCNSA called progressive-PCNSA that also detects "new" (untrained classes). Finally, we provide an experimental comparison of PCNSA and progressive PCNSA with SLDA and PCA and also with other classification algorithms-linear SVMs, kernel PCA, kernel discriminant analysis, and kernel SLDA, for object recognition and face recognition under large pose/expression variation. We also show applications of PCNSA to two classification problems in video--an action retrieval problem and abnormal activity detection.

Algorithms↗

Structure from planar motion.

Planar motion is arguably the most dominant type of motion in surveillance videos. The constraints on motion lead to a simplified factorization method for structure from planar motion when using a stationary perspective camera. Compared with methods for general motion, our approach has two major advantages: a measurement matrix that fully exploits the motion constraints is formed such that the new measurement matrix has a rank of at most 3, instead of 4; the measurement matrix needs similar scalings, but the estimation of fundamental matrices or epipoles is not needed. Experimental results show that the algorithm is accurate and fairly robust to noise and inaccurate calibration. As the new measurement matrix is a nonlinear function of the observed variables, a different method is introduced to deal with the directional uncertainty in the observed variables. Differences and the dual relationship between planar motion and planar object are also clarified. Based on our method, a fully automated vehicle reconstruction system has been designed.

Algorithms↗

Face verification across age progression.

Human faces undergo considerable amounts of varialions with aging. While face recognition systems have been proven to be sensitive to factors such as illumination and pose, their sensitivity to facial aging effects is yet to be studied. How does age progression affect the similarity between a pair of face images of an individual? What is the confidence associated with establishing the identity between a pair of age separated face images? In this paper, we develop a Bayesian age difference classifier that classifies face images of individuals based on age differences and performs face verification across age progression. Further, we study the similarity of faces across age progression. Since age separated face images invariably differ in illumination and pose, we propose preprocessing methods for minimizing such variations. Experimental results using a database comprising of pairs of face images that were retrieved from the passports of 465 individuals are presented. The verification system for faces separated by as many as nine years, attains an equal error rate of 8.5%.

Adult↗

From sample similarity to ensemble similarity: probabilistic distance measures in reproducing kernel Hilbert space.

This paper addresses the problem of characterizing ensemble similarity from sample similarity in a principled manner. Using reproducing kernel as a characterization of sample similarity, we suggest a probabilistic distance measure in the reproducing kernel Hilbert space (RKHS) as the ensemble similarity. Assuming normality in the RKHS, we derive analytic expressions for probabilistic distance measures that are commonly used in many applications, such as Chernoff distance (or the Bhattacharyya distance as its special case), Kullback-Leibler divergence, etc. Since the reproducing kernel implicitly embeds a nonlinear mapping, our approach presents a new way to study these distances whose feasibility and efficiency is demonstrated using experiments with synthetic and real examples. Further, we extend the ensemble similarity to the reproducing kernel for ensemble and study the ensemble similarity for more general data representations.

Algorithms↗

Bayesian algorithms for simultaneous structure from motion estimation of multiple independently moving objects.

In this paper, the problem of simultaneous structure from motion estimation for multiple independently moving objects from a monocular image sequence is addressed. Two Bayesian algorithms are presented for solving this problem using the sequential importance sampling (SIS) technique. The empirical posterior distribution of object motion and feature separation parameters is approximated by weighted samples. The first algorithm addresses the problem when only two moving objects are present. A singular value decomposition (SVD)-based sample clustering algorithm is shown to be capable of separating samples related to different objects. A pair of SIS procedures is used to track the posterior distribution of the motion parameters. In the second algorithm, a balancing step is added into the SIS procedure to preserve samples of low weights so that all objects have enough samples to propagate empirical motion distributions. By using the proposed algorithms, the relative motions of all the moving objects with respect to the camera can be simultaneously estimated. Both algorithms have been tested on synthetic and real-image sequences. Improved results have been achieved.

Algorithms↗

Background learning for robust face recognition with PCA in the presence of clutter.

We propose a new method within the framework of principal component analysis (PCA) to robustly recognize faces in the presence of clutter. The traditional eigenface recognition (EFR) method, which is based on PCA, works quite well when the input test patterns are faces. However, when confronted with the more general task of recognizing faces appearing against a background, the performance of the EFR method can be quite poor. It may miss faces completely or may wrongly associate many of the background image patterns to faces in the training set. In order to improve performance in the presence of background, we argue in favor of learning the distribution of background patterns and show how this can be done for a given test image. An eigenbackground space is constructed corresponding to the given test image and this space in conjunction with the eigenface space is used to impart robustness. A suitable classifier is derived to distinguish nonface patterns from faces. When tested on images depicting face recognition in real situations against cluttered background, the performance of the proposed method is quite good with fewer false alarms.

Algorithms↗

Statistical bias in 3-D reconstruction from a monocular video.

The present state-of-the-art in computing the error statistics in three-dimensional (3-D) reconstruction from video concentrates on estimating the error covariance. A different source of error which has not received much attention is the fact that the reconstruction estimates are often significantly statistically biased. In this paper, we derive a precise expression for the bias in the depth estimate, based on the continuous (differentiable) version of structure from motion (SfM). Many SfM algorithms, or certain portions of them, can be posed in a linear least-squares (LS) framework Ax = b. Examples include initialization procedures for bundle adjustment or algorithms that alternately estimate depth and camera motion. It is a well-known fact that the LS estimate is biased if the system matrix A is noisy. In SfM, the matrix A contains point correspondences, which are always difficult to obtain precisely; thus, it is expected that the structure and motion estimates in such a formulation of the problem would be biased. Existing results on the minimum achievable variance of the SfM estimator are extended by deriving a generalized Cramer-Rao lower bound. A detailed analysis of the effect of various camera motion parameters on the bias is presented. We conclude by presenting the effect of bias compensation on reconstructing 3-D face models from rendered images.

Algorithms↗

"Shape activity": a continuous-state HMM for moving/deforming shapes with application to abnormal activity detection.

The aim is to model "activity" performed by a group of moving and interacting objects (which can be people, cars, or different rigid components of the human body) and use the models for abnormal activity detection. Previous approaches to modeling group activity include co-occurrence statistics (individual and joint histograms) and dynamic Bayesian networks, neither of which is applicable when the number of interacting objects is large. We treat the objects as point objects (referred to as "landmarks") and propose to model their changing configuration as a moving and deforming "shape" (using Kendall's shape theory for discrete landmarks). A continuous-state hidden Markov model is defined for landmark shape dynamics in an activity. The configuration of landmarks at a given time forms the observation vector, and the corresponding shape and the scaled Euclidean motion parameters form the hidden-state vector. An abnormal activity is then defined as a change in the shape activity model, which could be slow or drastic and whose parameters are unknown. Results are shown on a real abnormal activity-detection problem involving multiple moving objects.

Algorithms↗

Matching shape sequences in video with applications in human movement analysis.

We present an approach for comparing two sequences of deforming shapes using both parametric models and nonparametric methods. In our approach, Kendall's definition of shape is used for feature extraction. Since the shape feature rests on a non-Euclidean manifold, we propose parametric models like the autoregressive model and autoregressive moving average model on the tangent space and demonstrate the ability of these models to capture the nature of shape deformations using experiments on gait-based human recognition. The nonparametric model is based on Dynamic Time-Warping. We suggest a modification of the Dynamic time-warping algorithm to include the nature of the non-Euclidean space in which the shape deformations take place. We also show the efficacy of this algorithm by its application to gait-based human recognition. We exploit the shape deformations of a person's silhouette as a discriminating feature and provide recognition results using the nonparametric model. Our analysis leads to some interesting observations on the role of shape and kinematics in automated gait-based person authentication.

Algorithms↗

Image-based face recognition under illumination and pose variations.

We present an image-based method for face recognition across different illuminations and poses, where the term image-based means that no explicit prior three-dimensional models are needed. As face recognition under illumination and pose variations involves three factors, namely, identity, illumination, and pose, generalizations in all these three factors are desired. We present a recognition approach that is able to generalize in the identity and illumination dimensions and handle a given set of poses. Specifically, the proposed approach derives an identity signature that is illumination- and pose-invariant, where the identity is tackled by means of subspace encoding, the illumination is characterized with a Lambertian reflectance model, and the given set of poses is treated as a whole. Experimental results using the Pose, Illumination, and Expression (PIE) database demonstrate the effectiveness of the proposed approach.

Algorithms↗

An information theoretic criterion for evaluating the quality of 3-D reconstructions from video.

Even though numerous algorithms exist for estimating the three-dimensional (3-D) structure of a scene from its video, the solutions obtained are often of unacceptable quality. To overcome some of the deficiencies, many application systems rely on processing more data than necessary, thus raising the question: how is the accuracy of the solution related to the amount of data processed by the algorithm? Can we automatically recognize situations where the quality of the data is so bad that even a large number of additional observations will not yield the desired solution? Previous efforts to answer this question have used statistical measures like second order moments. They are useful if the estimate of the structure is unbiased and the higher order statistical effects are negligible, which is often not the case. This paper introduces an alternative information-theoretic criterion for evaluating the quality of a 3-D reconstruction. The accuracy of the reconstruction is judged by considering the change in mutual information (MI) (termed as the incremental MI) between a scene and its reconstructions. An example of 3-D reconstruction from a video sequence using optical flow equations and known noise distribution is considered and it is shown how the MI can be computed from first principles. We present simulations on both synthetic and real data to demonstrate the effectiveness of the proposed criterion.

Algorithms↗

Identification of humans using gait.

We propose a view-based approach to recognize humans from their gait. Two different image features have been considered: the width of the outer contour of the binarized silhouette of the walking person and the entire binary silhouette itself. To obtain the observation vector from the image features, we employ two different methods. In the first method, referred to as the indirect approach, the high-dimensional image feature is transformed to a lower dimensional space by generating what we call the frame to exemplar (FED) distance. The FED vector captures both structural and dynamic traits of each individual. For compact and effective gait representation and recognition, the gait information in the FED vector sequences is captured in a hidden Markov model (HMM). In the second method, referred to as the direct approach, we work with the feature vector directly (as opposed to computing the FED) and train an HMM. We estimate the HMM parameters (specifically the observation probability B) based on the distance between the exemplars and the image features. In this way, we avoid learning high-dimensional probability density functions. The statistical nature of the HMM lends overall robustness to representation and recognition. The performance of the methods is illustrated using several databases.

Algorithms↗

Visual tracking and recognition using appearance-adaptive models in particle filters.

We present an approach that incorporates appearance-adaptive models in a particle filter to realize robust visual tracking and recognition algorithms. Tracking needs modeling interframe motion and appearance changes, whereas recognition needs modeling appearance changes between frames and gallery images. In conventional tracking algorithms, the appearance model is either fixed or rapidly changing, and the motion model is simply a random walk with fixed noise variance. Also, the number of particles is typically fixed. All these factors make the visual tracker unstable. To stabilize the tracker, we propose the following modifications: an observation model arising from an adaptive appearance model, an adaptive velocity motion model with adaptive noise variance, and an adaptive number of particles. The adaptive-velocity model is derived using a first-order linear predictor based on the appearance difference between the incoming observation and the previous particle configuration. Occlusion analysis is implemented using robust statistics. Experimental results on tracking visual objects in long outdoor and indoor video sequences demonstrate the effectiveness and robustness of our tracking algorithm. We then perform simultaneous tracking and recognition by embedding them in a particle filter. For recognition purposes, we model the appearance changes between frames and gallery images by constructing the intra- and extrapersonal spaces. Accurate recognition is achieved when confronted by pose and view variations.

Algorithms↗

Statistical-physical model for foliage clutter in ultra-wideband synthetic aperture radar images.

Analyzing foliage-penetrating (FOPEN) ultra-wideband synthetic aperture radar (SAR) images is a challenging problem owing to the noisy and impulsive nature of foliage clutter. Indeed, many target-detection algorithms for FOPEN SAR data are characterized by high false-alarm rates. In this work, a statistical-physical model for foliage clutter is proposed that explains the presence of outliers in the data and suggests the use of symmetric alpha-stable (SalphaS) distributions for accurate clutter modeling. Furthermore, with the use of general assumptions of the noise sources and propagation conditions, the proposed model relates the parameters of the SalphaS model to physical parameters such as the attenuation coefficient and foliage density.

Journal Article↗

Fast two-frame multiscale dense optical flow estimation using discrete wavelet filters.

A multiscale algorithm with complexity O(N) (where N is the number of pixels in one image) using wavelet filters is proposed to estimate dense optical flow from two frames. Hierarchical image representation by wavelet decomposition is integrated with differential techniques in a new multiscale framework. It is shown that if a compactly supported wavelet basis with one vanishing moment is carefully selected, hierarchical image, first-order derivative, and corner representations can be obtained from the wavelet decomposition. On the basis of this result, three of the four components of the wave let decomposition are employed to estimate dense optical flow with use of only two frames. This overcomes the "flattening-out" problem in traditional pyramid methods, which produce large errors when low-texture regions become flat at coarse levels as a result of blurring. A two-dimensional affine motion model is used to formulate the optical flow problem as a linear system, with all resolutions simultaneously (i.e., coarse-and-fine) rather than the traditional coarse-to-fine approach, which unavoidably propagates errors from the coarse level. This not only helps to improve the accuracy but also makes the hardware implementation of our algorithm simple. Experiments on different types of image sequences, together with quantitative and qualitative comparisons with several other optical flow methods, are given to demonstrate the effectiveness and the robustness of our algorithm.

Journal Article↗