Search PubMed⌕ Search

Biomedical subjects

Charles E Davidson

Publications and source records attributed to Charles E Davidson.

4 recordsLinked to original sources

Genetic algorithms for classification of olfactory stimulants.

We have developed and tested a genetic algorithm (GA) for pattern recognition, which identifies molecular descriptors that optimize the separation of the activity classes of olfactory stimulants in a plot of the two or three largest principal components of the data. Because principal components maximize variance, the bulk of the information encoded by these descriptors is about differences between olfactory classes in the dataset. In addition, the GA focuses on those classes and or samples that are difficult to classify as it trains using a form of boosting to modify the fitness landscape. Boosting minimizes the problem of convergence to a local optimum, because the fitness function of the GA is changing as the population is evolving toward a solution. Over time, compounds that consistently classify correctly are not as heavily weighted in the analysis as compounds that are difficult to classify. The pattern recognition GA learns its optimal parameters in a manner similar to a neural network. The algorithm integrates aspects of both strong and weak learning to yield a "smart" one-pass procedure for feature selection and classification.

Algorithms↗

Machine learning based pattern recognition applied to microarray data.

MOTIVATION: Microarrays have allowed the expression level of thousands of genes or proteins to be measured simultaneously. Data sets generated by these arrays consist of a small number of observations (e.g., 20-100 samples) on a very large number of variables (e.g., 10,000 genes or proteins). The observations in these data sets often have other attributes associated with them such as a class label denoting the pathology of the subject. Finding the genes or proteins that are correlated to these attributes is often a difficult task since most of the variables do not contain information about the pathology and as such can mask the identity of the relevant features. We describe a genetic algorithm (GA) that employs both supervised and unsupervised learning to mine gene expression and proteomic data. The pattern recognition GA selects features that increase clustering, while simultaneously searching for features that optimize the separation of the classes in a plot of the two or three largest principal components of the data. Because the largest principal components capture the bulk of the variance in the data, the features chosen by the GA contain information primarily about differences between classes in the data set. The principal component analysis routine embedded in the fitness function of the GA acts as an information filter, significantly reducing the size of the search space since it restricts the search to feature sets whose principal component plots show clustering on the basis of class. The algorithm integrates aspects of artificial intelligence and evolutionary computations to yield a smart one pass procedure for feature selection, clustering, classification, and prediction.

Algorithms↗

Electronic van der Waals surface property descriptors and genetic algorithms for developing structure-activity correlations in olfactory databases.

A methodology to facilitate the intelligent design of new odorants (e.g., musks) with specialized properties has been developed as part of an ongoing research effort in machine learning. In a traditional framework, the introduction of a new odorant is a lengthy, costly, and laborious discovery, development, and testing process. We propose to streamline this process utilizing large existing olfactory databases available through the open scientific literature as input for a new structure/activity correlation methodology. The first step in this process is to characterize each molecule in the database by an appropriate set of descriptors. To accomplish this task, an enhanced version of Breneman's Transferable Atom Equivalent (TAE) descriptor methodology will be used to create a large set of electron density derived shape/property hybrid (PEST), wavelet coefficient (WCD), and TAE histogram descriptors. We have chosen these molecular property descriptors to represent the problem because they have been shown to contain pertinent shape and electronic properties of the molecule and correlate with key modes of intermolecular interactions. Traditional QSAR methodologies, which employ fragment based descriptors, have been shown to be effective for QSAR development within homologous sets of molecules but are less effective when applied to data sets containing a great deal of structural variation. In contrast to previous attempts at SAR, our use of shape-aware electron density based molecular property descriptors has removed many of the limitations brought about by the use of descriptors based on substructure fragments, molecular surface properties, or other whole molecule descriptors. Another reason for the mixed success of past QSAR efforts can be traced to the nature of the underlying modeling problem, which is often quite complex. To meet these challenges, a genetic algorithm for pattern recognition analysis has been developed that selects descriptors which create class separation in a plot of the two largest principal components of the data while simultaneously searching for features that increase clustering of the data.

Journal Article↗

Spectral pattern recognition using self-organizing MAPS.

A Kohonen neural network is an iterative technique used to map multivariate data. The network is able to learn and display the topology of the data. Self-organizing maps have advantages as well as drawbacks when compared to principal component plots. One advantage is that data preprocessing is usually minimal. Another is that an outlier will only affect one map unit and its neighborhood. However, outliers can have a drastic and disproportionate effect on principal component plots. Removing them does not always solve the problem for as soon as the worst outliers are deleted, other data points may appear in this role. The advantage of using self-organizing maps for spectral pattern recognition is demonstrated by way of two studies recently completed in our laboratory. In the first study, Raman spectroscopy and self-organizing maps were used to differentiate six common household plastics by type for recycling purposes. The second study involves the development of a potential method to differentiate acceptable lots from unacceptable lots of avicel using diffuse reflectance near-infrared spectroscopy and self-organizing maps.

Journal Article↗