Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,225 records · Page 68Linked to original sources

Automatic identification of subcellular phenotypes on human cell arrays.

Light microscopic analysis of cell morphology provides a high-content readout of cell function and protein localization. Cell arrays and microwell transfection assays on cultured cells have made cell phenotype analysis accessible to high-throughput experiments. Both the localization of each protein in the proteome and the effect of RNAi knock-down of individual genes on cell morphology can be assayed by manual inspection of microscopic images. However, the use of morphological readouts for functional genomics requires fast and automatic identification of complex cellular phenotypes. Here, we present a fully automated platform for high-throughput cell phenotype screening combining human live cell arrays, screening microscopy, and machine-learning-based classification methods. Efficiency of this platform is demonstrated by classification of eleven subcellular patterns marked by GFP-tagged proteins. Our classification method can be adapted to virtually any microscopic assay based on cell morphology, opening a wide range of applications including large-scale RNAi screening in human cells.

Artifacts↗

Establishing glucose- and ABA-regulated transcription networks in Arabidopsis by microarray analysis and promoter classification using a Relevance Vector Machine.

Establishing transcriptional regulatory networks by analysis of gene expression data and promoter sequences shows great promise. We developed a novel promoter classification method using a Relevance Vector Machine (RVM) and Bayesian statistical principles to identify discriminatory features in the promoter sequences of genes that can correctly classify transcriptional responses. The method was applied to microarray data obtained from Arabidopsis seedlings treated with glucose or abscisic acid (ABA). Of those genes showing >2.5-fold changes in expression level, approximately 70% were correctly predicted as being up- or down-regulated (under 10-fold cross-validation), based on the presence or absence of a small set of discriminative promoter motifs. Many of these motifs have known regulatory functions in sugar- and ABA-mediated gene expression. One promoter motif that was not known to be involved in glucose-responsive gene expression was identified as the strongest classifier of glucose-up-regulated gene expression. We show it confers glucose-responsive gene expression in conjunction with another promoter motif, thus validating the classification method. We were able to establish a detailed model of glucose and ABA transcriptional regulatory networks and their interactions, which will help us to understand the mechanisms linking metabolism with growth in Arabidopsis. This study shows that machine learning strategies coupled to Bayesian statistical methods hold significant promise for identifying functionally significant promoter sequences.

Abscisic Acid↗

Subsystem identification through dimensionality reduction of large-scale gene expression data.

The availability of parallel, high-throughput biological experiments that simultaneously monitor thousands of cellular observables provides an opportunity for investigating cellular behavior in a highly quantitative manner at multiple levels of resolution. One challenge to more fully exploit new experimental advances is the need to develop algorithms to provide an analysis at each of the relevant levels of detail. Here, the data analysis method non-negative matrix factorization (NMF) has been applied to the analysis of gene array experiments. Whereas current algorithms identify relationships on the basis of large-scale similarity between expression patterns, NMF is a recently developed machine learning technique capable of recognizing similarity between subportions of the data corresponding to localized features in expression space. A large data set consisting of 300 genome-wide expression measurements of yeast was used as sample data to illustrate the performance of the new approach. Local features detected are shown to map well to functional cellular subsystems. Functional relationships predicted by the new analysis are compared with those predicted using standard approaches; validation using bioinformatic databases suggests predictions using the new approach may be up to twice as accurate as some conventional approaches.

Algorithms↗

Molecular scene analysis: the integration of direct-methods and artificial-intelligence strategies for solving protein crystal structure.

A knowledge-based approach to crystal structure determination is presented. The approach integrates direct-methods and artificial-intelligence strategies to rephrase the structure determination process as an exercise in scene analysis. A general joint probability distribution framework, which allows the incorporation of isomorphous replacement, anomalous scattering and a priori structural information, forms the basis of the direct-methods strategies. The accumulated knowledge on crystal and molecular structures is exploited through the use of artificial-intelligence strategies, which include techniques of knowledge representation, search and machine learning.

Journal Article↗

Automated protein classification using consensus decision.

We propose a novel technique for automatically generating the SCOP classification of a protein structure with high accuracy. High accuracy is achieved by combining the decisions of multiple methods using the consensus of a committee (or an ensemble) classifier. Our technique is rooted in machine learning which shows that by judicially employing component classifiers, an ensemble classifier can be constructed to outperform its components. We use two sequence- and three structure-comparison tools as component classifiers. Given a protein structure, using the joint hypothesis, we first determine if the protein belongs to an existing category (family, superfamily, fold) in the SCOP hierarchy. For the proteins that are predicted as members of the existing categories, we compute their family-, superfamily-, and fold-level classifications using the consensus classifier. We show that we can significantly improve the classification accuracy compared to the individual component classifiers. In particular, we achieve error rates that are 3-12 times less than the individual classifiers' error rates at the family level, 1.5-4.5 times less at the superfamily level, and 1.1-2.4 times less at the fold level.

Algorithms↗

A robust meta-classification strategy for cancer diagnosis from gene expression data.

One of the major challenges in cancer diagnosis from microarray data is to develop robust classification models which are independent of the analysis techniques used and can combine data from different laboratories. We propose a meta-classification scheme which uses a robust multivariate gene selection procedure and integrates the results of several machine learning tools trained on raw and pattern data. We validate our method by applying it to distinguish diffuse large B-cell lymphoma (DLBCL) from follicular lymphoma (FL) on two independent datasets: the HuGeneFL Affmetrixy dataset of Shipp et al. (www. genome.wi.mit.du/MPR /lymphoma) and the Hu95Av2 Affymetrix dataset (DallaFavera's laboratory, Columbia University). Our meta-classification technique achieves higher predictive accuracies than each of the individual classifiers trained on the same dataset and is robust against various data perturbations. We also find that combinations of p53 responsive genes (e.g., p53, PLK1 and CDK2) are highly predictive of the phenotype.

Algorithms↗

Spatio-spectral filters for improving the classification of single trial EEG.

Data recorded in electroencephalogram (EEG)-based brain-computer interface experiments is generally very noisy, non-stationary, and contaminated with artifacts that can deteriorate discrimination/classification methods. In this paper, we extend the common spatial pattern (CSP) algorithm with the aim to alleviate these adverse effects. In particular, we suggest an extension of CSP to the state space, which utilizes the method of time delay embedding. As we will show, this allows for individually tuned frequency filters at each electrode position and, thus, yields an improved and more robust machine learning procedure. The advantages of the proposed method over the original CSP method are verified in terms of an improved information transfer rate (bits per trial) on a set of EEG-recordings from experiments of imagined limb movements.

Brain↗

Associative clustering for exploring dependencies between functional genomics data sets.

High-throughput genomic measurements, interpreted as cooccurring data samples from multiple sources, open up a fresh problem for machine learning: What is in common in the different data sets, that is, what kind of statistical dependencies are there between the paired samples from the different sets? We introduce a clustering algorithm for exploring the dependencies. Samples within each data set are grouped such that the dependencies between groups of different sets capture as much of pairwise dependencies between the samples as possible. We formalize this problem in a novel probabilistic way, as optimization of a Bayes factor. The method is applied to reveal commonalities and exceptions in gene expression between organisms and to suggest regulatory interactions in the form of dependencies between gene expression profiles and regulator binding patterns.

Algorithms↗

The applicability of recurrent neural networks for biological sequence analysis.

Selection of machine learning techniques requires a certain sensitivity to the requirements of the problem. In particular, the problem can be made more tractable by deliberately using algorithms that are biased toward solutions of the requisite kind. In this paper, we argue that recurrent neural networks have a natural bias toward a problem domain of which biological sequence analysis tasks are a subset. We use experiments with synthetic data to illustrate this bias. We then demonstrate that this bias can be exploitable using a data set of protein sequences containing several classes of subcellular localization targeting peptides. The results show that, compared with feed forward, recurrent neural networks will generally perform better on sequence analysis tasks. Furthermore, as the patterns within the sequence become more ambiguous, the choice of specific recurrent architecture becomes more critical.

Algorithms↗

Efficiently mining gene expression data via a novel parameterless clustering method.

Clustering analysis has been an important research topic in the machine learning field due to the wide applications. In recent years, it has even become a valuable and useful tool for in-silico analysis of microarray or gene expression data. Although a number of clustering methods have been proposed, they are confronted with difficulties in meeting the requirements of automation, high quality, and high efficiency at the same time. In this paper, we propose a novel, parameterless and efficient clustering algorithm, namely, Correlation Search Technique (CST), which fits for analysis of gene expression data. The unique feature of CST is it incorporates the validation techniques into the clustering process so that high quality clustering results can be produced on the fly. Through experimental evaluation, CST is shown to outperform other clustering methods greatly in terms of clustering quality, efficiency, and automation on both of synthetic and real data sets.

Algorithms↗

Multitraining support vector machine for image retrieval.

Relevance feedback (RF) schemes based on support vector machines (SVMs) have been widely used in content-based image retrieval (CBIR). However, the performance of SVM-based RF approaches is often poor when the number of labeled feedback samples is small. This is mainly due to 1) the SVM classifier being unstable for small-size training sets because its optimal hyper plane is too sensitive to the training examples; and 2) the kernel method being ineffective because the feature dimension is much greater than the size of the training samples. In this paper, we develop a new machine learning technique, multitraining SVM (MTSVM), which combines the merits of the cotraining technique and a random sampling method in the feature space. Based on the proposed MTSVM algorithm, the above two problems can be mitigated. Experiments are carried out on a large image set of some 20,000 images, and the preliminary results demonstrate that the developed method consistently improves the performance over conventional SVM-based RFs in terms of precision and standard deviation, which are used to evaluate the effectiveness and robustness of a RF algorithm, respectively.

Algorithms↗

Multiple exemplar-based facial image retrieval using independent component analysis.

In this paper, we design a content-based image retrieval system where multiple query examples can be used to indicate the need to retrieve not only images similar to the individual examples, but also those images which actually represent a combination of the content of query images. We propose a scheme for representing content of an image as a combination of features from multiple examples. This scheme is exploited for developing a multiple example-based retrieval engine. We have explored the use of machine learning techniques for generating the most appropriate feature combination scheme for a given class of images. The combination scheme can be used for developing purposive query engines for specialized image databases. Here, we have considered facial image databases. The effectiveness of the image retrieval system is experimentally demonstrated on different databases.

Algorithms↗

Medical diagnosis with C4.5 Rule preceded by artificial neural network ensemble.

Comprehensibility is very important for any machine learning technique to be used in computer-aided medical diagnosis. Since an artificial neural network ensemble is composed of multiple artificial neural networks, its comprehensibility is worse than that of a single artificial neural network. In this paper, C4.5 Rule-PANE which combines artificial neural network ensemble with rule induction by regarding the former as a preprocess of the latter, is proposed. At first, an artificial neural network ensemble is trained. Then, a new training data set is generated by feeding the feature vectors of the original training instances to the trained ensemble and replacing the expected class labels of the original training instances with the class labels output from the ensemble. Additional training data may also be appended by randomly generating feature vectors and combining them with their corresponding class labels output from the ensemble. Finally, a specific rule induction approach, i.e., C4.5 Rule, is used to learn rules from the new training data set. Case studies on diabetes, hepatitis, and breast cancer show that C4.5 Rule-PANE could generate rules with strong generalization ability, which profits from artificial neural network ensemble, and strong comprehensibility, which profits from rule induction.

Diagnosis↗

The combined technique for detection of artifacts in clinical electroencephalograms of sleeping newborns.

In this paper, we describe a new method combining the polynomial neural network and decision tree techniques in order to derive comprehensible classification rules from clinical electroencephalograms (EEGs) recorded from sleeping newborns. These EEGs are heavily corrupted by cardiac, eye movement, muscle, and noise artifacts and, as a consequence, some EEG features are irrelevant to classification problems. Combining the polynomial network and decision tree techniques, we discover comprehensible classification rules while also attempting to keep their classification error down. This technique is shown to out-perform a number of commonly used machine learning technique applied to automatically recognize artifacts in the sleep EEGs.

Algorithms↗

An evaluation of contrast enhancement techniques for mammographic breast masses.

The main aim of this paper is to propose a novel set of metrics that measure the quality of the image enhancement of mammographic images in a computer-aided detection framework aimed at automatically finding masses using machine learning techniques. Our methodology includes a novel mechanism for the combination of the metrics proposed into a single quantitative measure. We have evaluated our methodology on 200 images from the publicly available digital database for screening mammograms. We show that the quantitative measures help us select the best suited image enhancement on a per mammogram basis, which improves the quality of subsequent image segmentation much better than using the same enhancement method for all mammograms.

Algorithms↗

Feature subset selection for improving the performance of false positive reduction in lung nodule CAD.

We propose a feature subset selection method based on genetic algorithms to improve the performance of false positive reduction in lung nodule computer-aided detection (CAD). It is coupled with a classifier based on support vector machines. The proposed approach determines automatically the optimal size of the feature set, and chooses the most relevant features from a feature pool. Its performance was tested using a lung nodule database (52 true nodules and 443 false ones) acquired by multislice CT scans. From 23 features calculated for each detected structure, the suggested method determined ten to be the optimal feature subset size, and selected the most relevant ten features. A support vector machine classifier trained with the optimal feature subset resulted in 100% sensitivity and 56.4% specificity using an independent validation set. Experiments show significant improvement achieved by a system incorporating the proposed method over a system without it. This approach can be also applied to other machine learning problems; e.g. computer-aided diagnosis of lung nodules.

Algorithms↗

Automated analysis of brachial ultrasound image sequences: early detection of cardiovascular disease via surrogates of endothelial function.

Early detection of cardiovascular disease would allow timely institution of preventive measures. Arterial endothelium play a primary role in processes leading to the development of atherosclerotic plaque and cardiovascular disease in general. Determination of flow-mediated dilatation (FMD) of brachial arteries from B-mode ultrasound image sequences offers a noninvasive surrogate index of endothelial function. A highly automated method for analysis of brachial ultrasound image sequences is reported and its performance assessed. The method overcomes the variability of brachial ultrasound images across subjects by incorporating machine learning and quality control steps. The automated method outperformed conventional manual analysis by providing a decreased analysis bias, increased reproducibility, and improved measurement accuracy. Consequently, it decreases inter- and intraobserver as well interinstitution variability. The method has been employed in a number of population studies with thousands of subjects analyzed.

Arterial Occlusive Diseases↗

Rule generation for protein secondary structure prediction with support vector machines and decision tree.

Support vector machines (SVMs) have shown strong generalization ability in a number of application areas, including protein structure prediction. However, the poor comprehensibility hinders the success of the SVM for protein structure prediction. The explanation of how a decision made is important for accepting the machine learning technology, especially for applications such as bioinformatics. The reasonable interpretation is not only useful to guide the "wet experiments," but also the extracted rules are helpful to integrate computational intelligence with symbolic AI systems for advanced deduction. On the other hand, a decision tree has good comprehensibility. In this paper, a novel approach to rule generation for protein secondary structure prediction by integrating merits of both the SVM and decision tree is presented. This approach combines the SVM with decision tree into a new algorithm called SVM_ DT, which proceeds in three steps. This algorithm first trains an SVM. Then, a new training set is generated through careful selection from the output of the SVM. Finally, the obtained training set is used to train a decision tree learning system and to extract the corresponding rule sets. The results of the experiments of protein secondary structure prediction on RS126 data set show that the comprehensibility of SVM_DT is much better than that of the SVM. Moreover, the generalization ability of SVM_DT is better than that of C4.5 decision trees and is similar to that of the SVM. Hence, SVM_DT can be used not only for prediction, but also for guiding biological experiments.

Algorithms↗