Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,279 records · Page 71Linked to original sources

DMP3: a dynamic multilayer perceptron construction algorithm.

This paper presents DMP3 (Dynamic Multilayer Perceptron 3), a multilayer perceptron (MLP) constructive training method that constructs MLPs by incrementally adding network elements of varying complexity to the network. DMP3 differs from other MLP construction techniques in several important ways, and the motivation for these differences are given. Information gain rather than error minimization is used to guide the growth of the network, which increases the utility of newly added network elements and decreases the likelihood that a premature dead end in the growth of the network will occur. The generalization performance of DMP3 is compared with that of several other well-known machine learning and neural network learning algorithms on nine real world data sets. Simulation results show that DMP3 performs better (on average) than any of the other algorithms on the data sets tested. The main reasons for this result are discussed in detail.

Algorithms↗

FrankSum: new feature selection method for protein function prediction.

In the study of in silico functional genomics, improving the performance of protein function prediction is the ultimate goal for identifying proteins associated with defined cellular functions. The classical prediction approach is to employ pairwise sequence alignments. However this method often faces difficulties when no statistically significant homologous sequences are identified. An alternative way is to predict protein function from sequence-derived features using machine learning. In this case the choice of possible features which can be derived from the sequence is of vital importance to ensure adequate discrimination to predict function. In this paper we have successfully selected biologically significant features for protein function prediction. This was performed using a new feature selection method (FrankSum) that avoids data distribution assumptions, uses a data independent measurement (p-value) within the feature, identifies redundancy between features and uses an appropriate ranking criterion for feature selection. We have shown that classifiers generated from features selected by FrankSum outperforms classifiers generated from full feature sets, randomly selected features and features selected from the Wrapper method. We have also shown the features are concordant across all species and top ranking features are biologically informative. We conclude that feature selection is vital for successful protein function prediction and FrankSum is one of the feature selection methods that can be applied successfully to such a domain.

Amino Acid Sequence↗

Data mining tools for biological sequences.

We describe a methodology, as well as some related data mining tools, for analyzing sequence data. The methodology comprises three steps: (a) generating candidate features from the sequences, (b) selecting relevant features from the candidates, and (c) integrating the selected features to build a system to recognize specific properties in sequence data. We also give relevant techniques for each of these three steps. For generating candidate features, we present various types of features based on the idea of k-grams. For selecting relevant features, we discuss signal-to-noise, t-statistics, and entropy measures, as well as a correlation-based feature selection method. For integrating selected features, we use machine learning methods, including C4.5, SVM, and Naive Bayes. We illustrate this methodology on the problem of recognizing translation initiation sites. We discuss how to generate and select features that are useful for understanding the distinction between ATG sites that are translation initiation sites and those that are not. We also discuss how to use such features to build reliable systems for recognizing translation initiation sites in DNA sequences.

Artificial Intelligence↗

Predicting risk of coronary artery disease from DNA microarray-based genotyping using neural networks and other statistical analysis tool.

This paper presents a novel approach for complex disease prediction that we have developed, exemplified by a study on risk of coronary artery disease (CAD). This multi-disciplinary approach straddles fields of microarray technology and genetics, neural networks (NN), data mining and machine learning, as well as traditional statistical analysis techniques, namely principal components analysis (PCA) and factor analysis (FA). A description of the biological background of the study is given, followed by a detailed description of how the problem has been modeled for analyses by neural networks and FA. A committee learning approach for NN has been used to improve generalization rates. We show that our NN approach is able to yield promising prediction results despite using only the most fundamental network structures. More interestingly, through the statistical analysis process, genes of similar biological functions have been clustered. In addition, a gene marker involved in breaking down lipids has been found to be the most correlated to CAD.

Algorithms↗

Hidden Markov Models, grammars, and biology: a tutorial.

Biological sequences and structures have been modelled using various machine learning techniques and abstract mathematical concepts. This article surveys methods using Hidden Markov Model and functional grammars for this purpose. We provide a formal introduction to Hidden Markov Model and grammars, stressing on a comprehensive mathematical description of the methods and their natural continuity. The basic algorithms and their application to analyzing biological sequences and modelling structures of bio-molecules like proteins and nucleic acids are discussed. A comparison of the different approaches is discussed, and possible areas of work and problems are highlighted. Related databases and softwares, available on the internet, are also mentioned.

Algorithms↗

Support vector machines for prediction and analysis of beta and gamma-turns in proteins.

Tight turns have long been recognized as one of the three important features of proteins, together with alpha-helix and beta-sheet. Tight turns play an important role in globular proteins from both the structural and functional points of view. More than 90% tight turns are beta-turns and most of the rest are gamma-turns. Analysis and prediction of beta-turns and gamma-turns is very useful for design of new molecules such as drugs, pesticides, and antigens. In this paper we investigated two aspects of applying support vector machine (SVM), a promising machine learning method for bioinformatics, to prediction and analysis of beta-turns and gamma-turns. First, we developed two SVM-based methods, called BTSVM and GTSVM, which predict beta-turns and gamma-turns in a protein from its sequence. When compared with other methods, BTSVM has a superior performance and GTSVM is competitive. Second, we used SVMs with a linear kernel to estimate the support of amino acids for the formation of beta-turns and gamma-turns depending on their position in a protein. Our analysis results are more comprehensive and easier to use than the previous results in designing turns in proteins.

Algorithms↗

Decision tree based information integration for automated protein classification.

We propose a novel technique for automatically generating the SCOP classification of a protein structure with high accuracy. We achieve accurate classification by combining the decisions of multiple methods using the consensus of a committee (or an ensemble) classifier. Our technique, based on decision trees, is rooted in machine learning which shows that by judicially employing component classifiers, an ensemble classifier can be constructed to outperform its components. We use two sequence- and three structure-comparison tools as component classifiers. Given a protein structure and using the joint hypothesis, we first determine if the protein belongs to an existing category (family, superfamily, fold) in the SCOP hierarchy. For the proteins that are predicted as members of the existing categories, we compute their family-, superfamily-, and fold-level classifications using the consensus classifier. We show that we can significantly improve the classification accuracy compared to the individual component classifiers. In particular, we achieve error rates that are 3-12 times less than the individual classifiers' error rates at the family level, 1.5-4.5 times less at the superfamily level, and 1.1-2.4 times less at the fold level.

Algorithms↗

PRIMA: peptide robust identification from MS/MS spectra.

In proteomics, tandem mass spectrometry is the key technology for peptide sequencing. However, partially due to the deficiency of peptide identification software, a large portion of the tandem mass spectra are discarded in almost all proteomics centers because they are not interpretable. The problem is more acute with the lower quality data from low end but more popular devices such as the ion trap instruments. In order to deal with the noisy and low quality data, this paper develops a systematic machine learning approach to construct a robust linear scoring function, whose coefficients are determined by a linear programming. A prototype, PRIMA, was implemented. When tested with large benchmarks of varying qualities, PRIMA consistently has higher accuracy than commonly used software MASCOT, SEQUEST and X! Tandem.

Amino Acid Sequence↗

Symmetric and asymmetric multi-modality biclustering analysis for microarray data matrix.

Machine learning techniques offer a viable approach to cluster discovery from microarray data, which involves identifying and classifying biologically relevant groups in genes and conditions. It has been recognized that genes (whether or not they belong to the same gene group) may be co-expressed via a variety of pathways. Therefore, they can be adequately described by a diversity of coherence models. In fact, it is known that a gene may participate in multiple pathways that may or may not be co-active under all conditions. It is therefore biologically meaningful to simultaneously divide genes into functional groups and conditions into co-active categories--leading to the so-called biclustering analysis. For this, we have proposed a comprehensive set of coherence models to cope with various plausible regulation processes. Furthermore, a multivariate biclustering analysis based on fusion of different coherence models appears to be promising because the expression level of genes from the same group may follow more than one coherence models. The simulation studies further confirm that the proposed framework enjoys the advantage of high prediction performance.

Algorithms↗

ANGLE: a sequencing errors resistant program for predicting protein coding regions in unfinished cDNA.

In the process of making full-length cDNA, predicting protein coding regions helps both in the preliminary analysis of genes and in any succeeding process. However, unfinished cDNA contains artifacts including many sequencing errors, which hinder the correct evaluation of coding sequences. Especially, predictions of short sequences are difficult because they provide little information for evaluating coding potential. In this paper, we describe ANGLE, a new program for predicting coding sequences in low quality cDNA. To achieve error-tolerant prediction, ANGLE uses a machine-learning approach, which makes better expression of coding sequence maximizing the use of limited information from input sequences. Our method utilizes not only codon usage, but also protein structure information which is difficult to be used for stochastic model-based algorithms, and optimizes limited information from a short segment when deciding coding potential, with the result that predictive accuracy does not depend on the length of an input sequence. The performance of ANGLE is compared with ESTSCAN on four dataset each of them having a different error rate (one frame-shift error or one substitution error per 200-500 nucleotides) and on one dataset which has no error. ANGLE outperforms ESTSCAN by 9.26% in average Matthews's correlation coefficient on short sequence dataset (< 1000 bases). On long sequence dataset, ANGLE achieves comparable performance.

Algorithms↗

New Insights into Genomic Variations and Mutational Events Associated with Plant-Pathogen Interactions.

Plant diseases threaten global food security, causing up to 40% crop yield losses and more than $220 billion in annual economic damage. This review synthesizes recent advances in understanding the genomic variations and mutational events underlying plant-pathogen interactions and durable plant disease resistance. Key insights into evolutionary dynamics, genetic variability, and coadaptive strategies reveal the complexity of host-pathogen relationships and the implications for developing durable disease resistance. Integrative approaches combining genome-wide association studies and functional genomics have uncovered the polygenic and epistatic architecture of quantitative resistance. Advances in pan-genomics and high-throughput sequencing have revealed extensive genetic variability in cultivated/elite germplasm and wild relatives. Emerging technologies, including gene editing, multi-omics, and machine learning, enable predictive modeling of resistance traits and support evolution that informs plant breeding strategies. Collectively, these advances provide a robust framework for developing durable resistance and sustainable crop protection in the face of global agricultural challenges.

Host-Pathogen Interactions↗

Automatic structuring of radiology free-text reports.

A natural language processor was developed that automatically structures the important medical information (eg, the existence, properties, location, and diagnostic interpretation of findings) contained in a radiology free-text document as a formal information model that can be interpreted by a computer program. The input to the system is a free-text report from a radiologic study. The system requires no reporting style changes on the part of the radiologist. Statistical and machine learning methods are used extensively throughout the system. A graphical user interface has been developed that allows the creation of hand-tagged training examples. Various aspects of the difficult problem of implementing an automated structured reporting system have been addressed, and the relevant technology is progressing well. Extensible Markup Language is emerging as the preferred syntactic standard for representing and distributing these structured reports within a clinical environment. Early successes hold out hope that similar statistically based models of language will allow deep understanding of textual reports. The success of these statistical methods will depend on the availability of large numbers of high-quality training examples for each radiologic subdomain. The acceptability of automated structured reporting systems will ultimately depend on the results of comprehensive evaluations.

Humans↗

A cell proliferation signature is a marker of extremely poor outcome in a subpopulation of breast cancer patients.

Breast cancer comprises a group of distinct subtypes that despite having similar histologic appearances, have very different metastatic potentials. Being able to identify the biological driving force, even for a subset of patients, is crucially important given the large population of women diagnosed with breast cancer. Here, we show that within a subset of patients characterized by relatively high estrogen receptor expression for their age, the occurrence of metastases is strongly predicted by a homogeneous gene expression pattern almost entirely consisting of cell cycle genes (5-year odds ratio of metastasis, 24.0; 95% confidence interval, 6.0-95.5). Overexpression of this set of genes is clearly associated with an extremely poor outcome, with the 10-year metastasis-free probability being only 24% for the poor group, compared with 85% for the good group. In contrast, this gene expression pattern is much less correlated with the outcome in other patient subpopulations. The methods described here also illustrate the value of combining clinical variables, biological insight, and machine-learning to dissect biological complexity. Our work presented here may contribute a crucial step towards rational design of personalized treatment.

Age Factors↗

Possible detection of pancreatic cancer by plasma protein profiling.

The survival rate of pancreatic cancer patients is the lowest among those with common solid tumors, and early detection is one of the most feasible means of improving outcomes. We compared plasma proteomes between pancreatic cancer patients and sex- and age-matched healthy controls using surface-enhanced laser desorption/ionization coupled with hybrid quadrupole time-of-flight mass spectrometry. Proteomic spectra were generated from a total of 245 plasma samples obtained from two institutes. A discriminating proteomic pattern was extracted from a training cohort (71 pancreatic cancer patients and 71 healthy controls) using a support vector machine learning algorithm and was applied to two validation cohorts. We recognized a set of four mass peaks at 8,766, 17,272, 28,080, and 14,779 m/z, whose mean intensities differed significantly (Mann-Whitney U test, P < 0.01), as most accurately discriminating cancer patients from healthy controls in the training cohort [sensitivity of 97.2% (69 of 71), specificity of 94.4% (67 of 71), and area under the curve value of 0.978]. This set discriminated cancer patients in the first validation cohort with a sensitivity of 90.9% (30 of 33) and a specificity of 91.1% (41 of 45), and its discriminating capacity was further validated in an independent cohort at a second institution. When combined with CA19-9, 100% (29 of 29 patients) of pancreatic cancers, including early-stage (stages I and II) tumors, were detected. Although a multi-institutional large-scale study will be necessary to confirm clinical significance, the biomarker set identified in this study may be applicable to using plasma samples to diagnose pancreatic cancer.

Biomarkers, Tumor↗

A multimarker model to predict outcome in tamoxifen-treated breast cancer patients.

PURPOSE: This study was designed to produce a model to predict outcome in tamoxifen-treated breast cancer patients based on clinicopathologic features and multiple molecular markers. EXPERIMENTAL DESIGN: This was a retrospective study of 324 stage I to III female breast cancer patients treated with tamoxifen for whom standard clinicopathologic data and tumor tissue microarrays were available. Nine molecular markers were studied by semiquantitative immunohistochemistry and/or fluorescence in situ hybridization. Cox proportional hazards analysis was used to determine the contributions of each variable to disease-specific and overall survival, and machine learning was used to produce a model to predict patient outcome. RESULTS: On a univariate basis, the following features were significantly associated with worse survival: high pathologic tumor or nodal class, histologic grade, epidermal growth factor receptor, ERBB2, MYC, or TP53; absent estrogen receptor (ER) or progesterone receptor; and low BCL2. CCND1 and CDKN1B did not reach statistical significance. On a multivariate basis, nodal class, ER, and MYC were statistically significant as independent factors for survival. However, the benefit of ER-positive status was moderated by BCL2, ERBB2, and progesterone receptor. BCL2 and TP53 also interacted as an independent risk factor. A kernel partial least squares polynomial model was developed with an area under the receiver operating characteristic curve of 0.90. CONCLUSIONS: Our data show the predictive value of BCL2, ERBB2, MYC, and TP53 in addition to the standard hormone receptors and clinicopathologic features, and they show the importance of conditional interpretation of certain molecular markers. Our multimarker predictive model performed significantly better than standard guidelines.

Aged↗

Predicting cancer drug response by proteomic profiling.

PURPOSE: Accurate prediction of an individual patient's drug response is an important prerequisite of personalized medicine. Recent pharmacogenomics research in chemosensitivity prediction has studied the gene-drug correlation based on transcriptional profiling. However, proteomic profiling will more directly solve the current functional and pharmacologic problems. We sought to determine whether proteomic signatures of untreated cells were sufficient for the prediction of drug response. EXPERIMENTAL DESIGN: In this study, a machine learning model system was developed to classify cell line chemosensitivity exclusively based on proteomic profiling. Using reverse-phase protein lysate microarrays, protein expression levels were measured by 52 antibodies in a panel of 60 human cancer cell (NCI-60) lines. The model system combined several well-known algorithms, including random forests, Relief, and the nearest neighbor methods, to construct the protein expression--based chemosensitivity classifiers. The classifiers were designed to be independent of the tissue origin of the cells. RESULTS: A total of 118 classifiers of the complete range of drug responses (sensitive, intermediate, and resistant) were generated for the evaluated anticancer drugs, one for each agent. The accuracy of chemosensitivity prediction of all the evaluated 118 agents was significantly higher (P < 0.02) than that of random prediction. Furthermore, our study found that the proteomic determinants for chemosensitivity of 5-fluorouracil were also potential diagnostic markers of colon cancer. CONCLUSIONS: The results showed that it was feasible to accurately predict chemosensitivity by proteomic approaches. This study provides a basis for the prediction of drug response based on protein markers in the untreated tumors.

Antineoplastic Agents↗

Statistical evaluation of local alignment features predicting allergenicity using supervised classification algorithms.

BACKGROUND: Recently, two promising alignment-based features predicting food allergenicity using the k nearest neighbor (kNN) classifier were reported. These features are the alignment score and alignment length of the best local alignment obtained in a database of known allergen sequences. METHODS: In the work reported here a much more comprehensive statistical evaluation of the potential of these features was performed, this time for the prediction of allergenicity in general. The evaluation consisted of the following four key components. (1) A new high quality database consisting of 318 carefully selected, non-redundant allergens and 1,007 sequences carefully selected to be non-allergens. (2) Three different supervised algorithms: the kNN classifier, the Bayesian linear Gaussian classifier, and the Bayesian quadratic Gaussian classifier. (3) A large set of local alignment procedures defined using the FASTA3 alignment program by means of a wide range of different parameter settings. (4) Novel performance curves, alternative to conventional receiver-operating characteristic curves, to display not only average behaviors but also statistical variations due to small data sets. RESULTS: The linear Gaussian classifier proved most useful among the tested supervised machine learning algorithms, closely followed by the quadratic Gaussian equivalent and kNN. The overall best classification results were obtained with a novel feature vector consisting of the combined alignment scores derived from local alignment procedures using different substitution matrices. CONCLUSIONS: The models reported here should be useful as a part of an integrated assessment scheme for potential protein allergenicity and for future comparisons with alternative bioinformatic approaches.

Algorithms↗

SNP-based analysis of genetic substructure in the German population.

OBJECTIVE: To evaluate the relevance and necessity to account for the effects of population substructure on association studies under a case-control design in central Europe, we analysed three samples drawn from different geographic areas of Germany. Two of the three samples, POPGEN (n = 720) and SHIP (n = 709), are from north and north-east Germany, respectively, and one sample, KORA (n = 730), is from southern Germany. METHODS: Population genetic differentiation was measured by classical F-statistics for different marker sets, either consisting of genome-wide selected coding SNPs located in functional genes, or consisting of selectively neutral SNPs from 'genomic deserts'. Quantitative estimates of the degree of stratification were performed comparing the genomic control approach [Devlin B, Roeder K: Biometrics 1999;55:997-1004], structured association [Pritchard JK, Stephens M, Donnelly P: Genetics 2000;155:945-959] and sophisticated methods like random forests [Breiman L: Machine Learning 2001;45:5-32]. RESULTS: F-statistics showed that there exists a low genetic differentiation between the samples along a north-south gradient within Germany (F(ST)(KORA/POPGEN): 1.7 . 10(-4); F(ST)(KORA/SHIP): 5.4 . 10(-4); F(ST)(POPGEN/SHIP): -1.3 . 10(-5)). CONCLUSION: Although the F(ST )-values are very small, indicating a minor degree of population structure, and are too low to be detectable from methods without using prior information of subpopulation membership, such as STRUCTURE [Pritchard JK, Stephens M, Donnelly P: Genetics 2000;155:945-959], they may be a possible source for confounding due to population stratification.

Case-Control Studies↗