Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Supervised Machine Learning”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Classifying vertical facial deformity using supervised and unsupervised learning.

OBJECTIVES: To evaluate the potential for machine learning techniques to identify objective criteria for classifying vertical facial deformity. METHODS: 19 parameters were determined from 131 lateral skull radiographs. Classifications were induced from raw data with simple visualisation, C5.0 and Kohonen feature maps; and using a Point Distribution Model (PDM) of shape templates comprising points taken from digitised radiographs. RESULTS: The induced decision trees enable a direct comparison of clinicians' idiosyncrasies in classification. Unsupervised algorithms induce models that are potentially more objective, but their blackbox nature makes them unsuitable for clinical application. The PDM methodology gives dramatic visualisations of two modes separating horizontal and vertical facial growth. Kohonen feature maps favour one clinician and PDM the other. Clinical response suggests that while Clinician 1 places greater weight on 5 of 6 parameters, Clinician 2 relies on more parameters that capture facial shape. CONCLUSIONS: While machine learning and statistical analyses classify subjects for vertical facial height, they have limited application in their present form. The supervised learning algorithm C5.0 is effective for generating rules for individual clinicians but its inherent bias invalidates its use for objective classification of facial form for research purposes. On the other hand, promising results from unsupervised strategies (especially the PDM) suggest a potential use for objective classification and further identification and analysis of ambiguous cases. At present, such methodologies may be unsuitable for clinical application because of the invisibility of their underlying processes. Further study is required with additional patient data and a wider group of clinicians.

Algorithms↗

Identifying marker genes in transcription profiling data using a mixture of feature relevance experts.

Transcription profiling experiments permit the expression levels of many genes to be measured simultaneously. Given profiling data from two types of samples, genes that most distinguish the samples (marker genes) are good candidates for subsequent in-depth experimental studies and developing decision support systems for diagnosis, prognosis, and monitoring. This work proposes a mixture of feature relevance experts as a method for identifying marker genes and illustrates the idea using published data from samples labeled as acute lymphoblastic and myeloid leukemia (ALL, AML). A feature relevance expert implements an algorithm that calculates how well a gene distinguishes samples, reorders genes according to this relevance measure, and uses a supervised learning method [here, support vector machines (SVMs)] to determine the generalization performances of different nested gene subsets. The mixture of three feature relevance experts examined implement two existing and one novel feature relevance measures. For each expert, a gene subset consisting of the top 50 genes distinguished ALL from AML samples as completely as all 7,070 genes. The 125 genes at the union of the top 50s are plausible markers for a prototype decision support system. Chromosomal aberration and other data support the prediction that the three genes at the intersection of the top 50s, cystatin C, azurocidin, and adipsin, are good targets for investigating the basic biology of ALL/AML. The same data were employed to identify markers that distinguish samples based on their labels of T cell/B cell, peripheral blood/bone marrow, and male/female. Selenoprotein W may discriminate T cells from B cells. Results from analysis of transcription profiling data from tumor/nontumor colon adenocarcinoma samples support the general utility of the aforementioned approach. Theoretical issues such as choosing SVM kernels and their parameters, training and evaluating feature relevance experts, and the impact of potentially mislabeled samples on marker identification (feature selection) are discussed.

Acute Disease↗

Machine learning to support diagnostics in the domain of asymptomatic liver disease.

Machine learning procedures, in unsupervised and supervised manner, can enable their users to achieve knowledge hardly comprehensible by even the best experts. This is true also if the clinical knowledge has been carefully assembled in a prospective way. A data set including 165 patients with elevated routine laboratory tests was extensively studied according to clinical history, laboratory profile and liver biopsy. Unsupervised learning by Kohonen feature map disclosed 4 groups of patients: the largest one with no or slight histopathological changes (116) and three smaller, more homogenous, with more diseased patients. Standardized histopathological scorings of the liver specimens defined patients into two groups. Fifty-eight of them were, according to the analysis, recommended for a liver biopsy, due to more severe degrees of inflammation and fibrosis. One-hundred and seven of the patients, in whom liver biopsy was retrospectively considered unnecessary, had only minor degrees of inflammation, fibrosis and/or steatosis. Supervised learning, using the inductive systems based on Quinlan's ID3 and CART algorithms, extracted knowledge in the form of decision trees. This approach could define a need for biopsy either with a very few significant findings or by pathways containing quotients and multiplications of the different basic items. These procedures were analyzed and compared for their theoretical and applicative performances. The cluster and Fischerian discriminant analyses were performed in order to compare the classification performance. The medical appropriateness of the obtained results is satisfying, therefore decision support systems, outlined in this study, should be evaluated in wider clinical practice. To achieve this goal, an example of a Medical Logical Module (MLM), based on the Arden Syntax, is given.

Adult↗

A weakly supervised deep learning-based recurrence prediction and risk stratification of lung adenocarcinoma from pathology whole-slide images.

BACKGROUND: Accurate prediction of postoperative recurrence in lung adenocarcinoma (LUAD) is essential for guiding clinical decision-making and improving patient outcomes. Although various predictive models have been developed, most rely on complex genomic analyses and high-dimensional clinical data. The complexity of these approaches substantially limits their feasibility for routine clinical use. To address this clinical challenge, this study aims to predict postoperative recurrence using routinely available hematoxylin and eosin (H&E)-stained images and characterize the associated biological features. METHODS: A total of 329 patients who underwent curative resection at the First Affiliated Hospital of Wenzhou Medical University (FHWMU) were retrospectively enrolled and randomly assigned to training and internal validation cohorts in a 7:3 ratio. An independent external validation cohort comprising 70 patients from the Clinical Proteomic Tumor Analysis Consortium (CPTAC) was included. Three patch-level feature extractors (Inception_V3, ResNet18, and DenseNet121) were evaluated within a weakly supervised multiple-instance learning (MIL) framework incorporating automated region-of-interest (ROI) detection on segmented whole-slide images (WSIs). Model performance was assessed using the area under the receiver operating characteristic curve (AUC), Kaplan-Meier (KM) survival analysis, and multivariable Cox proportional hazards regression. Transcriptomic profiling and gene set enrichment analysis (GSEA) were conducted to investigate biological differences between risk groups. RESULTS: The model achieved AUCs of 0.923 in the training cohort, 0.891 in the internal validation cohort, and 0.847 in the external validation cohort. The model effectively stratified patients into high- and low-risk groups with significantly different recurrence-free survival (RFS) across all cohorts (all P&#x2009;<&#x2009;0.001) and retained prognostic value within AJCC stages I-III. Transcriptomic analyses revealed consistent enrichment of cell cycle-related pathways and neutrophil extracellular trap (NET) formation in high-risk patients across both institutional and CPTAC cohorts, aligning with distinct biological profiles of the model-derived risk stratification. CONCLUSIONS: This weakly supervised deep learning framework enables accurate and externally validated prediction of postoperative recurrence in LUAD using routinely available histopathological images, and integration of histopathological features with molecular analyses enhances biological interpretability. This work provides a clinically accessible and cost-effective tool for postoperative risk assessment in LUAD patients.

Humans↗

Profile-based direct kernels for remote homology detection and fold recognition.

MOTIVATION: Protein remote homology detection is a central problem in computational biology. Supervised learning algorithms based on support vector machines are currently one of the most effective methods for remote homology detection. The performance of these methods depends on how the protein sequences are modeled and on the method used to compute the kernel function between them. RESULTS: We introduce two classes of kernel functions that are constructed by combining sequence profiles with new and existing approaches for determining the similarity between pairs of protein sequences. These kernels are constructed directly from these explicit protein similarity measures and employ effective profile-to-profile scoring schemes for measuring the similarity between pairs of proteins. Experiments with remote homology detection and fold recognition problems show that these kernels are capable of producing results that are substantially better than those produced by all of the existing state-of-the-art SVM-based methods. In addition, the experiments show that these kernels, even when used in the absence of profiles, produce results that are better than those produced by existing non-profile-based schemes. AVAILABILITY: The programs for computing the various kernel functions are available on request from the authors.

Algorithms↗

Genomic data sampling and its effect on classification performance assessment.

BACKGROUND: Supervised classification is fundamental in bioinformatics. Machine learning models, such as neural networks, have been applied to discover genes and expression patterns. This process is achieved by implementing training and test phases. In the training phase, a set of cases and their respective labels are used to build a classifier. During testing, the classifier is used to predict new cases. One approach to assessing its predictive quality is to estimate its accuracy during the test phase. Key limitations appear when dealing with small-data samples. This paper investigates the effect of data sampling techniques on the assessment of neural network classifiers. RESULTS: Three data sampling techniques were studied: Cross-validation, leave-one-out, and bootstrap. These methods are designed to reduce the bias and variance of small-sample estimations. Two prediction problems based on small-sample sets were considered: Classification of microarray data originating from a leukemia study and from small, round blue-cell tumours. A third problem, the prediction of splice-junctions, was analysed to perform comparisons. Different accuracy estimations were produced for each problem. The variations are accentuated in the small-data samples. The quality of the estimates depends on the number of train-test experiments and the amount of data used for training the networks. CONCLUSION: The predictive quality assessment of biomolecular data classifiers depends on the data size, sampling techniques and the number of train-test experiments. Conservative and optimistic accuracy estimations can be obtained by applying different methods. Guidelines are suggested to select a sampling technique according to the complexity of the prediction problem under consideration.

Computational Biology↗

Building multiclass classifiers for remote homology detection and fold recognition.

BACKGROUND: Protein remote homology detection and fold recognition are central problems in computational biology. Supervised learning algorithms based on support vector machines are currently one of the most effective methods for solving these problems. These methods are primarily used to solve binary classification problems and they have not been extensively used to solve the more general multiclass remote homology prediction and fold recognition problems. RESULTS: We present a comprehensive evaluation of a number of methods for building SVM-based multiclass classification schemes in the context of the SCOP protein classification. These methods include schemes that directly build an SVM-based multiclass model, schemes that employ a second-level learning approach to combine the predictions generated by a set of binary SVM-based classifiers, and schemes that build and combine binary classifiers for various levels of the SCOP hierarchy beyond those defining the target classes. CONCLUSION: Analyzing the performance achieved by the different approaches on four different datasets we show that most of the proposed multiclass SVM-based classification approaches are quite effective in solving the remote homology prediction and fold recognition problems and that the schemes that use predictions from binary models constructed for ancestral categories within the SCOP hierarchy tend to not only lead to lower error rates but also reduce the number of errors in which a superfamily is assigned to an entirely different fold and a fold is predicted as being from a different SCOP class. Our results also show that the limited size of the training data makes it hard to learn complex second-level models, and that models of moderate complexity lead to consistently better results.

Algorithms↗

Knowledge-based analysis of microarray gene expression data by using support vector machines.

We introduce a method of functionally classifying genes by using gene expression data from DNA microarray hybridization experiments. The method is based on the theory of support vector machines (SVMs). SVMs are considered a supervised computer learning method because they exploit prior knowledge of gene function to identify unknown genes of similar function from expression data. SVMs avoid several problems associated with unsupervised clustering methods, such as hierarchical clustering and self-organizing maps. SVMs have many mathematical features that make them attractive for gene expression analysis, including their flexibility in choosing a similarity function, sparseness of solution when dealing with large data sets, the ability to handle large feature spaces, and the ability to identify outliers. We test several SVMs that use different similarity metrics, as well as some other supervised learning methods, and find that the SVMs best identify sets of genes with a common function using expression data. Finally, we use SVMs to predict functional roles for uncharacterized yeast ORFs based on their expression data.

Algorithms↗

Regularized Least Squares Cancer classifiers from DNA microarray data.

BACKGROUND: The advent of the technology of DNA microarrays constitutes an epochal change in the classification and discovery of different types of cancer because the information provided by DNA microarrays allows an approach to the problem of cancer analysis from a quantitative rather than qualitative point of view. Cancer classification requires well founded mathematical methods which are able to predict the status of new specimens with high significance levels starting from a limited number of data. In this paper we assess the performances of Regularized Least Squares (RLS) classifiers, originally proposed in regularization theory, by comparing them with Support Vector Machines (SVM), the state-of-the-art supervised learning technique for cancer classification by DNA microarray data. The performances of both approaches have been also investigated with respect to the number of selected genes and different gene selection strategies. RESULTS: We show that RLS classifiers have performances comparable to those of SVM classifiers as the Leave-One-Out (LOO) error evaluated on three different data sets shows. The main advantage of RLS machines is that for solving a classification problem they use a linear system of order equal to either the number of features or the number of training examples. Moreover, RLS machines allow to get an exact measure of the LOO error with just one training. CONCLUSION: RLS classifiers are a valuable alternative to SVM classifiers for the problem of cancer classification by gene expression data, due to their simplicity and low computational complexity. Moreover, RLS classifiers show generalization ability comparable to the ones of SVM classifiers also in the case the classification of new specimens involves very few gene expression levels.

Algorithms↗

Prediction of the axillary lymph node status in mammary cancer on the basis of clinicopathological data and flow cytometry.

Axillary lymph node status is a major prognostic factor in mammary carcinoma. It is clinically desirable to predict the axillary lymph node status from data from the mammary cancer specimen. In the study, the axillary lymph node status, routine histological parameters and flow-cytometric data were retrospectively obtained from 1139 specimens of invasive mammary cancer. The ten variables: age, tumour type, tumour grade, tumour size, skin infiltration, lymphangiosis carcinomatosa, pT4 category, percentage of tumour cells in G2/M- and S-phases of the cell cycle, and ploidy index were considered as predictor variables, and the single variable lymph node metastasis pN (0 for pN0, or 1 for pN1 or pN2) was used as an output variable. A stepwise logistic regression analysis, with the axillary lymph node as a dependent variable, was used for feature selection. Only lymphangiosis carcinomatosa and tumour size proved to be significant as independent predictor variables; the other variables were non-contributory. Three paradigms with supervised learning rules (multilayer perceptron, learning vector quantisation and support vector machines) were used for the purpose of prediction. If any of these paradigms was used with the information from all ten input variables, 73% of cases could be correctly predicted, with specificity ranging from 82 to 84% and sensitivity ranging from 60 to 63%. If only the two significant input variables were used, lymphangiosis carcinomatosa and tumour diameter, the prediction accuracy was no worse. Nearly identical results were obtained by two different techniques of cross-validation (leave-one-out against ten-fold cross validation). It was concluded that: artificial neural networks can be used for risk stratification on the basis of routine data in individual cases of mammary cancer; and lymphangiosis carcinomatosa and tumour size are independent predictors of axillary lymph node metastasis in mammary cancer.

Algorithms↗

Analyzing microarray data using cluster analysis.

As pharmacogenetics researchers gather more detailed and complex data on gene polymorphisms that effect drug metabolizing enzymes, drug target receptors and drug transporters, they will need access to advanced statistical tools to mine that data. These tools include approaches from classical biostatistics, such as logistic regression or linear discriminant analysis, and supervised learning methods from computer science, such as support vector machines and artificial neural networks. In this review, we present an overview of another class of models, cluster analysis, which will likely be less familiar to pharmacogenetics researchers. Cluster analysis is used to analyze data that is not a priori known to contain any specific subgroups. The goal is to use the data itself to identify meaningful or informative subgroups. Specifically, we will focus on demonstrating the use of distance-based methods of hierarchical clustering to analyze gene expression data.

Cluster Analysis↗

Systematic learning of gene functional classes from DNA array expression data by using multilayer perceptrons.

Recent advances in microarray technology have opened new ways for functional annotation of previously uncharacterised genes on a genomic scale. This has been demonstrated by unsupervised clustering of co-expressed genes and, more importantly, by supervised learning algorithms. Using prior knowledge, these algorithms can assign functional annotations based on more complex expression signatures found in existing functional classes. Previously, support vector machines (SVMs) and other machine-learning methods have been applied to a limited number of functional classes for this purpose. Here we present, for the first time, the comprehensive application of supervised neural networks (SNNs) for functional annotation. Our study is novel in that we report systematic results for ~100 classes in the Munich Information Center for Protein Sequences (MIPS) functional catalog. We found that only ~10% of these are learnable (based on the rate of false negatives). A closer analysis reveals that false positives (and negatives) in a machine-learning context are not necessarily "false" in a biological sense. We show that the high degree of interconnections among functional classes confounds the signatures that ought to be learned for a unique class. We term this the "Borges effect" and introduce two new numerical indices for its quantification. Our analysis indicates that classification systems with a lower Borges effect are better suitable for machine learning. Furthermore, we introduce a learning procedure for combining false positives with the original class. We show that in a few iterations this process converges to a gene set that is learnable with considerably low rates of false positives and negatives and contains genes that are biologically related to the original class, allowing for a coarse reconstruction of the interactions between associated biological pathways. We exemplify this methodology using the well-studied tricarboxylic acid cycle.

Algorithms↗

An Integrated Machine Learning and Genomic Framework for Precise Detection of Gastric Cancer.

This study presents a novel integrative approach for the analysis of high-dimensional gene expression data, leveraging the complementary strengths of unsupervised clustering and supervised classification. Using K-means clustering, the data set is stratified into three distinct clusters, revealing intrinsic biological patterns and relationships. The resulting cluster assignments are subsequently used as pseudolabels to train machine learning models, including support vector machines, random forest, and a stacking ensemble classifier. To validate and enhance the robustness of clustering, complementary methods, such as hierarchical clustering and density-based spatial clustering of applications with noise (DBSCAN), are used, with results visualized through principal component analysis-driven dimensionality reduction. The high predictive accuracy achieved by the classifiers underlines the separability and reliability of the identified clusters. Furthermore, feature importance analysis highlighted key genetic determinants within each cluster, offering actionable insights into potential biomarkers and critical genomic features. This framework bridges the gap between exploratory unsupervised learning and predictive supervised modeling, providing a scalable and interpretable method for analyzing complex genomic data sets. Its applicability extends to biomarker discovery, patient stratification, and other precision medicine applications, emphasizing its utility in advancing genomic research and clinical practice.

Humans↗

Machine learning in bioinformatics.

This article reviews machine learning methods for bioinformatics. It presents modelling methods, such as supervised classification, clustering and probabilistic graphical models for knowledge discovery, as well as deterministic and stochastic heuristics for optimization. Applications in genomics, proteomics, systems biology, evolution and text mining are also shown.

Artificial Intelligence↗

In silico estimation of DMSO solubility of organic compounds for bioscreening.

Solubility of organic compounds in DMSO is an important issue for commercial and academic organizations handling large compound collections or performing biological screening. In particular, solubility data are critical for the optimization of storage conditions and for the selection of compounds for bioscreening compatible with the assay protocol. Solubility is largely determined by the solvation energy and the crystal disruption energy, and these molecular phenomena should be assessed in structure-solubility correlation studies. The authors summarize our long-term experimental observations and theoretical studies of physicochemical determinants of DMSO solubility of organic substances. They compiled a comprehensive reference database of proprietary data on compound solubility (55,277 compounds with good DMSO solubility and 10,223 compounds with poor DMSO solubility), calculated specific molecular descriptors (topological, electromagnetic, charge, and lipophilicity parameters), and applied an advanced machine-learning approach for training neural networks to address the solubility. Both supervised (feed-forward, back-propagated neural networks) and unsupervised (Kohonen neural networks) learning methods were used. The resulting neural network models were validated by successfully predicting DMSO solubility of compounds in independent test selections.

Dimethyl Sulfoxide↗

Reliable classification of two-class cancer data using evolutionary algorithms.

In the area of bioinformatics, the identification of gene subsets responsible for classifying available disease samples to two or more of its variants is an important task. Such problems have been solved in the past by means of unsupervised learning methods (hierarchical clustering, self-organizing maps, k-mean clustering, etc.) and supervised learning methods (weighted voting approach, k-nearest neighbor method, support vector machine method, etc.). Such problems can also be posed as optimization problems of minimizing gene subset size to achieve reliable and accurate classification. The main difficulties in solving the resulting optimization problem are the availability of only a few samples compared to the number of genes in the samples and the exorbitantly large search space of solutions. Although there exist a few applications of evolutionary algorithms (EAs) for this task, here we treat the problem as a multiobjective optimization problem of minimizing the gene subset size and minimizing the number of misclassified samples. Moreover, for a more reliable classification, we consider multiple training sets in evaluating a classifier. Contrary to the past studies, the use of a multiobjective EA (NSGA-II) has enabled us to discover a smaller gene subset size (such as four or five) to correctly classify 100% or near 100% samples for three cancer samples (Leukemia, Lymphoma, and Colon). We have also extended the NSGA-II to obtain multiple non-dominated solutions discovering as much as 352 different three-gene combinations providing a 100% correct classification to the Leukemia data. In order to have further confidence in the identification task, we have also introduced a prediction strength threshold for determining a sample's belonging to one class or the other. All simulation results show consistent gene subset identifications on three disease samples and exhibit the flexibilities and efficacies in using a multiobjective EA for the gene subset identification task.

Algorithms↗

A primer on gene expression and microarrays for machine learning researchers.

Data originating from biomedical experiments has provided machine learning researchers with an important source of motivation for developing and evaluating new algorithms. A new wave of algorithmic development has been initiated with the publication of gene expression data derived from microarrays. Microarray data analysis is particularly challenging given the large number of measurements (typically in the order of thousands) that are reported for relatively few samples (typically in the order of dozens). Many data sets are now available on the web. It is important that machine learning researchers understand how data are obtained and which assumptions are necessary in the analysis. Microarray data have the potential to cause significant impact in machine learning research, not just as a rich and realistic source of cases for testing new algorithms, as has been the UCI machine learning repository in the past decades, but also as a main motivation for their development. In this article, we briefly review the biology underlying microarrays, the process of obtaining gene expression measurements, and the rationale behind the common types of analyses involved in a microarray experiment. We outline the main challenges and reiterate critical considerations regarding the construction of supervised learning models that use this type of data. The goal of this article is to familiarize machine learning researchers with data originated from gene expression microarrays.

Algorithms↗

Machine learning approaches to supporting the identification of photoreceptor-enriched genes based on expression data.

BACKGROUND: Retinal photoreceptors are highly specialised cells, which detect light and are central to mammalian vision. Many retinal diseases occur as a result of inherited dysfunction of the rod and cone photoreceptor cells. Development and maintenance of photoreceptors requires appropriate regulation of the many genes specifically or highly expressed in these cells. Over the last decades, different experimental approaches have been developed to identify photoreceptor enriched genes. Recent progress in RNA analysis technology has generated large amounts of gene expression data relevant to retinal development. This paper assesses a machine learning methodology for supporting the identification of photoreceptor enriched genes based on expression data. RESULTS: Based on the analysis of publicly-available gene expression data from the developing mouse retina generated by serial analysis of gene expression (SAGE), this paper presents a predictive methodology comprising several in silico models for detecting key complex features and relationships encoded in the data, which may be useful to distinguish genes in terms of their functional roles. In order to understand temporal patterns of photoreceptor gene expression during retinal development, a two-way cluster analysis was firstly performed. By clustering SAGE libraries, a hierarchical tree reflecting relationships between developmental stages was obtained. By clustering SAGE tags, a more comprehensive expression profile for photoreceptor cells was revealed. To demonstrate the usefulness of machine learning-based models in predicting functional associations from the SAGE data, three supervised classification models were compared. The results indicated that a relatively simple instance-based model (KStar model) performed significantly better than relatively more complex algorithms, e.g. neural networks. To deal with the problem of functional class imbalance occurring in the dataset, two data re-sampling techniques were studied. A random over-sampling method supported the implementation of the most powerful prediction models. The KStar model was also able to achieve higher predictive sensitivities and specificities using random over-sampling techniques. CONCLUSION: The approaches assessed in this paper represent an efficient and relatively inexpensive in silico methodology for supporting large-scale analysis of photoreceptor gene expression by SAGE. They may be applied as complementary methodologies to support functional predictions before implementing more comprehensive, experimental prediction and validation methods. They may also be combined with other large-scale, data-driven methods to facilitate the inference of transcriptional regulatory networks in the developing retina. Furthermore, the methodology assessed may be applied to other data domains.

Animals↗