Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Supervised Machine Learning”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Supervised classification for gene network reconstruction.

One of the central problems of functional genomics is revealing gene expression networks - the relationships between genes that reflect observations of how the expression level of each gene affects those of others. Microarray data are currently a major source of information about the interplay of biochemical network participants in living cells. Various mathematical techniques, such as differential equations, Bayesian and Boolean models and several statistical methods, have been applied to expression data in attempts to extract the underlying knowledge. Unsupervised clustering methods are often considered as the necessary first step in visualization and analysis of the expression data. As for supervised classification, the problem mainly addressed so far has been how to find discriminative genes separating various samples or experimental conditions. Numerous methods have been applied to identify genes that help to predict treatment outcome or to confirm a diagnosis, as well as to identify primary elements of gene regulatory circuits. However, less attention has been devoted to using supervised learning to uncover relationships between genes and/or their products. To start filling this gap a machine-learning approach for gene networks reconstruction is described here. This approach is based on building classifiers--functions, which determine the state of a gene's transcription machinery through expression levels of other genes. The method can be applied to various cases where relationships between gene expression levels could be expected.

Genes↗

Machine learning in control of functional electrical stimulation systems for locomotion.

Two machine learning techniques were evaluated for automatic design of a rule-based control of functional electrical stimulation (FES) for locomotion of spinal cord injured humans. The task was to learn the invariant characteristics of the relationship between sensory information and the FES-control signal by using off-line supervised training. Sensory signals were recorded using pressure sensors installed in the insoles of a subject's shoes and goniometers attached across the joints of the affected leg. The FES-control consisted of pulses corresponding to time intervals when the subject pressed on the manual push-button to deliver the stimulation during FES-assisted ambulation. The machine learning techniques used were the adaptive logic network (ALN) [1] and the inductive learning algorithm (IL) [2]. Results to date suggest that, given the same training data, the IL learned faster than the ALN, while both performed the test rapidly. The generalization was estimated by measuring the test errors and it was better with an ALN, especially if past points were used to reflect the time dimension. Both techniques were able to predict future stimulation events. An advantage of the ALN over the IL was that ALN's can be retrained with new data without losing previously collected knowledge. The advantages of the IL over the ALN were that the IL produces small, explicit, comprehensible trees and that the relative importance of each sensory contribution can be quantified.

Adult↗

scFANCL: Dual contrastive learning with false-negative correction at cell level for single-cell RNA-seq clustering.

BACKGROUND: Single-cell RNA sequencing (scRNA-seq) enables cellular characterization at single-cell resolution. However, its high dimensionality, sparsity, and noise make clustering challenging. Approaches utilizing contrastive learning and data augmentation have been introduced to improve representation quality for scRNA-seq clustering. In particular, dual contrastive frameworks combining instance- and cluster-level objectives can capture both cell-cell similarities and inter-cluster variations. However, existing dual contrastive frameworks focus primarily on discrete cluster boundaries, neglecting the biological continuity inherent in scRNA-seq data. METHODS: We propose scFANCL, a dual contrastive framework designed to capture biological continuity in scRNA data. Rather than treating all non-augmented samples as negatives, scFANCL applies a cosine-similarity-based threshold to exclude cells of the same type from the negative pool, preserving continuous transcriptional relationships among them while maintaining inter-cluster separation. RESULTS: Extensive experiments across seven publicly available scRNA-seq datasets demonstrated that scFANCL achieves competitive clustering performance compared with existing baseline methods, consistently yielding high ARI and NMI scores across datasets of varying size and complexity. Ablation studies further confirmed the contribution of the false negative filtering component, showing measurable improvements over variants without filtering. Downstream analyses further suggest that the learned embeddings may reflect biologically meaningful transcriptional transitions, including continuous differentiation trajectories within related cell types. The source code is available at https://github.com/mjuailab/scFANCL . CONCLUSIONS: scFANCL addresses a key limitation of conventional contrastive learning by applying a cosine-similarity-based threshold to exclude cells of the same type from the negative pool, thereby preserving biological continuity within cell types while maintaining inter-cluster separation. Evaluations across seven benchmark scRNA-seq datasets demonstrate competitive clustering performance, with learned embeddings capturing biologically meaningful transcriptional structure and characteristics of rare cell populations.

Clustering Algorithms↗

Analysis of polarity information in medical text.

Knowing the polarity of clinical outcomes is important in answering questions posed by clinicians in patient treatment. We treat analysis of this information as a classification problem. Natural language processing and machine learning techniques are applied to detect four possibilities in medical text: no outcome, positive outcome, negative outcome, and neutral outcome. A supervised learning method is used to perform the classification at the sentence level. Five feature sets are constructed: unigrams, bigrams, change phrases, negations, and categories. The performance of different combinations of feature sets is compared. The results show that generalization using the category information in the domain knowledge base Unified Medical Language System is effective in the task. The effect of context information is significant. Combining linguistic features and domain knowledge leads to the highest accuracy.

Artificial Intelligence↗

The machine-learning classifier ALLCatchR2 identifies 20 T-ALL subtypes across cohorts and age groups.

T-cell acute lymphoblastic leukemia (T-ALL) comprises molecularly diverse subtypes, but robust cross-cohort validations and operational gene-expression definitions are lacking. To establish a gene-expression-anchored framework for T-ALL subtyping, we aggregated 2314 transcriptomes (15 cohorts, age: 0.8-90.8 years). An extended unsupervised approach defined 17 main clusters and 3 subclusters in samples with high blast fractions. Supervised analyses added an overarching immature T-ALL (early T cell precursor [ETP]-like) definition and resolved the LMO2 &#x3b3;&#x3b4;-like subtype. All clusters contained samples from at least two cohorts. Characteristic genomic driver enrichments were consistent across cohorts, while gene-expression clusters did not correspond exclusively to single driver events but also reflected developmental origins. A machine-learning classifier based on ALLCatchR, our B-cell acute lymphoblastic leukemia (B-ALL) classifier, identified these 20 transcriptomic subtypes and the immature T-ALL (ETP-like) signature with 0.995-1.0 accuracy in a validation set (n&#x2009;=&#x2009;203). Testing the classifier on a second hold-out data set (n&#x2009;=&#x2009;265 samples) showed that 92.7% of predictions matched with corresponding driver alterations. Across all samples, 83.2% of cases received high-confidence predictions, 7.3% candidate predictions, and 9.5% remained unclassified, largely because of low blast fractions. We identified a novel gene-expression cluster markedly enriched (P&#x2009;<&#x2009;0.001) for clonal hematopoiesis mutations (IDH2 R140Q, DNMT3A) and a stem-/progenitor cell-like gene expression. This novel clonal hematopoiesis-related T-ALL subtype was observed in six cohorts and accounted for 8.9% of adults and 39.5% of patients aged >50 years. We extended&#xa0;ALLCatchR into ALLCatchR2, a free R package that now enables B-/T-lineage separation, gene-expression subtyping, blast estimation, and developmental annotation to harmonize T-ALL classification across studies and clinical contexts.

Journal Article↗

Discovering hidden candidate plastic-degrading enzymes: Combined multi-omics and machine learning strategy.

Plastic pollution poses a major threat to the stability of natural ecosystems as well as human health. Microbial enzymes have long been considered a potential resource for targeted biodegradation but, except for a few successful cases, the discovery of efficient enzymes has proved challenging. Aiming to accelerate the process, we propose an approach combining metagenomics, metatranscriptomics and semi-supervised learning that selects promising plastic-degrading candidate enzymes from the proteome of relevant microorganisms. Tested on a dataset of over 10,000 microbial proteins, ranking models consistently prioritize known plastic-degrading enzymes, achieving an area under the cumulative distribution function curve above 0.96, with leave-one-family-out cross-validation indicating that performance is largely retained across protein families. As a case study, this work focuses on mixed microbial cultures exposed for extended periods to polyethylene, polyethylene terephthalate, and polyurethane substrates. The prevalent species after selective enrichment were functionally characterized, finding Rhodococcus aetherivorans as the most relevant species in two of the five cultures under investigation. Among the top-ranked proteins, several have high structural similarity with known enzymes despite not being identified by sequence similarity search. Moreover, according to metatranscriptomics results, several of these enzymes were found to be expressed at the same level or above that of annotated enzymes, suggesting that they may have functional relevance. Overall, this work highlights the potential of integrating multi-omics with data-driven methods for enzyme discovery and for accelerating the development of biotechnological solutions to plastic pollution.

Biodegradation, Environmental↗

A study on several machine-learning methods for classification of malignant and benign clustered microcalcifications.

In this paper, we investigate several state-of-the-art machine-learning methods for automated classification of clustered microcalcifications (MCs). The classifier is part of a computer-aided diagnosis (CADx) scheme that is aimed to assisting radiologists in making more accurate diagnoses of breast cancer on mammograms. The methods we considered were: support vector machine (SVM), kernel Fisher discriminant (KFD), relevance vector machine (RVM), and committee machines (ensemble averaging and AdaBoost), of which most have been developed recently in statistical learning theory. We formulated differentiation of malignant from benign MCs as a supervised learning problem, and applied these learning methods to develop the classification algorithm. As input, these methods used image features automatically extracted from clustered MCs. We tested these methods using a database of 697 clinical mammograms from 386 cases, which included a wide spectrum of difficult-to-classify cases. We analyzed the distribution of the cases in this database using the multidimensional scaling technique, which reveals that in the feature space the malignant cases are not trivially separable from the benign ones. We used receiver operating characteristic (ROC) analysis to evaluate and to compare classification performance by the different methods. In addition, we also investigated how to combine information from multiple-view mammograms of the same case so that the best decision can be made by a classifier. In our experiments, the kernel-based methods (i.e., SVM, KFD, and RVM) yielded the best performance (Az = 0.85, SVM), significantly outperforming a well-established, clinically-proven CADx approach that is based on neural network (Az = 0.80).

Algorithms↗

Metabolic fingerprinting of salt-stressed tomatoes.

The aim of this study was to adopt the approach of metabolic fingerprinting through the use of Fourier transform infrared (FT-IR) spectroscopy and chemometrics to study the effect of salinity on tomato fruit. Two varieties of tomato were studied, Edkawy and Simge F1. Salinity treatment significantly reduced the relative growth rate of Simge F1 but had no significant effect on that of Edkawy. In both tomato varieties salt-treatment significantly reduced mean fruit fresh weight and size class but had no significant affect on total fruit number. Marketable yield was however reduced in both varieties due to the occurrence of blossom end rot in response to salinity. Whole fruit flesh extracts from control and salt-grown tomatoes were analysed using FT-IR spectroscopy. Each sample spectrum contained 882 variables, absorbance values at different wavenumbers, making visual analysis difficult and therefore machine learning methods were applied. The unsupervised clustering method, principal component analysis (PCA) showed no discrimination between the control and salt-treated fruit for either variety. The supervised method, discriminant function analysis (DFA) was able to classify control and salt-treated fruit in both varieties. Genetic algorithms (GA) were applied to identify discriminatory regions within the FT-IR spectra important for fruit classification. The GA models were able to classify control and salt-treated fruit with a typical error, when classifying the whole data set, of 9% in Edkawy and 5% in Simge F1. Key regions were identified within the spectra corresponding to nitrile containing compounds and amino radicals. The application of GA enabled the identification of functional groups of potential importance in relation to the response of tomato to salinity.

Algorithms↗

Protein ranking by semi-supervised network propagation.

BACKGROUND: Biologists regularly search DNA or protein databases for sequences that share an evolutionary or functional relationship with a given query sequence. Traditional search methods, such as BLAST and PSI-BLAST, focus on detecting statistically significant pairwise sequence alignments and often miss more subtle sequence similarity. Recent work in the machine learning community has shown that exploiting the global structure of the network defined by these pairwise similarities can help detect more remote relationships than a purely local measure. METHODS: We review RankProp, a ranking algorithm that exploits the global network structure of similarity relationships among proteins in a database by performing a diffusion operation on a protein similarity network with weighted edges. The original RankProp algorithm is unsupervised. Here, we describe a semi-supervised version of the algorithm that uses labeled examples. Three possible ways of incorporating label information are considered: (i) as a validation set for model selection, (ii) to learn a new network, by choosing which transfer function to use for a given query, and (iii) to estimate edge weights, which measure the probability of inferring structural similarity. RESULTS: Benchmarked on a human-curated database of protein structures, the original RankProp algorithm provides significant improvement over local network search algorithms such as PSI-BLAST. Furthermore, we show here that labeled data can be used to learn a network without any need for estimating parameters of the transfer function, and that diffusion on this learned network produces better results than the original RankProp algorithm with a fixed network. CONCLUSION: In order to gain maximal information from a network, labeled and unlabeled data should be used to extract both local and global structure.

Algorithms↗

Relevance vector machine for automatic detection of clustered microcalcifications.

Clustered microcalcifications (MC) in mammograms can be an important early sign of breast cancer in women. Their accurate detection is important in computer-aided detection (CADe). In this paper, we propose the use of a recently developed machine-learning technique--relevance vector machine (RVM)--for detection of MCs in digital mammograms. RVM is based on Bayesian estimation theory, of which a distinctive feature is that it can yield a sparse decision function that is defined by only a very small number of so-called relevance vectors. By exploiting this sparse property of the RVM, we develop computerized detection algorithms that are not only accurate but also computationally efficient for MC detection in mammograms. We formulate MC detection as a supervised-learning problem, and apply RVM as a classifier to determine at each location in the mammogram if an MC object is present or not. To increase the computation speed further, we develop a two-stage classification network, in which a computationally much simpler linear RVM classifier is applied first to quickly eliminate the overwhelming majority, non-MC pixels in a mammogram from any further consideration. The proposed method is evaluated using a database of 141 clinical mammograms (all containing MCs), and compared with a well-tested support vector machine (SVM) classifier. The detection performance is evaluated using free-response receiver operating characteristic (FROC) curves. It is demonstrated in our experiments that the RVM classifier could greatly reduce the computational complexity of the SVM while maintaining its best detection accuracy. In particular, the two-stage RVM approach could reduce the detection time from 250 s for SVM to 7.26 s for a mammogram (nearly 35-fold reduction). Thus, the proposed RVM classifier is more advantageous for real-time processing of MC clusters in mammograms.

Algorithms↗

A support vector machine approach for detection of microcalcifications.

In this paper, we investigate an approach based on support vector machines (SVMs) for detection of microcalcification (MC) clusters in digital mammograms, and propose a successive enhancement learning scheme for improved performance. SVM is a machine-learning method, based on the principle of structural risk minimization, which performs well when applied to data outside the training set. We formulate MC detection as a supervised-learning problem and apply SVM to develop the detection algorithm. We use the SVM to detect at each location in the image whether an MC is present or not. We tested the proposed method using a database of 76 clinical mammograms containing 1120 MCs. We use free-response receiver operating characteristic curves to evaluate detection performance, and compare the proposed algorithm with several existing methods. In our experiments, the proposed SVM framework outperformed all the other methods tested. In particular, a sensitivity as high as 94% was achieved by the SVM method at an error rate of one false-positive cluster per image. The ability of SVM to out perform several well-known methods developed for the widely studied problem of MC detection suggests that SVM is a promising technique for object detection in a medical imaging application.

Algorithms↗

Feature selection and nearest centroid classification for protein mass spectrometry.

BACKGROUND: The use of mass spectrometry as a proteomics tool is poised to revolutionize early disease diagnosis and biomarker identification. Unfortunately, before standard supervised classification algorithms can be employed, the "curse of dimensionality" needs to be solved. Due to the sheer amount of information contained within the mass spectra, most standard machine learning techniques cannot be directly applied. Instead, feature selection techniques are used to first reduce the dimensionality of the input space and thus enable the subsequent use of classification algorithms. This paper examines feature selection techniques for proteomic mass spectrometry. RESULTS: This study examines the performance of the nearest centroid classifier coupled with the following feature selection algorithms. Student-t test, Kolmogorov-Smirnov test, and the P-test are univariate statistics used for filter-based feature ranking. From the wrapper approaches we tested sequential forward selection and a modified version of sequential backward selection. Embedded approaches included shrunken nearest centroid and a novel version of boosting based feature selection we developed. In addition, we tested several dimensionality reduction approaches, namely principal component analysis and principal component analysis coupled with linear discriminant analysis. To fairly assess each algorithm, evaluation was done using stratified cross validation with an internal leave-one-out cross-validation loop for automated feature selection. Comprehensive experiments, conducted on five popular cancer data sets, revealed that the less advocated sequential forward selection and boosted feature selection algorithms produce the most consistent results across all data sets. In contrast, the state-of-the-art performance reported on isolated data sets for several of the studied algorithms, does not hold across all data sets. CONCLUSION: This study tested a number of popular feature selection methods using the nearest centroid classifier and found that several reportedly state-of-the-art algorithms in fact perform rather poorly when tested via stratified cross-validation. The revealed inconsistencies provide clear evidence that algorithm evaluation should be performed on several data sets using a consistent (i.e., non-randomized, stratified) cross-validation procedure in order for the conclusions to be statistically sound.

Algorithms↗

Learning in brains and machines.

The problem of learning is arguably at the very core of the problem of intelligence, both biological and artificial. In this paper we sketch some of our work over the last ten years in the area of supervised learning, focusing on three interlinked directions of research: theory, engineering applications (that is, making intelligent software) and neuroscience (that is, understanding the brain's mechanisms of learning).

Brain↗

Dynamics and context dependence of visual category learning.

Visual category learning by humans is observed within a paradigm of supervised learning. Mental representations for recognition are reconstructed from the observed data structures by fitting to them predicted classification data obtained from similarity-based models of recognition on the one hand and machine vision systems for image understanding on the other hand. These principles are illustrated with examples concerning the dynamics and the dependence on context of processes of category learning.

Computer Simulation↗

Filter versus wrapper gene selection approaches in DNA microarray domains.

DNA microarray experiments generating thousands of gene expression measurements, are used to collect information from tissue and cell samples regarding gene expression differences that could be useful for diagnosis disease, distinction of the specific tumor type, etc. One important application of gene expression microarray data is the classification of samples into known categories. As DNA microarray technology measures the gene expression en masse, this has resulted in data with the number of features (genes) far exceeding the number of samples. As the predictive accuracy of supervised classifiers that try to discriminate between the classes of the problem decays with the existence of irrelevant and redundant features, the necessity of a dimensionality reduction process is essential. We propose the application of a gene selection process, which also enables the biology researcher to focus on promising gene candidates that actively contribute to classification in these large scale microarrays. Two basic approaches for feature selection appear in machine learning and pattern recognition literature: the filter and wrapper techniques. Filter procedures are used in most of the works in the area of DNA microarrays. In this work, a comparison between a group of different filter metrics and a wrapper sequential search procedure is carried out. The comparison is performed in two well-known DNA microarray datasets by the use of four classic supervised classifiers. The study is carried out over the original-continuous and three-intervals discretized gene expression data. While two well-known filter metrics are proposed for continuous data, four classic filter measures are used over discretized data. The same wrapper approach is used for both continuous and discretized data. The application of filter and wrapper gene selection procedures leads to considerably better accuracy results in comparison to the non-gene selection approach, coupled with interesting and notable dimensionality reductions. Although the wrapper approach mainly shows a more accurate behavior than filter metrics, this improvement is coupled with considerable computer-load necessities. We note that most of the genes selected by proposed filter and wrapper procedures in discrete and continuous microarray data appear in the lists of relevant-informative genes detected by previous studies over these datasets. The aim of this work is to make contributions in the field of the gene selection task in DNA microarray datasets. By an extensive comparison with more popular filter techniques, we would like to make contributions in the expansion and study of the wrapper approach in this type of domains.

Artificial Intelligence↗

Oligonucleotide microarray for prediction of early intrahepatic recurrence of hepatocellular carcinoma after curative resection.

BACKGROUND: Hepatocellular carcinoma has a poor prognosis because of the high intrahepatic recurrence rate. There are technological limitations to traditional methods such as TNM staging for accurate prediction of recurrence, suggesting that new techniques are needed. METHODS: We investigated mRNA expression profiles in tissue specimens from a training set, comprising 33 patients with hepatocellular carcinoma, with high-density oligonucleotide microarrays representing about 6000 genes. We used this training set in a supervised learning manner to construct a predictive system, consisting of 12 genes, with the Fisher linear classifier. We then compared the predictive performance of our system with that of a predictive system with a support vector machine (SVM-based system) on a blinded set of samples from 27 newly enrolled patients. FINDINGS: Early intrahepatic recurrence within 1 year after curative surgery occurred in 12 (36%) and eight (30%) patients in the training and blinded sets, respectively. Our system correctly predicted early intrahepatic recurrence or non-recurrence in 25 (93%) of 27 samples in the blinded set and had a positive predictive value of 88% and a negative predictive value of 95%. By contrast, the SVM-based system predicted early intrahepatic recurrence or non-recurrence correctly in only 16 (60%) individuals in the blinded set, and the result yielded a positive predictive value of only 38% and a negative predictive value of 79%. INTERPRETATION: Our system predicted early intrahepatic recurrence or non-recurrence for patients with hepatocellular carcinoma much more accurately than the SVM-based system, suggesting that our system could serve as a new method for characterising the metastatic potential of hepatocellular carcinoma.

Carcinoma, Hepatocellular↗

Enhancing text categorization with semantic-enriched representation and training data augmentation.

OBJECTIVE: Acquiring and representing biomedical knowledge is an increasingly important component of contemporary bioinformatics. A critical step of the process is to identify and retrieve relevant documents among the vast volume of modern biomedical literature efficiently. In the real world, many information retrieval tasks are difficult because of high data dimensionality and the lack of annotated examples to train a retrieval algorithm. Under such a scenario, the performance of information retrieval algorithms is often unsatisfactory, therefore improvements are needed. DESIGN: We studied two approaches that enhance the text categorization performance on sparse and high data dimensionality: (1) semantic-preserving dimension reduction by representing text with semantic-enriched features; and (2) augmenting training data with semi-supervised learning. A probabilistic topic model was applied to extract major semantic topics from a corpus of text of interest. The representation of documents was projected from the high-dimensional vocabulary space onto a semantic topic space with reduced dimensionality. A semi-supervised learning algorithm based on graph theory was applied to identify potential positive training cases, which were further used to augment training data. The effects of data transformation and augmentation on text categorization by support vector machine (SVM) were evaluated. RESULTS AND CONCLUSION: Semantic-enriched data transformation and the pseudo-positive-cases augmented training data enhance the efficiency and performance of text categorization by SVM.

Algorithms↗

Deep learning-based multimodal pathogenomics integration for precision cancer prognosis.

BACKGROUND: Recent studies have revealed valuable prognostic insights in haematoxylin and eosin (H&E)-stained histological sections and transcriptomic profiles, suggesting potential applications in machine learning. However, existing methods lack sufficient intra- and inter-modal interactions, and face challenges in clinical validation due to incomplete multimodal data. METHODS: We proposed PathoGems (PathoGenomics-based integrative survival prediction), a weakly-supervised, interpretable multimodal learning framework that integrates histology and genomic profiles for precise cancer prognosis prediction. To evaluate the robustness of PathoGems, we initially curated a dataset of 1965 cases across four cohorts from The Cancer Genome Atlas (TCGA), including breast, colorectal, glioblastoma, and esophageal cancers. For external validation, PathoGems was further evaluated on four independent cohorts, consisting of 76 breast cancer and 41 esophageal squamous cell carcinoma cases from Zhejiang Cancer Hospital, as well as 102 colorectal cancer and 58 glioblastoma cases from the Clinical Proteomic Tumor Analysis Consortium (CPTAC). RESULTS: PathoGems effectively stratified patients into favorable and unfavorable risk groups, revealing significant differences in histological patterns, genomic features, and overall survival (log-rank test, p&#x2009;<&#x2009;0.05). Moreover, the model&#x2019;s predictions are further supported by visualization and transcriptomic analysis, enhancing interpretability and reliability. CONCLUSIONS: By fusing histological and clinicogenomic multimodal models, PathoGems will provide a solid foundation for developing an innovative tool that aids clinicians in making informed decisions and selection personalized treatment strategies for cancer patients.

Humans↗