Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Supervised Machine Learning”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

142 records · Page 8Linked to original sources

Data processing and classification analysis of proteomic changes: a case study of oil pollution in the mussel, Mytilus edulis.

BACKGROUND: Proteomics may help to detect subtle pollution-related changes, such as responses to mixture pollution at low concentrations, where clear signs of toxicity are absent. The challenges associated with the analysis of large-scale multivariate proteomic datasets have been widely discussed in medical research and biomarker discovery. This concept has been introduced to ecotoxicology only recently, so data processing and classification analysis need to be refined before they can be readily applied in biomarker discovery and monitoring studies. RESULTS: Data sets obtained from a case study of oil pollution in the Blue mussel were investigated for differential protein expression by retentate chromatography-mass spectrometry and decision tree classification. Different tissues and different settings were used to evaluate classifiers towards their discriminatory power. It was found that, due the intrinsic variability of the data sets, reliable classification of unknown samples could only be achieved on a broad statistical basis (n > 60) with the observed expression changes comprising high statistical significance and sufficient amplitude. The application of stringent criteria to guard against overfitting of the models eventually allowed satisfactory classification for only one of the investigated data sets and settings. CONCLUSION: Machine learning techniques provide a promising approach to process and extract informative expression signatures from high-dimensional mass-spectrometry data. Even though characterisation of the proteins forming the expression signatures would be ideal, knowledge of the specific proteins is not mandatory for effective class discrimination. This may constitute a new biomarker approach in ecotoxicology, where working with organisms, which do not have sequenced genomes render protein identification by database searching problematic. However, data processing has to be critically evaluated and statistical constraints have to be considered before supervised classification algorithms are employed.

Journal Article↗

ASGCL: Adaptive Sparse Mapping-based graph contrastive learning network for cancer drug response prediction.

Personalized cancer drug treatment is emerging as a frontier issue in modern medical research. Considering the genomic differences among cancer patients, determining the most effective drug treatment plan is a complex and crucial task. In response to these challenges, this study introduces the Adaptive Sparse Graph Contrastive Learning Network (ASGCL), an innovative approach to unraveling latent interactions in the complex context of cancer cell lines and drugs. The core of ASGCL is the GraphMorpher module, an innovative component that enhances the input graph structure via strategic node attribute masking and topological pruning. By contrasting the augmented graph with the original input, the model delineates distinct positive and negative sample sets at both node and graph levels. This dual-level contrastive approach significantly amplifies the model's discriminatory prowess in identifying nuanced drug responses. Leveraging a synergistic combination of supervised and contrastive loss, ASGCL accomplishes end-to-end learning of feature representations, substantially outperforming existing methodologies. Comprehensive ablation studies underscore the efficacy of each component, corroborating the model's robustness. Experimental evaluations further illuminate ASGCL's proficiency in predicting drug responses, offering a potent tool for guiding clinical decision-making in cancer therapy.

Humans↗

[Interview in medicine].

The very earliest myths relating to the art of healing give great weight to the vital importance of human dialogue. Today's technology, however, supplies the physician with so many diagnostic and therapeutic tools that the one-on-one encounter with the patient is steadily losing its significance. And yet, even today the dialogue is still one of the most important components of the healing process. The dialogue between physician and patient has three generally acknowledged objectives: gaining information about the illness, getting a clear understanding of the patient as a person, and the therapeutic effect. Taking the patient's history largely serves the first objective. At this stage, the physician--depending on his orientation--already chooses from a wide spectrum of approaches: from conducting a cursory, purpose-oriented interview to engaging in an open compassionate dialogue. The last approach, anchored in a spirit attuned to the patient's psyche and mind-set, will not only yield important clues as to the systemic interrelationships underlying the disease, but it also provides the foundation for the therapeutic dimension of the dialogue. By engaging in a dialogue and by his very readiness to do so, the physician unconsciously reveals a lot about his own person. In order to achieve a true dialogue, it is necessary that he abandon the role-playing so common to physician-patient relationships and that he meet his patient on a person-to-person basis. In the process he will reveal his self-perception, his relationship to fellow human beings and, not least, his idea of what it means to be a healer. This idea depends completely on his perception of human beings in general, on whether he sees people as complex machines controlled by the brain computer or as spiritual beings whose bodies function as the screen onto which their thoughts, ideas and convictions are projected. A true dialogue between physician and patient provides the foundation for healing to take place. The complete trust of the patient, his willingness to cooperate, his compliance, and, in the end, the very healing process itself depend on the physician's ability to engage in dialogue. The art of meaningful communication can be taught. Medical students, the physicians-to-be, can learn a lot from the example set by their mentors. When university professors, chiefs of service and supervising attending physicians are appointed, their ability to demonstrate not only a gift of dialogue, but also the awareness of its vital importance in a doctor-physician relationship should be a decisive factor.

Communication↗

Dynamic On-line Clustering and State Extraction: An Approach to Symbolic Learning.

Although recurrent neural nets have been moderately successful in learning to emulate finite-state machines (FSMs), the continuous internal state dynamics of a neural net are not well matched to the discrete behavior of an FSM. We describe an architecture, called DOLCE, that allows discrete states to evolve in a net as learning progresses. DOLCE consists of a standard recurrent neural net trained by gradient descent and an adaptive clustering technique that quantizes the state space. We describe two implementations of DOLCE. The first implementation, called DOLCE(u), uses an adaptive clustering scheme in an unsupervised mode to determine both the number of clusters and the partitioning of the state space as learning progresses. The second model, DOLCE(s), uses a Gaussian Mixture Model in a supervised learning framework to infer the states of an FSM. DOLCE(s) is based on the assumption that a finite set of discrete internal states is required for the task, and that the actual network state belongs to this set but has been corrupted by noise due to inaccuracy in the weights. DOLCE(s) learns to recover the discrete state with maximum a posteriori probability from the noisy state. Simulations show that both implementations of DOLCE lead to a significant improvement in generalization performance over earlier neural net approaches to FSM induction. The idea of adaptive quantization is not just applicable to DOLCE but can be applied to other domains as well.

Journal Article↗

A unified benchmark of supervised and retrieval-based methods for viral genomic sequence classification.

The rapid growth of genomic sequencing demands fast, accurate, and scalable analysis methods. In viral genomic classification, expanding labeled reference collections can make supervised models costly to update and dependent on fixed label sets, motivating retrieval-based genomic classification as a simpler, more flexible alternative. We present a unified benchmark of supervised and retrieval-based methods for viral genomic sequence classification across three viral classification tasks: hepatitis C virus (HCV) genotyping, COVID-19 discrimination, and human papillomavirus (HPV) genotyping. We compare standard sequence encodings (one-hot, k-mers, FCGR) with dense embeddings (dna2vec, DNABERT). For each representation, we evaluate supervised classifiers (Random Forest, Decision Tree, XGBoost) and retrieval-based classification, where sequence vectors are indexed with FAISS and labels are assigned via similarity-weighted k-NN. Furthermore, we benchmark multiple FAISS index types (Flat, IVF, HNSW, IVFPQ, OPQ) to characterize accuracy-speed-memory trade-offs at scale. The results show that XGBoost and retrieval using Flat or IVF indexes achieve strong classification performance under different computational profiles. Compressed indexes such as IVFPQ and OPQ substantially reduce memory usage, although their accuracy loss depends on the dataset and representation. Overall, supervised XGBoost provides a favorable accuracy-size trade-off, while retrieval-based classification remains competitive and allows labeled reference sequences to be incorporated without retraining a global classifier. This benchmark provides practical guidance for selecting sequence representations, classifiers, and vector-search indexes under different accuracy, memory, and update requirements.

Genome, Viral↗

Evaluation of different biological data and computational classification methods for use in protein interaction prediction.

Protein-protein interactions play a key role in many biological systems. High-throughput methods can directly detect the set of interacting proteins in yeast, but the results are often incomplete and exhibit high false-positive and false-negative rates. Recently, many different research groups independently suggested using supervised learning methods to integrate direct and indirect biological data sources for the protein interaction prediction task. However, the data sources, approaches, and implementations varied. Furthermore, the protein interaction prediction task itself can be subdivided into prediction of (1) physical interaction, (2) co-complex relationship, and (3) pathway co-membership. To investigate systematically the utility of different data sources and the way the data is encoded as features for predicting each of these types of protein interactions, we assembled a large set of biological features and varied their encoding for use in each of the three prediction tasks. Six different classifiers were used to assess the accuracy in predicting interactions, Random Forest (RF), RF similarity-based k-Nearest-Neighbor, Naïve Bayes, Decision Tree, Logistic Regression, and Support Vector Machine. For all classifiers, the three prediction tasks had different success rates, and co-complex prediction appears to be an easier task than the other two. Independently of prediction task, however, the RF classifier consistently ranked as one of the top two classifiers for all combinations of feature sets. Therefore, we used this classifier to study the importance of different biological datasets. First, we used the splitting function of the RF tree structure, the Gini index, to estimate feature importance. Second, we determined classification accuracy when only the top-ranking features were used as an input in the classifier. We find that the importance of different features depends on the specific prediction task and the way they are encoded. Strikingly, gene expression is consistently the most important feature for all three prediction tasks, while the protein interactions identified using the yeast-2-hybrid system were not among the top-ranking features under any condition.

Computational Biology↗

A joint health and social services initiative for children with disabilities.

The children's disability team in Cambridge provides an integrated health and social care service for children with complex learning and physical disabilities and their families. The team uses a multidisciplinary and multi-agency teamwork approach to care provision. The effectiveness of the team was evaluated using a cooperative review of its functions, in which all the 'subjects' were active participants in defining and delivering the evaluation. This was combined with individual questionnaires regarding the team's perceived strengths and weaknesses. Particular implications for training and supervision emerged from the findings. This article discusses the ways in which the team has successfully refined its practice of collaborative working in a developmental way between 1992-1998.

Child↗

Profile-based string kernels for remote homology detection and motif extraction.

We introduce novel profile-based string kernels for use with support vector machines (SVMs) for the problems of protein classification and remote homology detection. These kernels use probabilistic profiles, such as those produced by the PSI-BLAST algorithm, to define position-dependent mutation neighborhoods along protein sequences for inexact matching of k-length subsequences ("k-mers") in the data. By use of an efficient data structure, the kernels are fast to compute once the profiles have been obtained. For example, the time needed to run PSI-BLAST in order to build the pro- files is significantly longer than both the kernel computation time and the SVM training time. We present remote homology detection experiments based on the SCOP database where we show that profile-based string kernels used with SVM classifiers strongly outperform all recently presented supervised SVM methods. We also show how we can use the learned SVM classifier to extract "discriminative sequence motifs" -- short regions of the original profile that contribute almost all the weight of the SVM classification score -- and show that these discriminative motifs correspond to meaningful structural features in the protein data. The use of PSI-BLAST profiles can be seen as a semi-supervised learning technique, since PSI-BLAST leverages unlabeled data from a large sequence database to build more informative profiles. Recently presented "cluster kernels" give general semi-supervised methods for improving SVM protein classification performance. We show that our profile kernel results are comparable to cluster kernels while providing much better scalability to large datasets.

Algorithms↗

Feature subset selection for splice site prediction.

MOTIVATION: The large amount of available annotated Arabidopsis thaliana sequences allows the induction of splice site prediction models with supervised learning algorithms (see Haussler (1998) for a review and references). These algorithms need information sources or features from which the models can be computed. For splice site prediction, the features we consider in this study are the presence or absence of certain nucleotides in close proximity to the splice site. Since it is not known how many and which nucleotides are relevant for splice site prediction, the set of features is chosen large enough such that the probability that all relevant information sources are in the set is very high. Using only those features that are relevant for constructing a splice site prediction system might improve the system and might also provide us with useful biological knowledge. Using fewer features will of course also improve the prediction speed of the system. RESULTS: A wrapper-based feature subset selection algorithm using a support vector machine or a naive Bayes prediction method was evaluated against the traditional method for selecting features relevant for splice site prediction. Our results show that this wrapper approach selects features that improve the performance against the use of all features and against the use of the features selected by the traditional method. AVAILABILITY: The data and additional interactive graphs on the selected feature subsets are available at http://www.psb.rug.ac.be/gps

Arabidopsis↗

Best harmony, unified RPCL and automated model selection for unsupervised and supervised learning on Gaussian mixtures, three-layer nets and ME-RBF-SVM models.

After introducing the fundamentals of BYY system and harmony learning, which has been developed in past several years as a unified statistical framework for parameter learning, regularization and model selection, we systematically discuss this BYY harmony learning on systems with discrete inner-representations. First, we shown that one special case leads to unsupervised learning on Gaussian mixture. We show how harmony learning not only leads us to the EM algorithm for maximum likelihood (ML) learning and the corresponding extended KMEAN algorithms for Mahalanobis clustering with criteria for selecting the number of Gaussians or clusters, but also provides us two new regularization techniques and a unified scheme that includes the previous rival penalized competitive learning (RPCL) as well as its various variants and extensions that performs model selection automatically during parameter learning. Moreover, as a by-product, we also get a new approach for determining a set of 'supporting vectors' for Parzen window density estimation. Second, we shown that other special cases lead to three typical supervised learning models with several new results. On three layer net, we get (i) a new regularized ML learning, (ii) a new criterion for selecting the number of hidden units, and (iii) a family of EM-like algorithms that combines harmony learning with new techniques of regularization. On the original and alternative models of mixture-of-expert (ME) as well as radial basis function (RBF) nets, we get not only a new type of criteria for selecting the number of experts or basis functions but also a new type of the EM-like algorithms that combines regularization techniques and RPCL learning for parameter learning with either least complexity nature on the original ME model or automated model selection on the alternative ME model and RBF nets. Moreover, all the results for the alternative ME model are also applied to other two popular nonparametric statistical approaches, namely kernel regression and supporting vector machine. Particularly, not only we get an easily implemented approach for determining the smoothing parameter in kernel regression, but also we get an alternative approach for deciding the set of supporting vectors in supporting vector machine.

Algorithms↗

Profile-based string kernels for remote homology detection and motif extraction.

We introduce novel profile-based string kernels for use with support vector machines (SVMs) for the problems of protein classification and remote homology detection. These kernels use probabilistic profiles, such as those produced by the PSI-BLAST algorithm, to define position-dependent mutation neighborhoods along protein sequences for inexact matching of k-length subsequences ("k-mers") in the data. By use of an efficient data structure, the kernels are fast to compute once the profiles have been obtained. For example, the time needed to run PSI-BLAST in order to build the profiles is significantly longer than both the kernel computation time and the SVM training time. We present remote homology detection experiments based on the SCOP database where we show that profile-based string kernels used with SVM classifiers strongly outperform all recently presented supervised SVM methods. We further examine how to incorporate predicted secondary structure information into the profile kernel to obtain a small but significant performance improvement. We also show how we can use the learned SVM classifier to extract "discriminative sequence motifs"--short regions of the original profile that contribute almost all the weight of the SVM classification score--and show that these discriminative motifs correspond to meaningful structural features in the protein data. The use of PSI-BLAST profiles can be seen as a semi-supervised learning technique, since PSI-BLAST leverages unlabeled data from a large sequence database to build more informative profiles. Recently presented "cluster kernels" give general semi-supervised methods for improving SVM protein classification performance. We show that our profile kernel results also outperform cluster kernels while providing much better scalability to large datasets.

Algorithms↗

PharaCon: a new framework for identifying bacteriophages via conditional representation learning.

MOTIVATION: Identifying bacteriophages (phages) within metagenomic sequences is essential for understanding microbial community dynamics. Transformer-based foundation models have been successfully employed to address various biological challenges. However, these models are typically pre-trained with self-supervised tasks that do not consider label variance in the pre-training data. This presents a challenge for phage identification as pre-training on mixed bacterial and phage data may lead to information bias due to the imbalance between bacterial and phage samples. RESULTS: To overcome this limitation, we proposed a novel conditional BERT framework that incorporates label classes as special tokens during pre-training. Specifically, our conditional BERT model attaches labels directly during tokenization, introducing label constraints into the model's input. Additionally, we introduced a new fine-tuning scheme that enables the conditional BERT to be effectively utilized for classification tasks. This framework allows the BERT model to acquire label-specific contextual representations from mixed sequence data during pre-training and applies the conditional BERT as a classifier during fine-tuning, and we named the fine-tuned model as PharaCon. We evaluated PharaCon against several existing methods on both simulated sequence datasets and real metagenomic contig datasets. The results demonstrate PharaCon's effectiveness and efficiency in phage identification, highlighting the advantages of incorporating label information during both pre-training and fine-tuning. AVAILABILITY AND IMPLEMENTATION: The source code and associated data can be accessed at https://github.com/Celestial-Bai/PharaCon.

Bacteriophages↗

[Increase in strength after active therapy in chronic low back pain (CLBP) patients: muscular adaptations and clinical relevance].

INTRODUCTION: Active treatments are advocated for the management of non-specific chronic low back pain (CLBP), although few studies have documented the relative efficacy of differing types of programme. A number of the available treatments comprise exercise routines on specially designed training machines, which are ostensibly better disposed to reverse the compromised trunk muscle function displayed by these patients than are 'free exercise' programmes. However, in using these muscle-training programmes, the physiological or anatomical adaptations that might account for the improved performance are rarely investigated, let alone identified. This is an important issue, because if the 'newly-acquired strength' is mostly specific to performance on the devices on which the patient has trained and been tested, and reflects the skill in executing these particular tasks, this will not necessarily assist the patient during performance of his/her everyday activities. The aims of the present study were (1) to quantify the changes in back muscle performance in chronic LBP patients following 3 months active therapy, and (2) to analyse the corresponding changes in activation and cross-sectional area of the paraspinal muscles. METHODS: 148 individuals (57% women) with CLBP (age 45.0+/-10.0 years; duration of LBP 10.9+/-9.5 years) were randomised to a treatment which they attended 2/week for 3 months: active physiotherapy, muscle reconditioning on training devices, or low-impact aerobics. Pre- and post-therapy, assessments were made of isometric trunk muscle strength in each plane of movement and of erector spinae activation (using surface electromyography) during back extension. In a sub-group of 56 patients, the cross-sectional area of the paravertebral muscles was determined using magnetic resonance imaging (MRI). In all patients, self-rated pain intensity, pain frequency and disability were assessed before and after therapy. RESULTS: 132/148 patients completed the therapy. Isometric strength in each movement plane increased significantly in all groups post-therapy. Apart from trunk extension, the changes were significantly greater in the devices group than in the other two groups (Fig 1). Activation of the paraspinal muscles during back extension also increased significantly in all groups (Fig 2) and was weakly, but significantly (r = 0.37; p = 0.0001) correlated with increased strength in back extension. Although, at baseline, highly significant correlations were observed between the size of the paraspinal muscles (at L3/4 and at L4/5) and isometric back extension strength (r=0.75; p< 0.0001), post-training increases in strength were not accompanied by corresponding changes in muscle size. None of the improvements in strength showed any relationship with the clinical changes in pain and disability, regardless of whether the latter were examined on an individual basis or in relation to 'outcome groups'. CONCLUSION: The superior trunk strength shown by the devices group post-therapy was considered to be attributable, in part, to a 'learning effect', of the type often seen when training and testing are carried out on the same machines. These gains are considered to be mostly 'task-specific'. However, part of the improvement in strength after active therapy (in all groups) also appeared to be due to an increased neural activation of the trunk muscles. These positive effects should be transferable to the performance of everyday activities for which the same muscles are employed, although the percentage improvement is probably not as high as the measured increase in strength might suggest. Possible roles for improved co-ordination and changes in motivation and/or pain tolerance after therapy cannot be excluded. No differences in the clinical outcome were observed between the three therapy groups, and the changes in physical performance after therapy did not correlate with the clinical outcome. It is therefore questionable whether strength measurements have any clinical significance in documenting the success of rehabilitation programmes, other than on a motivational basis. The results of the present study suggest that the value of supervised active therapy programmes does not reside in the reversal of specific muscular deficiencies, but rather in the provision of a source of confirmation/encouragement for the patient, that movement is not harmful, and a foundation upon which to further build. Whether the utilisation of specific training devices, or individual instruction, is necessary to elicit these particular effects is questionable.

Adaptation, Physiological↗

A Meta-learning-driven strategy for adulteration detection in sweet potato starch and vermicelli using Raman spectroscopy.

To address the widespread adulteration of sweet potato starch and its vermicelli with cheaper starches and overcome conventional supervised learning's dependency on large labeled datasets, this study developed a few-shot discrimination method integrating Raman spectroscopy with meta-learning. We constructed a meta-learning framework using cassava- and wheat-adulterated sweet potato starch as the source domain for training, with potato-adulterated sweet potato starch and cassava-adulterated sweet potato vermicelli as two target domains for testing. Raman spectra showed high consistency between sweet potato vermicelli and its raw starch, laying the foundation for cross-domain detection. Testing yielded comprehensive classification accuracies of 95.33% and 98.00% for the two target domains, significantly outperforming SVM, RF, and CNN (max. 85.24%). This approach effectively identifies subtle starch variety differences in complex adulteration, providing novel food quality inspection solutions and verifying the feasibility of raw material-to-finished product cross-domain detection.

Ipomoea batatas↗

Learning optimized features for hierarchical models of invariant object recognition.

There is an ongoing debate over the capabilities of hierarchical neural feedforward architectures for performing real-world invariant object recognition. Although a variety of hierarchical models exists, appropriate supervised and unsupervised learning methods are still an issue of intense research. We propose a feedforward model for recognition that shares components like weight sharing, pooling stages, and competitive nonlinearities with earlier approaches but focuses on new methods for learning optimal feature-detecting cells in intermediate stages of the hierarchical network. We show that principles of sparse coding, which were previously mostly applied to the initial feature detection stages, can also be employed to obtain optimized intermediate complex features. We suggest a new approach to optimize the learning of sparse features under the constraints of a weight-sharing or convolutional architecture that uses pooling operations to achieve gradual invariance in the feature hierarchy. The approach explicitly enforces symmetry constraints like translation invariance on the feature set. This leads to a dimension reduction in the search space of optimal features and allows determining more efficiently the basis representatives, which achieve a sparse decomposition of the input. We analyze the quality of the learned feature representation by investigating the recognition performance of the resulting hierarchical network on object and face databases. We show that a hierarchy with features learned on a single object data set can also be applied to face recognition without parameter changes and is competitive with other recent machine learning recognition approaches. To investigate the effect of the interplay between sparse coding and processing nonlinearities, we also consider alternative feedforward pooling nonlinearities such as presynaptic maximum selection and sum-of-squares integration. The comparison shows that a combination of strong competitive nonlinearities with sparse coding offers the best recognition performance in the difficult scenario of segmentation-free recognition in cluttered surround. We demonstrate that for both learning and recognition, a precise segmentation of the objects is not necessary.

Learning↗

Organ-delimited gene regulatory networks provide high accuracy in candidate transcription factor selection across diverse processes.

Organ-specific gene expression datasets that include hundreds to thousands of experiments allow the reconstruction of organ-level gene regulatory networks (GRNs). However, creating such datasets is greatly hampered by the requirements of extensive and tedious manual curation. Here, we trained a supervised classification model that can accurately classify the organ-of-origin for a plant transcriptome. This K-Nearest Neighbor-based multiclass classifier was used to create organ-specific gene expression datasets for the leaf, root, shoot, flower, and seed in Arabidopsis thaliana. A GRN inference approach was used to determine the: i. influential transcription factors (TFs) in each organ and, ii. most influential TFs for specific biological processes in that organ. These genome-wide, organ-delimited GRNs (OD-GRNs), recalled many known regulators of organ development and processes operating in those organs. Importantly, many previously unknown TF regulators were uncovered as potential regulators of these processes. As a proof-of-concept, we focused on experimentally validating the predicted TF regulators of lipid biosynthesis in seeds, an important food and biofuel trait. Of the top 20 predicted TFs, eight are known regulators of seed oil content, e.g., WRI1, LEC1, FUS3. Importantly, we validated our prediction of MybS2, TGA4, SPL12, AGL18, and DiV2 as regulators of seed lipid biosynthesis. We elucidated the molecular mechanism of MybS2 and show that it induces purple acid phosphatase family genes and lipid synthesis genes to enhance seed lipid content. This general approach has the potential to be extended to any species with sufficiently large gene expression datasets to find unique regulators of any trait-of-interest.

Arabidopsis↗