Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

Symmetric and asymmetric multi-modality biclustering analysis for microarray data matrix.

Machine learning techniques offer a viable approach to cluster discovery from microarray data, which involves identifying and classifying biologically relevant groups in genes and conditions. It has been recognized that genes (whether or not they belong to the same gene group) may be co-expressed via a variety of pathways. Therefore, they can be adequately described by a diversity of coherence models. In fact, it is known that a gene may participate in multiple pathways that may or may not be co-active under all conditions. It is therefore biologically meaningful to simultaneously divide genes into functional groups and conditions into co-active categories--leading to the so-called biclustering analysis. For this, we have proposed a comprehensive set of coherence models to cope with various plausible regulation processes. Furthermore, a multivariate biclustering analysis based on fusion of different coherence models appears to be promising because the expression level of genes from the same group may follow more than one coherence models. The simulation studies further confirm that the proposed framework enjoys the advantage of high prediction performance.

Algorithms↗

Classification and selection of biomarkers in genomic data using LASSO.

High-throughput gene expression technologies such as microarrays have been utilized in a variety of scientific applications. Most of the work has been done on assessing univariate associations between gene expression profiles with clinical outcome (variable selection) or on developing classification procedures with gene expression data (supervised learning). We consider a hybrid variable selection/classification approach that is based on linear combinations of the gene expression profiles that maximize an accuracy measure summarized using the receiver operating characteristic curve. Under a specific probability model, this leads to the consideration of linear discriminant functions. We incorporate an automated variable selection approach using LASSO. An equivalence between LASSO estimation with support vector machines allows for model fitting using standard software. We apply the proposed method to simulated data as well as data from a recently published prostate cancer study.

Journal Article↗

A machine learning approach to the analysis of time-frequency maps, and its application to neural dynamics.

The statistical analysis of experimentally recorded brain activity patterns may require comparisons between large sets of complex signals in order to find meaningful similarities and differences between signals with large variability. High-level representations such as time-frequency maps convey a wealth of useful information, but they involve a large number of parameters that make statistical investigations of many signals difficult at present. In this paper, we describe a method that performs drastic reduction in the complexity of time-frequency representations through a modelling of the maps by elementary functions. The method is validated on artificial signals and subsequently applied to electrophysiological brain signals (local field potential) recorded from the olfactory bulb of rats while they are trained to recognize odours. From hundreds of experimental recordings, reproducible time-frequency events are detected, and relevant features are extracted, which allow further information processing, such as automatic classification.

Algorithms↗

Warmr: a data mining tool for chemical data.

Data mining techniques are becoming increasingly important in chemistry as databases become too large to examine manually. Data mining methods from the field of Inductive Logic Programming (ILP) have potential advantages for structural chemical data. In this paper we present Warmr, the first ILP data mining algorithm to be applied to chemoinformatic data. We illustrate the value of Warmr by applying it to a well studied database of chemical compounds tested for carcinogenicity in rodents. Data mining was used to find all frequent substructures in the database, and knowledge of these frequent substructures is shown to add value to the database. One use of the frequent substructures was to convert them into probabilistic prediction rules relating compound description to carcinogenesis. These rules were found to be accurate on test data, and to give some insight into the relationship between structure and activity in carcinogenesis. The substructures were also used to prove that there existed no accurate rule, based purely on atom-bond substructure with less than seven conditions, that could predict carcinogenicity. This results put a lower bound on the complexity of the relationship between chemical structure and carcinogenicity. Only by using a data mining algorithm, and by doing a complete search, is it possible to prove such a result. Finally the frequent substructures were shown to add value by increasing the accuracy of statistical and machine learning programs that were trained to predict chemical carcinogenicity. We conclude that Warmr, and ILP data mining methods generally, are an important new tool for analysing chemical databases.

Algorithms↗

Locally linear discriminant analysis for multimodally distributed classes for face recognition with a single model image.

We present a novel method of nonlinear discriminant analysis involving a set of locally linear transformations called "Locally Linear Discriminant Analysis (LLDA)." The underlying idea is that global nonlinear data structures are locally linear and local structures can be linearly aligned. Input vectors are projected into each local feature space by linear transformations found to yield locally linearly transformed classes that maximize the between-class covariance while minimizing the within-class covariance. In face recognition, linear discriminant analysis (LDA) has been widely adopted owing to its efficiency, but it does not capture nonlinear manifolds of faces which exhibit pose variations. Conventional nonlinear classification methods based on kernels such as generalized discriminant analysis (GDA) and support vector machine (SVM) have been developed to overcome the shortcomings of the linear method, but they have the drawback of high computational cost of classification and overfitting. Our method is for multiclass nonlinear discrimination and it is computationally highly efficient as compared to GDA. The method does not suffer from overfitting by virtue of the linear base structure of the solution. A novel gradient-based learning algorithm is proposed for finding the optimal set of local linear bases. The optimization does not exhibit a local-maxima problem. The transformation functions facilitate robust face recognition in a low-dimensional subspace, under pose variations, using a single model image. The classification results are given for both synthetic and real face data.

Algorithms↗

Integrating histology and spatial transcriptomics via multimodal transformers and contrastive representation learning for accurate gene expression prediction.

Predicting spatial gene expression from Histological images is a fundamental task in understanding tissue organization and molecular phenotypes. However, existing methods often rely on single-model representations or lack effective alignment between image and transcriptomic features. To address these limitations, we propose a unified multimodal learning framework that integrates histological imaging and spatial transcriptomics through a shared latent representation space. Specifically, histological H&E images are encoded by a ResNet50-based convolutional stem and a MobileViT Transformer backbone to extract hierarchical visual representations. Both modalities are projected into a shared latent space via linear-GELU-dropout transformation blocks, enabling cross-modal alignment through a contrastive learning objective that maximizes agreement between the corresponding image and the spot embeddings. Experimental results on the 10x Genomics Visium dataset of human liver tissue demonstrate that MViTGene achieves significantly higher prediction accuracy than existing methods across multiple gene subsets, with improvements of 20%, 33%, and 12% in predicting marker genes, highly expressed genes, and highly variable genes, respectively. The significant improvement in relevance indicates that the model can more accurately capture the true correspondence between tissue morphology and gene expression, therefore enabling more reliable biological interpretation. It provides a computational tool for high-throughput spatial gene expression prediction that balances performance and interpretability.

Humans↗

Global survey of organ and organelle protein expression in mouse: combined proteomic and transcriptomic profiling.

Organs and organelles represent core biological systems in mammals, but the diversity in protein composition remains unclear. Here, we combine subcellular fractionation with exhaustive tandem mass spectrometry-based shotgun sequencing to examine the protein content of four major organellar compartments (cytosol, membranes [microsomes], mitochondria, and nuclei) in six organs (brain, heart, kidney, liver, lung, and placenta) of the laboratory mouse, Mus musculus. Using rigorous statistical filtering and machine-learning methods, the subcellular localization of 3274 of the 4768 proteins identified was determined with high confidence, including 1503 previously uncharacterized factors, while tissue selectivity was evaluated by comparison to previously reported mRNA expression patterns. This molecular compendium, fully accessible via a searchable web-browser interface, serves as a reliable reference of the expressed tissue and organelle proteomes of a leading model mammal.

Animals↗

Fuzzy support vector machines for adaptive Morse code recognition.

Morse code is now being harnessed for use in rehabilitation applications of augmentative-alternative communication and assistive technology, facilitating mobility, environmental control and adapted worksite access. In this paper, Morse code is selected as a communication adaptive device for persons who suffer from muscle atrophy, cerebral palsy or other severe handicaps. A stable typing rate is strictly required for Morse code to be effective as a communication tool. Therefore, an adaptive automatic recognition method with a high recognition rate is needed. The proposed system uses both fuzzy support vector machines and the variable-degree variable-step-size least-mean-square algorithm to achieve these objectives. We apply fuzzy memberships to each point, and provide different contributions to the decision learning function for support vector machines. Statistical analyses demonstrated that the proposed method elicited a higher recognition rate than other algorithms in the literature.

Algorithms↗

Multidimensional signal exploration using multiple correspondence analysis. An example of a load lifting study.

Most empirical studies concerning rehabilitation yield numerous multidimensional signals (dozens of time variables are obtained for dozens of empirical situations). The purpose of this paper is to suggest a statistical analysis procedure based on: 1) space-time fuzzy windowing; 2) signal behavior characterization within the windows using membership value averages (MVA); and 3) MVA analysis using the multiple correspondence analysis (MCA). A load lifting study provided an example of 78 multidimensional signals including 89 time variables (forces, energy indicators, linear and angular positions, speeds, and accelerations). The main goal of MCA was to compare and contrast biomechanical signals from two lifting modes: "free" and "isokinetic." In the first mode, three loads were tested--light, medium, and heavy. In the second, three speeds were tested--slow, medium, and fast. Thirteen male individuals without disabilities participated in this study. The MCA showed that most of the free load-lifting strategies cannot be used in isokinetic lifting because the constraints of the subject and the environment are different. In addition, as the level of difficulty increases, free lifting became more economical while isokinetic lifting became less economical. These results would appear to indicate that movement strategies used for free lifting cannot be learned using an isokinetic machine during rehabilitation sessions for chronic low back pain. MCA was also suggested as a tool for comparing patients with control individuals. To achieve this aim, the notion of "supplementary data" was introduced.

Adult↗

PicSOr: an objective test of perceptual skill that predicts laparoscopic technical skill in three initial studies of laparoscopic performance.

BACKGROUND: Laparoscopic surgery requires surgeons to infer the shape of 3-D structures, such as the internal organs of patients, from 2-D displays on a video monitor. Recent evidence indicates that the issue is not resolved by the use of contemporary 3-D camera systems. It is therefore crucial to find ways of measuring differences in aptitude for recovering 3-D structure from 2-D images, and assessing its impact on performance. Our aim was to test empirically for a relationship between laparoscopic ability and the perceptual skill of recovering information about 3-D structures from 2-D monitor displays. METHODS: Participants in three studies completed a simulated laparoscopic cutting task as well as the Pictorial Surface Orientation (PicSOr)3 Test. In studies 1 (n = 48) and 2 (n = 32) both groups were laparoscopic novices, and in study 3 (n = 34) 18 of the participants were experienced laparoscopic surgeons. FINDINGS: All three studies showed that PicSOr consistently predicted the laparoscopic performance of participants on the laparoscopic cutting task (study 1, r = 0.5, p < 0.0003; study 2, r = 0.5, p < 0.004; and study 3, r = 0.42, p = 0.017). Furthermore, it was also a significant predictor of laparoscopic surgeons' performance (r = 0.54, p = 0.047). INTERPRETATIONS: This is the first objective perceptual psychometric test to reliably predict laparoscopic technical skills. PicSOr provides a tool for assessing which trainees have the potential to learn minimal access surgery.

Adult↗

Order of Search in Fuzzy ART and Fuzzy ARTMAP: Effect of the Choice Parameter.

This paper focuses on two ART architectures, the Fuzzy ART and the Fuzzy ARTMAP. Fuzzy ART is a pattern clustering machine, while Fuzzy ARTMAP is a pattern classification machine. Our study concentrates on the order according to which categories in Fuzzy ART, or the ART(a) model of Fuzzy ARTMAP are chosen. Our work provides a geometrical, and clearer understanding of why, and in what order, these categories are chosen for various ranges of the choice parameter of the Fuzzy ART module. This understanding serves as a powerful tool in developing properties of learning pertaining to these neural network architectures; to strengthen this argument, it is worth mentioning that the order according to which categories are chosen in ART 1 and ARTMAP provided a valuable tool in proving important properties about these architectures. Copyright 1996 Elsevier Science Ltd.

Journal Article↗

Sex-specific associations of the plasma-proteome with incident coronary artery disease.

AIMS: The etiology of coronary artery Disease (CAD) appears different for men and women, yet insights into underlying sex-specific biological mechanisms are limited. We integrated genomic and proteomic analyses to investigate sex-specific associations of the plasma-proteome with CAD. METHODS AND RESULTS: In 40,829 UK Biobank participants (free-of-CAD, baseline-365 days thereafter; 55% women; mean age 56.9&#x2009;&#xb1;&#x2009;8.1 years), we examined associations between 2,922 plasma proteins and incident CAD over a median follow-up of 13.7 years (IQR 13.1-14.4) using multivariable-adjusted Cox proportional hazards models. Sex-specific analyses identified 440 female exclusive and 32 male exclusive proteins associated with incident CAD (FDR-corrected p&#x2009;<&#x2009;0.05), revealing distinct pathway enrichments, including innate immune response in women and angiogenesis in men. Causality was assessed through combined and sex-stratified two-sample Mendelian randomization (MR) using inverse-variance-weighted analyses with genome wide association summary statistics from 422,108 men (61,969 cases) and 521,695 women (27,128 cases) (UK Biobank, FinnGen freeze 9). Integration of direct sex-protein interaction analyses with sex-combined MR identified 59 proteins with evidence for sex-specific causal effects. Four proteins demonstrated concordant directionality in sex-stratified MR analyses (n&#x2009;=&#x2009;943,803) and multivariable regression models, namely CDKN2D, MYH9, and SKAP2 (women), and CTSH (men). To assess translational relevance, prioritized targets were further evaluated in secondary major adverse cardiovascular events among carotid endarterectomy patients (MACE; Athero-Express) and acute myocardial infarction (AMI; MISSION!) using plasma proteomics and ELISA. After further top-target identification in the context of MACE and AMI, clinical drug candidates were identified through a machine learning framework, including CTSH (men), and TNFRSF4 (both sexes). CONCLUSIONS: We identified sex-specific associations of proteins and biological pathways with incident CAD. Whereas the majority of proteins had consistent associations in both men and women, our findings suggest a degree of sex-specific pathogenesis with evidence for potential causality, opening new alleys for tailored prevention strategies and clinical cardiovascular risk management.

Journal Article↗

Placenta-derived Exosomes Mitigate Hypoxia-Induced Trophoblast Apoptosis and Inflammatory Progression via SASH1.

SASH1 is a signal adaptor protein involved in cell growth, apoptosis, and immune regulation, and has been increasingly studied in tumor and immune cells. Emerging evidence suggests that SASH1 plays an important role in inflammatory responses and cellular homeostasis, processes that are closely associated with the development of PE. This study aimed to determine whether SASH1 contributes to trophoblast apoptosis and inflammatory responses in PE and whether P-EXOS exerts protective effects through SASH1 regulation. In this study, three PE-related transcriptomic datasets (GSE75010, GSE10588, and GSE60438) were analyzed to identify shared differentially expressed genes (DEGs), followed by Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses. Machine learning algorithms were further applied to screen key candidate genes, and single-cell RNA sequencing data were used to characterize cellular heterogeneity in placental tissue and to determine cell type-specific expression patterns. SASH1 was identified as a consensus candidate gene and was significantly upregulated in trophoblast cells from PE samples. In vitro, a hypoxia-treated HTR-8/SVneo trophoblast cell model was established, combined with SASH1 knockdown, SASH1 overexpression, and co-culture with P-EXOS. Functional experiments showed that knockdown of SASH1 significantly suppressed hypoxia-induced trophoblast apoptosis and reduced the secretion of pro-inflammatory cytokines, including IL-6, IL-1&#x3b2;, and TNF-&#x3b1;, whereas SASH1 overexpression promoted apoptosis and inflammatory responses. In addition, P-EXOS treatment markedly reduced SASH1 expression at both mRNA and protein levels and attenuated hypoxia-induced trophoblast injury, while SASH1 overexpression largely abolished these protective effects. Taken together, these findings indicate that SASH1 plays a critical role in trophoblast apoptosis and inflammatory responses in PE. P-EXOS may alleviate hypoxia-induced trophoblastic injury by suppressing SASH1 expression, providing new insights into the molecular mechanisms and potential therapeutic targets for PE.

Trophoblasts↗

CAKL: Commutative algebra k-mer learning of genomics.

Despite the availability of various sequence analysis models, comparative genomic analysis remains a challenge in genomics, genetics, and phylogenetics. Commutative algebra, a fundamental tool in algebraic geometry and number theory, has rarely been used in data and biological sciences. In this study, we introduce commutative algebra k-mer learning (CAKL) as the first-ever nonlinear algebraic framework for analyzing genomic sequences. CAKL bridges between commutative algebra, algebraic topology, combinatorics, and machine learning to establish a new mathematical paradigm for comparative genomic analysis. We evaluate its effectiveness on three tasks-genetic variant identification, phylogenetic tree analysis, and viral genome classification-typically requiring alignment-based, alignment-free, and machine-learning approaches, respectively. Across eleven datasets, CAKL outperforms five state-of-the-art sequence analysis methods, particularly in viral classification, and maintains stable predictive accuracy as dataset size increases, underscoring its scalability and robustness. This work ushers in a new era in commutative algebraic data analysis and learning.

Journal Article↗

Estimation of the stapes-bone thickness in the stapedotomy surgical procedure using a machine-learning technique.

Stapedotomy is a surgical procedure aimed at the treatment of hearing impairment due to otosclerosis. The treatment consists of drilling a hole through the stapes bone in the inner ear in order to insert a prosthesis. Safety precautions require knowledge of the nonmeasurable stapes thickness. The technical goal herein has been the design of high-level controls for an intelligent mechatronics drilling tool in order to enable the estimation of stapes thickness from measurable drilling data. The goal has been met by learning a map between drilling features, hence no model of the physical system has been necessary. Learning has been achieved as explained in this paper by a scheme, namely the d-sigma Fuzzy Lattice Neurocomputing (d sigma-FLN) scheme for classification, within the framework of fuzzy lattices. The successful application of the d sigma-FLN scheme is demonstrated in estimating the thickness of a stapes bone "on-line" using drilling data obtained experimentally in the laboratory.

Deafness↗

A machine learning strategy to identify candidate binding sites in human protein-coding sequence.

BACKGROUND: The splicing of RNA transcripts is thought to be partly promoted and regulated by sequences embedded within exons. Known sequences include binding sites for SR proteins, which are thought to mediate interactions between splicing factors bound to the 5' and 3' splice sites. It would be useful to identify further candidate sequences, however identifying them computationally is hard since exon sequences are also constrained by their functional role in coding for proteins. RESULTS: This strategy identified a collection of motifs including several previously reported splice enhancer elements. Although only trained on coding exons, the model discriminates both coding and non-coding exons from intragenic sequence. CONCLUSION: We have trained a computational model able to detect signals in coding exons which seem to be orthogonal to the sequences' primary function of coding for proteins. We believe that many of the motifs detected here represent binding sites for both previously unrecognized proteins which influence RNA splicing as well as other regulatory elements.

Algorithms↗

Predicting carcinoid heart disease with the noisy-threshold classifier.

OBJECTIVE: To predict the development of carcinoid heart disease (CHD), which is a life-threatening complication of certain neuroendocrine tumors. To this end, a novel type of Bayesian classifier, known as the noisy-threshold classifier, is applied. MATERIALS AND METHODS: Fifty-four cases of patients that suffered from a low-grade midgut carcinoid tumor, of which 22 patients developed CHD, were obtained from the Netherlands Cancer Institute (NKI). Eleven attributes that are known at admission have been used to classify whether the patient develops CHD. Classification accuracy and area under the receiver operating characteristics (ROC) curve of the noisy-threshold classifier are compared with those of the naive-Bayes classifier, logistic regression, the decision-tree learning algorithm C4.5, and a decision rule, as formulated by an expert physician. RESULTS: The noisy-threshold classifier showed the best classification accuracy of 72% correctly classified cases, although differences were significant only for logistic regression and C4.5. An area under the ROC curve of 0.66 was attained for the noisy-threshold classifier, and equaled that of the physician's decision-rule. CONCLUSIONS: The noisy-threshold classifier performed favorably to other state-of-the-art classification algorithms, and equally well as a decision-rule that was formulated by the physician. Furthermore, the semantics of the noisy-threshold classifier make it a useful machine learning technique in domains where multiple causes influence a common effect.

Algorithms↗

Classification of non-coding RNA using graph representations of secondary structure.

Some genes produce transcripts that function directly in regulatory, catalytic, or structural roles in the cell. These non-coding RNAs are prevalent in all living organisms, and methods that aid the understanding of their functional roles are essential. RNA secondary structure, the pattern of base-pairing, contains the critical information for determining the three dimensional structure and function of the molecule. In this work we examine whether the basic geometric and topological properties of secondary structure are sufficient to distinguish between RNA families in a learning framework. First, we develop a labeled dual graph representation of RNA secondary structure by adding biologically meaningful labels to the dual graphs proposed by Gan et al [1]. Next, we define a similarity measure directly on the labeled dual graphs using the recently developed marginalized kernels [2]. Using this similarity measure, we were able to train Support Vector Machine classifiers to distinguish RNAs of known families from random RNAs with similar statistics. For 22 of the 25 families tested, the classifier achieved better than 70% accuracy, with much higher accuracy rates for some families. Training a set of classifiers to automatically assign family labels to RNAs using a one vs. all multi-class scheme also yielded encouraging results. From these initial learning experiments, we suggest that the labeled dual graph representation, together with kernel machine methods, has potential for use in automated analysis and classification of uncharacterized RNA molecules or efficient genome-wide screens for RNA molecules from existing families.

Base Sequence↗