Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “machine learning prediction model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

619 records · Page 35Linked to original sources

Collateral missing value imputation: a new robust missing value estimation algorithm for microarray data.

MOTIVATION: Microarray data are used in a range of application areas in biology, although often it contains considerable numbers of missing values. These missing values can significantly affect subsequent statistical analysis and machine learning algorithms so there is a strong motivation to estimate these values as accurately as possible before using these algorithms. While many imputation algorithms have been proposed, more robust techniques need to be developed so that further analysis of biological data can be accurately undertaken. In this paper, an innovative missing value imputation algorithm called collateral missing value estimation (CMVE) is presented which uses multiple covariance-based imputation matrices for the final prediction of missing values. The matrices are computed and optimized using least square regression and linear programming methods. RESULTS: The new CMVE algorithm has been compared with existing estimation techniques including Bayesian principal component analysis imputation (BPCA), least square impute (LSImpute) and K-nearest neighbour (KNN). All these methods were rigorously tested to estimate missing values in three separate non-time series (ovarian cancer based) and one time series (yeast sporulation) dataset. Each method was quantitatively analyzed using the normalized root mean square (NRMS) error measure, covering a wide range of randomly introduced missing value probabilities from 0.01 to 0.2. Experiments were also undertaken on the yeast dataset, which comprised 1.7% actual missing values, to test the hypothesis that CMVE performed better not only for randomly occurring but also for a real distribution of missing values. The results confirmed that CMVE consistently demonstrated superior and robust estimation capability of missing values compared with other methods for both series types of data, for the same order of computational complexity. A concise theoretical framework has also been formulated to validate the improved performance of the CMVE algorithm. AVAILABILITY: The CMVE software is available upon request from the authors.

Algorithms↗

Automatic discovery of cross-family sequence features associated with protein function.

BACKGROUND: Methods for predicting protein function directly from amino acid sequences are useful tools in the study of uncharacterized protein families and in comparative genomics. Until now, this problem has been approached using machine learning techniques that attempt to predict membership, or otherwise, to predefined functional categories or subcellular locations. A potential drawback of this approach is that the human-designated functional classes may not accurately reflect the underlying biology, and consequently important sequence-to-function relationships may be missed. RESULTS: We show that a self-supervised data mining approach is able to find relationships between sequence features and functional annotations. No preconceived ideas about functional categories are required, and the training data is simply a set of protein sequences and their UniProt/Swiss-Prot annotations. The main technical aspect of the approach is the co-evolution of amino acid-based regular expressions and keyword-based logical expressions with genetic programming. Our experiments on a strictly non-redundant set of eukaryotic proteins reveal that the strongest and most easily detected sequence-to-function relationships are concerned with targeting to various cellular compartments, which is an area already well studied both experimentally and computationally. Of more interest are a number of broad functional roles which can also be correlated with sequence features. These include inhibition, biosynthesis, transcription and defence against bacteria. Despite substantial overlaps between these functions and their corresponding cellular compartments, we find clear differences in the sequence motifs used to predict some of these functions. For example, the presence of polyglutamine repeats appears to be linked more strongly to the "transcription" function than to the general "nuclear" function/location. CONCLUSION: We have developed a novel and useful approach for knowledge discovery in annotated sequence data. The technique is able to identify functionally important sequence features and does not require expert knowledge. By viewing protein function from a sequence perspective, the approach is also suitable for discovering unexpected links between biological processes, such as the recently discovered role of ubiquitination in transcription.

Algorithms↗

Identification of Biomarkers for Right Ventricular Dysfunction in Idiopathic Dilated Cardiomyopathy Via Urinary Proteomics and Machine Learning.

BACKGROUND: Right ventricular dysfunction (RVD) is a common complication of idiopathic dilated cardiomyopathy linked to poor outcomes. However, reliable noninvasive biomarkers for RVD remain lacking. This study aimed to identify urinary proteomic markers using mass spectrometry and machine learning. METHODS: In this prospective cohort, patients with idiopathic dilated cardiomyopathy were classified by cardiac magnetic resonance imaging into groups with RVD (RV ejection fraction <45%) and without RVD groups. Baseline urine samples were profiled by data-independent acquisition mass spectrometry. Differentially expressed proteins were identified and selected by least absolute shrinkage and selection operator regression to build a diagnostic model, developed in a training set, and validated in a test set. The primary end point was a composite of cardiovascular death, heart failure rehospitalization, left ventricular assist device implantation, or heart transplantation. RESULTS: The study enrolled 147 patients with idiopathic dilated cardiomyopathy (64 with RVD, 83 without), with a median follow-up of 19.3&#x2009;months. Of 3579 quantified urinary proteins, 46 were differentially expressed between groups. A 3-protein panel (RARRES1 [retinoic acid receptor responder protein 1], MVB12B [multivesicular body subunit 12B], GSK3A [glycogen synthase kinase 3 alpha]) was identified and showed excellent diagnostic accuracy (training area under the curve 0.946; validation area under the curve0.935), outperforming both NT-proBNP (N-terminal pro-brain natriuretic peptide) and tricuspid annular plane systolic excursion. The risk score derived from this panel effectively stratified patients, with the high-risk group exhibiting significantly worse outcomes than the low-risk group (hazard ratio, 3.24 [95% CI, 1.56-6.71], P=0.002). CONCLUSIONS: The urinary proteomic panel developed in this study demonstrates diagnostic and prognostic potential for identifying RVD in idiopathic dilated cardiomyopathy, providing a promising noninvasive tool for precise detection and clinical risk stratification.

Humans↗

Uncovering the genetic architecture of ME/CFS: a precision approach reveals impact of rare monogenic variation.

BACKGROUND: Myalgic encephalomyelitis/chronic fatigue syndrome (ME/CFS) is a disabling and heterogeneous disorder lacking validated biomarkers or targeted therapies. Clinical variability and elusive pathophysiology hinder progress toward effective diagnostics and treatment. Core symptoms include persistent fatigue, post-exertional malaise, unrefreshing sleep, cognitive dysfunction, and pain. We tested whether an individualized, &#x201c;n-of-1&#x201d; genomic and transcriptomic framework combined with comprehensive, participant-informed phenotyping could reveal molecular signatures unique to each patient. METHODS: Clinical-grade whole-genome sequencing was conducted in 31 affected individuals from 25 families, with RNA-seq performed on a subset (16 affected, 7 unaffected) using blood samples. Machine-learning assisted variant triage, transcript-aware damage prediction, and expert review identified pathogenic or likely pathogenic variants in 8 of 25 probands (32%) and 12 of 31 affected individuals (39%). RESULTS: Findings revealed marked genetic heterogeneity, including large-effect rare and more common variants. Implicated pathways included ATP generation, oxidative phosphorylation, fatty acid oxidation; regulation of glycolysis, amino acid and lipid turnover; ion and solute homeostasis; synaptic signaling, excitability, oxygen transport, and muscle integrity, resilience, and post-exertional recovery; previously implicated processes. Plausible modifiers influencing disease onset, severity, and relapsing&#x2013;remitting patterns and possibly explaining intrafamilial variability and inconsistent findings across studies, were also identified. Despite gene-level diversity, downstream effects converged on impaired energy production, reduced stress resilience, and vulnerability to post-exertional metabolic failure; disruptions consistent with core ME/CFS symptoms of exertional intolerance, cognitive fog, and fatigue. CONCLUSIONS: Our findings support the hypothesis that at least a subset of ME/CFS cases represent distinct molecular disorders that converge on shared physiological pathways. Validation in larger, more diverse cohorts will be essential to test this hypothesis and establish generalizability, but increase size alone is unlikely to resolve causation in a disorder defined by rarity, heterogeneity, and molecular complexity. We suggest that progress will require experimental designs that integrate individual-level genomic data with deep, participant-informed deep phenotyping, capturing the combined effects of rare and common variants and environmental modifiers on disease expression and progression. We believe that an individualized precision medicine framework will uncover molecular drivers and modifiers of ME/CFS previously obscured by heterogeneity, enabling biologically informed stratification, improved trial design, biomarker discovery, and targeted interventions in this historically neglected condition.

Humans↗

Predicting DNA-binding sites of proteins from amino acid sequence.

BACKGROUND: Understanding the molecular details of protein-DNA interactions is critical for deciphering the mechanisms of gene regulation. We present a machine learning approach for the identification of amino acid residues involved in protein-DNA interactions. RESULTS: We start with a Naïve Bayes classifier trained to predict whether a given amino acid residue is a DNA-binding residue based on its identity and the identities of its sequence neighbors. The input to the classifier consists of the identities of the target residue and 4 sequence neighbors on each side of the target residue. The classifier is trained and evaluated (using leave-one-out cross-validation) on a non-redundant set of 171 proteins. Our results indicate the feasibility of identifying interface residues based on local sequence information. The classifier achieves 71% overall accuracy with a correlation coefficient of 0.24, 35% specificity and 53% sensitivity in identifying interface residues as evaluated by leave-one-out cross-validation. We show that the performance of the classifier is improved by using sequence entropy of the target residue (the entropy of the corresponding column in multiple alignment obtained by aligning the target sequence with its sequence homologs) as additional input. The classifier achieves 78% overall accuracy with a correlation coefficient of 0.28, 44% specificity and 41% sensitivity in identifying interface residues. Examination of the predictions in the context of 3-dimensional structures of proteins demonstrates the effectiveness of this method in identifying DNA-binding sites from sequence information. In 33% (56 out of 171) of the proteins, the classifier identifies the interaction sites by correctly recognizing at least half of the interface residues. In 87% (149 out of 171) of the proteins, the classifier correctly identifies at least 20% of the interface residues. This suggests the possibility of using such classifiers to identify potential DNA-binding motifs and to gain potentially useful insights into sequence correlates of protein-DNA interactions. CONCLUSION: Naïve Bayes classifiers trained to identify DNA-binding residues using sequence information offer a computationally efficient approach to identifying putative DNA-binding sites in DNA-binding proteins and recognizing potential DNA-binding motifs.

Algorithms↗

The Continuity Trap in Data Science Health Research.

Secondary use is now the ordinary condition of data science health research rather than an exception to it. Electronic health records collected for clinical care become prediction tools and inputs for generative AI; imaging archives become foundation-model corpora; genomic datasets become resources for polygenic risk scores; and legacy biospecimens become renewable, indefinitely distributable cell lines. Governance has responded by emphasizing verifiable instruments such as provenance logs, repository approvals, broad-consent forms, data-use agreements, model cards, records of processing, and locality-preserving architectures. These instruments are necessary, and they answer real questions about lineage, privacy, institutional responsibility, and accountability, but they are not sufficient to establish that a present use remains ethically justified. We define ethical continuity as the persistence of normatively relevant relationships between the original conditions of data generation or material collection and subsequent downstream uses, such that current uses remain justifiable in light of the expectations, permissions, meanings, and relational obligations present at entrustment. We then define the Continuity Trap as a review-stage governance error in which a salient signal of continuity in one domain is treated as sufficient evidence of ethical continuity overall, causing inquiry into the remaining domains to close prematurely. The trap is not ordinary noncompliance, ethics creep, or a demand for universal rereview; it is a cross-domain inference error that can arise even in careful, good-faith review. We distinguish it from proxy closure, of which it is a continuity-specific subtype, and from Goodhart's and Campbell's laws, which describe how measures degrade once they become targets. We operationalize ethical continuity across 4 domains: provenance, semantics, authorization, and relational standing, developed in our Representational Veracity framework, and we show that these domains can diverge as data are linked, transformed, modeled, and redeployed. We identify the institutional mechanisms-provenance privilege, descriptor sedimentation, authorization fossilization, and community effacement-that cause auditable signals to be overread, and we examine how the US Health Insurance Portability and Accountability Act (HIPAA) of 1996, the General Data Protection Regulation, the European Health Data Space, US Food and Drug Administration guidance, the US National Institute of Standards and Technology (NIST) AI Risk Management Framework, and federated-learning governance can reduce risk while still inducing continuity traps. We apply the framework to consent and nonconsent settings, including public health, immunization, syndromic, and wastewater surveillance, polygenic risk scores, induced pluripotent stem cells, federated learning, and health-related large language models. The policy implication is trigger-based continuity review: rather than rereviewing every reuse, investigators and reviewers should identify the weakest continuity domain at the present data stage and impose a domain-matched safeguard, recorded in a short continuity statement. This reframing is intended for the committees, repositories, funders, and governance bodies that decide whether reuse may proceed, and it matters most in cross-border and low-resource settings. Provenance should begin ethical review; it should not end it.

Data Science↗

Transcriptome Analysis and Experimental Validation of Palmitoylation- Related Biomarkers in Atherosclerosis.

INTRODUCTION: Protein palmitoylation contributes to membrane localisation, signal transduction, and cell-fate regulation. It is closely associated with lipid metabolic dysfunction, immune inflammation, and vascular remodelling in atherosclerosis (AS). However, key palmitoylation-related transcriptomic markers and their potential causal associations with AS remain incompletely defined. METHODS: The Gene Expression Omnibus (GEO) dataset GSE100927 was used as the training cohort, and GSE43292 was used as an external validation cohort. Differentially expressed genes were identified using limma and intersected with palmitoylation-related genes to obtain palmitoylation-related differentially expressed genes (PRDEGs). Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses were then performed using clusterProfiler. Two-sample Mendelian randomisation was used to evaluate potential causal relationships between characteristic genes and AS. Feature selection was conducted using random forest and support vector machine recursive feature elimination (SVM-RFE), and the overlapping genes selected by both methods were retained. Receiver operating characteristic (ROC) curves were used to assess diagnostic performance. A five-gene nomogram was constructed, and its clinical utility was evaluated using calibration curves and decision curve analysis (DCA). Gene set variation analysis (GSVA) was applied to compare pathway activity between high- and low-expression groups for each core gene. Single-cell analysis using Seurat and expression-based cell-cell communication analysis using CellChat were conducted with GSE159677, and upstream transcription factors were predicted using NetworkAnalyst. For in vivo validation, an AS model was established in ApoE&#x2078;/&#x2078; mice fed a high-fat diet, and aortic gene and protein expression were assessed by RT-qPCR and western blotting. RESULTS: In GSE100927, 51 PRDEGs were identified. GO and KEGG enrichment analyses highlighted pathways associated with regulation of monoatomic ion transport, sarcomere and myofibril organisation, and immune inflammation. Mendelian randomisation suggested a potential protective causal association between SLC7A7 and AS. By integrating MR with random forest and SVM-RFE feature selection, we prioritised five core genes: PLCB2, GMIP, NEXN, PLN, and SLC7A7. These genes showed good diagnostic performance in GSE43292. The resulting nomogram was well calibrated and demonstrated stable net benefit in decision curve and clinical impact curve analyses. Single-gene GSVA identified consistently activated pathways across multiple genes, including innate and adaptive immune recognition, calcium signalling and myocardial contraction/cardiomyopathy, extracellular matrix-receptor interaction, cell junction pathways, autophagy-lysosome pathways, and several metabolic programmes. At the single-cell level, PLCB2 and GMIP were predominantly expressed in T cells and macrophages, NEXN and PLN were enriched in vascular smooth muscle cells, and SLC7A7 was mainly expressed in macrophages. CellChat analysis indicated increased signals for immune-related ligand-receptor interactions. In ApoE&#x2078;/&#x2078; mice fed a high-fat diet, PLCB2, GMIP, and SLC7A7 were upregulated, whereas NEXN and PLN were downregulated; protein-level changes were concordant with the transcriptomic trends. DISCUSSION: These findings indicate that palmitoylation-related dysregulation in AS converges on immune inflammation, calcium signalling/contractile programmes, ECM remodelling, and autophagy-linked metabolism. The five-gene panel is supported by external validation, single-cell localisation to immune and vascular compartments, and concordant results in ApoE&#x2078;/&#x2078; mice. CONCLUSION: This study identified and validated five palmitoylation-related genes associated with AS. SLC7A7 showed a potential protective causal signal in MR analysis. The enriched pathway patterns linked these genes to immune inflammation, calcium signalling-contraction coupling, ECM remodelling, cell adhesion, and autophagy- associated metabolic reprogramming. The five-gene nomogram showed potential utility for diagnostic classification and decision support, nominating candidate biomarkers and pathway targets for AS molecular subtyping, diagnosis, and mechanistic investigation.

Atherosclerosis (AS)↗