Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

Natural discriminant analysis using interactive Potts models.

Natural discriminant analysis based on interactive Potts models is developed in this work. A generative model composed of piece-wise multivariate gaussian distributions is used to characterize the input space, exploring the embedded clustering and mixing structures and developing proper internal representations of input parameters. The maximization of a log-likelihood function measuring the fitness of all input parameters to the generative model, and the minimization of a design cost summing up square errors between posterior outputs and desired outputs constitutes a mathematical framework for discriminant analysis. We apply a hybrid of the mean-field annealing and the gradient-descent methods to the optimization of this framework and obtain multiple sets of interactive dynamics, which realize coupled Potts models for discriminant analysis. The new learning process is a whole process of component analysis, clustering analysis, and labeling analysis. Its major improvement compared to the radial basis function and the support vector machine is described by using some artificial examples and a real-world application to breast cancer diagnosis.

Journal Article↗

Machine learning on multiple epigenetic features reveals H3K27Ac as a driver of gene expression prediction across patients with glioblastoma.

Epigenetic mechanisms play a crucial role in driving transcript expression and shaping the phenotypic plasticity of glioblastoma stem cells (GSCs), contributing to tumor heterogeneity and therapeutic resistance. These mechanisms dynamically regulate the expression of key oncogenic and stemness-associated genes, enabling GSCs to adapt to environmental cues and evade targeted therapies. Importantly, epigenetic reprogramming allows GSCs to transition between cellular states, including therapy-resistant mesenchymal-like phenotypes, underscoring the need for epigenetic-targeting strategies to disrupt these adaptive processes. Understanding these epigenetic drivers of gene expression provides a foundation for novel therapeutic interventions aimed at eradicating GSCs and improving glioblastoma outcomes. Using machine learning (ML), we employ cross-patient prediction of transcript expression in GSCs by combining epigenetic features from various sources, including ATAC-seq, CTCF ChIP-seq, RNAPII ChIP-seq, H3K27Ac ChIP-seq, and RNA-seq. We investigate different ML and deep learning (DL) models for this task and ultimately build our final pipeline using XGBoost. The model trained on one patient generalizes to other 11 patients with high performance. Notably, H3K27Ac alone from a single patient is sufficient to predict gene expression in all 11 patients. Furthermore, the distribution of H3K27Ac peaks across the genomes of all patients is remarkably similar. These findings suggest that GSCs share a common distributional pattern of enhancer activity characterized by H3K27Ac, which can be utilized to predict gene expression in GSCs across patients. In summary, while GSCs are known for their transcriptomic and phenotypic heterogeneity, we propose that they share a common epigenetic pattern of enhancer activation that defines their underlying transcriptomic expression pattern. This pattern can predict gene expression across patient samples, providing valuable insights into the biology of GSCs.

Glioblastoma↗

An empirical comparison of back propagation and the RDSE algorithm on continuously valued real world data.

The ability of a neural network to generalise is dependent on how representative the training patterns were of the whole data domain, and how smoothly the network has fitted to these patterns [Sethi, I.K. (1990). IEEE International Joint Conference on Neural Networks, Seattle, WA, Vol. 2, pp. 219-224]. In non-scaled continuous data domains, training examples will lie at differing distances from each other, making the fitting problem more difficult and varied. This paper introduces a new neuron with an adaptive steepness parameter, implemented as an extra internal connection, which is altered to better interpolate between the data points that its hyperplane divides. Networks of the new neuronal model are trained using a new paradigm entitled the random directed search by entropy algorithm (RDSE). This involves constructing a network by training one neuron at a time and freezing the weights. Each neuron is trained using directed random search [Baba (1989). Neural Networks, 2, 367-373] to find a hyperplane that separates examples by minimising an entropy measure [Quinlan (1986). Induction of Decision Trees, Machine Learning, Vol. 1, pp. 81-106]. This training paradigm solves the problem of pre-defining a network topology, has few problems with local minima, can handle unscaled continuous input data and can be fully trained in a relatively short time scale when compared with other methods, e.g. back propagation (BP).An example benchmark problem is used to illustrate the effects of the new neuronal model, and results for two real world data domains are given which display an improved classification rate when compared against networks with a constant steepness value for every neuron. An empirical comparison between BP and RDSE for the two data sets are also given. These results display improved training times, robustness and classification rates by RDSE when compared against BP.

Journal Article↗

Feedback error learning neural network for trans-femoral prosthesis.

Feedback-error learning (FEL) neural network was developed for control of a powered trans-femoral prosthesis. Nonlinearities and time-variations of the dynamics of the plant, in addition to redundancy and dynamic uncertainty during the double support phase of walking, makes conventional control methods very difficult to use. Rule-based control, which uses a knowledge base determined by machine learning and finite automata method is limited since it does not respond well to perturbations and environmental changes. FEL can be regarded as a hybrid control, because it combines nonparametric identification with parametric modeling and control. This paper presents simulation of a powered trans-femoral prosthesis controlled by a FEL neural network. Results suggest that FEL can be used to identify inverse dynamics of an arbitrary trans-femoral prosthesis during simple single joint movements (e.g., sinusoidal oscillations). The identified inverse dynamics then allows the tracking of an arbitrary trajectory such as a desired walking pattern within a multijoint structure. Simulation shows that the identified controller responds correctly when the leg motion is exposed to a perturbation such as a frequent change of the ground reaction force or the hip joint torque generated by the user. FEL eliminates the need for precise, tedious, and complex identification of model parameters.

Activities of Daily Living↗

Honest assessments of automatic learning algorithm performance.

OBJECTIVE: To compare methods of evaluating probabilistic predictors in systems that learn from examples. STUDY DESIGN: The performance of four automatic learning algorithms, representing current machine learning technology, were assessed using four methodologies in the task of separating normal squamous intermediate cervical cells from all other segmented objects in digital images. Two of the methodologies were carefully constructed to model sources of variation associated with the choice of training and test sets. These assessments were statistically compared with assessments using both standard and a modified version of cross-validation. RESULTS: The investigation illustrates the tradeoffs involved in obtaining statistical rigor as compared with the cost of collecting data. While cross-validation makes frugal use of data, it can produce misleading assessments of algorithm performance in terms of both bias and variance. The modified version produces more reliable assessments but in some cases may also be misleading. CONCLUSION: We suggest that users of learning algorithms should exercise judicious care in evaluating learning algorithm performance in order to avoid unnecessary bias and large variance in their assessments.

Algorithms↗

CORA--a knowledge-based system for the analysis of case-control studies.

Carrying out a statistical analysis, the researcher is concerned with the problem of choosing an appropriate statistical technique from a large number of competing methods. Most common statistical software offer different methods for analysing the data without giving any support regarding the adequacy of a method for a particular data set. This paper outlines the main features of the computer system CORA which provides a statistical analysis of stratified contingency tables and additionally supports the researcher at the different steps of this analysis. Here, the support given by the system consists of two different aspects. On the one hand, the help system of CORA contains general information on the implemented statistical methods which can be obtained on request. On the other hand, an advice tool recommends an adequate statistical method which depends on the actual empirical case-control data to be analysed. To build up the advice tool, a set of rules being discovered by machine learning from simulation studies is integrated into the system CORA.

Case-Control Studies↗

Genome-wide association, polygenic risk scores, and machine learning for chronic post-surgical pain risk stratification: A UK biobank study.

Chronic post-surgical pain is a prevalent and debilitating complication following surgery, representing a clinical challenge. Despite the established heritability of pain phenotypes, large-scale genetic studies remain limited. This study aimed to identify genetic variants associated with chronic post-surgical pain, develop polygenic risk scores, and integrate these with clinical features for risk prediction. UK Biobank data from 47,836 participants (2490 cases and 45,346 controls) were split into training (80%; n = 38,268) and validation (20%; n = 9568) sets prior to analysis. A genome-wide association study was conducted on the training set only, across 19 million variants, and polygenic risk scores were constructed and integrated with clinical features in a logistic regression framework. Two close, rare, imputed signals crossed the genome-wide significance threshold but lacked local linkage-disequilibrium support, while 220 variants crossed the suggestive threshold. In the held-out validation set, cases had higher mean polygenic risk scores than controls (0.138 vs. -0.021; Cohen's d = 0.16, p < 0.001). A logistic regression model integrating clinical features and polygenic risk scores achieved an area under the curve of 0.639 (95% CI: 0.583-0.693), higher than models using either feature set alone. The polygenic risk score for chronic post-surgical pain was among the most important predictors. Risk stratification revealed the top quartile had 3.84-fold higher odds of chronic post-surgical pain than the bottom quartile (95% CI: 2.00-7.37). These findings suggest a possible modest genetic contribution to chronic post-surgical pain. Polygenic risk scores may complement clinical factors in surgical risk stratification. PERSPECTIVE: Chronic post-surgical pain may have a modest genetic contribution. This UK Biobank study identified over 220 variants at suggestive significance and constructed a polygenic risk score that was significantly elevated in cases. A combined clinical-genomic model achieved a 3.84-fold difference in odds across predicted-risk quartiles.

Chronic post-surgical pain↗

Landscape of essential growth and fluconazole-resistance genes in the human fungal pathogen Cryptococcus neoformans.

Fungi can cause devastating invasive infections, typically in immunocompromised patients. Treatment is complicated both by the evolutionary similarity between humans and fungi and by the frequent emergence of drug resistance. Studies in fungal pathogens have long been slowed by a lack of high-throughput tools and community resources that are common in model organisms. Here we demonstrate a high-throughput transposon mutagenesis and sequencing (TN-seq) system in Cryptococcus neoformans that enables genome-wide determination of gene essentiality. We employed a random forest machine learning approach to classify the C. neoformans genome as essential or nonessential, predicting 1,465 essential genes, including 302 that lack human orthologs. These genes are ideal targets for new antifungal drug development. TN-seq also enables genome-wide measurement of the fitness contribution of genes to phenotypes of interest. As proof of principle, we demonstrate the genome-wide contribution of genes to growth in fluconazole, a clinically used antifungal. We show a novel role for the well-studied RIM101 pathway in fluconazole susceptibility. We also show that insertions of transposons into the 5' upstream region can drive sensitization of essential genes, enabling screenlike assays of both essential and nonessential components of the genome. Using this approach, we demonstrate a role for mitochondrial function in fluconazole sensitivity, such that tuning down many essential mitochondrial genes via 5' insertions can drive resistance to fluconazole. Our assay system will be valuable in future studies of C. neoformans, particularly in examining the consequences of genotypic diversity.

Cryptococcus neoformans↗

Machine learning approaches for cancer prognosis and diagnosis via non-coding RNA: a comprehensive review.

Non-coding RNAs (ncRNAs), once considered genomic dark matter, are now established as key regulators of gene expression with widespread roles in cellular homeostasis and disease. In cancer, ncRNA expression is frequently and systematically dysregulated, and many of these molecules circulate in stable, protected form within biofluids, offering a compelling basis for non-invasive or minimally invasive diagnostic strategies. However, their clinical translation remains substantially hindered to date due to biological complexity, technical noise, and high dimensionality inherent to ncRNA expression datasets. In this context, machine learning (ML) has emerged as a powerful analytical tool to address these challenges, enabling the identification of subtle, reproducible ncRNA signatures predictive of diverse malignancies. This review critically evaluates ML-driven frameworks for cancer diagnosis and prognosis across four ncRNA subclasses, namely miRNAs, lncRNAs, circRNAs, and piRNAs, while also acknowledging the biophysical and thermodynamic models that reinforce ncRNA bioinformatics. Despite substantial methodological progress in ML-based cancer diagnosis and prognosis, key challenges persist, including tumor biological heterogeneity, limited multicenter validation, and the lack of widely adopted standardized protocols for preprocessing, normalization, and reporting workflows. Furthermore, many current ML models lack interpretability in biological or clinical context, constraining their translational utility. By synthesizing recent advances and identifying unresolved barriers, this review charts a roadmap for developing a robust, clinically actionable ncRNA biomarker platform for cancer detection. With global cancer incidence projected to exceed 35 million annual cases by 2050, validated ncRNA-ML-driven frameworks hold potential to revolutionize early-stage detection and personalized therapeutic strategies, thereby reducing the escalating socio-economic burden of cancer worldwide.

Humans↗

N6-methyladenine identification using deep learning and discriminative feature integration.

N6-methyladenine (6&#xa0;mA) is a pivotal DNA modification that plays a crucial role in epigenetic regulation, gene expression, and various biological processes. With advancements in sequencing technologies and computational biology, there is an increasing focus on developing accurate methods for 6&#xa0;mA site identification to enhance early detection and understand its biological significance. Despite the rapid progress of machine learning in bioinformatics, accurately detecting 6&#xa0;mA sites remains a challenge due to the limited generalizability and efficiency of existing approaches. In this study, we present Deep-N6mA, a novel Deep Neural Network (DNN) model incorporating optimal hybrid features for precise 6&#xa0;mA site identification. The proposed framework captures complex patterns from DNA sequences through a comprehensive feature extraction process, leveraging k-mer, Dinucleotide-based Cross Covariance (DCC), Trinucleotide-based Auto Covariance (TAC), Pseudo Single Nucleotide Composition (PseSNC), Pseudo Dinucleotide Composition (PseDNC), and Pseudo Trinucleotide Composition (PseTNC). To optimize computational efficiency and eliminate irrelevant or noisy features, an unsupervised Principal Component Analysis (PCA) algorithm is employed, ensuring the selection of the most informative features. A multilayer DNN serves as the classification algorithm to identify N6-methyladenine sites accurately. The robustness and generalizability of Deep-N6mA were rigorously validated using fivefold cross-validation on two benchmark datasets. Experimental results reveal that Deep-N6mA achieves an average accuracy of 97.70% on the F. vesca dataset and 95.75% on the R. chinensis dataset, outperforming existing methods by 4.12% and 4.55%, respectively. These findings underscore the effectiveness of Deep-N6mA as a reliable tool for early 6&#xa0;mA site detection, contributing to epigenetic research and advancing the field of computational biology.

Deep Learning↗

Patient-specific modeling identifies metabolic interventions for reversing glucose use reprogramming in alcohol-associated hepatitis.

Alcoholic hepatitis (AH) is an acute form of alcohol-associated liver disease with very few treatment options. Recent studies highlighted liver metabolic reprogramming in AH as an indicator of severity. We aim at identifying new intervention points to reverse liver metabolic dysregulation across varying degrees of AH. We develop 89 personalized genome-scale metabolic models by integrating a generic human cellular metabolic model with liver transcriptomics data from AH patients with varying disease severity and healthy controls. We grade the AH patients based on the model-predicted level of glycolysis reprogramming and validate the results using published metabolomics data. We test in silico gene knockdown interventions to reverse the aberrant metabolic reprogramming in AH. Knockdown of two glycolytic genes, Hkdc1 and Pkm, significantly rebalance the metabolic fluxes toward a healthy liver metabolic phenotype. We use machine learning on the glycolysis fluxes to develop a quantitative glucose use reprogramming score, which correlates with AH severity and patient-specific responses to in silico gene knockdown interventions. The score was independently validated using a published AH liver transcriptomics dataset. We propose a cellular metabolism-based therapy targeting Hkdc1 and Pkm in the glycolysis pathway as a potential treatment for reversing the aberrant glucose metabolism in AH.

Humans↗

Analysis of molecular profile data using generative and discriminative methods.

A modular framework is proposed for modeling and understanding the relationships between molecular profile data and other domain knowledge using a combination of generative (here, graphical models) and discriminative [Support Vector Machines (SVMs)] methods. As illustration, naive Bayes models, simple graphical models, and SVMs were applied to published transcription profile data for 1,988 genes in 62 colon adenocarcinoma tissue specimens labeled as tumor or nontumor. These unsupervised and supervised learning methods identified three classes or subtypes of specimens, assigned tumor or nontumor labels to new specimens and detected six potentially mislabeled specimens. The probability parameters of the three classes were utilized to develop a novel gene relevance, ranking, and selection method. SVMs trained to discriminate nontumor from tumor specimens using only the 50-200 top-ranked genes had the same or better generalization performance than the full repertoire of 1,988 genes. Approximately 90 marker genes were pinpointed for use in understanding the basic biology of colon adenocarcinoma, defining targets for therapeutic intervention and developing diagnostic tools. These potential markers highlight the importance of tissue biology in the etiology of cancer. Comparative analysis of molecular profile data is proposed as a mechanism for predicting the physiological function of genes in instances when comparative sequence analysis proves uninformative, such as with human and yeast translationally controlled tumour protein. Graphical models and SVMs hold promise as the foundations for developing decision support systems for diagnosis, prognosis, and monitoring as well as inferring biological networks.

Bayes Theorem↗

Benchmark of biomarker identification and prognostic modeling methods on diverse censored data.

The practices of identifying biomarkers and developing prognostic models using genomic data has become increasingly prevalent. Such data often features characteristics that make these practices difficult, namely high dimensionality, correlations between predictors, and sparsity. Many modern methods have been developed to address these problematic characteristics while performing feature selection and prognostic modeling, but a large-scale comparison of their performances in these tasks on diverse right-censored time to event data (aka survival time data) is much needed. We have compiled many existing methods, including some machine learning methods, several which have performed well in previous benchmarks, primarily for comparison in regards to variable selection capability, and secondarily for survival time prediction on many synthetic datasets with varying levels of sparsity, correlation between predictors, and signal strength of informative predictors. For illustration, we have also performed multiple analyses on a publicly available and widely used cancer cohort from The Cancer Genome Atlas using these methods. We evaluated the methods through extensive simulation studies in terms of the false discovery rate, F1-score, concordance index, Brier score, root mean square error, and computation time. Of the methods compared, CoxBoost and the Adaptive LASSO performed well in all metrics, and the LASSO and elastic net excelled when evaluating concordance index and F1-score. The Benjamini-Hoschberg and q-value procedures showed volatile performances in controlling the false discovery rate. Some methods' performances were greatly affected by differences in the data characteristics. With our extensive numerical study, we have identified the best performing methods for a plethora of data characteristics using informative metrics. This will help cancer researchers in choosing the best approach for their needs when working with genomic data.

Humans↗

GiantHunter: accurate detection of giant virus in metagenomic data using reinforcement-learning and Monte Carlo tree search.

MOTIVATION: Nucleocytoplasmic large DNA viruses (NCLDVs) are notable for their large genomes and extensive gene repertoires, which contribute to their widespread environmental presence and critical roles in processes such as host metabolic reprogramming and nutrient cycling. Metagenomic sequencing has emerged as a powerful tool for uncovering novel NCLDVs in environmental samples. However, identifying NCLDV sequences in metagenomic data remains challenging due to their high genomic diversity, limited reference genomes, and shared regions with other microbes. Existing alignment-based and machine learning methods struggle with achieving optimal trade-offs between sensitivity and precision. RESULTS: In this work, we present GiantHunter, a reinforcement learning-based tool for identifying NCLDVs from metagenomic data. By employing a Monte Carlo tree search strategy, GiantHunter dynamically selects representative non-NCLDV sequences as the negative training data, enabling the model to establish a robust decision boundary. Benchmarking on rigorously designed experiments shows that GiantHunter achieves high precision while maintaining competitive sensitivity, improving the F1-score by 10% and reducing computational cost by 90% compared to the second-best method. To demonstrate its real-world utility, we applied GiantHunter to 60 metagenomic datasets collected from six cities along the Yangtze River, located both upstream and downstream of the Three Gorges Dam. The results reveal significant differences in NCLDV diversity correlated with proximity to the dam, likely influenced by reduced flow velocity caused by the dam. These findings highlight GiantHunter's potential to advance our understanding of NCLDVs and their ecological roles in diverse environments. AVAILABILITY AND IMPLEMENTATION: The source code of GiantHunter is available via: https://github.com/FuchuanQu/GiantHunter.

Metagenomics↗

Machine learning-assisted Mn-N-C nanozyme colorimetric sensor array for trace-level detection of biogenic amines in meat.

Accurate detection of biogenic amines (BAs) in meat remains challenging due to their high structural similarity and co-occurrence. Herein, an Mn-N-C nanozyme was synthesized via a metal-organic framework confined pyrolysis strategy, possessing excellent oxidase (OXD)- and peroxidase (POD)-like activities. The dual enzyme-like activity showed Km values of 0.1584&#xa0;mM (OXD) and 0.1498&#xa0;mM (POD), respectively, in detection system. Leveraging these properties, a colorimetric sensor array was constructed, enabling the detection of four representative BAs within a concentration range of 2-10&#xa0;ppm with 100% classification accuracy. In addition, a concentration independent recognition model based on an artificial neural network was developed to address signal nonlinearity interference in meat. The integrated system achieved accurate trace-level identification of BAs in perishable fish, pork, and chicken, demonstrating its applicability for early-stage BAs monitoring and quality deterioration warning during storage and transportation.

Biogenic Amines↗

Anatomic models and phantoms for diagnostic ultrasound instruction.

The preparation and application of anatomic models and phantoms to facilitate learning diagnostic ultrasound is described. Imaging with diagnostic ultrasound requires mastery of many skills, along with knowledge of sound-tissue interactions which contribute to the formation of diagnostic images and artifacts. Understanding the genesis of artifacts encountered during ultrasound scanning can avoid misinterpretation and aid diagnosis. In addition, development of machine related knowledge and skills, including manipulation of the transducer and the selection of correct settings for variables such as gain, power, time-gain compensation, and transducer type, is dependent on an understanding of how these factors affect the image. The normal appearance of an organ relates to both its echogenicity and morphologic characteristics, and confirmation of the nature of an abnormality often requires ultrasound guided biopsy. The use of anatomic models and phantoms in ultrasound instruction allows principles to be demonstrated, knowledge acquired, and biopsy procedures practiced and mastered in a controlled setting. This can minimize live animal use, and enhance the knowledge base and skills of the clinician prior to applying this diagnostic technique to the clinical patient.

Animals↗

Normativity in 18th century discourse on speech.

Eighteenth century phoneticians, such as Dodart, Ferrein, and Hellwag, extended the taxonomy of visible articulatory processes into the realm of the invisible, notably with the exploration of the voicing mechanism. Remedial initiatives were not simply confined to consideration of the outward manifestations of speech and its disorders: The work of Haller, Kuestner, and Morgagni shows an acute awareness of the nervous organization underlying verbal behavior. There was a characteristic preoccupation with mechanical models of speech, which led to the attempts of Kempelen and other investigators to construct actual "speaking machines." Eighteenth century scholars regarded language as not only an innate capacity peculiar to human nature, but also as a bodily habit learned by experience. The function of the orthoepist was to teach the right speech habits, and the upward mobility of the bourgeoisie created a demand for his services.

Europe↗

Conserved HSFA1-dependent chromatin dynamics drive heat stress responses in plants.

Eukaryotic organisms remodel chromatin landscapes to regulate gene expression in response to environmental stress. In plants, heat stress (HS) induces widespread chromatin changes, yet the role of heat shock transcription factors (HSFs) in chromatin remodeling and their evolutionary conservation remains unclear. Using Marchantia polymorpha Mphsf mutants and Arabidopsis thaliana Athsfa1s mutants, we identify HSFA1 as a key regulator of HS-induced cis-regulatory element (CRE) accessibility, a mechanism conserved across land plants, mice, and humans. Gene regulatory network modeling reveals parallel transcription factor subnetworks, with MpWRKY10 and MpABI5B acting as indirect and negative HS regulators. We further showed that ABA modulates gene expression in an HSFA1-dependent manner without inducing chromatin remodeling. Finally, we develop a machine learning framework integrating chromatin accessibility and CRE information to predict gene expression across species, revealing stress-responsive regulatory logic at the transcriptional level. These findings provide insights into how TFs coordinate chromatin architecture to drive stress adaptation.

Heat-Shock Response↗