Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

Identification and Classification of Expressed Orphan Genes, Spurious Orphan Genes, and Conserved Genes in the Human Gut Microbiome.

Orphan genes (OGs)-genes lacking detectable homologs outside a species-are widespread in microbial genomes and are thought to contribute to their adaptation and molecular innovation. However, not all predicted OGs may represent novel functional coding sequences. False positive OGs, also called spurious OGs, can arise from gene prediction errors. We reason that OGs lacking detectable expression are more likely to be spurious. To test this, we combined large-scale metatranscriptomic profiling of the human gut microbiome with machine learning to distinguish expressed OGs from spurious ones and compare them with conserved genes (CGs) found in multiple species. Using nearly 5,000 metatranscriptome libraries, we identified ∼218,000 OGs supported by expression evidence, while ∼330,000 predicted OGs lacked detectable expression and were classified as spurious. We extracted 154 features for sequence, structural, and evolutionary properties for each gene and trained XGBoost classifiers while accounting for genomic representation. The models achieved an area under the receiver operating characteristic curve (AUC) of 0.82 in distinguishing expressed OGs from spurious OGs and an AUC of 0.93 in distinguishing expressed OGs from CGs. Interpretation based on SHAP (SHapley Additive exPlanations) revealed clear biological signals. Particularly, expressed orphans were present in more genomes than spurious ones, and expressed OGs were shorter than CGs. This work improves OG discovery and suggests that expressed OGs differ systematically from CGs and spurious OGs in sequence composition, structural constraints, and evolutionary signals.

Humans↗

CCNA2 orchestrates the PI3K/AKT signaling axis to propel prostate cancer metastasis.

BACKGROUND: Prostate cancer (PCa) remains one of the most common malignancies in men, posing a persistent global burden in terms of both public health and socioeconomic costs. Although early detection is essential for improving patient outcomes, existing clinical tools, including prostate-specific antigen (PSA) screening, digital rectal examination, and transrectal ultrasound-guided biopsy, are hampered by suboptimal specificity and positive predictive value, resulting in frequent overdiagnosis and overtreatment of indolent lesions while missing a subset of aggressive tumors at an early stage. In this context, the rapid advancement of high-throughput omics technologies, coupled with sophisticated machine learning (ML) algorithms, provides a powerful computational framework to dissect high-dimensional genomic data, uncover latent gene expression signatures, and identify candidate biomarkers with superior discriminative performance over conventional clinicopathological parameters. Therefore, in this study, we sought to screen for crucial ML-based biomarkers associated with PCa, with a particular focus on systematically assessing the diagnostic and prognostic value of CCNA2. Leveraging large-scale transcriptomic cohorts from public repositories, we employed an ensemble of ML approaches to prioritize candidate genes and subsequently evaluated the diagnostic performance of CCNA2 through receiver operating characteristic curve analysis, as well as its prognostic utility via Kaplan-Meier survival estimation and multivariate Cox proportional hazards modeling. Our findings are anticipated to elucidate the molecular landscape of PCa and offer a promising biomarker candidate for early detection and risk stratification. METHODS: This study integrated single-cell RNA sequencing, bulk transcriptomic data from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) repositories, immunofluorescence, and multiple ML algorithms with in vitro functional assays to evaluate CCNA2 expression, clinical relevance, and biological behavior in PCa. RESULTS: CCNA2 was linked to metastasis and poor prognosis. High CCNA2 expression significantly correlated with adverse survival outcomes, and knockdown of CCNA2 suppressed proliferation, migration, and invasion in PCa cell lines. Mechanistically, CCNA2 modulated the PI3K/AKT signaling pathway. An ML-based diagnostic model incorporating CCNA2 demonstrated high predictive accuracy across multiple validation cohorts. CONCLUSIONS: CCNA2 serves as a promising prognostic biomarker and therapeutic target in prostate adenocarcinoma, driving tumor progression potentially via the PI3K/AKT axis.

CCNA2↗

Causal associations between hormone replacement therapy and brain structure: Evidence from large-scale Mendelian randomization and double machine learning.

BACKGROUND: Hormone replacement therapy (HRT) is widely prescribed for the management of hormone deficiency, particularly during menopause, yet its causal effects on human brain structure remain incompletely understood. Observational studies have reported heterogeneous associations, underscoring the need for robust causal inference. METHODS: We applied an integrated causal framework combining two-sample Mendelian Randomization (MR) and Double Machine Learning (DML) to evaluate the effects of four HRT-related exposures-age at initiation, age at cessation, ever-use of HRT, and a composite medication-based phenotype-on 1366 brain imaging-derived phenotypes from the UK Biobank. Genetic instruments were derived from large-scale GWAS summary statistics, and causal estimates were validated using non-parametric DML models with cross-fitting and performance evaluation. RESULTS: Genetic instruments for age at HRT initiation, age at cessation, and ever-use of HRT were strong (median F-statistics 16.29-36.66). MR analyses identified a causal association between later initiation of HRT and lower orientation dispersion in the right inferior cerebellar peduncle (ubm-a-542; primary finding, no pleiotropy detected). An additional association with the left tapetum FA (ubm-a-243) was identified but exhibited significant directional horizontal pleiotropy (MR-Egger intercept P = 0.001) and is excluded from primary conclusions (Supplementary Note S2). Later cessation of HRT was associated with increased cortical thickness in the left middle occipital gyrus, reduced surface area in the left frontopolar cortex, and increased orientation dispersion in the splenium of the corpus callosum. Ever-use of HRT was causally linked to larger volumes of the right inferior frontal gyrus and right nucleus accumbens. These associations were corroborated by independent DML validation, which provided causally debiased estimates robust to high-dimensional confounding. Results for ukb-b-8080 (median F = 1.45) are provided in Supplementary Note S1 only; weak-instrument bias precludes causal inference. CONCLUSIONS: This study provides genetic-instrument-based and machine-learning-validated evidence for causal associations between HRT exposure-particularly its timing and lifetime use-and specific features of human brain structure, including white-matter microarchitecture, cortical thickness, and regional brain volume. These findings are FDR-controlled within exposures and independently replicated by DML, but require replication in external neuroimaging GWAS cohorts to establish definitive causal conclusions. They highlight the neurobiological relevance of sex steroid exposure and inform future research on brain aging and personalized hormone-based interventions.

Humans↗

A data-mining approach to spacer oligonucleotide typing of Mycobacterium tuberculosis.

MOTIVATION: The Direct Repeat (DR) locus of Mycobacterium tuberculosis is a suitable model to study (i) molecular epidemiology and (ii) the evolutionary genetics of tuberculosis. This is achieved by a DNA analysis technique (genotyping), called sp acer oligo nucleotide typing (spoligotyping ). In this paper, we investigated data analysis methods to discover intelligible knowledge rules from spoligotyping, that has not yet been applied on such representation. This processing was achieved by applying the C4.5 induction algorithm and knowledge rules were produced. Finally, a Prototype Selection (PS) procedure was applied to eliminate noisy data. This both simplified decision rules, as well as the number of spacers to be tested to solve classification tasks. In the second part of this paper, the contribution of 25 new additional spacers and the knowledge rules inferred were studied from a machine learning point of view. From a statistical point of view, the correlations between spacers were analyzed and suggested that both negative and positive ones may be related to potential structural constraints within the DR locus that may shape its evolution directly or indirectly. RESULTS: By generating knowledge rules induced from decision trees, it was shown that not only the expert knowledge may be modeled but also improved and simplified to solve automatic classification tasks on unknown patterns. A practical consequence of this study may be a simplification of the spoligotyping technique, resulting in a reduction of the experimental constraints and an increase in the number of samples processed.

Algorithms↗

EDAmame: interactive exploratory data analyses with explainable models.

SUMMARY: Complex tabular datasets comprising many diverse features can require specific expertise to interpret, posing a barrier to researchers with minimal data science experience. EDAmame is an interactive tool that simplifies initial analysis and visualization of these datasets, providing insights into data quality and feature relationships. By leveraging open-source machine learning frameworks in R, EDAmame allows researchers to perform effective exploratory data analysis without command-line or coding requirements. AVAILABILITY AND IMPLEMENTATION: A limited online version can be accessed at https://edamame.org.au/ or can be downloaded from https://doi.org/10.5281/zenodo.15356492. The app is developed in R Shiny and implements tidyverse and tidymodels packages.

Machine Learning↗

Improving the efficiency of a user-driven learning system with reconfigurable hardware. Application to DNA splicing.

This paper describes a new approach to problem solving by splitting up problem component parts between software and hardware. Our main idea arises from the combination of two previously published works. The first one proposed a conceptual environment of concept modelling in which the machine and the human expert interact. The second one reported an algorithm based on reconfigurable hardware system which outperforms any kind of previously published genetic data base scanning hardware or algorithms. Here we show how efficient the interaction between the machine and the expert is when the concept modelling is based on reconfigurable hardware system. Their cooperation is thus achieved with an real time interaction speed. The designed system has been partially applied to the recognition of primate splice junctions sites in genetic sequences.

Algorithms↗

Statistical mechanics of learning with soft margin classifiers.

We study the typical learning properties of the recently introduced soft margin classifiers (SMCs), learning realizable and unrealizable tasks, with the tools of statistical mechanics. We derive analytically the behavior of the learning curves in the regime of very large training sets. We obtain exponential and power laws for the decay of the generalization error towards the asymptotic value, depending on the task and on general characteristics of the distribution of stabilities of the patterns to be learned. The optimal learning curves of the SMCs, which give the minimal generalization error, are obtained by tuning the coefficient controlling the trade-off between the error and the regularization terms in the cost function. If the task is realizable by the SMC, the optimal performance is better than that of a hard margin support vector machine and is very close to that of a Bayesian classifier.

Algorithms↗

Potential assessment of the "support vector machine" method in forecasting ambient air pollutant trends.

Monitoring and forecasting of air quality parameters are popular and important topics of atmospheric and environmental research today due to the health impact caused by exposing to air pollutants existing in urban air. The accurate models for air pollutant prediction are needed because such models would allow forecasting and diagnosing potential compliance or non-compliance in both short- and long-term aspects. Artificial neural networks (ANN) are regarded as reliable and cost-effective method to achieve such tasks and have produced some promising results to date. Although ANN has addressed more attentions to environmental researchers, its inherent drawbacks, e.g., local minima, over-fitting training, poor generalization performance, determination of the appropriate network architecture, etc., impede the practical application of ANN. Support vector machine (SVM), a novel type of learning machine based on statistical learning theory, can be used for regression and time series prediction and have been reported to perform well by some promising results. The work presented in this paper aims to examine the feasibility of applying SVM to predict air pollutant levels in advancing time series based on the monitored air pollutant database in Hong Kong downtown area. At the same time, the functional characteristics of SVM are investigated in the study. The experimental comparisons between the SVM model and the classical radial basis function (RBF) network demonstrate that the SVM is superior to the conventional RBF network in predicting air quality parameters with different time series and of better generalization performance than the RBF model.

Air Pollutants↗

Machine learning in quantitative histopathology.

The role of expert systems functioning as process controllers in learning image understanding systems is discussed. Numeric learning systems already have found a number of applications in cytologic and histopathologic diagnosis. Depending on the required capabilities, systems of increasing complexity are needed. Expert systems to guide scene segmentation in histopathologic imagery require model-based reasoning. Diagnostic image interpretation with learning capability demands a full model of the human expert's competence, including a considerable variety of knowledge representation schemes and inference strategies, coordinated by a meta-process controller.

Artificial Intelligence↗

Multi-cohort integration and machine learning identify CPVL as a novel oncogenic driver in gastric cancer.

BACKGROUND: Gastric cancer (GC) remains a leading cause of cancer-related mortality worldwide, and the prognosis of advanced GC remains poor. Systematic identification of robust biomarkers through multi-cohort integration and computational prioritization may facilitate the discovery of novel therapeutic targets. AIM: To identify key genes associated with gastric cancer progression through integrative multi-omics analysis and to elucidate the biological functions and molecular mechanisms of the top-prioritized candidate gene. METHODS: Comprehensive bioinformatics analyses integrating The Cancer Genome Atlas (TCGA), Genotype-Tissue Expression (GTEx), and Gene Expression Omnibus (GEO) datasets were performed using differential expression analysis, weighted gene co-expression network analysis (WGCNA), Cox regression, and eight machine-learning algorithms to systematically identify and prioritize GC-associated hub genes. Among the identified candidates, CPVL was selected for further validation based on its diagnostic and prognostic performance. CPVL expression and clinical relevance were validated by independent datasets and immunohistochemistry. Lentiviral constructs were used to overexpress or silence CPVL in GC cell lines. Functional assays were performed, including CCK-8, colony formation, EdU incorporation, and flow cytometry, to assess cell proliferation and cell-cycle distribution. Western blotting and JAK2 inhibitor (AZD1480) rescue experiments were performed to elucidate the underlying mechanisms, and a nude mouse xenograft model was used to evaluate tumorigenicity in vivo. RESULTS: Multi-cohort screening identified five hub genes (CPVL, AADAC, BCAT1, CPXM1, and FBN1). Among them, CPVL exhibited the highest diagnostic accuracy (AUC = 0.895) and the strongest correlation with poor overall survival, and was therefore selected for mechanistic investigation. CPVL expression was markedly upregulated in GC tissues and cell lines. Functional assays demonstrated that CPVL promotes GC cell proliferation and accelerates G1/S-phase transition. Mechanistically, CPVL activated the JAK2/STAT3 signaling pathway, upregulating Cyclin D1 and CDK4 while downregulating p27. Treatment with the JAK2 inhibitor AZD1480 partially reversed these effects. In vivo, CPVL knockdown significantly inhibited tumor growth. CONCLUSION: Through systematic multi-cohort integration and machine-learning prioritization, CPVL was identified as a novel oncogenic driver in gastric cancer. CPVL promotes tumor growth via activation of the JAK2/STAT3 pathway and regulation of the Cyclin D1/CDK4/p27 axis, highlighting its potential as a diagnostic biomarker and therapeutic target.

Biomarker↗

Generative topographic mapping applied to clustering and visualization of motor unit action potentials.

The identification and visualization of clusters formed by motor unit action potentials (MUAPs) is an essential step in investigations seeking to explain the control of the neuromuscular system. This work introduces the generative topographic mapping (GTM), a novel machine learning tool, for clustering of MUAPs, and also it extends the GTM technique to provide a way of visualizing MUAPs. The performance of GTM was compared to that of three other clustering methods: the self-organizing map (SOM), a Gaussian mixture model (GMM), and the neural-gas network (NGN). The results, based on the study of experimental MUAPs, showed that the rate of success of both GTM and SOM outperformed that of GMM and NGN, and also that GTM may in practice be used as a principled alternative to the SOM in the study of MUAPs. A visualization tool, which we called GTM grid, was devised for visualization of MUAPs lying in a high-dimensional space. The visualization provided by the GTM grid was compared to that obtained from principal component analysis (PCA).

Action Potentials↗

A time-resolved single-cell roadmap of the logic driving anterior neural crest diversification from neural border to migration stages.

Neural crest cells exemplify cellular diversification from a multipotent progenitor population. However, the full sequence of early molecular choices orchestrating the emergence of neural crest heterogeneity from the embryonic ectoderm remains elusive. Gene-regulatory-networks (GRN) govern early development and cell specification toward definitive neural crest. Here, we combine ultradense single-cell transcriptomes with machine-learning and large-scale transcriptomic and epigenomic experimental validation of selected trajectories, to provide the general principles and highlight specific features of the GRN underlying neural crest fate diversification from induction to early migration stages using Xenopus frog embryos as a model. During gastrulation, a transient neural border zone state precedes the choice between neural crest and placodes which includes multiple converging gene programs. During neurulation, transcription factor connectome, and bifurcation analyses demonstrate the early emergence of neural crest fates at the neural plate stage, alongside an unbiased multipotent-like lineage persisting until epithelial-mesenchymal transition stage. We also decipher circuits driving cranial and vagal neural crest formation and provide a broadly applicable high-throughput validation strategy for investigating single-cell transcriptomes in vertebrate GRNs in development, evolution, and disease.

Animals↗

GANN: genetic algorithm neural networks for the detection of conserved combinations of features in DNA.

BACKGROUND: The multitude of motif detection algorithms developed to date have largely focused on the detection of patterns in primary sequence. Since sequence-dependent DNA structure and flexibility may also play a role in protein-DNA interactions, the simultaneous exploration of sequence- and structure-based hypotheses about the composition of binding sites and the ordering of features in a regulatory region should be considered as well. The consideration of structural features requires the development of new detection tools that can deal with data types other than primary sequence. RESULTS: GANN (available at http://bioinformatics.org.au/gann) is a machine learning tool for the detection of conserved features in DNA. The software suite contains programs to extract different regions of genomic DNA from flat files and convert these sequences to indices that reflect sequence and structural composition or the presence of specific protein binding sites. The machine learning component allows the classification of different types of sequences based on subsamples of these indices, and can identify the best combinations of indices and machine learning architecture for sequence discrimination. Another key feature of GANN is the replicated splitting of data into training and test sets, and the implementation of negative controls. In validation experiments, GANN successfully merged important sequence and structural features to yield good predictive models for synthetic and real regulatory regions. CONCLUSION: GANN is a flexible tool that can search through large sets of sequence and structural feature combinations to identify those that best characterize a set of sequences.

Algorithms↗

Using rough sets, neural networks, and logistic regression to predict compliance with cholesterol guidelines goals in patients with coronary artery disease.

Coronary artery disease is a leading cause of death and disability in the United States and throughout the developed world. Results from large randomized, blinded, placebo-controlled trials have demonstrated clearly the benefit of lowering LDL cholesterol in lowering the risk for coronary artery disease. Unfortunately, despite the quantity of evidence, and the availability of medications that can efficiently lower LDL cholesterol with few side effects, not everyone who could benefit from cholesterol lowering interventions actually receives them. Despite the dissemination of national care guidelines for the evaluation and treatment of cholesterol levels (NCEP - National Cholesterol Education Program), compliance with such guidelines is suboptimal. There clearly is room for improvement in narrowing the gap between evidence based guidelines and actual clinical practice. The ability to classify those patients who are or will likely to be noncompliant on the basis of patient data routinely collected during patient care could be potentially useful by enabling the focusing of limited health care resources to those who are or will be at high risk of being under treated. In order to explore this possibility further, we attempted to create such classifiers of cholesterol guideline compliance. To do this, we obtained data from an ambulatory electronic medical record system at use at the MGH adult primary care practices for over 20 years. We obtained the data from this hierarchically-structured EMR using its own native query language, called MQL (Medical Query Language). Next, we applied to the collected data the machine learning techniques of rough set theory, neural networks (feed forward backpropagation nets), and logistic regression. We did this by using commonly available software that for the most part is freely available via the internet. We then compared the accuracy of the classifier models using the receiver operating characteristic (ROC) area and C-index summary metrics.

Cholesterol↗

A model for setting up interdisciplinary collaborative working in groups: lessons from an experience of action learning.

There is a current policy emphasis within health services on collaborative interdisciplinary groupworking between professionals, exemplified by the increasing use of action learning sets in a health context. Most of the evaluative research into this type of group collaboration has concentrated on evaluating outcomes, whereas this project aimed to evaluate qualitatively the experiences of the professional group members. The research uses a grounded theory methodology to investigate their perceptions and to analyse the data collected through interview methods. The research addresses an emergent theoretical model that could be of use when planning multidisciplinary task groups. It aims to enhance the success of these groups using a theory based on the concept of project momentum.

Cooperative Behavior↗

Support-vector-machine classification of linear functional motifs in proteins.

Our algorithm predicts short linear functional motifs in proteins using only sequence information. Statistical models for short linear functional motifs in proteins are built using the database of short sequence fragments taken from proteins in the current release of the Swiss-Prot database. Those segments are confirmed by experiments to have single-residue post-translational modification. The sensitivities of the classification for various types of short linear motifs are in the range of 70%. The query protein sequence is dissected into short overlapping fragments. All segments are represented as vectors. Each vector is then classified by a machine learning algorithm (Support Vector Machine) as potentially modifiable or not. The resulting list of plausible post-translational sites in the query protein is returned to the user. We also present a study of the human protein kinase C family as a biological application of our method.

Databases, Genetic↗

The automatic discovery of structural principles describing protein fold space.

The study of protein structure has been driven largely by the careful inspection of experimental data by human experts. However, the rapid determination of protein structures from structural-genomics projects will make it increasingly difficult to analyse (and determine the principles responsible for) the distribution of proteins in fold space by inspection alone. Here, we demonstrate a machine-learning strategy that automatically determines the structural principles describing 45 folds. The rules learnt were shown to be both statistically significant and meaningful to protein experts. With the increasing emphasis on high-throughput experimental initiatives, machine-learning and other automated methods of analysis will become increasingly important for many biological problems.

Algorithms↗

A case-acquisition and decision-support system for the analysis of group-average lactation curves.

A case-acquisition and decision-support system was developed to support the analysis of group-average lactation curves and to acquire example cases from domain specialists. This software was developed through several iterations of a three-step approach involving 1) problem analysis and formulation in consultation with two dairy nutrition specialists; 2) development of a case-acquisition and decision-support prototype by the system developer; and 3) use of the prototype by the domain specialists to analyze and classify milk-recording data from example herds. The overall problem was decomposed into three subproblems: removal of outlier tests and lactation curves of individual cows; interpretation of group-average lactation curves; and diagnosis of detected abnormalities at the herd level through the identification of potential management deficiencies. For each subproblem, a software module was developed allowing the user to analyze both graphical and numerical performance representations and classify these representations using predefined linguistic descriptors. The example-based method for the development of the program proved to be very useful, facilitating the communication between system developer and domain specialists, and allowing the specialists to explore the appropriateness of the various prototypes developed. The resulting software represents a formalization of the approach to group-average lactation curve analysis, elicited from the two domain specialists. In future research, the case-acquisition and decision-support system will be complemented with knowledge to automate identified classification tasks, which will be captured through the application of machine-learning techniques to example cases, acquired from domain specialists using the software.

Animals↗