Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

Medical diagnostic system using Fuzzy Coloured Petri Nets under uncertainty.

We propose a medical diagnostic system using Fuzzy Coloured Petri Nets (FCPN) in this paper. For complex real-world knowledge Fuzzy Petri Net (FPN) models have been proposed to perform fuzzy reasoning automatically. However, in the Petri Net we have to represent all kinds of processes by separate subnets even though the process has the same behavior of other one. Real-world knowledge often contains many parts which are similar, but not identical. This means that the total PTN becomes very large. The kind of problems may be annoying for a small system, and it may be catastrophic for the description of large-scale system. To avoid this kind of problems we propose a learning and reasoning method using FCPNs under uncertainty. On the other hand to correct the rules of knowledge-based system hand-built classifier and empirical learning method both based on domain theory have been proposed as machine learning methods, where there is a significant gap between the knowledge-intensive approach in the former and the virtually knowledge-free approach in the later. To resolve such problems simultaneously we propose a hybrid learning method which is built on the top of knowledge-based FCPN and Genetic Algorithms (GA). To verify the validity and the effectiveness of the proposed system, we have successfully applied it to the diagnosis of intervertebral diseases.

Algorithms↗

CAKR: commutative algebra k-mer representations for genomics.

Despite the availability of various sequence analysis models, comparative genomic analysis remains a challenge in genomics, genetics, and phylogenetics. Commutative algebra, a fundamental tool in algebraic geometry and number theory, has rarely been used in data and biological sciences. In this study, we introduce commutative algebra k-mer representations as a nonlinear algebraic framework for analyzing genomic sequences. This representation bridges commutative algebra, algebraic topology, combinatorics, and machine learning to establish a mathematical framework for comparative genomic analysis. We evaluate its effectiveness on three tasks including genetic variant classification, phylogenetic tree reconstruction, and viral classification, typically requiring alignment-based, alignment-free, and machine-learning approaches, respectively. In this work, we show that commutative algebra k-mer representations outperform five state-of-the-art sequence analysis methods across twelve primary datasets, with two additional supplementary fragment-placement benchmarks, especially in viral classification, and maintain relatively stable predictive accuracy as dataset size increases, underscoring scalability and robustness.

Genomics↗

Warmr: a data mining tool for chemical data.

Data mining techniques are becoming increasingly important in chemistry as databases become too large to examine manually. Data mining methods from the field of Inductive Logic Programming (ILP) have potential advantages for structural chemical data. In this paper we present Warmr, the first ILP data mining algorithm to be applied to chemoinformatic data. We illustrate the value of Warmr by applying it to a well studied database of chemical compounds tested for carcinogenicity in rodents. Data mining was used to find all frequent substructures in the database, and knowledge of these frequent substructures is shown to add value to the database. One use of the frequent substructures was to convert them into probabilistic prediction rules relating compound description to carcinogenesis. These rules were found to be accurate on test data, and to give some insight into the relationship between structure and activity in carcinogenesis. The substructures were also used to prove that there existed no accurate rule, based purely on atom-bond substructure with less than seven conditions, that could predict carcinogenicity. This results put a lower bound on the complexity of the relationship between chemical structure and carcinogenicity. Only by using a data mining algorithm, and by doing a complete search, is it possible to prove such a result. Finally the frequent substructures were shown to add value by increasing the accuracy of statistical and machine learning programs that were trained to predict chemical carcinogenicity. We conclude that Warmr, and ILP data mining methods generally, are an important new tool for analysing chemical databases.

Algorithms↗

Integrating histology and spatial transcriptomics via multimodal transformers and contrastive representation learning for accurate gene expression prediction.

Predicting spatial gene expression from Histological images is a fundamental task in understanding tissue organization and molecular phenotypes. However, existing methods often rely on single-model representations or lack effective alignment between image and transcriptomic features. To address these limitations, we propose a unified multimodal learning framework that integrates histological imaging and spatial transcriptomics through a shared latent representation space. Specifically, histological H&E images are encoded by a ResNet50-based convolutional stem and a MobileViT Transformer backbone to extract hierarchical visual representations. Both modalities are projected into a shared latent space via linear-GELU-dropout transformation blocks, enabling cross-modal alignment through a contrastive learning objective that maximizes agreement between the corresponding image and the spot embeddings. Experimental results on the 10x Genomics Visium dataset of human liver tissue demonstrate that MViTGene achieves significantly higher prediction accuracy than existing methods across multiple gene subsets, with improvements of 20%, 33%, and 12% in predicting marker genes, highly expressed genes, and highly variable genes, respectively. The significant improvement in relevance indicates that the model can more accurately capture the true correspondence between tissue morphology and gene expression, therefore enabling more reliable biological interpretation. It provides a computational tool for high-throughput spatial gene expression prediction that balances performance and interpretability.

Humans↗

Multidimensional signal exploration using multiple correspondence analysis. An example of a load lifting study.

Most empirical studies concerning rehabilitation yield numerous multidimensional signals (dozens of time variables are obtained for dozens of empirical situations). The purpose of this paper is to suggest a statistical analysis procedure based on: 1) space-time fuzzy windowing; 2) signal behavior characterization within the windows using membership value averages (MVA); and 3) MVA analysis using the multiple correspondence analysis (MCA). A load lifting study provided an example of 78 multidimensional signals including 89 time variables (forces, energy indicators, linear and angular positions, speeds, and accelerations). The main goal of MCA was to compare and contrast biomechanical signals from two lifting modes: "free" and "isokinetic." In the first mode, three loads were tested--light, medium, and heavy. In the second, three speeds were tested--slow, medium, and fast. Thirteen male individuals without disabilities participated in this study. The MCA showed that most of the free load-lifting strategies cannot be used in isokinetic lifting because the constraints of the subject and the environment are different. In addition, as the level of difficulty increases, free lifting became more economical while isokinetic lifting became less economical. These results would appear to indicate that movement strategies used for free lifting cannot be learned using an isokinetic machine during rehabilitation sessions for chronic low back pain. MCA was also suggested as a tool for comparing patients with control individuals. To achieve this aim, the notion of "supplementary data" was introduced.

Adult↗

PicSOr: an objective test of perceptual skill that predicts laparoscopic technical skill in three initial studies of laparoscopic performance.

BACKGROUND: Laparoscopic surgery requires surgeons to infer the shape of 3-D structures, such as the internal organs of patients, from 2-D displays on a video monitor. Recent evidence indicates that the issue is not resolved by the use of contemporary 3-D camera systems. It is therefore crucial to find ways of measuring differences in aptitude for recovering 3-D structure from 2-D images, and assessing its impact on performance. Our aim was to test empirically for a relationship between laparoscopic ability and the perceptual skill of recovering information about 3-D structures from 2-D monitor displays. METHODS: Participants in three studies completed a simulated laparoscopic cutting task as well as the Pictorial Surface Orientation (PicSOr)3 Test. In studies 1 (n = 48) and 2 (n = 32) both groups were laparoscopic novices, and in study 3 (n = 34) 18 of the participants were experienced laparoscopic surgeons. FINDINGS: All three studies showed that PicSOr consistently predicted the laparoscopic performance of participants on the laparoscopic cutting task (study 1, r = 0.5, p < 0.0003; study 2, r = 0.5, p < 0.004; and study 3, r = 0.42, p = 0.017). Furthermore, it was also a significant predictor of laparoscopic surgeons' performance (r = 0.54, p = 0.047). INTERPRETATIONS: This is the first objective perceptual psychometric test to reliably predict laparoscopic technical skills. PicSOr provides a tool for assessing which trainees have the potential to learn minimal access surgery.

Adult↗

Order of Search in Fuzzy ART and Fuzzy ARTMAP: Effect of the Choice Parameter.

This paper focuses on two ART architectures, the Fuzzy ART and the Fuzzy ARTMAP. Fuzzy ART is a pattern clustering machine, while Fuzzy ARTMAP is a pattern classification machine. Our study concentrates on the order according to which categories in Fuzzy ART, or the ART(a) model of Fuzzy ARTMAP are chosen. Our work provides a geometrical, and clearer understanding of why, and in what order, these categories are chosen for various ranges of the choice parameter of the Fuzzy ART module. This understanding serves as a powerful tool in developing properties of learning pertaining to these neural network architectures; to strengthen this argument, it is worth mentioning that the order according to which categories are chosen in ART 1 and ARTMAP provided a valuable tool in proving important properties about these architectures. Copyright 1996 Elsevier Science Ltd.

Journal Article↗

Sex-specific associations of the plasma-proteome with incident coronary artery disease.

AIMS: The etiology of coronary artery Disease (CAD) appears different for men and women, yet insights into underlying sex-specific biological mechanisms are limited. We integrated genomic and proteomic analyses to investigate sex-specific associations of the plasma-proteome with CAD. METHODS AND RESULTS: In 40,829 UK Biobank participants (free-of-CAD, baseline-365 days thereafter; 55% women; mean age 56.9&#x2009;&#xb1;&#x2009;8.1 years), we examined associations between 2,922 plasma proteins and incident CAD over a median follow-up of 13.7 years (IQR 13.1-14.4) using multivariable-adjusted Cox proportional hazards models. Sex-specific analyses identified 440 female exclusive and 32 male exclusive proteins associated with incident CAD (FDR-corrected p&#x2009;<&#x2009;0.05), revealing distinct pathway enrichments, including innate immune response in women and angiogenesis in men. Causality was assessed through combined and sex-stratified two-sample Mendelian randomization (MR) using inverse-variance-weighted analyses with genome wide association summary statistics from 422,108 men (61,969 cases) and 521,695 women (27,128 cases) (UK Biobank, FinnGen freeze 9). Integration of direct sex-protein interaction analyses with sex-combined MR identified 59 proteins with evidence for sex-specific causal effects. Four proteins demonstrated concordant directionality in sex-stratified MR analyses (n&#x2009;=&#x2009;943,803) and multivariable regression models, namely CDKN2D, MYH9, and SKAP2 (women), and CTSH (men). To assess translational relevance, prioritized targets were further evaluated in secondary major adverse cardiovascular events among carotid endarterectomy patients (MACE; Athero-Express) and acute myocardial infarction (AMI; MISSION!) using plasma proteomics and ELISA. After further top-target identification in the context of MACE and AMI, clinical drug candidates were identified through a machine learning framework, including CTSH (men), and TNFRSF4 (both sexes). CONCLUSIONS: We identified sex-specific associations of proteins and biological pathways with incident CAD. Whereas the majority of proteins had consistent associations in both men and women, our findings suggest a degree of sex-specific pathogenesis with evidence for potential causality, opening new alleys for tailored prevention strategies and clinical cardiovascular risk management.

Journal Article↗

Placenta-derived Exosomes Mitigate Hypoxia-Induced Trophoblast Apoptosis and Inflammatory Progression via SASH1.

SASH1 is a signal adaptor protein involved in cell growth, apoptosis, and immune regulation, and has been increasingly studied in tumor and immune cells. Emerging evidence suggests that SASH1 plays an important role in inflammatory responses and cellular homeostasis, processes that are closely associated with the development of PE. This study aimed to determine whether SASH1 contributes to trophoblast apoptosis and inflammatory responses in PE and whether P-EXOS exerts protective effects through SASH1 regulation. In this study, three PE-related transcriptomic datasets (GSE75010, GSE10588, and GSE60438) were analyzed to identify shared differentially expressed genes (DEGs), followed by Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses. Machine learning algorithms were further applied to screen key candidate genes, and single-cell RNA sequencing data were used to characterize cellular heterogeneity in placental tissue and to determine cell type-specific expression patterns. SASH1 was identified as a consensus candidate gene and was significantly upregulated in trophoblast cells from PE samples. In vitro, a hypoxia-treated HTR-8/SVneo trophoblast cell model was established, combined with SASH1 knockdown, SASH1 overexpression, and co-culture with P-EXOS. Functional experiments showed that knockdown of SASH1 significantly suppressed hypoxia-induced trophoblast apoptosis and reduced the secretion of pro-inflammatory cytokines, including IL-6, IL-1&#x3b2;, and TNF-&#x3b1;, whereas SASH1 overexpression promoted apoptosis and inflammatory responses. In addition, P-EXOS treatment markedly reduced SASH1 expression at both mRNA and protein levels and attenuated hypoxia-induced trophoblast injury, while SASH1 overexpression largely abolished these protective effects. Taken together, these findings indicate that SASH1 plays a critical role in trophoblast apoptosis and inflammatory responses in PE. P-EXOS may alleviate hypoxia-induced trophoblastic injury by suppressing SASH1 expression, providing new insights into the molecular mechanisms and potential therapeutic targets for PE.

Trophoblasts↗

CAKL: Commutative algebra k-mer learning of genomics.

Despite the availability of various sequence analysis models, comparative genomic analysis remains a challenge in genomics, genetics, and phylogenetics. Commutative algebra, a fundamental tool in algebraic geometry and number theory, has rarely been used in data and biological sciences. In this study, we introduce commutative algebra k-mer learning (CAKL) as the first-ever nonlinear algebraic framework for analyzing genomic sequences. CAKL bridges between commutative algebra, algebraic topology, combinatorics, and machine learning to establish a new mathematical paradigm for comparative genomic analysis. We evaluate its effectiveness on three tasks-genetic variant identification, phylogenetic tree analysis, and viral genome classification-typically requiring alignment-based, alignment-free, and machine-learning approaches, respectively. Across eleven datasets, CAKL outperforms five state-of-the-art sequence analysis methods, particularly in viral classification, and maintains stable predictive accuracy as dataset size increases, underscoring its scalability and robustness. This work ushers in a new era in commutative algebraic data analysis and learning.

Journal Article↗

Estimation of the stapes-bone thickness in the stapedotomy surgical procedure using a machine-learning technique.

Stapedotomy is a surgical procedure aimed at the treatment of hearing impairment due to otosclerosis. The treatment consists of drilling a hole through the stapes bone in the inner ear in order to insert a prosthesis. Safety precautions require knowledge of the nonmeasurable stapes thickness. The technical goal herein has been the design of high-level controls for an intelligent mechatronics drilling tool in order to enable the estimation of stapes thickness from measurable drilling data. The goal has been met by learning a map between drilling features, hence no model of the physical system has been necessary. Learning has been achieved as explained in this paper by a scheme, namely the d-sigma Fuzzy Lattice Neurocomputing (d sigma-FLN) scheme for classification, within the framework of fuzzy lattices. The successful application of the d sigma-FLN scheme is demonstrated in estimating the thickness of a stapes bone "on-line" using drilling data obtained experimentally in the laboratory.

Deafness↗

Artificial intelligence for anticancer drug discovery from natural products of macroalgae and sponges: A systematic review.

Marine natural products (MNPs) from macroalgae and marine sponges have inspired clinically important anticancer agents, including the cytarabine pharmacophore and the eribulin scaffold, while cyanobacterial dolastatin chemistry supplies the auristatin payloads of several marine-inspired antibody-drug conjugates (ADCs) such as brentuximab vedotin. Artificial intelligence (AI) methods, encompassing both classical machine learning (ML) with hand-engineered features and modern deep learning (DL) with many-layered neural networks, are increasingly supporting key decisions in natural-product anticancer drug discovery, including bioactivity prediction, target identification, absorption, distribution, metabolism, excretion and toxicity (ADMET) filtering, generative analogue design, and the selection of preclinical candidates. DL architectures relevant to this field include graph neural networks, transformer-based molecular generators, diffusion models for protein-ligand docking, and convolutional networks for mass spectrometry, while classical ML contributes interpretable fingerprint-based bioactivity models and molecular networking for dereplication. This review follows a systematic literature review methodology to organize the landscape of AI methods now applied to MNP anticancer discovery, distinguishing ML and DL approaches where relevant, situating them within the chemical context of macroalgal and sponge-derived oncology leads, and critically examining published case studies, including validation level (computational, in vitro, in vivo, clinical). The principal bottleneck for medical translation has shifted partly from algorithmic capability toward data infrastructure and experimental validation. Sparse, heterogeneous, and taxonomically biased bioactivity records limit what current models can learn and reduce the reliability of AI-prioritized candidates entering the preclinical pipeline. A roadmap is proposed that prioritizes open MNP-specific benchmarks, symbiont-aware modeling, and active learning loops with synthesizability and ADMET constraints. These AI workflows may accelerate the prioritization of marine-derived anticancer leads and support earlier, more evidence-based translational decisions in oncology drug development.

Biological Products↗

Metabolic fingerprinting of salt-stressed tomatoes.

The aim of this study was to adopt the approach of metabolic fingerprinting through the use of Fourier transform infrared (FT-IR) spectroscopy and chemometrics to study the effect of salinity on tomato fruit. Two varieties of tomato were studied, Edkawy and Simge F1. Salinity treatment significantly reduced the relative growth rate of Simge F1 but had no significant effect on that of Edkawy. In both tomato varieties salt-treatment significantly reduced mean fruit fresh weight and size class but had no significant affect on total fruit number. Marketable yield was however reduced in both varieties due to the occurrence of blossom end rot in response to salinity. Whole fruit flesh extracts from control and salt-grown tomatoes were analysed using FT-IR spectroscopy. Each sample spectrum contained 882 variables, absorbance values at different wavenumbers, making visual analysis difficult and therefore machine learning methods were applied. The unsupervised clustering method, principal component analysis (PCA) showed no discrimination between the control and salt-treated fruit for either variety. The supervised method, discriminant function analysis (DFA) was able to classify control and salt-treated fruit in both varieties. Genetic algorithms (GA) were applied to identify discriminatory regions within the FT-IR spectra important for fruit classification. The GA models were able to classify control and salt-treated fruit with a typical error, when classifying the whole data set, of 9% in Edkawy and 5% in Simge F1. Key regions were identified within the spectra corresponding to nitrile containing compounds and amino radicals. The application of GA enabled the identification of functional groups of potential importance in relation to the response of tomato to salinity.

Algorithms↗

Identification and ranking of genetic and laboratory environment factors influencing a behavioral trait, thermal nociception, via computational analysis of a large data archive.

Laboratory conditions in biobehavioral experiments are commonly assumed to be 'controlled', having little impact on the outcome. However, recent studies have illustrated that the laboratory environment has a robust effect on behavioral traits. Given that environmental factors can interact with trait-relevant genes, some have questioned the reliability and generalizability of behavior genetic research designed to identify those genes. This problem might be alleviated by the identification of the most relevant environmental factors, but the task is hindered by the large number of factors that typically vary between and within laboratories. We used a computational approach to retrospectively identify and rank sources of variability in nociceptive responses as they occurred in a typical research laboratory over several years. A machine-learning algorithm was applied to an archival data set of 8034 independent observations of baseline thermal nociceptive sensitivity. This analysis revealed that a factor even more important than mouse genotype was the experimenter performing the test, and that nociception can be affected by many additional laboratory factors including season/humidity, cage density, time of day, sex and within-cage order of testing. The results were confirmed by linear modeling in a subset of the data, and in confirmatory experiments, in which we were able to partition the variance of this complex trait among genetic (27%), environmental (42%) and genetic x environmental (18%) sources.

Animals↗

A molecular map of mesenchymal tumors.

BACKGROUND: Bone and soft tissue tumors represent a diverse group of neoplasms thought to derive from cells of the mesenchyme or neural crest. Histological diagnosis is challenging due to the poor or heterogenous differentiation of many tumors, resulting in uncertainty over prognosis and appropriate therapy. RESULTS: We have undertaken a broad and comprehensive study of the gene expression profile of 96 tumors with representatives of all mesenchymal tissues, including several problem diagnostic groups. Using machine learning methods adapted to this problem we identify molecular fingerprints for most tumors, which are pathognomonic (decisive) and biologically revealing. CONCLUSION: We demonstrate the utility of gene expression profiles and machine learning for a complex clinical problem, and identify putative origins for certain mesenchymal tumors.

Gene Expression Profiling↗

NanoSSL: attention mechanism-based self-supervised learning method for protein identification using nanopores.

MOTIVATION: Nanopores are cutting-edge interdisciplinary tools that can analyze biomolecules at the single-molecule level for many applications, e.g. DNA sequencing. Efforts are underway to extend nanopores to proteomics, including the development of machine learning algorithms for protein sequencing and identification. However, single-molecule data are intrinsically noisy and hard to process. Moreover, the development and performance of machine learning for nanopore is jeopardized by data scarcity. Self-supervised learning is an emerging method that may yield advantages in nanopore scenarios. RESULTS: We propose and experimentally validate Nanopore analysis using Self-Supervised Learning (NanoSSL), a generative self-supervised learning framework based on attention mechanisms for the identification of protein signals from nanopores. Leveraging a two-step approach consisting of self-supervised pre-training and supervised fine-tuning, NanoSSL learns useful feature representations from empirical data to facilitate downstream classification tasks. Inspired by the concept of fragmentation in conventional protein sequencing technologies, during pretraining each translocation event is split into multiple non-overlapping fragments of equal size, some of which are randomly masked and reconstructed using a masked autoencoder. Learning the feature representations of the reconstructed nanopore events facilitates molecular identification in fine-tuning. In this study, we retested a publicly available nanopore multiplexed protein sensing dataset for model iteration, and subsequently measured Alzheimer's disease biomarker A&#x3b2;1-42 using homemade solid-state nanopores. Empirical results indicated NanoSSL achieved an unprecedented performance across four metrics: accuracy, precision, recall, and F1 score, when classifying two mutated A&#x3b2;1-42, E22G and G37R. The self-supervised learning and attention mechanism were verified as the source of performance gains. AVAILABILITY AND IMPLEMENTATION: The main program is available at https://doi.org/10.5281/zenodo.17172822.

Nanopores↗

The immune system as a model for pattern recognition and classification.

OBJECTIVE: To design a pattern recognition engine based on concepts derived from mammalian immune systems. DESIGN: A supervised learning system (Immunos-81) was created using software abstractions of T cells, B cells, antibodies, and their interactions. Artificial T cells control the creation of B-cell populations (clones), which compete for recognition of "unknowns." The B-cell clone with the "simple highest avidity" (SHA) or "relative highest avidity" (RHA) is considered to have successfully classified the unknown. MEASUREMENT: Two standard machine learning data sets, consisting of eight nominal and six continuous variables, were used to test the recognition capabilities of Immunos-81. The first set (Cleveland), consisting of 303 cases of patients with suspected coronary artery disease, was used to perform a ten-way cross-validation. After completing the validation runs, the Cleveland data set was used as a training set prior to presentation of the second data set, consisting of 200 unknown cases. RESULTS: For cross-validation runs, correct recognition using SHA ranged from a high of 96 percent to a low of 63.2 percent. The average correct classification for all runs was 83.2 percent. Using the RHA metric, 11.2 percent were labeled "too close to determine" and no further attempt was made to classify them. Of the remaining cases, 85.5 percent were correctly classified. When the second data set was presented, correct classification occurred in 73.5 percent of cases when SHA was used and in 80.3 percent of cases when RHA was used. CONCLUSIONS: The immune system offers a viable paradigm for the design of pattern recognition systems. Additional research is required to fully exploit the nuances of immune computation.

Algorithms↗

Machine psychology: autonomous behavior, perceptual categorization and conditioning in a brain-based device.

In studying brain activity during the behavior of living animals, it is not possible simultaneously to analyze all levels of control from molecular events to motor responses. To provide insights into how levels of control interact, we have carried out synthetic neural modeling using a brain-based real-world device. We describe here the design and performance of such a device, designated Darwin VII, which is guided by computer-simulated analogues of cortical and subcortical structures. All levels of Darwin VII's neural architecture can be examined simultaneously as the device behaves in a real environment. Analysis of its neural activity during perceptual categorization and conditioned behavior suggests neural mechanisms for invariant object recognition, experience-dependent perceptual categorization, first-order and second-order conditioning, and the effects of different learning rates on responses to appetitive and aversive events. While individual Darwin VII exemplars developed similar categorical responses that depended on exploration of the environment and sensorimotor adaptation, each showed highly individual patterns of changes in synaptic strengths. By allowing exhaustive analysis and manipulation of neuroanatomy and large-scale neural dynamics, such brain-based devices provide valuable heuristics for understanding cortical interactions. These devices also provide the groundwork for the development of intelligent machines that follow neurobiological rather than computational principles in their construction.

Animals↗