Search PubMedSearch

SEARCH · Search PubMed

Results for “computer sciences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Artificial intelligence-driven advancements in agricultural biotechnology.

The need for faster and more informative data processing for better decision-making is driving the adoption of artificial intelligence (AI) in the agricultural sector. Thanks to recent advancements in computer science and the increase in computational powers of modern computers, AI is not only augmenting traditional solutions, but also helping in developing novel solutions to existing challenging matters. AI-driven models have an exceptional ability to identify patterns and combine a diverse collection of data together and make inference. The increasing pressure on farmlands posed by the growing global population and climate change is lessening growth, yield, and productivity ultimately posing risk to food security worldwide. Incorporation of AI in agriculture has the potential to drive farming efficiency to new heights. This comprehensive review critically evaluates the evolution of AI in agricultural biotechnology from a theoretical concept to a global phenomenon. A comprehensive literature search was performed using major scientific databases, including PubMed, Web of Science, Embase, Scopus, Lens and the Cochrane Library. In this review, we empirically demonstrate the fields advancement toward more capable AI systems and discuss the current applications of AI across crop improvement and precision agriculture such as crop improvement and genetic engineering, genomic selection and plant breeding, pest and disease detection, precision agriculture and smart farming, soil health and nutrient management, climate resilient crop development, livestock biotechnology, challenges and ethical considerations in AI based agricultural biotechnology. Furthermore, this review addresses the exponential growth of commercial intellectual property in the field and contrast it with academic publication outputs. Finally, we critically assess the ethical challenges impeding equitable adoption of AI including data sovereignty and digital divide, while projecting future frontiers involving quantum computing. This review will help build sustainable agricultural systems capable of adapting to climate change, contribute to the development of climate-resilient and high-yielding crops, and address global food security challenges.

Agriculture

Large language models in bioinformatics: a comprehensive survey.

The emergence of foundation models with trillion-level parameters has redefined the landscape of artificial intelligence. Various fields are developing their own large-scale models, which can solve many problems within the field and improve work efficiency. Biological large-scale models are a cross-disciplinary research field that combines mathematics, computer science, and biology, aiming to simulate and understand the structure, function, and dynamic changes of biological systems through the establishment of complex computational models. This field covers multiple levels such as biological pathways, population dynamics, protein folding, etc., providing us with tools for deep exploration of the mysteries of life and applications in medicine, ecology, and other fields. This article reviews the background and research status of biological large-scale models, and discusses future directions. Large language models (LLMs) and other large-scale foundation models have rapidly advanced in recent years, enabling powerful representation learning and generation across text, sequences, and multimodal data. In bioinformatics and biomedicine, these models are increasingly used to analyze genomic sequences, infer protein properties and structures, support drug discovery, and integrate heterogeneous biomedical evidence. This survey reviews the basic principles of LLMs and summarizes representative applications in (i) gene and genome sequence analysis, (ii) protein structure and function prediction, and (iii) drug design, including virtual screening and personalized medicine. We also discuss emerging multi-model modeling approaches, as well as key challenges such as data quality and privacy, interpretability, generalization to new organisms and tasks, and responsible deployment in health-related settings. Finally, we outline future directions for developing reliable, scalable, and explainable bioinformatics foundation models.

bioinformatics

Polyploidy Arithmetic.

Polyploidy occurs in plants and animals, and is an important force in speciation and genome evolution. The main focus of this paper is the following fundamental question that was recently posed by Huber and Maher: Given the ploidy numbers of a collection of extant species, or their ploidy profile, what is the smallest number of hybridizations needed in any evolutionary history for these species to completely represent these numbers? In this paper, we shall show that this question can be rephrased in terms of addition chains and the closely related addition sequences, which have been studied for over a century in mathematics and computer science. These are sequences of natural numbers that start with 1, so that each number in the sequence larger than 1 is the sum of two other numbers arising earlier in the sequence. In our first main result, we show that finding the smallest number of hybridization events to explain a ploidy profile, or the hybrid number, is equivalent to solving the so-called addition sequence problem. This immediately implies that computing the hybridization number is computationally intractable. Even so, it also leads to new connections to representing polyploid evolution using networks. More specifically, in our second main result we show that ploidy profiles representable by tree-child networks are exactly the addition chains, implying a polynomial-time algorithm for identifying these profiles. We then consider beaded tree-child networks, which permit the representation of autopolyploidy events, and in our third main result we provide a greedy polynomial-time algorithm to decide whether a given profile can be realized by such a network. We expect that our results can be leveraged in future work through, for example, making use of known algorithms for computing short addition sequences to give bounds for the hybrid number, and in guiding network reconstruction for polyploid species.

Polyploidy

CSGL: chemical synthesis graph learning for molecule representation.

MOTIVATION: Molecule representation learning (MRL) translates molecules into a real vector space, serving as input to downstream tasks in biology, chemistry, and computer science. This article introduces a chemical synthesis graph learning (CSGL) framework, which enhances MRL by considering both the atomic structures of molecules and their roles in chemical reactions through a hierarchical graph representation. Specifically, molecules are first modeled based on their molecular graphs, which capture atomic-level structural information. They are then further refined using a chemical synthesis graph, where nodes represent reactant and product molecule sets, and edges encode chemical transformations between reactants and products (e.g. changes in molecular structures). CSGL optimizes molecular embeddings of reactant and product nodes in a fashion that ensures the embeddings conform to a chemical balance constraint. RESULTS: Experimental results show that our method CSGL achieves strong performance on a variety of tasks, including product prediction, reaction classification, and molecular property prediction. AVAILABILITY AND IMPLEMENTATION: https://github.com/li-2023/CSGL.

Machine Learning

Protocol to perform integrative analysis of high-dimensional single-cell multimodal data using an interpretable deep learning technique.

The advent of single-cell multi-omics sequencing technology makes it possible for researchers to leverage multiple modalities for individual cells. Here, we present a protocol to perform integrative analysis of high-dimensional single-cell multimodal data using an interpretable deep learning technique called moETM. We describe steps for data preprocessing, multi-omics integration, inclusion of prior pathway knowledge, and cross-omics imputation. As a demonstration, we used the single-cell multi-omics data collected from bone marrow mononuclear cells (GSE194122) as in our original study. For complete details on the use and execution of this protocol, please refer to Zhou et al.1.

Deep Learning

Using cancer profiles to identify synthetic lethal therapeutic targets and predictive biomarkers in cancer gene dependency data.

MOTIVATION: Large scale loss-of-function screens utilising CRISPR or siRNA can provide profound insights into the importance of individual genes for the survival of a cancer cell and can drive the identification of therapeutic targets and biomarkers, and the development of targeted drugs. However, the analysis of these data and the substantial bodies of metadata that relate to them, is technically challenging and typically requires substantial expertise in data science and computer coding. RESULTS: To facilitate the analysis of cancer gene dependency data by cancer biologists and clinical scientists, we have developed DepMine-a computational toolkit providing a powerful system for framing complex queries relating cancer gene dependency to the underlying genetic changes that occur in cancer cells. DepMine identifies synthetic lethal relationships between putative target genes and complex 'cancer profiles' built from user-specified combinations of mutations, copy-number variation, and expression levels, and can refine these to optimal biomarker definitions for target dependency. AVAILABILITY: The Python implementation of DepMine and associated data files can be obtained at https://github.com/UOSbioinformaticslab/depmine and is free to academics and Not-For-Profit organisations. The DepMine release referenced in this paper is archived as DOI: 10.5281/zenodo.19570601.

Humans

Challenges and Opportunities in Analyzing Cancer-Associated Microbiomes.

The study of cancer-associated microbiomes has gained significant attention in recent years, spurred by advances in high-throughput sequencing and metagenomic analysis. Microbiome research holds promise for identifying noninvasive biomarkers and possibly new paradigms for cancer treatment. In this review, we explore the key computational challenges and opportunities in analyzing cancer-associated microbiomes (in tumor/normal tissues and other body sites, e.g., gut, oral, and skin), focusing on sequencing-driven strategies and associated considerations for taxonomic and functional characterization. The discussion covers the strengths and limitations of current analysis tools for identifying contamination, determining compositional bias, and resolving species and strains, as well as the statistical, metabolic, and network inferences that are essential to uncover host-microbiome interactions. Several key considerations are required to guide the choice of databases used for metagenomic analysis in such studies. Recent advances in spatial and single-cell technologies have provided insights into cancer-associated microbiomes, and Artificial Intelligence-driven protein function prediction might enable rapid advances in this field. Finally, we provide a perspective on how the field can evolve to manage the ever-growing size of datasets and generate robust and testable hypotheses. This article is part of a special series: Driving Cancer Discoveries with Computational Research, Data Science, and Machine Learning/AI .

Humans

Integrative Multiomics and Drug Sensitivity Profiling Reveal Potential Biomarkers and Therapeutic Strategies in Pediatric Solid Tumors.

UNLABELLED: Cure rates for childhood malignancies using established therapy protocols have increased to an average of 80% but have reached a plateau. Moreover, survival rates are particularly low for some pediatric tumors-such as high-risk group 3 medulloblastomas, osteosarcomas, Ewing sarcomas, high-risk neuroblastomas, and high-grade gliomas-and dismal for patients with relapsed malignancies. A functional drug response profiling platform for pediatric solid and brain tumors has been established within the INFORM program to identify patient-specific vulnerabilities and biomarkers and to unravel molecular mechanisms associated with drug response profiles for clinical translation. In this study, we performed a multiomics analysis using drug sensitivity profiles, as well as genomic and transcriptomic data, of 81 pediatric solid tumor samples. The integrative analysis suggested two multiomics signatures associated with drug sensitivity. One signature distinguished neuroblastoma samples with sensitivity to navitoclax, a BCL2 family inhibitor. A second signature was specific to a subset of Wilms tumors harboring the SIX1 (Q177R) hotspot mutation that displayed high expression of MGAM, PTPN14, STAT4, and KDM2B and high sensitivity to MEK inhibitors. A patient-specific causal interaction network analysis suggested possible molecular interactions between MEK inhibitors and the SIX1 mutation in Wilms tumor samples. In conclusion, the integration of drug sensitivity profiling and multiomics data revealed potential biomarkers that may be associated with drug sensitivity in pediatric solid tumors. Patient-specific causal interaction network analysis further elucidated the interaction between inhibitors and signature biomarkers, providing insights that may inform clinical translation. SIGNIFICANCE: The combination of multiomics analysis and drug sensitivity profiling identified two signatures related to drug sensitivity in pediatric solid tumors, contributing to the advancement of functional precision medicine and personalized treatment strategies. This article is part of a special series: Driving Cancer Discoveries with Computational Research, Data Science, and Machine Learning/AI .

Humans

Modeling Early-Onset Cancer Kinetics Reveals Changes in Underlying Risk and the Impact of Population Screening.

UNLABELLED: Recent studies have reported increases in early-onset cancer cases (diagnosed less than 50 years of age) and raised questions about whether the increase is related to earlier diagnosis from nonspecific medical tests as reflected by decreasing tumor-size-at-diagnosis (apparent effects) or actual increases in underlying cancer risk (true effects), or both. The classic Multistage Clonal Expansion (MSCE) model assumes cancer detection at the first malignant cell's emergence, although later modifications have included lag-times or stochasticity in detection to represent the delay in tumor detection. In this study, we introduced an approach to explicitly incorporate tumor-size-at-diagnosis in the MSCE framework accounting for improvements in cancer detection over time to distinguish between apparent and true increases in early-onset cancer incidence. The model was structurally identifiable and provided better parameter estimation than the classic model. The model was applied to colorectal, breast, and thyroid cancers to examine changes in cancer risk while accounting for detection improvements over time in three representative birth cohorts (1950-1954, 1965-1969, and 1980-1984). The analyses suggested accelerated carcinogenic events and shorter mean sojourn times (the average time from the first malignant cell emergence to cancer detection) in more recent cohorts. Furthermore, using this model to examine the screening impact on the incidence of breast and colorectal cancers, for which both have established screening protocols, provided results that align with well-documented differences in screening effects between these cancers. These findings underscore the importance of incorporating tumor-size-at-diagnosis in cancer modeling and support true increases in early-onset cancer risk in recent years for breast, colorectal, and thyroid cancers. SIGNIFICANCE: A model of early-onset cancer trends that distinguishes true risk from detection effects accurately captures cancer kinetics, trends in cancer progression, and the impact of screening, which could inform cancer prevention strategies. This article is part of a special series: Driving Cancer Discoveries with Computational Research, Data Science, and Machine Learning/AI .

Humans

Practicing Data Science in Interactive Notebooks.

The Jupyter Notebook is a platform for interactive computing that displays code and results in the same browser, making it valuable for teaching, prototyping, data analysis, and collaboration. Its explicit and transparent structure greatly reproducibility while its backend server supports flexible deployment. In the past few years, Jupyter notebooks and similar tools have become increasingly popular. In this chapter, we will review key aspects of data analysis in a cloud environment and demonstrate common tasks for analyzing metabolomics data using template notebooks. This is an accompaniment to the basic bioinformatics tools and essential data science toolkit introduced in the first edition.

Software

Transcriptome-wide analysis reveals sequence selection to avoid mRNA aggregation in E. coli.

The stability of RNA base pairing and its limited four-letter code create an intrinsic potential for promiscuous RNA-RNA interactions. In vitro, such interactions drive RNA to self-assemble into aggregates. This raises a fundamental unanswered question: within a confined cellular volume at physiological mRNA abundances, how much aggregation would arise from sequence-encoded chemistry alone? Here, we establish this baseline with large-scale kinetic simulations of the E. coli transcriptome. Our simulations reveal that sequence-encoded base-pairing energetics is sufficient to generate a dynamic network of large aggregates, organized by long, multivalent mRNA hubs. Strikingly, evolutionary analysis shows that native E. coli sequences exhibit clear signatures of selection to counteract this propensity: they fold more stably, minimize unstructured regions, and form weaker intermolecular contacts than dinucleotide-preserving controls. These findings demonstrate that maintaining transcriptome solubility has been a significant, previously unrecognized constraint shaping genome evolution, and provide a new lens to interpret cellular RNA management.

Biological Sciences (Biophysics and Computational

Generative AI Models in Time-Varying Biomedical Data: Scoping Review.

BACKGROUND: Trajectory modeling is a long-standing challenge in the application of computational methods to health care. In the age of big data, traditional statistical and machine learning methods do not achieve satisfactory results as they often fail to capture the complex underlying distributions of multimodal health data and long-term dependencies throughout medical histories. Recent advances in generative artificial intelligence (AI) have provided powerful tools to represent complex distributions and patterns with minimal underlying assumptions, with major impact in fields such as finance and environmental sciences, prompting researchers to apply these methods for disease modeling in health care. OBJECTIVE: While AI methods have proven powerful, their application in clinical practice remains limited due to their highly complex nature. The proliferation of AI algorithms also poses a significant challenge for nondevelopers to track and incorporate these advances into clinical research and application. In this paper, we introduce basic concepts in generative AI and discuss current algorithms and how they can be applied to health care for practitioners with little background in computer science. METHODS: We surveyed peer-reviewed papers on generative AI models with specific applications to time-series health data. Our search included single- and multimodal generative AI models that operated over structured and unstructured data, physiological waveforms, medical imaging, and multi-omics data. We introduce current generative AI methods, review their applications, and discuss their limitations and future directions in each data modality. RESULTS: We followed the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines and reviewed 155 articles on generative AI applications to time-series health care data across modalities. Furthermore, we offer a systematic framework for clinicians to easily identify suitable AI methods for their data and task at hand. CONCLUSIONS: We reviewed and critiqued existing applications of generative AI to time-series health data with the aim of bridging the gap between computational methods and clinical application. We also identified the shortcomings of existing approaches and highlighted recent advances in generative AI that represent promising directions for health care modeling.

Artificial Intelligence

Algorithms and tools for data-driven omics integration to achieve multilayer biological insights: a narrative review.

Systems biology is a holistic approach to biological sciences that combines experimental and computational strategies, aimed at integrating information from different scales of biological processes to unravel pathophysiological mechanisms and behaviours. In this scenario, high-throughput technologies have been playing a major role in providing huge amounts of omics data, whose integration would offer unprecedented possibilities in gaining insights on diseases and identifying potential biomarkers. In the present review, we focus on strategies that have been applied in literature to integrate genomics, transcriptomics, proteomics, and metabolomics in the year range 2018-2024. Integration approaches were divided into three main categories: statistical-based approaches, multivariate methods, and machine learning/artificial intelligence techniques. Among them, statistical approaches (mainly based on correlation) were the ones with a slightly higher prevalence, followed by multivariate approaches, and machine learning techniques. Integrating multiple biological layers has shown great potential in uncovering molecular mechanisms, identifying putative biomarkers, and aid classification, most of the time resulting in better performances when compared to single omics analyses. However, significant challenges remain. The high-throughput nature of omics platforms introduces issues such as variable data quality, missing values, collinearity, and dimensionality. These challenges further increase when combining multiple omics datasets, as the complexity and heterogeneity of the data increase with integration. We report different strategies that have been found in literature to cope with these challenges, but some open issues still remain and should be addressed to disclose the full potential of omics integration.

Algorithms

FusionTarget: Computational framework for drug repurposing against modeled fusion protein structures from genomic breakpoints.

Many fusion genes have been recognized as biomarkers and therapeutic targets. However, the lack of knowledge on protein structures and targeting approaches made it challenging to develop effective targeting therapeutics. To fill this, we developed a computational pipeline, FusionTarget, which annotates the genomic DNA breakage to RNA and protein sequences, predicts the 3D structures of fusion proteins, and performs comparative virtual screening, comparative molecular dynamics simulation, and quantitative analyses to identify the fusion protein-selective small molecules by selecting drugs with consistent high-fold binding affinity between fusion and wild-type proteins in multiple isoforms. We applied our pipeline to EWSR1::FLI1 in Ewing sarcoma and KMT2A::AFF1 in infant acute lymphoblastic leukemia. Further cell assay experiments confirmed that cells expressing individual fusion genes were more sensitive to the suggested drugs, and the key downstream genes were affected by our drugs. FusionTarget provides a unique foundation for developing therapeutics targeting fusion proteins.

applied computing in medical science

Fantastic microbes and where to find them: evaluating learning-by-doing outcomes in a crowdfunded metagenomics workshop.

Metagenomics offers a powerful framework for authentic, interdisciplinary learning, yet it remains underrepresented in undergraduate education due to technical and infrastructural barriers. We hypothesized that a research-based, learning-by-doing metagenomics workshop supported by accessible bioinformatics tools could enhance students' perceived skills, self-efficacy, and conceptual understanding of metagenomic analysis. To test this hypothesis, we designed and evaluated a hybrid hands-on workshop in which undergraduate and postgraduate students analyzed real environmental shotgun metagenomic datasets generated from soil samples collected during a citizen science initiative. Using the graphical workflow platform KBase, participants completed an end-to-end metagenomic analysis, from quality control and assembly to genome reconstruction, taxonomic classification, functional annotation, and scientific presentation of results. Educational outcomes were assessed through validated retrospective pre-post questionnaires, self-efficacy scales, and an open-ended conceptual understanding task. Participants showed significant increases in perceived metagenomic skills and confidence in performing metagenomic analyses, while gains in perceived learning showed a positive trend. Conceptual understanding improved across educational levels, particularly among participants with limited prior experience. Together, these findings demonstrate that authentic, data-driven metagenomics activities can effectively lower barriers to computational biology and foster meaningful learning through hands-on research experiences.

Metagenomics