Search PubMedSearch

SEARCH · Search PubMed

Results for “bioinformatic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

SPOP expression is associated with tumor-infiltrating lymphocytes in pancreatic cancer.

BACKGROUND: Speckle Type POZ Protein (SPOP), despite its tumor type-dependent role in tumorigenesis, primarily as a tumor suppressor gene is associated with a variety of different cancers. However, its function in pancreatic cancer remains uncertain. METHODS: SPOP expression and the association between its expression and patient prognosis and immune function were evaluated using The Cancer Genome Atlas (TCGA), Genotype-Tissue Expression (GTEx), The Tumor Immune Estimation Resource 2.0 (TIMER2.0) database, cBioportal, and various bioinformatic databases. Enrichment analysis of SPOP and the association between SPOP expression with clinical stage and grade were analyzed using the R software package. Then immunohistochemistry (IHC) was used to estimate the correlation between SPOP and tumor-infiltrating lymphocytes (TILs) in patients with pancreatic cancer. RESULTS: As part of our study, we assessed that SPOP was anomalously expressed in kinds of cancers, associated with clinical stage and outcomes. Meanwhile, SPOP also played a crucial role in the tumor microenvironment (TME). The expression level of SPOP was significantly correlated to tumor-infiltrating immune cells (TICs) in pancreatic cancer. CONCLUSIONS: Our study uncovered the potential corrections in SPOP with TICs, suggesting that SPOP may act as a biomarker for immunotherapy in pancreatic cancer.

Humans

MYH16 upregulation is associated with lung adenocarcinoma aggressiveness and immune infiltration.

Myosin heavy chain 16 (MYH16) may significantly affect cell cycle progression. Nevertheless, there is a lack of evidence about the clinical relevance of MYH16 upregulation in pan cancers, including lung adenocarcinoma (LUAD). MYH16 expression patterns were evaluated in various bioinformatics databases using The Cancer Genome Atlas data set. Clinical and pathological factor data were employed to risk-stratify patients. The Kaplan-Meier plotter approach was used to estimate survival rates. Tumor immune infiltration was explored via the TIMER tool, and gene set enrichment analysis (GSEA) was used to identify the pathways involved in MYH16 upregulation. The results showed that MYH16 was abnormally upregulated in pan cancers, including LUAD. MYH16 expression induction in LUAD was found to be related to the tumor stage. Furthermore, MYH16 upregulation was correlated with LUAD development and worse overall survival, particularly in women. Notably, MYH16 overexpression in LUAD tissues corresponded to the amount of immune infiltration in the tumor. Additionally, univariate Cox hazard regression analysis revealed that MYH16 may be an independent prognostic indicator for LUAD. Furthermore, a nomogram was constructed according to MYH16 expression and clinical characteristics. BMP6 expression deficiency may be a key factor contributing to MYH16 upregulation in LUAD. Finally, GSEA demonstrated that MYH16 might mediate meiosis and gene silencing through RNA signaling pathways. This study, for the first time, showed that MYH16 upregulation in LUAD is associated with various risk factors, increased cancer aggressiveness, enhanced infiltration of tumor immune cells, and reduced survival rates.

Female

Identification of a prognostic signature consisting of three macrophage-related genes for glioblastoma based on bulk and single-cell transcriptomes analyses.

BACKGROUND: Tumor-associated macrophages have been implicated in the progression and treatment resistance of glioblastoma (GBM). This study aimed to identify macrophage-related genes associated with prognosis and therapeutic response in GBM. MATERIALS AND METHODS: Bulk RNA-seq data from 533 patients with GBM were downloaded from the Cancer Genome Atlas (TCGA) and Chinese Glioma Genome Atlas (CGGA) databases. Bioinformatic tools were used to detect the co-expression gene modules associated with the infiltration of immune cells, identify a prognostic macrophage-related gene signature, and explore their association with sensitivity to chemotherapeutic drugs and immune checkpoint blockade. Single-cell RNA-seq data and multiplexed immunofluorescence were used to validate ISG20 expression (a member of the identified gene signature) in macrophages. RESULTS: We detected gene modules associated with macrophages and identified a signature consisting of three macrophage-related genes (ISG20, PARP12 and IFIT5) in the discovery set (TCGA-GBM, n = 159), and validated its prognostic value in the validation set (CGGA-GBM, n = 374). This gene signature demonstrated favorable accuracy in predicting prognosis and resistance of immuno- and chemo-therapy. The co-expression of ISG20 and PD-1 in macrophages was verified by single-cell RNA-seq data and multiplex immunofluorescence. CONCLUSIONS: This study presents a macrophage-related gene signature to predict prognosis and therapeutic response in GBM. ISG20, PARP12 and IFIT5 are interferon-stimulated genes, and further investigations may provide new insights into the interplay between macrophages and interferon signaling in GBM.

Humans

Active components and potential mechanisms of Wuzhuyu decoction in the treatment of ethanol-induced acute gastric mucosal injury: a network pharmacology and experimental verification.

OBJECTIVE: To investigate the underlying mechanisms and active components of Wuzhuyu decoction (, WD) in alleviating ethanol-induced acute gastric mucosal injury (GMI) using an integrated approach of network pharmacology and experimental verification. METHODS: Sprague-Dawley rats were randomly divided into six groups: control (Con), model (Mod), bismuth potassium citrate (BPC), WD at low (WD-L), medium (WD-M), and high (WD-H) doses. Following seven days of continuous intragastric administration of the respective treatments, an ethanol-induced gastric mucosal injury model was established in all groups except the control group by oral gavage of anhydrous ethanol. The gastric mucosal injury index was evaluated, and pathological changes were assessed viahematoxylin and eosin (HE) staining. Levels of tumor necrosis factor-alpha (TNF-α), interleukin-1 beta (IL-1β), malondialdehyde (MDA), superoxide dismutase (SOD), and glutathione peroxidase (GSH-Px) were measured by enzyme-linked immunosorbent assay (ELISA). The chemical composition was identified by ultra-performance liquid chromatography-tandem mass spectrometry. Active compounds were screened using the Swiss-absorption, distribution, metabolism, and excretion database, and their potential targets were predicted using the Swiss Target Prediction database and bioinformatics annotation database for molecular mechanism. Simultaneously, disease targets related to GMI were retrieved from the online mendelian inheritance in man and GeneCards databases. A protein-protein interaction (PPI) network was constructed, and functional enrichment analyses of gene ontology (GO) and Kyoto encyclopedia of genes and genomes (KEGG) enrichment analyses were performed using the Metascape database. Key predictions from the network pharmacology analysis were subsequently verified through animal experiments. Protein expression levels of B-cell lymphoma-2 (Bcl-2), Bcl-2-associated X protein (Bax), Cleaved Caspase-3, and Cleaved Caspase-9 were analyzed by Western blot. Finally, molecular docking was performed using AutoDock Vina to investigate the interactions between the active components and core targets. RESULTS: WD treatment significantly reduced the gastric mucosal injury index and the levels of TNF-α, IL-1β, MDA, while it increased the activities of SOD and GSH-Px. Histopathological examination revealed marked improvement in gastric tissue morphology. A total of 145 compounds were identified in WD. Network pharmacology analysis identified 440 overlapping targets between WD and GMI. GO and KEGG enrichment analyses highlighted the apoptosis signaling pathway as a key mechanism for WD's protective effect against ethanol-induced GMI. Experimental validation demonstrated that WD treatment reduced the apoptosis of gastric mucosal epithelial cells, promoted the expression of Bcl-2, and inhibited the expression of Bax, Cleaved Caspase-3 and Cleaved Caspase-9. Molecular docking results indicated that dehydroevodiamine, rutaecarpine, evodiamine, hexahydrocurcumin, and isorhamnetin are potential active components in WD that contribute to the inhibition of apoptosis. CONCLUSIONS: WD alleviates ethanol-induced acute GMI, at least in part, by inhibiting the apoptosis. The primary active components responsible for this effect are dehydroevodiamine, rutaecarpine, evodiamine, hexahydrocurcumin, and isorhamnetin.

Drugs, Chinese Herbal

KSHVbook: An Information-Sharing Database for Kaposi's Sarcoma-Associated Herpesvirus.

Kaposi's sarcoma-associated herpesvirus (KSHV) is a double-stranded DNA virus belonging to the γ-herpesvirus subfamily. KSHV is the causative agent of Kaposi's sarcoma (KS), primary effusion lymphoma (PEL), multicentric Castleman's disease (MCD), and KSHV inflammatory cytokine syndrome (KICS). Since its discovery, research on KSHV has rapidly progressed, but existing information platforms relatively lack comprehensiveness and do not provide efficient analysis tools tailored for KSHV. To further promote the research on KSHV more effectively, we have developed KSHVbook (http://www.kshvbook.com), a specialized information-sharing database dedicated to KSHV. This platform offers extensive information on genes, coding sequences, proteins, and the gene regulatory region. Besides, the KSHVbook includes about 35 010 transcription factor binding sites (TFBSs), 342 010 pairs of KSHV miRNA-host target gene relationships, protein structures predicted by AlphaFold3, qPCR primers, and so on. We also develop analytical tools for viral genome regions, TFBSs, and KSHV miRNA target genes to discover previously unknown biological functions of KSHV. These analytical tools can effectively identify the potential regulatory relationships between host transcription factors and viral genes. Overall, this platform provides a centralized data resource for KSHV research by integrating multiple databases, offering accessible analysis tools, and simplifying data acquisition. The KSHVbook will continue to be updated, and more features can be found on the website.

Herpesvirus 8, Human

Functional Analysis of MS-Based Proteomics Data: From Protein Groups to Networks.

Mass spectrometry-based proteomics allows the quantification of thousands of proteins, protein variants, and their modifications, in many biological samples. These are derived from the measurement of peptide relative quantities, and it is not always possible to distinguish proteins with similar sequences due to the absence of protein-specific peptides. In such cases, peptide signals are reported in protein groups that can correspond to several genes. Here, we show that multi-gene protein groups have a limited impact on GO-term enrichment, but selecting only one gene per group affects network analysis. We thus present the Cytoscape app Proteo Visualizer (https://apps.cytoscape.org/apps/ProteoVisualizer) that is designed for retrieving protein interaction networks from STRING using protein groups as input and thus allows visualization and network analysis of bottom-up MS-based proteomics data sets.

Proteomics

Targeting SUV4-20H2-mediated H4K20 methylation restrains growth and migration in pediatric high-grade astrocytomas.

Pediatric astrocytomas are characterized by increased molecular and clinical heterogeneity with epigenetic alterations contributing to aggressiveness and therapy resistance. The repressive histone mark H4K20 trimethylation (H4K20me3) and the methyltransferase SUV4-20H2 (KMT5C) are critical regulators of chromatin integrity and genome stability, with limited investigation in pediatric astrocytomas. KMT5C mRNA levels were evaluated in a publicly available pediatric gliomas database using bioinformatic analysis. Investigation of SUV4-20H2 and H4K20me3 expression was performed in a cohort of 43 pediatric astrocytoma tissues by immunohistochemistry. Their functional role and mechanism of action was investigated in pediatric glioma cell lines by using the substrate-competitive inhibitor of SUV4-20, A-196. Cell viability, apoptosis and migration were assessed using XTT, cleaved PARP, and wound healing assays, respectively. Effects of treatment on H4K20 methylation, DNA damage, mitotic stress [Polo-like kinase (PLK1) expression], and invasion markers (N-cadherin, β-catenin expression) were examined by western immunoblotting. KMT5C mRNA was significantly enriched in pediatric high-grade astrocytomas compared to low-grade tumors. A significant elevation of SUV4-20H2 and H4K20me3 expression was detected in astrocytoma tissues indicating epigenetic dysregulation contributing to malignancy. Treatment with A-196 reduced cell proliferation of pediatric glioma cell lines and induced apoptosis in a dose-dependent manner. It further impaired cell migration, accompanied by reduced N-cadherin and β-catenin expression. Mechanistically, inhibition of SUV4-20 depleted H4K20me3, inducing chromatin destabilization, replication-associated DNA damage and was associated with increased PLK1 expression, consistent with activation of a mitotic stress response. Our findings indicate that SUV4-20H2-mediated H4K20 activity in pediatric high-grade astrocytomas maintains their growth and migratory potential by regulating chromatin integrity and may serve as potential therapeutic target.

H4K20me2/3

Programmatic access to ICTV virus taxonomy through a public ontology API.

BACKGROUND: The International Committee on Taxonomy of Viruses (ICTV) is responsible for developing and maintaining a universal virus taxonomy. As the reference framework for organising the viral world, it is essential for virology and related fields. Despite its widespread use in research and public health, programmatic access to ICTV taxonomy has remained limited, posing challenges for integration, versioning, and interoperability across databases and bioinformatics resources requiring up-to-date virus taxonomy. FINDINGS: To address this, we developed a public and sustainable solution leveraging ontology-based APIs. All available ICTV Master Species List (MSL) releases, from MSL1 to MSL41, were transformed into a unified, semantically structured ontology comprising more than 195,000 current and historical entities and deployed through the Ontology Lookup Service (OLS). The ontology is automatically rebuilt and republished whenever a new MSL release becomes available. Complementary ICTV-NCBI mappings and helper libraries support integration into downstream systems. CONCLUSIONS: Together, these resources enable, for the first time, public programmatic retrieval of current and historical ICTV taxon names, taxonomic relationships, metadata, and persistent identifiers through stable endpoints, including resolution of former taxonomic terms to their current accepted taxon or taxa and retrieval of taxon histories across releases. More broadly, this work illustrates a general strategy for transforming structured biological datasets into semantically enriched graph resources exposed through scalable public APIs. These developments enhance interoperability, reduce manual curation, and support FAIR-aligned taxonomic data management in virology and pandemic preparedness.

API

Programmatic access to ICTV virus taxonomy through a public ontology API.

The International Committee on Taxonomy of Viruses (ICTV) is responsible for developing and maintaining a universal virus taxonomy. As the reference framework for organising the viral world, it is essential for virology and related fields. Despite its widespread use in research and public health, programmatic access to ICTV taxonomy has remained limited, posing challenges for integration, versioning, and interoperability across databases and bioinformatics resources requiring up-to-date virus taxonomy. To address this, we developed a public and sustainable solution leveraging ontology-based APIs. Successive ICTV Master Species List (MSL) releases were transformed into a structured ontology and deployed as a unified representation through the Ontology Lookup Service (OLS). The framework also provides ICTV-NCBI mappings and helper libraries for integration into downstream systems. This enables, for the first time, public programmatic retrieval of current and historical virological taxon names, taxonomic relationships, metadata, and persistent identifiers through stable endpoints. More broadly, this work illustrates a general strategy for transforming structured biological datasets into semantically enriched graph resources exposed through scalable public APIs. These developments enhance interoperability, reduce manual curation, and support FAIR-aligned taxonomic data management in virology and pandemic preparedness.

API

Integrative proteomics and bioinformatics pipelines for PTM profiling.

Post-translational modifications (PTMs) regulate protein function across all life forms and allow plants to respond rapidly to biotic and abiotic stress. Over 450 PTM types have been described across organisms, of which 23-33 have been experimentally confirmed in plants, including phosphorylation, acetylation, methylation, glycosylation, ubiquitination, and sumoylation. These modifications are highly dynamic and often reversible, and frequently act in combination, or "crosstalk," to fine-tune cellular processes. Advances in high-resolution mass spectrometry and large-scale genome sequencing continue to expand the catalogue of known PTM sites, while machine learning and deep learning approaches increasingly support prediction of PTM site localization and function. Unlike broader surveys of plant PTMs, this review focuses specifically on O-phosphorylation and Lys-N(ε)-acetylation, the two best-characterized and most extensively crosstalking PTMs in plants, and integrates four perspectives: the historical development of proteomic and bioinformatics approaches to these modifications; current mass spectrometry-based workflows and enrichment strategies; the bioinformatics tools and databases available for their analysis; and the technical and species-related challenges, particularly in non-model plants, that currently limit their study. We close by outlining priority directions for future research, including multi-omics integration, AI-based prediction, and the translation of PTM knowledge into crop stress resilience and breeding applications.

Protein Processing, Post-Translational

Integrated Bioinformatics Analysis Revealing that the NSDHL Gene Might Be Associated with the Progression of Western HFD/SW-Induced Hepatocellular Carcinoma.

BACKGROUND AND OBJECTIVE: Hepatocellular carcinoma (HCC) remains a significant global health concern. However, the etiology and pathogenesis of HCC have yet to be fully elucidated. Previous studies have indicated a close association between obesity and the occurrence and progression of HCC. The objective of this study was to employ bioinformatics strategies in order to explore key genes associated with the clinical diagnosis and prognosis of HCC induced by a Western high-fat diet and sugar water (HFD/SW). MATERIALS AND METHODS: We obtained the expression profile chip data GSE197884 from the Gene Expression Omnibus (GEO) database. Subsequently, “DESeq” and “Limma” R packages were employed to identify differentially expressed genes (DEGs) while constructing a co-expressed gene network using weighted gene co-expression analysis (WGCNA). Functional enrichment analyses were then carried out, followed by the construction of a protein-protein interaction (PPI) network to uncover core genes. The core genes were confirmed through data retrieved from The Cancer Genome Atlas (TCGA) database in order to determine their status as hub genes. Finally, survival and tumor immune infiltration analyses were performed to unveil the prognostic significance of these hub genes. RESULTS: In total, 126 intersection targets were retrieved through the Venn diagram. Gene ontology (GO) enrichment and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analyses revealed that the DEGs were primarily related to the proliferation and apoptosis of HCC cells, the digestion and metabolism of liver cells, the HCC tumor microenvironment, and immune response. The PPI network analysis identified 11 core targets, among which seven hub genes, including NSDHL, MVK, SQLW, GCAT, ALAS2, GLDC, and AGXT, were obtained after TCGA database validation. Furthermore, it was found that NSDHL was closely associated with the clinical diagnosis and prognosis of HCC induced by HFD/SW and also affected the cellular immune infiltration in the HCC tumor microenvironment. CONCLUSION: The present study demonstrated a significantly elevated expression of NSDHL in HCC tissues, suggesting its potential as a specific biomarker for precise clinical diagnosis and prognosis assessment of HCC induced by HFD/SW.

Computational Biology

ANXA3 hypomethylation as a prognostic biomarker in hepatitis B virus-related acute-on-chronic liver failure.

BACKGROUND: Hepatitis B virus-related acute-on-chronic liver failure (HBV-ACLF) is associated with a poor prognosis. This research aimed to characterize the expression pattern and clinical value of Annexin A3 (ANXA3) in HBV-ACLF patients. METHODS: First of all, ACLF-related datasets were downloaded from the Gene Expression Omnibus (GEO) database to carry out bioinformatics analyses. RT-qPCR, ELISA, and Methylight were used to measure ANXA3 gene expression and promoter methylation levels. A validation cohort was leveraged to further validate the results. RESULTS: Transcriptome analysis showed that ANXA3 was among the most differentially expressed genes when comparing dead patients with HBV-ACLF to those with survivors. The mRNA and serum levels of ANXA3 were elevated, and methylation levels were decreased in HBV-ACLF patients. The PMR value of ANXA3 in patients with HBV-ACLF was negatively correlated with inflammation-related cytokines IL-6, TNF-&#x3b1;, and IL-1&#x3b2;, as well as quantitative clinical parameters AST, TBIL, PT, INR, NEUT%, and MELD score, and positively correlated with PTA (all p&#x2009;<&#x2009;0.05). In HBV-ACLF patients, ANXA3 was considered to be an independent influence factor for the 90-day mortality. It was also found that ANXA3, especially hypomethylation, was associated with 28- and 90-day overall survival in patients with HBV-ACLF based on receiver operating characteristic (ROC) analysis, decision curve analysis (DCA), and Kaplan-Meier curves. CONCLUSIONS: ANXA3 hypomethylation has a prominent predictive value for short-term mortality in patients with HBV-ACLF and may serve as a promising biomarker of HBV-ACLF prognosis.

Humans

GBRAP: A Comprehensive Database and Tool for Exploring Genomic Diversity Across All Domains of Life.

Evolutionary studies require extensive examination of genomic information across all domains of life. Despite the availability of a large number of genomes through GenBank, the effective visualization or comparison of the information they contain is challenging due to many reasons, including their size. We introduce genome-based retrieval and analysis parser, a comprehensive software tool to analyze genome files, and an online database housing an extensive collection of carefully curated, high-quality genome statistics for all the organisms available in the RefSeq database of National Center for Biotechnology Information. Users can either directly search, or select from precategorized groups, the organisms of their choice and retrieve data, and the output is generated as tables containing more than 200 columns of useful genomic information (base counts, GC content, Shannon entropy, codon usage, etc.) separately calculated for different genomic elements (e.g. coding sequences, introns, transfer RNA, ribosomal RNA, noncoding RNA, etc.). The data are independently displayed (if applicable) for each chromosomal, mitochondrial, plastid, or plasmid sequence. All the data can be visualized on the database or downloaded as comma-separated value or Excel files. The genome-based retrieval and analysis parser database is free to access without any registration and is publicly available at http://tacclab.org/gbrap/.

Software

A novel start-loss mutation of the SLC29A3 gene in a consanguineous family with H syndrome: clinical characteristics, in silico analysis and literature review.

BACKGROUND: The SLC29A3 gene, which encodes a nucleoside transporter protein, is primarily located in intracellular membranes. The mutations in this gene can give rise to various clinical manifestations, including H syndrome, dysosteosclerosis, Faisalabad histiocytosis, and pigmented hypertrichosis with insulin-dependent diabetes. The aim of this study is to present two Iranian patients with H syndrome and to describe a novel start-loss mutation in SLC29A3 gene. METHODS: In this study, we employed whole-exome sequencing (WES) as a method to identify genetic variations that contribute to the development of H syndrome in a 16-year-old girl and her 8-year-old brother. These siblings were part of an Iranian family with consanguineous parents. To confirmed the pathogenicity of the identified variant, we utilized in-silico tools and cross-referenced various databases to confirm its novelty. Additionally, we conducted a co-segregation study and verified the presence of the variant in the parents of the affected patients through Sanger sequencing. RESULTS: In our study, we identified a novel start-loss mutation (c.2T&#x2009;>&#x2009;A, p.Met1Lys) in the SLC29A3 gene, which was found in both of two patients. Co-segregation analysis using Sanger sequencing confirmed that this variant was inherited from the parents. To evaluate the potential pathogenicity and novelty of this mutation, we consulted various databases. Additionally, we employed bioinformatics tools to predict the three-dimensional structure of the mutant SLC29A3 protein. These analyses were conducted with the aim of providing valuable insights into the functional implications of the identified mutation on the structure and function of the SLC29A3 protein. CONCLUSION: Our study contributes to the expanding body of evidence supporting the association between mutations in the SLC29A3 gene and H syndrome. The molecular analysis of diseases related to SLC29A3 is crucial in understanding the range of variability and raising awareness of H syndrome, with the ultimate goal of facilitating early diagnosis and appropriate treatment. The discovery of this novel biallelic variant in the probands further underscores the significance of utilizing genetic testing approaches, such as WES, as dependable diagnostic tools for individuals with this particular condition.

Humans

PICRUSt2-SC: an update to the reference database used for functional prediction within PICRUSt2.

SUMMARY: PICRUSt2 is a bioinformatic tool that predicts microbial functions in amplicon sequencing data using a database of annotated reference genomes. We have constructed an updated database for PICRUSt2 that has substantially increased the number of bacterial (19,493 to 26,868) and archaeal (406 to 1,002) genomes as well as the number of functional annotations present. The previous PICRUSt2 database relied on many timely and computationally intensive manual processes that made it difficult to update. We constructed a new streamlined process to allow regular upgrades to the PICRUSt2 database on an ongoing basis, and used this process to create a new database, PICRUSt2-SC (Sugar-Coated). Additionally, we have shown that this updated database contains genomes that more closely match study sequences from a range of different environments. The genomes contained in the database therefore better represent these environments and this leads to an improvement in the predicted functional annotations obtained from PICRUSt2. AVAILABILITY AND IMPLEMENTATION: PICRUSt2 source code is freely available at https://github.com/picrust/picrust2 and at https://anaconda.org/bioconda/picrust2. The latest version of PICRUSt2 at the time of writing is also archived: https://doi.org/10.5281/zenodo.15119781. The PICRUSt2-SC database comes pre-installed with PICRUSt2 from version 2.6.0 onwards. Step-by-step instructions for making the updated database are at https://github.com/picrust/picrust2/wiki/Updating-the-PICRUSt2-database. All code used for the analyses and figures in this manuscript is at https://github.com/R-Wright-1/PICRUSt2-SC_application_note and https://doi.org/10.5281/zenodo.15119770.

Software

Unveiling the Molecular Secrets of Seaweeds: A Comprehensive Review of Bioinformatics Applications in Algal Research.

Recent advances in high-throughput sequencing, bioinformatics, and multi-omics technologies have transformed seaweed research by overcoming long-standing challenges associated with complex genomes, diverse life cycles, and limited genomic resources. This review provides a comprehensive overview of bioinformatics approaches used to investigate seaweed genomics, transcriptomics, proteomics, metabolomics, microbiomes, and functional genomics, with emphasis on the computational tools and databases that support these analyses. Applications of bioinformatics in phylogenetics, drug discovery, microbiome characterization, and the development of biofuels, nutraceuticals, pharmaceuticals, and sustainable agriculture are also discussed. Particular attention is given to emerging strategies involving multi-omics integration, genome editing, artificial intelligence, machine learning, and synthetic biology that are reshaping seaweed research. The review further examines current challenges, including incomplete genomic resources, data standardization, and the need for experimental validation of computational predictions. Collectively, these advances highlight the growing role of bioinformatics in enabling systems-level understanding of seaweed biology and accelerating their translation into sustainable biotechnological and marine bioeconomy applications.

macroalgal genomics

Mining thermophile photosynthesis genes: a synthetic operon expressing Chloroflexota species reaction center genes in Rhodobacter sphaeroides.

Photosynthesis is the foundation of the vast majority of life systems, and therefore the most important bioenergetic process on earth, and the greatest diversity in photosynthetic systems are found in microorganisms. However, understanding of the biophysical and biochemical processes that transduce light to chemical energy has derived from the relatively small subset of proteins from microbes that are amenable to cultivation, in contrast to the huge number of microbial DNA sequences encoding proteins that catalyze the initial photochemical reactions that has been deposited in databases, such as from metagenomics. We describe the use of a Rhodobacter sphaeroides laboratory strain for expression of heterologous photosynthesis genes to demonstrate the feasibility of mining this resource, focusing on hot spring Chloroflexota gene sequences. Using a synthetic operon of genes, we produced a photochemically active complex of reaction center proteins in our biological system. We also present bioinformatic analyses of anoxygenic type II reaction center sequences from metagenomic samples collected from hot (42-90&#xb0; C) springs available through the JGI IMG database, to generate a resource of diverse sequences that potentially are adapted to photosynthesis at such temperatures. These data provide a view into the natural diversity of anoxygenic photosynthesis, through a lens focused on high-temperature environments. The approach we took to express such genes can be applied for potential biotechnology purposes as well as for studies of fundamental catalytic properties of these heretofore inaccessible protein complexes.

Chloroflexota

Distinct mutational landscapes for germline and somatic cancer variants in forty tumor suppressor genes.

Germline and somatic cancer variants in tumor suppressor genes (TSGs) share loss-of-function mechanisms, but studies of a few genes (DICER1 and CEBPA) have demonstrated differences in variant consequence and location. To systematically assess whether TSGs display distinct mutational patterns, we leveraged large public genetic databases and compared 32,941 high-quality pathogenic/likely pathogenic (P/LP) germline variants in ClinVar, with 12,907 oncogenic/likely oncogenic (O/LO) somatic tumor variants from cBioPortal across 40 TSGs. Only 3,863 (9.2%) variants were shared. Eighteen TSGs showed significantly different distributions of variant occurrences by molecular consequence, replicated with non-overlapping somatic data from the COSMIC database (chi-squared tests, false discovery rate = 5%). DICER1, TP53, and SMAD4 displayed excess somatic missense events, while nine TSGs (e.g., RB1 and APC) contained excess somatic stop-gain events throughout the coding sequence. Analysis by tumor type revealed excess stop-gain events in tissues exposed to environmental mutagens with corresponding mutation signatures. For several TSGs (WT1), germline variants predispose to tumors (Wilms' tumor) distinct from the majority source of somatic data (myeloid leukemia). Germline and somatic events are also distributed unevenly across cDNA locations, with 103 regions of preferential clustering in 39 TSGs (78 somatic and 25 germline). Twenty somatic clusters contained recurring frameshifts in homopolymer runs, many in tumors with microsatellite instability. Germline clusters contain more germline-exclusive variants, some driving non-cancer phenotypes reflecting genetic pleiotropy. Altogether, germline and somatic variants of TSGs represent unique sets with substantially different patterns shaped by selection pressures from gene-specific and somatic mutational mechanisms. Characterizing these distinctions enables more accurate clinical interpretation of TSG variants.

Humans