Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Deep learning-based annotation of plant abiotic stress resistance genes for crops.

The declining costs of DNA sequencing have expanded genomic data, crucial for understanding plant abiotic stress responses and crop improvement. However, accurate gene annotation remains challenging. To address this limitation, we propose the PASRGA, a deep learning approach that leverages transfer learning and contrastive learning to annotate genes related to drought, salt, cold, and UV resistance. PASRGA achieves high F1-scores, area under the receiver operating characteristic (AUROC), area under the precision-recall curve (AUPRC), and Matthews correlation coefficient (MCC) in annotating stress resistance genes, significantly outperforming the general protein annotation model CLEAN, the plant phosphatase gene annotation model PF-NET, the top-ranked model in the CAFA5 challenge NetGO 4.0, and four traditional machine learning methods. Its effectiveness was further validated with a salt stress treatment experiment in Eutrema salsugineum. To facilitate crop breeding practices, we utilized PASRGA to annotate the genomes of 17 major crops. To improve accessibility and utility, we incorporated both manually curated and PASRGA-predicted gene data, together with the PASRGA tool, into the PlantASRG database (https://bioinfor.nefu.edu.cn/PlantASRG/). This comprehensive resource aims to support crop breeding initiatives and ensure food security.

Crops, Agricultural↗

Tackling non-canonical splicing in arrhythmogenic cardiomyopathy to reduce the uncertain significance variants burden.

BACKGROUND: Splice-altering variants (SAVs), particularly those outside canonical splice sites, are an underappreciated contributor to inherited cardiovascular diseases. In arrhythmogenic cardiomyopathy (ACM), these variants frequently remain classified as of uncertain significance (VUS) due to limited predictive power and lack of transcript-level evidence, constraining genetic yield and clinical management. Our study aimed to determine the functional impact of SAVs in ACM genes and refine their classification using ACMG/AMP and ClinGen SVI criteria. METHODS: SAVs identified in 200 ACM probands underwent SpliceAI prediction, GTEx cardiac exon-usage annotation, and functional assessment using pSPL3-based minigene assays. Aberrant transcripts were quantified using Percent Splicing Alteration (PSA). Segregation data and ACMG/AMP criteria refined by ClinGen SVI were applied to integrate functional and clinical evidence for classification. RESULTS: Aberrant splicing was confirmed in 9/20 variants (45%), including synonymous, missense, and non-canonical intronic changes. SpliceAI scores correlated strongly with PSA values (R²=0.86). Case-control burden testing revealed significant enrichment of splice-altering variants in DSP, DSG2, DSC2 and FLNC. Integrating predictive algorithms with experimental validation and segregation analysis markedly enhances reclassification of 16/20 variants (80%). CONCLUSION: Splicing defects beyond canonical sites significantly shape ACM genetic landscape. Integrating predictive models with experimental validation clarifies uncertain variants bridging the gap between genomic uncertainty and clinical decision-making.

Humans↗

Predicting functions from protein sequences--where are the bottlenecks?

The exponential growth of sequence data does not necessarily lead to an increase in knowledge about the functions of genes and their products. Prediction of function using comparative sequence analysis is extremely powerful but, if not performed appropriately, may also lead to the creation and propagation of assignment errors. While current homology detection methods can cope with the data flow, the identification, verification and annotation of functional features need to be drastically improved.

Amino Acid Sequence↗

From fold to function.

A number of recent advances have been made in deriving function information from protein structure. A fold relationship to an already characterized protein will often allow general information about function to be deduced. More detailed information can be obtained using sequence relationships to already studied proteins. Methods of deducing function directly from structure, without the use of evolutionary relationships, are developing rapidly. All such methods may be used with models of protein structure, rather than with experimentally determined ones, but model accuracy imposes limitations. The rapid expansion of the structural genomics field has created a new urgency for improved methods of structure-based annotation of function.

Animals↗

CDACHIE: chromatin domain annotation by integrating chromatin interaction and epigenomic data with contrastive learning.

MOTIVATION: Chromatin domain annotation identifies functional genomic regions, such as active and inactive zones, based on epigenomic features like histone modifications, DNA methylation, and chromatin accessibility. While recent methods have utilized both chromatin interaction data (e.g. Hi-C) and epigenomic data, they often overlook the direct relationship between these data types. RESULTS: In this study, we introduce Chromatin Domain Annotation using Contrastive Learning for Hi-C and Epigenomic Data (CDACHIE), a method for identifying chromatin domains from Hi-C and epigenomic data. Our approach leverages contrastive learning to generate aligned representative vectors for both data types at each genomic bin. The concatenated vectors are then clustered using K-means to classify distinct chromatin domain types. CDACHIE achieves superior performance in Variance Explained, evaluated across gene expression, replication timing, and ChIA-PET data. This highlights its robust ability to integrate semantic associations between Hi-C and epigenomic features within the embedding space. AVAILABILITY AND IMPLEMENTATION: The source code is available at GitHub: https://github.com/maruyama-lab-design/CDACHIE. An archival snapshot of the code used in this study is available on Zenodo: https://doi.org/10.5281/zenodo.15751780.

Chromatin↗

Fishing for a reelGene: evaluating gene models with evolution and machine learning.

Assembled genomes and their associated annotations have transformed our study of gene function. However, each new annotated assembly generates new gene models. Inconsistencies between annotations likely arise from biological and technical causes, including pseudogene misclassification, transposon activity, and intron retention from sequencing of unspliced transcripts. To evaluate gene model predictions, we developed reelGene, a pipeline of machine learning models focused on (1) transcription boundaries, (2) mRNA integrity, and (3) protein structure. The first two models leverage sequence characteristics and evolutionary conservation across related taxa to learn the grammar of conserved transcription boundaries and mRNA sequences, while the third uses the conserved evolutionary grammar of protein sequences to predict whether a gene can produce a protein. Evaluating 1.8 million transcript models in Zea mays ssp. mays (maize), reelGene classified 28% as incorrectly annotated or non-functional. We find that reelGene classifies 92.2% of genes in the maize proteome and 99.2% of genes within the maize classical gene list as functional. reelGene also provides a way to further investigate genome biology- for instance, reelGene indicates that 10.3% of dispensable genes in B73 are functional, and within retained duplicate genes, reelGene identifies a 30% bias toward the retention of the M1 subgenome when one copy is functional and the other is non-functional. As an annotation-evaluating tool, reelGene is directly applicable to species of the Andropogoneae tribe, including other important crops like sorghum and miscanthus. As a community resource, reelGene has been integrated onto MaizeGDB both as a browser track and as an individual Shiny App, allowing researchers to evaluate gene model accuracy and further investigate genome biology.

Machine Learning↗

StrainMake: reproducible hybrid metagenomics with MAG recovery and strain-level resolution.

SUMMARY: Metagenomic workflows involve complex multi-step analyses, from quality control and assembly to binning, annotation, and strain-level profiling. Few existing metagenomic pipelines achieve the combination of flexibility, reproducibility, and hybrid assembly support within a unified workflow. We present StrainMake, a Snakemake-based workflow for de novo metagenomic analysis from short, long, or hybrid sequencing data. StrainMake integrates widely used tools across all major steps-quality control, assembly, binning, dereplication, taxonomic and functional annotation-while also providing non-redundant gene catalogues, community-scale metabolic models, and strain-level microdiversity metrics. The modular design enables the use of alternative tools, scalable execution on HPC systems, and full reproducibility through Snakemake and Conda. RESULTS: Applied to the CAMI II strain-madness dataset, StrainMake produced high-quality assemblies and metagenome-assembled genomes (MAGs), while enabling strain-resolved comparisons across samples. Hybrid assemblies improved contiguity, whereas short-read assemblies offered faster runtimes, illustrating the workflow's benchmarking capacity. AVAILABILITY AND IMPLEMENTATION: StrainMake is open source and available at https://github.com/UMMISCO/strainmake, together with comprehensive documentation. Generated data are deposited in Zenodo (doi: 10.5281/zenodo.16950162).

Metagenomics↗

Transcriptomic responses of Porphyrophora sophorae larvae during licorice root colonization reveal coordinated remodeling of translation, mitochondrial energy metabolism and defense-related genes.

BACKGROUND: Porphyrophora sophorae is a subterranean piercing-sucking scale insect that damages licorice (Glycyrrhiza uralensis) roots, but the molecular responses associated with larval root colonization remain insufficiently defined. METHODS: We compared non-parasitic larvae (NP) and root-colonizing larvae (RC) using six RNA-seq libraries, de novo transcriptome assembly, DESeq2-based differential expression analysis, GO/KEGG enrichment, annotation-based candidate gene screening, and RT-qPCR validation of selected genes. RESULTS: Sequencing yielded 260.91 million clean reads, and de novo assembly produced 60,794 non-redundant transcripts. DESeq2 identified 703 FDR-significant DEGs, including 49 upregulated and 654 downregulated genes in RC larvae. Upregulated genes were mainly associated with translation- and ribosome-related processes, whereas downregulated genes were enriched in mitochondrial, oxidation-reduction, energy metabolism, and oxidative phosphorylation-related functions. Annotation-based screening identified 75 FDR-significant candidate genes associated with chemosensation, defense-related responses, and energy metabolism, with mitochondrial energy metabolism-related genes forming the largest module. RT-qPCR validation based on the raw Ct data showed concordant expression directions for ten selected transcript targets. CONCLUSIONS: Root colonization in P. sophorae larvae was associated with coordinated transcriptional remodeling involving selective activation of translation-related processes, adjustment of mitochondrial energy metabolism, and changes in defense-related gene expression. These results provide candidate molecular targets for future functional studies of host contact, feeding establishment, and physiological adjustment in this subterranean scale insect.

Animals↗

Bioinformatics analysis of miR-2861 and miR-5011-5p that function as potential tumor suppressors in colorectal carcinogenesis.

BACKGROUND: The study aimed to was to investigate the relationship between miR-2861, miR-5011-5p, and colorectal carcinogenesis. METHOD: In the present study, it was isolated RNA from both the tumor and non-tumor tissue of a total of 80 CRC patients and after synthesizing the cDNA, it was performed qRT-PCR to determine the expression levels of miR‑2861 and miR‑5011-5p. In addition, it was predicted that dysregulated miRNAs targets, pathways and functional gene annotations that may be important in colorectal carcinogenesis using KEGG pathway and GO analysis. RESULTS: The resulting data revealed that both expression levels of miR-2861 and miR-5011-5p were significantly decreased in tumor tissues compared with non-tumor tissues of CRC patients. The GO and KEGG pathway analysis showed that miR-2861 and miR-5011-5p may participate in multiple the biological process, cellular components, and molecular function subcategories such as mitotic cell cycle, regulation of small GTPase mediated signal transduction, cell death, and acid binding transcription factor activity. It was also revealed that target genes of miRNAs can be found in signaling pathways such as TGF-beta, Rap1, Ras, cAMP, Wnt, mTOR and, PI3K-Akt signaling pathways. CONCLUSION: These findings imply that miR-2861 and miR-5011-5p might function as tumor suppressors in the development of CRC.

MicroRNAs↗

Structural genomics sheds light on protein functions and remote homologs across the insect tree of life.

Protein structure bridges the sequence-function relationship, enabling deep exploration of biological processes across diverse organisms. Insects, the most diverse animal lineage, accounting for over 50% of all described animal species, provide an exceptional system for exploring sequence-structure-function relationships. Here, we reconstructed a comprehensive and well-resolved phylogeny of 4854 insects, spanning all orders. Leveraging this framework, we created an atlas of 13.29 million predicted protein structures from 824 representative species, including 11.63 million newly predicted structures. Structural clustering revealed that proteins with divergent sequences but similar structures could be effectively grouped together. Structural similarity searches against proteins with well-characterized functions yielded annotations for 7.61 million insect proteins, including up to 14% of previously unannotated proteins. We further identified 750 million remote homologs between insect proteins, many of which trace back to ancient branches of the insect phylogeny. Remarkably, despite extensive sequence divergence, cGAS-like receptors (cGLRs) were structurally conserved across all 824 insects. Experimental assays demonstrated that these structurally identified cGLRs play a crucial role in antiviral defense in the yellow fever mosquito. Our findings highlight the significance of structural genomics for understanding protein function and evolution across the tree of life.

Animals↗

Chromosome-level genome assembly and annotation of Pterygoplichthys pardalis.

Suckermouth catfishes, with their evolved powerful features, have become notorious invasive species, causing significant damage to aquatic ecosystems. However, the lack of high-quality genomes severely restricts research on this group within the field. In this study, we de novo assembled the chromosome-level genome assembly of Pterygoplichthys pardalis using multiple platforms of sequencing data, including Illumina short reads, Nanopore long reads, and Hi-C sequencing reads, resulting in a 1.51 Gb genome assembly. Multiple evaluations, including read mapping ratio (98.52%), transcript mapping ratio (99.61%), conserved BUSCO gene set (98.8%), and N50 score (49.47 Mb), indicated the high continuity and accuracy of the genome assembly we generated. Genome annotation found that 0.97 Gb of genome sequences are repetitive sequences, accounting for 64.47% of the genome assembly. Further, 23,859 protein-coding genes were successfully predicted, 92.92% of which could be annotated in functional databases. This high-quality genome assembly of P. pardalis provides a valuable resource for understanding the genetic underpinnings of P. pardalis's invasive success and offers critical data for future fisheries research and management.

Animals↗

PICRUSt2-SC: an update to the reference database used for functional prediction within PICRUSt2.

SUMMARY: PICRUSt2 is a bioinformatic tool that predicts microbial functions in amplicon sequencing data using a database of annotated reference genomes. We have constructed an updated database for PICRUSt2 that has substantially increased the number of bacterial (19,493 to 26,868) and archaeal (406 to 1,002) genomes as well as the number of functional annotations present. The previous PICRUSt2 database relied on many timely and computationally intensive manual processes that made it difficult to update. We constructed a new streamlined process to allow regular upgrades to the PICRUSt2 database on an ongoing basis, and used this process to create a new database, PICRUSt2-SC (Sugar-Coated). Additionally, we have shown that this updated database contains genomes that more closely match study sequences from a range of different environments. The genomes contained in the database therefore better represent these environments and this leads to an improvement in the predicted functional annotations obtained from PICRUSt2. AVAILABILITY AND IMPLEMENTATION: PICRUSt2 source code is freely available at https://github.com/picrust/picrust2 and at https://anaconda.org/bioconda/picrust2. The latest version of PICRUSt2 at the time of writing is also archived: https://doi.org/10.5281/zenodo.15119781. The PICRUSt2-SC database comes pre-installed with PICRUSt2 from version 2.6.0 onwards. Step-by-step instructions for making the updated database are at https://github.com/picrust/picrust2/wiki/Updating-the-PICRUSt2-database. All code used for the analyses and figures in this manuscript is at https://github.com/R-Wright-1/PICRUSt2-SC_application_note and https://doi.org/10.5281/zenodo.15119770.

Software↗

Assessment of genomic prediction capabilities of transcriptome data in a barley multi-parent RIL population.

Low-cost and high-throughput RNA sequencing data for barley RILs achieved GP performance comparable to or better than traditional SNP array datasets when combined with parental whole-genome sequencing SNP data. The field of genomic selection (GS) is advancing rapidly on many fronts including the utilization of multi-omics datasets with the goal of increasing prediction ability and becoming an integral part of an increasing number of breeding programs ensuring future food security. In this study, we used RNA sequencing (RNA-Seq) data to perform genomic prediction (GP) on three related barley RIL populations. We investigated the potential of increasing prediction ability by combining genomic and transcriptomic datasets, adding whole-genome sequencing (WGS) SNP data, functional annotation-based filtering, and empirical quality filtering. Our RNA-Seq data were generated cost-efficiently using small-footprint plant cultivation, high-throughput RNA extraction, and Library preparation miniaturization. We also examined sequencing depth reduction as an additional cost-saving measure. We used fivefold cross-validation to evaluate the prediction ability of the gene expression dataset, the RNA-Seq SNP dataset, and the consensus SNP dataset between the RNA-Seq and parental WGS data, resulting in prediction abilities between 0.73 and 0.78. The consensus SNP dataset performed best, with five out of eight traits performing significantly better compared to a 50K SNP array, which served as a benchmark. The advantage of the consensus SNP dataset was most prominent in the inter-population predictions, in which the training and validation sets originated from different RIL sub-populations. We were therefore able to not only show that RNA-Seq data alone are able to predict various complex traits in barley using RILs, but also that the performance can be further increased with WGS data for which the public availability will steadily increase.

Hordeum↗

A multi-model genome-wide association study identifies genetic variants underlying resistance to Largemouth Bass Ranavirus (LMBV) in Micropterus salmoides.

Largemouth bass (Micropterus salmoides) is an economically important freshwater aquaculture species, yet recurrent outbreaks of Largemouth Bass Ranavirus (LMBV) continue to impair production and cause substantial losses. The genetic basis of host variation in LMBV resistance remains insufficiently characterized. Here, we applied a multi-model genome-wide association study (GWAS) to identify loci associated with resistance following a controlled challenge with the LMBV-23PY strain. Whole-genome resequencing was performed for 146 phenotyped fish, including 72 susceptible and 74 resistant individuals. After stringent quality control, 877,262 high-quality variants were retained and tested using six GWAS models. Across binary survival status and survival time phenotypes, 32 shared suggestive variants were consistently detected across models, representing suggestive loci for LMBV-23PY resistance. Genes within ±50 kb of these loci were annotated, and functional enrichment highlighted immune- and redox-related biological processes. Three prioritized candidates-GSTT3L (glutathione S-transferase theta-3-like), CGRP2 (calcitonin gene-related peptide 2), and NPPC (natriuretic peptide C)-were associated with pathways involved in oxidative stress responses and immune regulation. Collectively, these results provide insight into the genetic architecture of LMBV-23PY resistance in largemouth bass and identify suggestive variants and associated candidate genes for downstream validation, functional interrogation, and the development of marker-assisted and genome-enabled breeding strategies.

Animals↗

Cholesterol Metabolism-related Characteristics Predict Therapeutic Response and Survival in Esophageal Cancer.

INTRODUCTION: Cholesterol homeostasis has been identified as an essential downstream pathway of mutations in TP53. Esophageal cancer is one of the most prevalent malignancies exhibiting the mutation. OBJECTIVES: To explore the significance of cholesterol metabolism-related characteristics in tumor phenotype and treatment outcomes of esophageal cancer. METHODS: We established a cholesterol metabolism-related gene set (CMGs) and performed Lasso-Cox analysis to identify prognostic signatures. Nomogram-based risk scores and clinical stages afterwards were constructed and evaluated. We simultaneously identified two metabolic subtypes based on the distinct features of the CMGs. We annotated the functional and pathway characteristics of differentially expressed genes between the clusters and compared the differences in clinical and immune characteristics. Finally, we assessed the prognostic value of signatures in the GSE53625 and two clinical cohorts using whole-exon sequencing and multiplex immunofluorescence. RESULTS: Our study identified five cholesterol prognosis-related genes (CRGs) that demonstrated superior prognostic efficacy in the training set compared to clinical staging, validated in independent public databases and two clinical cohorts. According to the different expression patterns of the signatures, patients were divided into two subtypes. The C1 group demonstrated poorer overall survival, response to immunotherapy, and downregulation of the p53 pathway. In the immune correlation analysis, we found that the risk score based on 5-signature model was significantly positively correlated with the abundance of suppressive immune cells and the immune checkpoints. Finally, we explored the impact of expression and genomic polymorphism of the signatures on the prognosis at the pan-cancer level. CONCLUSIONS: Our findings underscore the distinct expression patterns of CRGs in esophageal cancer. These signatures are efficient to serve as prognostic indicators and assess the effectiveness of immunotherapy. They may also represent promising targets in other TP53 mutant malignancies.

Humans↗

Unraveling the genomic blueprint of the Indian black soldier fly: From genome assembly to evolutionary insights.

The black soldier fly (BSF) (Hermetia illucens) has been renowned for its sustainable bioconversion capabilities, resulting in smart protein production with wide applications in animal feed, bioenergy, and biofertilizer. However, the genetic mechanisms underlying efficient bioconversion and productivity remain poorly understood. To advance strain-specific applications and strengthen genetic resource availability, we present the whole genome sequencing (WGS) data for an Indian isolate of black soldier fly. The assembled genome was 1.46 Gb with a scaffold N50 of 172.7 Mb, and a GC content of 42.6%. Furthermore, 64.17% of genomic sequences were masked as repeated, and 14,317 protein-coding sequences were identified. Variant analysis against the reference genome identified 34.44 million variants (∼33.25 million SNPs and ∼ 1.18 million INDELs), with the majority (99.3%) classified as MODIFIER, 0.54% as LOW impact, 0.14% as MODERATE, and only 0.003% as HIGH impact. Comparative genomic analysis with other related species revealed expansions of gene families in BSF associated with Immune effector (Antimicrobial peptides (AMPs), Lysozymes, and Peptidoglycan Recognition Protein (PGRP) and Detoxification (cytochrome P450 enzymes). Notably, AMPs in the Indian isolate showed enhanced copy number variation in defensin (27) and PGRP (40) compared to reference BSF, suggesting potential regional adaptations to pathogen exposure. Collectively, this genomic data provides an improved resource for evolutionary studies, functional genomics, and targeted genetic improvement of BSF for sustainable bioconversion applications.

Comparative genomics↗

Selection system for genes encoding nuclear-targeted proteins.

Nuclear proteins have essential roles in cell proliferation and differentiation. We have developed a yeast selection system-the nuclear transportation trap (NTT)-to identify genes encoding nuclear transport signals. Both unknown and previously identified nuclear localization signals were identified from a human fetal brain cDNA library. The majority (75%) of the unknown proteins examined were exclusively localized to the nucleus in COS-7 cells. We propose that NTT is an efficient method for isolating cDNAs that encode nuclear targeted proteins that can be applied to the retrieval of novel nuclear proteins and to annotate gene function.

Amino Acid Sequence↗

OpTiles: an R package for adaptive tiling and methylation variability profiling.

SUMMARY: OpTiles is an R package that dynamically defines tiling windows based on the distribution of sequenced CpGs, addressing the limitations of traditional fixed-tiling approaches in targeted methylation datasets. By integrating CpG density with intra-region methylation variability, it provides a reliability metric and extended functionality for annotating, prioritizing, and interpreting complex methylation data. AVAILABILITY AND IMPLEMENTATION: OpTiles is implemented in R and source code is freely available at https://github.com/fhaive/OpTiles. Data are available on Zenodo at https://doi.org/10.5281/zenodo.16961292.

DNA Methylation↗