Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Unraveling the genomic blueprint of the Indian black soldier fly: From genome assembly to evolutionary insights.

The black soldier fly (BSF) (Hermetia illucens) has been renowned for its sustainable bioconversion capabilities, resulting in smart protein production with wide applications in animal feed, bioenergy, and biofertilizer. However, the genetic mechanisms underlying efficient bioconversion and productivity remain poorly understood. To advance strain-specific applications and strengthen genetic resource availability, we present the whole genome sequencing (WGS) data for an Indian isolate of black soldier fly. The assembled genome was 1.46 Gb with a scaffold N50 of 172.7 Mb, and a GC content of 42.6%. Furthermore, 64.17% of genomic sequences were masked as repeated, and 14,317 protein-coding sequences were identified. Variant analysis against the reference genome identified 34.44 million variants (∼33.25 million SNPs and ∼ 1.18 million INDELs), with the majority (99.3%) classified as MODIFIER, 0.54% as LOW impact, 0.14% as MODERATE, and only 0.003% as HIGH impact. Comparative genomic analysis with other related species revealed expansions of gene families in BSF associated with Immune effector (Antimicrobial peptides (AMPs), Lysozymes, and Peptidoglycan Recognition Protein (PGRP) and Detoxification (cytochrome P450 enzymes). Notably, AMPs in the Indian isolate showed enhanced copy number variation in defensin (27) and PGRP (40) compared to reference BSF, suggesting potential regional adaptations to pathogen exposure. Collectively, this genomic data provides an improved resource for evolutionary studies, functional genomics, and targeted genetic improvement of BSF for sustainable bioconversion applications.

Comparative genomics↗

Selection system for genes encoding nuclear-targeted proteins.

Nuclear proteins have essential roles in cell proliferation and differentiation. We have developed a yeast selection system-the nuclear transportation trap (NTT)-to identify genes encoding nuclear transport signals. Both unknown and previously identified nuclear localization signals were identified from a human fetal brain cDNA library. The majority (75%) of the unknown proteins examined were exclusively localized to the nucleus in COS-7 cells. We propose that NTT is an efficient method for isolating cDNAs that encode nuclear targeted proteins that can be applied to the retrieval of novel nuclear proteins and to annotate gene function.

Amino Acid Sequence↗

OpTiles: an R package for adaptive tiling and methylation variability profiling.

SUMMARY: OpTiles is an R package that dynamically defines tiling windows based on the distribution of sequenced CpGs, addressing the limitations of traditional fixed-tiling approaches in targeted methylation datasets. By integrating CpG density with intra-region methylation variability, it provides a reliability metric and extended functionality for annotating, prioritizing, and interpreting complex methylation data. AVAILABILITY AND IMPLEMENTATION: OpTiles is implemented in R and source code is freely available at https://github.com/fhaive/OpTiles. Data are available on Zenodo at https://doi.org/10.5281/zenodo.16961292.

DNA Methylation↗

The SBASE protein domain library, release 7.0: a collection of annotated protein sequence segments.

SBASE 7.0 is the seventh release of the SBASE protein domain library sequences that contains 237 937 annotated structural, functional, ligand-binding and topogenic segments of proteins, cross-referenced to all major sequence databases and sequence pattern collections. The entries are clustered into over 1811 groups and are provided with two WWW-based search facilities for on-line use. SBASE 7.0 is freely available by anonymous 'ftp' file transfer from ftp.icgeb. trieste.it. Automated searching of SBASE with BLAST can be carried out with the WWW servers http://www.icgeb.trieste.it/sbase/and http://sbase.abc.hu/sbase/

Amino Acid Sequence↗

The InterPro database, an integrated documentation resource for protein families, domains and functional sites.

Signature databases are vital tools for identifying distant relationships in novel sequences and hence for inferring protein function. InterPro is an integrated documentation resource for protein families, domains and functional sites, which amalgamates the efforts of the PROSITE, PRINTS, Pfam and ProDom database projects. Each InterPro entry includes a functional description, annotation, literature references and links back to the relevant member database(s). Release 2.0 of InterPro (October 2000) contains over 3000 entries, representing families, domains, repeats and sites of post-translational modification encoded by a total of 6804 different regular expressions, profiles, fingerprints and Hidden Markov Models. Each InterPro entry lists all the matches against SWISS-PROT and TrEMBL (more than 1,000,000 hits from 462,500 proteins in SWISS-PROT and TrEMBL). The database is accessible for text- and sequence-based searches at http://www.ebi.ac.uk/interpro/. Questions can be emailed to interhelp@ebi.ac.uk.

Databases, Factual↗

The SBASE protein domain library, release 8.0: a collection of annotated protein sequence segments.

SBASE 8.0 is the eighth release of the SBASE library of protein domain sequences that contains 294 898 annotated structural, functional, ligand-binding and topogenic segments of proteins, cross-referenced to most major sequence databases and sequence pattern collections. The entries are clustered into over 2005 statistically validated domain groups (SBASE-A) and 595 non-validated groups (SBASE-B), provided with several WWW-based search and browsing facilities for online use. A domain-search facility was developed, based on non-parametric pattern recognition methods, including artificial neural networks. SBASE 8.0 is freely available by anonymous 'ftp' file transfer from ftp.icgeb.trieste.it. Automated searching of SBASE can be carried out with the WWW servers http://www.icgeb.trieste.it/sbase/ and http://sbase.abc. hu/sbase/.

Binding Sites↗

VicMAG, an open-source tool for visualizing circular metagenome-assembled genomes highlighting bacterial virulence and antimicrobial resistance.

Bacterial pathogens spread in clinical and environmental settings, and mobile genetic elements (MGEs), such as plasmids and phages, mediate the transfer of virulence factor genes (VFGs) and antimicrobial resistance genes (ARGs) among bacterial communities. Metagenomic analysis of environmental and wastewater samples using highly accurate long-read sequencing technologies, such as Pacific Biosciences (PacBio) HiFi sequencing, provides valuable insights into monitoring the regional spread of VFGs and ARGs, including dissemination mediated by MGEs. No visualization tool is currently available for the comprehensive display of numerous resulting circular metagenome-assembled genomes (cMAGs) with functional gene annotations. Here, we developed visualization of circular metagenome-assembled genome (VicMAG), a visualization tool for highly complex cMAGs derived from long-read metagenome assemblies annotated using updated databases of VFGs, ARGs, and MGEs. Using 353 cMAGs from PacBio HiFi sequencing of a wastewater sample, we demonstrated the utility of VicMAG for metagenome visualization. VicMAG provides comprehensive, size-aware visualization of cMAGs representing bacterial chromosomes and plasmids, annotated with VFGs, ARGs, and phages. By simultaneously visualizing all cMAGs in a framework, VicMAG facilitates a holistic understanding of the distribution and genomic context of VFGs and ARGs across complex microbial communities. This tool supports integrated surveillance of bacteria associated with virulence and antimicrobial resistance across clinical, environmental, and One Health contexts.

Metagenome↗

Gene and pathway analysis of genome-wide genetic associations of bladder cancer.

BACKGROUND: Although genetic variants associated with bladder cancer (BCa) risk have been identified through hypothesis-driven and genome-wide association studies, a systematic understanding of BCa genetic susceptibility at the gene and pathway levels remains to be achieved. MATERIALS AND METHODS: In this 2-stage functional genomics study, we used 5 independent tools for genome-wide gene mapping and ranking based on BCa genome-wide association studies summary statistics, followed by a meta-analysis of gene-level significance p values, to obtain a consensus gene ranking in terms of association with BCa. Subsequently, we performed preranked gene-set enrichment analysis to identify the functional pathways involved in BCa genetic susceptibility. Joint analysis with gene-set enrichment analysis, based on somatic alteration frequency, was performed to explore the pathway-level relationships between genetic susceptibility and somatic alterations in BCa. RESULTS: Other than the well-known BCa genes (such as FGFR3, MYC, TERT, CCNE1, and TP63), we additionally prioritized a set of novel genes likely to be genetically implicated in BCa development, including SETD2, a possible tumor suppressor gene involved in chromatin remodeling. We further demonstrated convergence between genetic associations and somatic alterations at both the gene (eg, FGFR3 and TERT) and pathway levels (eg, cell cycle and chromatin modification), as well as functional ontologies specifically implicated in germline predisposition to BCa (eg, CD8/TCR signaling, immune checkpoints, and cytokine signaling). CONCLUSIONS: We identified several novel genes associated with BCa and demonstrated that genetic variants contribute to the development of BCa by affecting antitumor immunity, response to toxic exposure, and RNA and protein homeostasis and synergizing with somatic alterations in various cancer-related pathways.

Bladder cancer↗

Moderate expression and activity of flocculins underlie the characteristic flocculation phenotype of Saccharomyces pastorianus.

Flocculation is a key technological trait in lager brewing, governing fermentation performance, yeast recovery, and beer quality. In the allo-aneuploid hybrid yeast Saccharomyces pastorianus, the genetic basis of flocculation remains poorly resolved due to its complex dual sub-genome architecture. Here, we systematically re-annotated and functionally characterized the complete FLO gene repertoire of the Group II strain CBS 1483. Thirteen FLO genes were identified, including allelic variants and a previously uncharacterized adhesin, Flo12, containing a Hyphal_reg_CWP domain instead of the canonical PA14 lectin-binding domain. Structural modeling revealed strong conservation of Ca²+-binding residues in PA14 domains, alongside repeat-region diversification likely contributing to functional variability. Using optogenetic expression in a FLO-null background, we demonstrated that SpcI-FLO9-1 and SpcI-FLO9-2_1 are the strongest drivers of flocculation, exhibiting NewFlo-like sugar sensitivity. Transcriptomic analysis during 17°P wort fermentation showed dynamic induction of these genes coinciding with flocculation onset. Surprisingly, deletion of both loci in CBS 1483 did not abolish but only delayed sedimentation in wort, accompanied by improved maltose utilization and attenuation. These findings reveal functional redundancy and compensatory mechanisms within the FLO network of lager yeast, highlighting the genetic complexity underlying flocculation, and providing a molecular framework to inform yeast selection, strain development, and optimization of the lager fermentation processes.IMPORTANCEFlocculation, the process by which yeast cells aggregate and settle, is essential for producing clear, high-quality lager beer, and for efficient yeast recovery during brewing. However, the genetic basis of this trait in lager yeast has remained poorly understood because these strains possess unusually complex hybrid genomes. In this study, we systematically identified and characterized the complete set of flocculation genes in the industrial lager yeast Saccharomyces pastorianus CBS 1483. We demonstrated that lager yeast flocculation is not controlled by a single dominant gene, but instead emerges from the combined action of several moderately active adhesion proteins that are expressed at low levels during fermentation. Surprisingly, deleting the two strongest candidate genes only delayed, rather than eliminated, sedimentation, revealing a robust compensatory network that preserves brewing performance. These findings refine the current understanding of yeast flocculation and provide a molecular framework for developing brewing strains with improved fermentation efficiency, product consistency, and flavor quality.

Saccharomyces pastorianus↗

Multi-omics analysis identifies key genes and functional loci affecting teat number in American Large White and Landrace pigs and their application in optimizing genomic selection models.

BACKGROUND: Teat number is a crucial economic trait in pigs. It directly affects the ability of sows to lactate, which in turn influences the survival and health of piglets. The teat number of French Large White pigs is close to 16, while the teat number of American Large White and Landrace pigs is about 14. In order to improve the teat number of American Landrace and Large White pigs through molecular approaches and precise breeding techniques, we genotyped 2,131 American Landrace and 4,564 American Large White with teat number phenotype using a 50 K SNP chip. Then, the SNP-chip data was imputed to the level of whole-genome sequencing (iWGS). Based on iWGS data, we conducted GWAS to identify novel, significant SNPs associated with teat number and to incorporate them into genomic selection. RESULTS: In Landrace pigs, significant SNPs for TTN mapped to SSC2, SSC7, SSC8, and SSC14; the SSC8 and SSC14 effects are novel. LTN mapped to SSC7, RTN to SSC7 and SSC8. The lead SSC7 SNP explained 2.60% of TTN phenotypic variance. In Large White pigs, significant SNPs were detected on SSC7 and SSC10 for TTN; SSC7, SSC10, and SSC12 for LTN; and SSC7 and SSC10 for RTN. The most significant locus on SSC7 accounted for 2.99% of the phenotypic variance in TTN. Additionally, a multi-population meta-analysis detected significant novel SNPs for LTN on SSC1 and SSC8. By utilizing Bayesian fine mapping, the most precise QTL confidence interval on SSC7 for both TTN and RTN in Large White pigs was reduced to 40 kb. By integrating functional gene annotation with RNA-seq and ATAC-seq data from Erhualian and Bamaxiang pigs mammary placodes at embryonic day 26, we prioritized PTPN13, TRPV3, ZDHHC13, and BRD2 as novel candidate genes for teat number. We then incorporated the significant SNPs to GBLUP and benchmarked genomic-selection accuracy. In both breeds, fitting the top SNP as fixed maximized prediction for TTN and RTN, whereas treating all significant loci as an additional random effect optimized LTN. CONCLUSIONS: Our findings provide a theoretical basis for dissecting new key genes affecting teat number and for advancing molecular breeding of teat number in pigs.

Animals↗

Genomic distribution characteristics and interspecific differences of microsatellite landscapes in Felidae.

BACKGROUND: Microsatellites within genomes play crucial roles in regulating gene expression, DNA replication, and chromosomal structure and function. Analyzing the composition and distribution patterns of microsatellites in closely related species not only reveals their evolutionary dynamics and adaptive mechanisms but also provides essential technical support for applications in genetic breeding, species conservation, and disease research. As one of the world's most captivating animal groups, the landscape patterns of microsatellites across feline genomes remain to be systematically characterized. RESULTS: This study utilized high-quality genomic data to conduct a systematic comparative analysis of microsatellite landscape distribution patterns across the genomes of 13 felid species. The findings revealed that microsatellite abundance and distribution exhibit species-specific characteristics, with a non-random genomic distribution and a negative correlation between microsatellite abundance and repeat length. The predominant distribution pattern followed the sequence: single > double > quadruple > triple > quintuple > sextuple nucleotide repeats. Microsatellite abundance peaked in intergenic regions, whereas trinucleotide repeats were more prevalent within exons. Coding regions showed a marked preference for trinucleotide and hexanucleotide repeats. Enrichment analysis of GO and KEGG pathways indicated that coding sequences containing microsatellites were primarily involved in transcription and translation processes. CONCLUSIONS: Our study elucidates the distribution patterns and characteristics of microsatellites across diverse feline species, providing significant insights into their evolutionary mechanisms and functional roles. Furthermore, these findings establish a valuable reference and foundational dataset for the future development of high-quality, species-specific microsatellite markers in felids.

Animals↗

eQTM (expression quantitative trait methylation) Atlas: a comprehensive resource of over 11 million DNA methylation-gene expression associations through across 11 tissues and 4 diseases.

MOTIVATION: Epigenome-wide association studies (EWAS) have identified numerous DNA methylation (DNAm) CpG sites associated with complex traits and diseases, but interpretation of those CpG sites remains challenging because in EWAS, CpGs are mostly linked to nearby genes based only on genomic proximity. Expression quantitative trait methylation (eQTM) analyses connect DNAm CpGs with statistically associated gene expression levels. However, a comprehensive, searchable resource integrating eQTMs across diverse tissues and disease contexts has been lacking. RESULTS: We developed the eQTM Atlas, a web-based resource that manually curates more than 11 million DNAm-gene expression associations from eight cohorts, covering 11 tissue types, four broad disease contexts, 173,886 unique CpG probes and 20,231 unique genes. The Atlas supports gene- or CpG- searches by tissue or disease type and finding associated CpG or genes, visualization of cis- and trans-eQTMs through genome browser, heatmap interfaces across various tissues, and cohort-level data downloads. By integrating eQTM results with EWAS resources, the eQTM Atlas enables users to connect disease- or trait-associated CpGs to statistically associated genes rather than relying solely on proximity-based gene annotation, supporting functional interpretation of EWAS findings and generation of disease-specific regulatory hypotheses. AVAILABILITY AND IMPLEMENTATION: The eQTM Atlas is freely available at https://shiny.crc.pitt.edu/eqtm_browser/. The web interface is implemented in R Shiny and hosted through the University of Pittsburgh Center for Research Computing (CRC). Source code is available at https://github.com/ads303/eQTM-Atlas.

DNA methylation↗

Re-annotating the Mycoplasma pneumoniae genome sequence: adding value, function and reading frames.

Four years after the original sequence submission, we have re-annotated the genome of Mycoplasma pneumoniae to incorporate novel data. The total number of ORFss has been increased from 677 to 688 (10 new proteins were predicted in intergenic regions, two further were newly identified by mass spectrometry and one protein ORF was dismissed) and the number of RNAs from 39 to 42 genes. For 19 of the now 35 tRNAs and for six other functional RNAs the exact genome positions were re-annotated and two new tRNA(Leu) and a small 200 nt RNA were identified. Sixteen protein reading frames were extended and eight shortened. For each ORF a consistent annotation vocabulary has been introduced. Annotation reasoning, annotation categories and comparisons to other published data on M.pneumoniae functional assignments are given. Experimental evidence includes 2-dimensional gel electrophoresis in combination with mass spectrometry as well as gene expression data from this study. Compared to the original annotation, we increased the number of proteins with predicted functional features from 349 to 458. The increase includes 36 new predictions and 73 protein assignments confirmed by the published literature. Furthermore, there are 23 reductions and 30 additions with respect to the previous annotation. mRNA expression data support transcription of 184 of the functionally unassigned reading frames.

Amino Acid Sequence↗

Shared genetic architecture of obesity and gastroesophageal reflux disease.

Obesity is identified as a risk factor of gastroesophageal reflux disease (GERD). This study aims to elucidate the shared genetic architecture of obesity-related phenotypes and GERD. Based on the publicly available genome-wide association studies' datasets, this genome-wide pleiotropic association study was conducted with various genetic approaches (including linkage disequilibrium score regression, high-definition likelihood inference for genetic correlations, pleiotropic analysis under composite null hypothesis, Functional Mapping and Annotation, Bayesian colocalization, summary-based Mendelian randomization, and multi-marker analysis of genomic annotation analysis) sequentially to unravel the genetic associations from single-nucleotide polymorphism to gene levels, and to reveal the underlying shared genetic architecture between obesity-related phenotypes and GERD. This study discovered shared genetic mechanisms between GERD and several obesity-related phenotypes, including arm fat percentage (left), arm fat percentage (right), leg fat percentage (left), leg fat percentage (right), trunk fat percentage, waist-to-hip ratio, and body mass index. Significant genetic correlations were observed by linkage disequilibrium score regression and high-definition likelihood inference for genetic correlations, with multiple associated pleiotropic loci and their mapped genes identified by pleiotropic analysis under composite null hypothesis, Functional Mapping and Annotation, Bayesian colocalization, summary-based Mendelian randomization, and multi-marker analysis of genomic annotation analysis. Additionally, several brain tissues were identified to be linked to both obesity and GERD by multi-marker analysis of genomic annotation. This research provided strong evidence of genetic correlations and brought novel insights into the underlying genetic connections and shared genetic architectures of obesity and GERD.

Humans↗

META-DIFF: a k-mer-based pipeline that detects differentially abundant sequences in metagenomics whole genome sequencing.

Traditional case-control metagenomic studies are constrained by their dependence on taxonomic and functional databases. Because annotation occurs before differential analysis, they are limited to known elements and keep function and taxonomy separate. Although binning strategies have emerged to reconstruct genomes and mitigate this issue, they still require an assembly step, preventing the use of all available sequencing data. Here, we introduce META-DIFF, a pipeline based on differentially abundant k-mers independently of any prior annotation. From those k-mers, it reconstructs longer sequences and provides biological context, as well as the best set of unitigs to discriminate between conditions. Across both taxonomy-centric and functionally-centric benchmarks, it showed robust performance and displayed great reproducibility. It also behaved more conservatively than did other univariate methodologies, i.e. it maintained a high precision at the expense of recall, particularly in conditions of low fold-change and limited sequencing depth. The efficacy of META-DIFF was further validated through its application to a real-world colorectal cancer dataset, which produced both confirmatory and novel results compared with those of previous publications. The pipeline is able to exploit all reads and identify differentially abundant elements, including unknown DNA, prior to annotation. With the guidelines provided, META-DIFF provides users with great exploratory power to unravel microbiome changes.

Metagenomics↗

Intrinsic errors in genome annotation.

Genome sequencing is usually followed by routine annotation of protein function based on the assumption that similar sequences will have similar functions. Here, we introduce a simple calculation to estimate the magnitude of any possible annotation errors. We counted the number of discrepancies in the annotation of well-established sets of similar proteins and extrapolated these values to the pairs of similar sequences used for the annotation of different microbial genomes. We conclude that the number of potential errors in the prediction of detailed functions is higher than is usually believed.

Binding Sites↗

The CATH Dictionary of Homologous Superfamilies (DHS): a consensus approach for identifying distant structural homologues.

A consensus approach has been developed for identifying distant structural homologues. This is based on the CATH Dictionary of Homologous Superfamilies (DHS), a database of validated multiple structural alignments annotated with consensus functional information for evolutionary protein superfamilies (URL: http://www. biochem.ucl.ac.uk/bsm/dhs). Multiple structural alignments have been generated for 362 well-populated superfamilies in the CATH structural domain database and annotated with secondary structure, physicochemical properties, functional sequence patterns and protein-ligand interaction data. Consensus functional information for each superfamily includes descriptions and keywords extracted from SWISS-PROT and the ENZYME database. The Dictionary provides a powerful resource to validate, examine and visualize key structural and functional features of each homologous superfamily. The value of the DHS, for assessing functional variability and identifying distant evolutionary relationships, is illustrated using the pyridoxal-5'-phosphate (PLP) binding aspartate aminotransferase superfamily. The DHS also provides a tool for examining sequence-structure relationships for proteins within each fold group.

Amino Acid Sequence↗

Assigning genomic sequences to CATH.

We report the latest release (version 1.6) of the CATH protein domains database (http://www.biochem.ucl. ac.uk/bsm/cath ). This is a hierarchical classification of 18 577 domains into evolutionary families and structural groupings. We have identified 1028 homo-logous superfamilies in which the proteins have both structural, and sequence or functional similarity. These can be further clustered into 672 fold groups and 35 distinct architectures. Recent developments of the database include the generation of 3D templates for recognising structural relatives in each fold group, which has led to significant improvements in the speed and accuracy of updating the database and also means that less manual validation is required. We also report the establishment of the CATH-PFDB (Protein Family Database), which associates 1D sequences with the 3D homologous superfamilies. Sequences showing identifiable homology to entries in CATH have been extracted from GenBank using PSI-BLAST. A CATH-PSIBLAST server has been established, which allows you to scan a new sequence against the database. The CATH Dictionary of Homologous Superfamilies (DHS), which contains validated multiple structural alignments annotated with consensus functional information for evolutionary protein superfamilies, has been updated to include annotations associated with sequence relatives identified in GenBank. The DHS is a powerful tool for considering the variation of functional properties within a given CATH superfamily and in deciding what functional properties may be reliably inherited by a newly identified relative.

Amino Acid Sequence↗