Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Assigning genomic sequences to CATH.

We report the latest release (version 1.6) of the CATH protein domains database (http://www.biochem.ucl. ac.uk/bsm/cath ). This is a hierarchical classification of 18 577 domains into evolutionary families and structural groupings. We have identified 1028 homo-logous superfamilies in which the proteins have both structural, and sequence or functional similarity. These can be further clustered into 672 fold groups and 35 distinct architectures. Recent developments of the database include the generation of 3D templates for recognising structural relatives in each fold group, which has led to significant improvements in the speed and accuracy of updating the database and also means that less manual validation is required. We also report the establishment of the CATH-PFDB (Protein Family Database), which associates 1D sequences with the 3D homologous superfamilies. Sequences showing identifiable homology to entries in CATH have been extracted from GenBank using PSI-BLAST. A CATH-PSIBLAST server has been established, which allows you to scan a new sequence against the database. The CATH Dictionary of Homologous Superfamilies (DHS), which contains validated multiple structural alignments annotated with consensus functional information for evolutionary protein superfamilies, has been updated to include annotations associated with sequence relatives identified in GenBank. The DHS is a powerful tool for considering the variation of functional properties within a given CATH superfamily and in deciding what functional properties may be reliably inherited by a newly identified relative.

Amino Acid Sequence↗

Cis-regulatory elements: systematic identification and horticultural applications.

Cis-regulatory elements (CREs) are the genetic DNA fragments bound by transcription factors (TFs). CREs function as molecular switches that precisely modulate the dosage and spatiotemporal patterns of gene expression. The systematic identification of CREs not only facilitates the annotation of the functional non-coding genome but also provides essential insights into the architecture of gene regulatory networks and sheds light on an accurate selection of the target sites for genetic engineering of crops. In this review, we summarize the current high-throughput methodologies used for identifying CREs, illustrate the associations between CREs and agronomic traits in horticultural crops, and discuss how CREs can be exploited to facilitate crop breeding.

Breeding↗

KEGG: kyoto encyclopedia of genes and genomes.

KEGG (Kyoto Encyclopedia of Genes and Genomes) is a knowledge base for systematic analysis of gene functions, linking genomic information with higher order functional information. The genomic information is stored in the GENES database, which is a collection of gene catalogs for all the completely sequenced genomes and some partial genomes with up-to-date annotation of gene functions. The higher order functional information is stored in the PATHWAY database, which contains graphical representations of cellular processes, such as metabolism, membrane transport, signal transduction and cell cycle. The PATHWAY database is supplemented by a set of ortholog group tables for the information about conserved subpathways (pathway motifs), which are often encoded by positionally coupled genes on the chromosome and which are especially useful in predicting gene functions. A third database in KEGG is LIGAND for the information about chemical compounds, enzyme molecules and enzymatic reactions. KEGG provides Java graphics tools for browsing genome maps, comparing two genome maps and manipulating expression maps, as well as computational tools for sequence comparison, graph comparison and path computation. The KEGG databases are daily updated and made freely available (http://www. genome.ad.jp/kegg/).

Animals↗

A functional update of the Escherichia coli K-12 genome.

BACKGROUND: Since the genome of Escherichia coli K-12 was initially annotated in 1997, additional functional information based on biological characterization and functions of sequence-similar proteins has become available. On the basis of this new information, an updated version of the annotated chromosome has been generated. RESULTS: The E. coli K-12 chromosome is currently represented by 4,401 genes encoding 116 RNAs and 4,285 proteins. The boundaries of the genes identified in the GenBank Accession U00096 were used. Some protein-coding sequences are compound and encode multimodular proteins. The coding sequences (CDSs) are represented by modules (protein elements of at least 100 amino acids with biological activity and independent evolutionary history). There are 4,616 identified modules in the 4,285 proteins. Of these, 48.9% have been characterized, 29.5% have an imputed function, 2.1% have a phenotype and 19.5% have no function assignment. Only 7% of the modules appear unique to E. coli, and this number is expected to be reduced as more genome data becomes available. The imputed functions were assigned on the basis of manual evaluation of functions predicted by BLAST and DARWIN analyses and by the MAGPIE genome annotation system. CONCLUSIONS: Much knowledge has been gained about functions encoded by the E. coli K-12 genome since the 1997 annotation was published. The data presented here should be useful for analysis of E. coli gene products as well as gene products encoded by other genomes.

Bacterial Proteins↗

Phylogeny of related functions: the case of polyamine biosynthetic enzymes.

Genome annotation requires explicit identification of gene function. This task frequently uses protein sequence alignments with examples having a known function. Genetic drift, co-evolution of subunits in protein complexes and a variety of other constraints interfere with the relevance of alignments. Using a specific class of proteins, it is shown that a simple data analysis approach can help solve some of the problems posed. The origin of ureohydrolases has been explored by comparing sequence similarity trees, maximizing amino acid alignment conservation. The trees separate agmatinases from arginases but suggest the presence of unknown biases responsible for unexpected positions of some enzymes. Using factorial correspondence analysis, a distance tree between sequences was established, comparing regions with gaps in the alignments. The gap tree gives a consistent picture of functional kinship, perhaps reflecting some aspects of phylogeny, with a clear domain of enzymes encoding two types of ureohydrolases (agmatinases and arginases) and activities related to, but different from ureohydrolases. Several annotated genes appeared to correspond to a wrong assignment if the trees were significant. They were cloned and their products expressed and identified biochemically. This substantiated the validity of the gap tree. Its organization suggests a very ancient origin of ureohydrolases. Some enzymes of eukaryotic origin are spread throughout the arginase part of the trees: they might have been derived from the genes found in the early symbiotic bacteria that became the organelles. They were transferred to the nucleus when symbiotic genes had to escape Muller's ratchet. This work also shows that arginases and agmatinases share the same two manganese-ion-binding sites and exhibit only subtle differences that can be accounted for knowing the three-dimensional structure of arginases. In the absence of explicit biochemical data, extreme caution is needed when annotating genes having similarities to ureohydrolases.

Amino Acid Sequence↗

Large-scale benchmarking of prokaryotic annotation tools across thousands of species.

BACKGROUND: Genome annotation is an important step in deriving functional meaning from prokaryotic sequencing data, yet systematic evaluations guiding tool selection are lacking. We present the first large-scale investigation of four prominent open-source annotation tools (Prokka, Bakta, EggNOG-mapper, and PGAP) across 156,033 diverse genomes. This includes Escherichia coli strains for baseline performance, thousands of archaea and bacteria genomes, as well as frameshifted and metagenome-assembled genomes. RESULTS: Bakta excels in annotating high-quality bacterial genomes, while PGAP was better for archaeal genomes and challenging bacterial assemblies, including metagenome-assembled, fragmented, or contaminated samples. For Gene Ontology annotation, PGAP consistently provides broader term coverage, whereas EggNOG-mapper offers more terms per feature. CONCLUSIONS: Our findings highlight tool-specific strengths crucial for selecting optimal solutions based on genome quality, taxonomy, and origin (e.g. MAGs). This study provides an evidence-based guide for users and informs future tool development.

Molecular Sequence Annotation↗

Identifying the 3'-terminal exon in human DNA.

MOTIVATION: We present JTEF, a new program for finding 3' terminal exons in human DNA sequences. This program is based on quadratic discriminant analysis, a standard non-linear statistical pattern recognition method. The quadratic discriminant functions used for building the algorithm were trained on a set of 3' terminal exons of type 3tuexon (those containing the true STOP codon). RESULTS: We showed that the average predictive accuracy of JTEF is higher than the presently available best programs (GenScan and Genemark.hmm) based on a test set of 65 human DNA sequences with 121 genes. In particular JTEF performs well on larger genomic contigs containing multiple genes and significant amounts of intergenic DNA. It will become a valuable tool for genome annotation and gene functional studies. AVAILABILITY: JTEF is available free for academic users on request from ftp://cshl.org/pub/science/mzhanglab/JTEF and will be made available through the World Wide Web (http://argon.cshl.org/).

Algorithms↗

Gene expression changes triggered by exposure of Haemophilus influenzae to novobiocin or ciprofloxacin: combined transcription and translation analysis.

The responses of Haemophilus influenzae to DNA gyrase inhibitors were analyzed at the transcriptional and the translational level. High-density microarrays based on the genomic sequence were used to monitor the expression levels of >80% of the genes in this bacterium. In parallel the proteins were analyzed by two-dimensional electrophoresis. DNA gyrase inhibitors of two different functional classes were used. Novobiocin, as a representative of one class, inhibits the ATPase activity of the enzyme, thereby indirectly changing the degree of DNA supercoiling. Ciprofloxacin, a representative of the second class, obstructs supercoiling by inhibiting the DNA cleavage-resealing reaction. Our results clearly show that different responses can be observed. Treatment with the ATPase inhibitor Novobiocin changed the expression rates of many genes, reflecting the fact that the initiation of transcription for many genes is sensitive to DNA supercoiling. Ciprofloxacin mainly stimulated the expression of DNA repair systems as a response to the DNA damage caused by the stable ternary complexes. In addition, changed expression levels were also observed for some genes coding for proteins either annotated as "unknown function" or "hypothetical" or for proteins not directly involved in DNA topology or repair.

Bacterial Proteins↗

metaExpertPro: A Computational Workflow for Metaproteomics Spectral Library Construction and Data-Independent Acquisition Mass Spectrometry Data Analysis.

Analysis of large-scale data-independent acquisition mass spectrometry metaproteomics data remains a computational challenge. Here, we present a computational pipeline called metaExpertPro for metaproteomics data analysis. This pipeline encompasses spectral library generation using data-dependent acquisition MS, protein identification and quantification using data-independent acquisition mass spectrometry, functional and taxonomic annotation, as well as quantitative matrix generation for both microbiota and hosts. By integrating FragPipe and DIA-NN, metaExpertPro offers compatibility with both Orbitrap and timsTOF MS instruments. To evaluate the depth and accuracy of identification and quantification, we conducted extensive assessments using human fecal samples and benchmark tests. Performance tests conducted on human fecal samples indicated that metaExpertPro quantified an average of 45,000 peptides in a 60-min diaPASEF injection. Notably, metaExpertPro outperformed three existing software tools by characterizing a higher number of peptides and proteins. Importantly, metaExpertPro maintained a low factual false discovery rate of approximately 5% for protein groups across four benchmark tests. Applying a filter of five peptides per genus, metaExpertPro achieved relatively high accuracy (F-score = 0.67-0.90) in genus diversity and showed a high correlation (rSpearman = 0.73-0.82) between the measured and true genus relative abundance in benchmark tests. Additionally, the quantitative results at the protein, taxonomy, and function levels exhibited high reproducibility and consistency across the commonly adopted public human gut microbial protein databases IGC and UHGP. In a metaproteomic analysis of dyslipidemia patients, metaExpertPro revealed characteristic alterations in microbial functions and potential interactions between the microbiota and the host.

Proteomics↗

NAVIP: Unraveling the influence of neighboring small sequence variants on functional impact prediction.

Once a suitable reference sequence has been generated, intra-species variation is often assessed by re-sequencing. Variant calling processes can reveal all differences between strains, accessions, genotypes, or individuals. These variants can be enriched with predictions about their functional implications based on available structural annotations, i.e., gene models. Although these functional impact predictions on a per-variant basis are often accurate, some challenging cases require the simultaneous incorporation of multiple adjacent variants into this prediction process. Examples include neighboring variants which modify each other's functional impact. The Neighborhood-Aware Variant Impact Predictor (NAVIP) considers all variants within a given protein coding sequence when predicting the effect. As a proof of concept, variants between the Arabidopsis thaliana accessions Columbia-0 and Niederzenz-1 were annotated. NAVIP is freely available on GitHub (https://github.com/bpucker/NAVIP) and accessible through a web server (https://pbb-tools.de).

Arabidopsis↗

Beyond Exons: Linking Noncoding Heritability and Polygenicity across Complex Human Traits and Disorders.

The genetic architecture of complex traits spans a continuum of polygenicity, yet it remains unclear how differences in polygenicity relate to the functional localization of SNP heritability across the genome. We use a MiXeR-based framework to partition heritability across exonic, intronic, and intergenic regions for 34 traits and introduce a likelihood-based annotation contribution score that quantifies annotation-specific impact on heritability. Exons explain a minority of heritability, and their contribution decreases with increasing polygenicity, from an average of 22% in less polygenic somatic diseases and biomarkers to 13% in highly polygenic psychiatric and cognitive phenotypes. Intergenic fractions show the opposite trend, whereas intronic fractions remain relatively stable. Analysis of a broader set of functional annotations reveals systematic differences along the polygenicity axis: highly polygenic traits show stronger contributions from comparative genomics and variant-effect scores, whereas less polygenic traits show stronger contributions in promoter, transcription, and chromatin annotations. Together, these results indicate that the functional partitioning of heritability systematically varies with polygenicity, pointing to a shift from gene-proximal regulatory architectures to architectures shaped by numerous dispersed regulatory effects as a key determinant of differences in polygenicity across traits.

Journal Article↗

Integration of genome mining and HiTES reveals secondary metabolic potential in marine-derived Aspergillus sp. WHUF0304.

AIMS: Marine-derived Aspergillus species are prolific producers of bioactive secondary metabolites, yet the majority of their biosynthetic gene clusters (BGCs) remain silent. This study aimed to integrate genome mining with high-throughput elicitor screening (HiTES) to unlock the metabolic potential of Aspergillus sp. WHUF0304 and identify elicitors that promote the accumulation of previously undetected metabolites. METHODS AND RESULTS: A high-quality genome of Aspergillus sp. WHUF0304 was assembled and annotated using multiple functional databases, revealing substantial secondary metabolic potential. antiSMASH analysis identified diverse BGCs, including NRPS/indole-related clusters potentially associated with indole diketopiperazine biosynthesis. A HiTES-inspired elicitor screening strategy was then applied to evaluate 42 small molecules for their ability to alter the metabolite profile of this strain. Among the tested elicitors, fluconazole was identified as the optimal inducer, triggering the production of several indole diketopiperazine-related differential metabolites. Subsequent activity-guided isolation led to the identification of a bioactive indole diketopiperazine dimer, cristatumin E, which exhibited antibacterial activity against Escherichia coli and Bacillus subtilis with minimum inhibitory concentrations (MICs) of 32 µg mL-1 and 256 µg mL-1, respectively. CONCLUSIONS: These findings demonstrate that integrating genomic and functional approaches effectively activates silent BGCs in marine fungi. The fluconazole-associated accumulation and subsequent isolation of cristatumin E, a bioactive indole diketopiperazine dimer, highlight the potential of elicitor-mediated activation to expand the detectable metabolite profile of Aspergillus sp. WHUF0304.

Aspergillus↗

The prostate expression database (PEDB): status and enhancements in 2000.

The Prostate Expression Database (PEDB) is an online resource designed to access and analyze gene expression information derived from the human prostate. PEDB archives >55 000 expressed sequence tags (ESTs) from 43 cDNA libraries in a curated relational database that provides detailed library information including tissue source, library construction methods, sequence diversity and sequence abundance. The differential expression of each EST species can be viewed across all libraries using a Virtual Expression Analysis Tool (VEAT), a graphical user interface written in Java for intra- and inter-library species comparisons. Recent enhancements to PEDB include: (i) the functional categorization of annotated EST assemblies using a classification scheme developed at The Institute for Genome Research; (ii) catalogs of expressed genes in specific prostate tissue sources designated as transcriptomes; and (iii) the addition of prostate proteome information derived from two-dimensional electrophoreses and mass spectrometry of prostate cancer cell lines. PEDB may be accessed via the WWW at http://www.mbt.washington.edu/PEDB/

Databases, Factual↗

The COG database: a tool for genome-scale analysis of protein functions and evolution.

Rational classification of proteins encoded in sequenced genomes is critical for making the genome sequences maximally useful for functional and evolutionary studies. The database of Clusters of Orthologous Groups of proteins (COGs) is an attempt on a phylogenetic classification of the proteins encoded in 21 complete genomes of bacteria, archaea and eukaryotes (http://www. ncbi.nlm. nih.gov/COG). The COGs were constructed by applying the criterion of consistency of genome-specific best hits to the results of an exhaustive comparison of all protein sequences from these genomes. The database comprises 2091 COGs that include 56-83% of the gene products from each of the complete bacterial and archaeal genomes and approximately 35% of those from the yeast Saccharomyces cerevisiae genome. The COG database is accompanied by the COGNITOR program that is used to fit new proteins into the COGs and can be applied to functional and phylogenetic annotation of newly sequenced genomes.

Database Management Systems↗

On the spatial disposition of the fifth transmembrane helix and the structural integrity of the transmembrane binding site in the opioid and ORL1 G protein-coupled receptor family.

Evidence from statistical cluster analyses of a multiple sequence alignment of G protein-coupled receptor seven-helix folds supports the existence of structurally conserved transmembrane (TM) ligand binding sites in the opioid/opioid receptor-like (ORL1) and amine receptor families. Based on the expectation that functionally conserved regions in homologous proteins will display locally higher levels of sequence identity compared with global sequence similarities that pertain to the overall fold, this approach may have wider applications in functional genomics to annotate sequence data. Binding sites in models of the kappa-opioid receptor seven-helix bundle built from the rhodopsin templates of Baldwin et al. (1997) [J. Mol. Biol., 272, 144-164] and Herzyk and Hubbard (1998) [J. Mol. Biol., 281, 742-751] are compared. The Herzyk and Hubbard template is found to be in better accord with experimental studies of amine, opioid and rhodopsin receptors owing to the reduced physical separation of the extracellular parts of TM helices V and VI and differences in the rotational orientation of the N-terminal of helix V that reveal side chain accessibilities in the Baldwin et al. structure to be out of phase with relative alkylation rates of engineered cysteine residues in the TM binding site of the alpha(2A)-adrenergic receptor. TM helix V in the Baldwin et al. template has been remodelled with a different proline kink to satisfy experimental constraints. A recent proposal that rotation of helix V is associated with receptor activation is critically discussed.

Binding Sites↗

Patterns of Genomic Divergence and Introgression in Two Primulina Hybrid Zones.

Hybrid zones have long been promoted as natural laboratories for understanding the mechanisms of speciation. Multiple or replicated hybrid zones are particularly informative, as they allow for assessing the consistency of genomic divergence and introgression across different environmental contexts and demographic histories, thereby improving our understanding of the factors that drive or hinder speciation on a broader scale. Here, using whole-genome resequencing data, we compare the patterns of genomic divergence and introgression in two Primulina hybrid zones. We found that genomic divergence in both hybrid zones is largely shaped by neutral processes, with only a few genomic regions showing signatures of balancing or lineage-specific selection. Genomic cline analyses identified numerous SNPs that showed significantly steeper clines and biased centres than the genome-wide expectation in both hybrid zones, consistent with the existence of reproductive barriers. Within regions of restricted gene flow, we identified 21 genes shared between the two hybrid zones. Annotation of gene function revealed that several genes are involved in reproductive processes. In addition, many zone-specific outlier loci were linked to genes associated with pollen and flower development, suggesting that these barriers may contribute to reproductive isolation under localised ecological conditions. Overall, these findings suggest that while certain reproductive barriers remain consistent across independent hybrid zones, others may be contingent on local environmental contexts. Our results demonstrate that both general and zone-specific mechanisms contribute to reproductive isolation in Primulina, providing empirical evidence that some genomic barriers recur across independent hybrid zones while others arise through localised adaptation.

Lamiales↗

Hypertrophic cardiomyopathy: a genome-wide association meta-analysis and polygenic risk score.

BACKGROUND: Hypertrophic cardiomyopathy (HCM) is a heritable trait with marked variability in expression and outcomes. Our aims were to discover new genetic loci associated with HCM and to test the effect of a new polygenic risk score (PRS) on incidence, phenotype and outcomes stratified by genotype status. METHODS: A discovery genome-wide association study (GWAS) was performed on 2284 HCM cases and 4525 controls. Two fixed-effects meta-analyses combined our discovery GWAS with single-trait and multi-trait results from a published study. Discovered loci underwent comprehensive bioinformatic analysis including functional and druggability annotations. A PRS using loci from the two meta-analyses was evaluated for association with HCM diagnosis in 411 213 individuals from UK Biobank (UKBB); imaging phenotypes in individuals without HCM; a composite endpoint (including all-cause mortality and transplantation); and sudden cardiac death (SCD) in 1756 HCM cases. PRS analyses were stratified by genotype status. RESULTS: Three loci were found in the discovery GWAS (BAG3, FHOD3 and novel locus PPP1R3A). In the meta-analyses, 70 unique loci were identified, four novel (MYPN, YWHAE, NOS1AP and OBSCN). Bioinformatic analyses identified NOS1AP as a candidate HCM gene. A new PRS was significantly associated with HCM diagnosis (HR=3.19, 95% CI 2.46 to 4.14 for top 5% vs lower 95%; HR=1.88, 95% CI 1.72 to 2.06 per SD increase). Significant associations were found between PRS and greater left ventricular (LV) wall thickness and higher LV ejection fraction in UKBB participants without HCM. Genotype-negative HCM cases in the top 20% of the PRS distribution had an increased risk of SCD (HR=2.72, 95% CI 1.03 to 7.17). CONCLUSIONS: We report novel HCM loci. A new PRS predicted the risk of HCM development and associated imaging characteristics in the UKBB and outcomes in an HCM cohort.

Cardiomyopathies↗

Genetic insights into the relationship between age at menarche and mental health-related phenotypes.

BACKGROUND: Multiple observational studies have reported associations between age at menarche (AAM) and mental health problems, yet their shared genetic architecture remains poorly characterized. METHODS: We leveraged genome-wide association study summary statistics for AAM and 15 mental health-related phenotypes. We conducted a multi-method integrative analysis encompassing linkage disequilibrium score regression, pleiotropic analysis under the composite null hypothesis, functional mapping and annotation, multi-marker analysis of genomic annotation, pathway enrichment, and bidirectional two-sample Mendelian randomization (MR) to explore shared genetic architecture and potential causal relationships. RESULTS: Our study identified significant genetic correlations between AAM and eight mental health-related phenotypes (miserableness, fed-up feelings, nervous feelings, ever thought that life is not worth living, ever self-harmed, depression, ever smoker, and age started smoking in former smokers). A total of 155 pleiotropic loci, 18 colocalized loci (e.g., 6q16.3), and 203 pleiotropic genes (e.g., LIN28B) were identified. These genes are expressed in multiple regions, including the cerebral cortex and hypothalamus, and are involved in various biological processes and signaling pathways. Additionally, MR analysis revealed causal associations between AAM and 5 mental health-related phenotypes (mood swings, miserableness, fed-up feelings, and age at which smokers started smoking in former/current smokers). CONCLUSIONS: Our study revealed extensive genetic associations between AAM and mental health-related phenotypes, and further explored the potential causal relationships between them. These findings enhance our understanding of the relationship from a genetic perspective and establish a foundation for future research to explore the biological pathways and environmental interactions contributing to these associations.

Genome-Wide Association Study↗