Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

META-DIFF: a k-mer-based pipeline that detects differentially abundant sequences in metagenomics whole genome sequencing.

Traditional case-control metagenomic studies are constrained by their dependence on taxonomic and functional databases. Because annotation occurs before differential analysis, they are limited to known elements and keep function and taxonomy separate. Although binning strategies have emerged to reconstruct genomes and mitigate this issue, they still require an assembly step, preventing the use of all available sequencing data. Here, we introduce META-DIFF, a pipeline based on differentially abundant k-mers independently of any prior annotation. From those k-mers, it reconstructs longer sequences and provides biological context, as well as the best set of unitigs to discriminate between conditions. Across both taxonomy-centric and functionally-centric benchmarks, it showed robust performance and displayed great reproducibility. It also behaved more conservatively than did other univariate methodologies, i.e. it maintained a high precision at the expense of recall, particularly in conditions of low fold-change and limited sequencing depth. The efficacy of META-DIFF was further validated through its application to a real-world colorectal cancer dataset, which produced both confirmatory and novel results compared with those of previous publications. The pipeline is able to exploit all reads and identify differentially abundant elements, including unknown DNA, prior to annotation. With the guidelines provided, META-DIFF provides users with great exploratory power to unravel microbiome changes.

Metagenomics↗

Intrinsic errors in genome annotation.

Genome sequencing is usually followed by routine annotation of protein function based on the assumption that similar sequences will have similar functions. Here, we introduce a simple calculation to estimate the magnitude of any possible annotation errors. We counted the number of discrepancies in the annotation of well-established sets of similar proteins and extrapolated these values to the pairs of similar sequences used for the annotation of different microbial genomes. We conclude that the number of potential errors in the prediction of detailed functions is higher than is usually believed.

Binding Sites↗

Associating genes with gene ontology codes using a maximum entropy analysis of biomedical literature.

Functional characterizations of thousands of gene products from many species are described in the published literature. These discussions are extremely valuable for characterizing the functions not only of these gene products, but also of their homologs in other organisms. The Gene Ontology (GO) is an effort to create a controlled terminology for labeling gene functions in a more precise, reliable, computer-readable manner. Currently, the best annotations of gene function with the GO are performed by highly trained biologists who read the literature and select appropriate codes. In this study, we explored the possibility that statistical natural language processing techniques can be used to assign GO codes. We compared three document classification methods (maximum entropy modeling, naïve Bayes classification, and nearest-neighbor classification) to the problem of associating a set of GO codes (for biological process) to literature abstracts and thus to the genes associated with the abstracts. We showed that maximum entropy modeling outperforms the other methods and achieves an accuracy of 72% when ascertaining the function discussed within an abstract. The maximum entropy method provides confidence measures that correlate well with performance. We conclude that statistical methods may be used to assign GO codes and may be useful for the difficult task of reassignment as terminology standards evolve over time.

Algorithms↗

The CATH Dictionary of Homologous Superfamilies (DHS): a consensus approach for identifying distant structural homologues.

A consensus approach has been developed for identifying distant structural homologues. This is based on the CATH Dictionary of Homologous Superfamilies (DHS), a database of validated multiple structural alignments annotated with consensus functional information for evolutionary protein superfamilies (URL: http://www. biochem.ucl.ac.uk/bsm/dhs). Multiple structural alignments have been generated for 362 well-populated superfamilies in the CATH structural domain database and annotated with secondary structure, physicochemical properties, functional sequence patterns and protein-ligand interaction data. Consensus functional information for each superfamily includes descriptions and keywords extracted from SWISS-PROT and the ENZYME database. The Dictionary provides a powerful resource to validate, examine and visualize key structural and functional features of each homologous superfamily. The value of the DHS, for assessing functional variability and identifying distant evolutionary relationships, is illustrated using the pyridoxal-5'-phosphate (PLP) binding aspartate aminotransferase superfamily. The DHS also provides a tool for examining sequence-structure relationships for proteins within each fold group.

Amino Acid Sequence↗

Assigning genomic sequences to CATH.

We report the latest release (version 1.6) of the CATH protein domains database (http://www.biochem.ucl. ac.uk/bsm/cath ). This is a hierarchical classification of 18 577 domains into evolutionary families and structural groupings. We have identified 1028 homo-logous superfamilies in which the proteins have both structural, and sequence or functional similarity. These can be further clustered into 672 fold groups and 35 distinct architectures. Recent developments of the database include the generation of 3D templates for recognising structural relatives in each fold group, which has led to significant improvements in the speed and accuracy of updating the database and also means that less manual validation is required. We also report the establishment of the CATH-PFDB (Protein Family Database), which associates 1D sequences with the 3D homologous superfamilies. Sequences showing identifiable homology to entries in CATH have been extracted from GenBank using PSI-BLAST. A CATH-PSIBLAST server has been established, which allows you to scan a new sequence against the database. The CATH Dictionary of Homologous Superfamilies (DHS), which contains validated multiple structural alignments annotated with consensus functional information for evolutionary protein superfamilies, has been updated to include annotations associated with sequence relatives identified in GenBank. The DHS is a powerful tool for considering the variation of functional properties within a given CATH superfamily and in deciding what functional properties may be reliably inherited by a newly identified relative.

Amino Acid Sequence↗

Cis-regulatory elements: systematic identification and horticultural applications.

Cis-regulatory elements (CREs) are the genetic DNA fragments bound by transcription factors (TFs). CREs function as molecular switches that precisely modulate the dosage and spatiotemporal patterns of gene expression. The systematic identification of CREs not only facilitates the annotation of the functional non-coding genome but also provides essential insights into the architecture of gene regulatory networks and sheds light on an accurate selection of the target sites for genetic engineering of crops. In this review, we summarize the current high-throughput methodologies used for identifying CREs, illustrate the associations between CREs and agronomic traits in horticultural crops, and discuss how CREs can be exploited to facilitate crop breeding.

Breeding↗

KEGG: kyoto encyclopedia of genes and genomes.

KEGG (Kyoto Encyclopedia of Genes and Genomes) is a knowledge base for systematic analysis of gene functions, linking genomic information with higher order functional information. The genomic information is stored in the GENES database, which is a collection of gene catalogs for all the completely sequenced genomes and some partial genomes with up-to-date annotation of gene functions. The higher order functional information is stored in the PATHWAY database, which contains graphical representations of cellular processes, such as metabolism, membrane transport, signal transduction and cell cycle. The PATHWAY database is supplemented by a set of ortholog group tables for the information about conserved subpathways (pathway motifs), which are often encoded by positionally coupled genes on the chromosome and which are especially useful in predicting gene functions. A third database in KEGG is LIGAND for the information about chemical compounds, enzyme molecules and enzymatic reactions. KEGG provides Java graphics tools for browsing genome maps, comparing two genome maps and manipulating expression maps, as well as computational tools for sequence comparison, graph comparison and path computation. The KEGG databases are daily updated and made freely available (http://www. genome.ad.jp/kegg/).

Animals↗

A functional update of the Escherichia coli K-12 genome.

BACKGROUND: Since the genome of Escherichia coli K-12 was initially annotated in 1997, additional functional information based on biological characterization and functions of sequence-similar proteins has become available. On the basis of this new information, an updated version of the annotated chromosome has been generated. RESULTS: The E. coli K-12 chromosome is currently represented by 4,401 genes encoding 116 RNAs and 4,285 proteins. The boundaries of the genes identified in the GenBank Accession U00096 were used. Some protein-coding sequences are compound and encode multimodular proteins. The coding sequences (CDSs) are represented by modules (protein elements of at least 100 amino acids with biological activity and independent evolutionary history). There are 4,616 identified modules in the 4,285 proteins. Of these, 48.9% have been characterized, 29.5% have an imputed function, 2.1% have a phenotype and 19.5% have no function assignment. Only 7% of the modules appear unique to E. coli, and this number is expected to be reduced as more genome data becomes available. The imputed functions were assigned on the basis of manual evaluation of functions predicted by BLAST and DARWIN analyses and by the MAGPIE genome annotation system. CONCLUSIONS: Much knowledge has been gained about functions encoded by the E. coli K-12 genome since the 1997 annotation was published. The data presented here should be useful for analysis of E. coli gene products as well as gene products encoded by other genomes.

Bacterial Proteins↗

Phylogeny of related functions: the case of polyamine biosynthetic enzymes.

Genome annotation requires explicit identification of gene function. This task frequently uses protein sequence alignments with examples having a known function. Genetic drift, co-evolution of subunits in protein complexes and a variety of other constraints interfere with the relevance of alignments. Using a specific class of proteins, it is shown that a simple data analysis approach can help solve some of the problems posed. The origin of ureohydrolases has been explored by comparing sequence similarity trees, maximizing amino acid alignment conservation. The trees separate agmatinases from arginases but suggest the presence of unknown biases responsible for unexpected positions of some enzymes. Using factorial correspondence analysis, a distance tree between sequences was established, comparing regions with gaps in the alignments. The gap tree gives a consistent picture of functional kinship, perhaps reflecting some aspects of phylogeny, with a clear domain of enzymes encoding two types of ureohydrolases (agmatinases and arginases) and activities related to, but different from ureohydrolases. Several annotated genes appeared to correspond to a wrong assignment if the trees were significant. They were cloned and their products expressed and identified biochemically. This substantiated the validity of the gap tree. Its organization suggests a very ancient origin of ureohydrolases. Some enzymes of eukaryotic origin are spread throughout the arginase part of the trees: they might have been derived from the genes found in the early symbiotic bacteria that became the organelles. They were transferred to the nucleus when symbiotic genes had to escape Muller's ratchet. This work also shows that arginases and agmatinases share the same two manganese-ion-binding sites and exhibit only subtle differences that can be accounted for knowing the three-dimensional structure of arginases. In the absence of explicit biochemical data, extreme caution is needed when annotating genes having similarities to ureohydrolases.

Amino Acid Sequence↗

Large-scale benchmarking of prokaryotic annotation tools across thousands of species.

BACKGROUND: Genome annotation is an important step in deriving functional meaning from prokaryotic sequencing data, yet systematic evaluations guiding tool selection are lacking. We present the first large-scale investigation of four prominent open-source annotation tools (Prokka, Bakta, EggNOG-mapper, and PGAP) across 156,033 diverse genomes. This includes Escherichia coli strains for baseline performance, thousands of archaea and bacteria genomes, as well as frameshifted and metagenome-assembled genomes. RESULTS: Bakta excels in annotating high-quality bacterial genomes, while PGAP was better for archaeal genomes and challenging bacterial assemblies, including metagenome-assembled, fragmented, or contaminated samples. For Gene Ontology annotation, PGAP consistently provides broader term coverage, whereas EggNOG-mapper offers more terms per feature. CONCLUSIONS: Our findings highlight tool-specific strengths crucial for selecting optimal solutions based on genome quality, taxonomy, and origin (e.g. MAGs). This study provides an evidence-based guide for users and informs future tool development.

Molecular Sequence Annotation↗

Identifying the 3'-terminal exon in human DNA.

MOTIVATION: We present JTEF, a new program for finding 3' terminal exons in human DNA sequences. This program is based on quadratic discriminant analysis, a standard non-linear statistical pattern recognition method. The quadratic discriminant functions used for building the algorithm were trained on a set of 3' terminal exons of type 3tuexon (those containing the true STOP codon). RESULTS: We showed that the average predictive accuracy of JTEF is higher than the presently available best programs (GenScan and Genemark.hmm) based on a test set of 65 human DNA sequences with 121 genes. In particular JTEF performs well on larger genomic contigs containing multiple genes and significant amounts of intergenic DNA. It will become a valuable tool for genome annotation and gene functional studies. AVAILABILITY: JTEF is available free for academic users on request from ftp://cshl.org/pub/science/mzhanglab/JTEF and will be made available through the World Wide Web (http://argon.cshl.org/).

Algorithms↗

Gene expression changes triggered by exposure of Haemophilus influenzae to novobiocin or ciprofloxacin: combined transcription and translation analysis.

The responses of Haemophilus influenzae to DNA gyrase inhibitors were analyzed at the transcriptional and the translational level. High-density microarrays based on the genomic sequence were used to monitor the expression levels of >80% of the genes in this bacterium. In parallel the proteins were analyzed by two-dimensional electrophoresis. DNA gyrase inhibitors of two different functional classes were used. Novobiocin, as a representative of one class, inhibits the ATPase activity of the enzyme, thereby indirectly changing the degree of DNA supercoiling. Ciprofloxacin, a representative of the second class, obstructs supercoiling by inhibiting the DNA cleavage-resealing reaction. Our results clearly show that different responses can be observed. Treatment with the ATPase inhibitor Novobiocin changed the expression rates of many genes, reflecting the fact that the initiation of transcription for many genes is sensitive to DNA supercoiling. Ciprofloxacin mainly stimulated the expression of DNA repair systems as a response to the DNA damage caused by the stable ternary complexes. In addition, changed expression levels were also observed for some genes coding for proteins either annotated as "unknown function" or "hypothetical" or for proteins not directly involved in DNA topology or repair.

Bacterial Proteins↗

metaExpertPro: A Computational Workflow for Metaproteomics Spectral Library Construction and Data-Independent Acquisition Mass Spectrometry Data Analysis.

Analysis of large-scale data-independent acquisition mass spectrometry metaproteomics data remains a computational challenge. Here, we present a computational pipeline called metaExpertPro for metaproteomics data analysis. This pipeline encompasses spectral library generation using data-dependent acquisition MS, protein identification and quantification using data-independent acquisition mass spectrometry, functional and taxonomic annotation, as well as quantitative matrix generation for both microbiota and hosts. By integrating FragPipe and DIA-NN, metaExpertPro offers compatibility with both Orbitrap and timsTOF MS instruments. To evaluate the depth and accuracy of identification and quantification, we conducted extensive assessments using human fecal samples and benchmark tests. Performance tests conducted on human fecal samples indicated that metaExpertPro quantified an average of 45,000 peptides in a 60-min diaPASEF injection. Notably, metaExpertPro outperformed three existing software tools by characterizing a higher number of peptides and proteins. Importantly, metaExpertPro maintained a low factual false discovery rate of approximately 5% for protein groups across four benchmark tests. Applying a filter of five peptides per genus, metaExpertPro achieved relatively high accuracy (F-score = 0.67-0.90) in genus diversity and showed a high correlation (rSpearman = 0.73-0.82) between the measured and true genus relative abundance in benchmark tests. Additionally, the quantitative results at the protein, taxonomy, and function levels exhibited high reproducibility and consistency across the commonly adopted public human gut microbial protein databases IGC and UHGP. In a metaproteomic analysis of dyslipidemia patients, metaExpertPro revealed characteristic alterations in microbial functions and potential interactions between the microbiota and the host.

Proteomics↗

NAVIP: Unraveling the influence of neighboring small sequence variants on functional impact prediction.

Once a suitable reference sequence has been generated, intra-species variation is often assessed by re-sequencing. Variant calling processes can reveal all differences between strains, accessions, genotypes, or individuals. These variants can be enriched with predictions about their functional implications based on available structural annotations, i.e., gene models. Although these functional impact predictions on a per-variant basis are often accurate, some challenging cases require the simultaneous incorporation of multiple adjacent variants into this prediction process. Examples include neighboring variants which modify each other's functional impact. The Neighborhood-Aware Variant Impact Predictor (NAVIP) considers all variants within a given protein coding sequence when predicting the effect. As a proof of concept, variants between the Arabidopsis thaliana accessions Columbia-0 and Niederzenz-1 were annotated. NAVIP is freely available on GitHub (https://github.com/bpucker/NAVIP) and accessible through a web server (https://pbb-tools.de).

Arabidopsis↗

Beyond Exons: Linking Noncoding Heritability and Polygenicity across Complex Human Traits and Disorders.

The genetic architecture of complex traits spans a continuum of polygenicity, yet it remains unclear how differences in polygenicity relate to the functional localization of SNP heritability across the genome. We use a MiXeR-based framework to partition heritability across exonic, intronic, and intergenic regions for 34 traits and introduce a likelihood-based annotation contribution score that quantifies annotation-specific impact on heritability. Exons explain a minority of heritability, and their contribution decreases with increasing polygenicity, from an average of 22% in less polygenic somatic diseases and biomarkers to 13% in highly polygenic psychiatric and cognitive phenotypes. Intergenic fractions show the opposite trend, whereas intronic fractions remain relatively stable. Analysis of a broader set of functional annotations reveals systematic differences along the polygenicity axis: highly polygenic traits show stronger contributions from comparative genomics and variant-effect scores, whereas less polygenic traits show stronger contributions in promoter, transcription, and chromatin annotations. Together, these results indicate that the functional partitioning of heritability systematically varies with polygenicity, pointing to a shift from gene-proximal regulatory architectures to architectures shaped by numerous dispersed regulatory effects as a key determinant of differences in polygenicity across traits.

Journal Article↗

Integration of genome mining and HiTES reveals secondary metabolic potential in marine-derived Aspergillus sp. WHUF0304.

AIMS: Marine-derived Aspergillus species are prolific producers of bioactive secondary metabolites, yet the majority of their biosynthetic gene clusters (BGCs) remain silent. This study aimed to integrate genome mining with high-throughput elicitor screening (HiTES) to unlock the metabolic potential of Aspergillus sp. WHUF0304 and identify elicitors that promote the accumulation of previously undetected metabolites. METHODS AND RESULTS: A high-quality genome of Aspergillus sp. WHUF0304 was assembled and annotated using multiple functional databases, revealing substantial secondary metabolic potential. antiSMASH analysis identified diverse BGCs, including NRPS/indole-related clusters potentially associated with indole diketopiperazine biosynthesis. A HiTES-inspired elicitor screening strategy was then applied to evaluate 42 small molecules for their ability to alter the metabolite profile of this strain. Among the tested elicitors, fluconazole was identified as the optimal inducer, triggering the production of several indole diketopiperazine-related differential metabolites. Subsequent activity-guided isolation led to the identification of a bioactive indole diketopiperazine dimer, cristatumin E, which exhibited antibacterial activity against Escherichia coli and Bacillus subtilis with minimum inhibitory concentrations (MICs) of 32 µg mL-1 and 256 µg mL-1, respectively. CONCLUSIONS: These findings demonstrate that integrating genomic and functional approaches effectively activates silent BGCs in marine fungi. The fluconazole-associated accumulation and subsequent isolation of cristatumin E, a bioactive indole diketopiperazine dimer, highlight the potential of elicitor-mediated activation to expand the detectable metabolite profile of Aspergillus sp. WHUF0304.

Aspergillus↗

The prostate expression database (PEDB): status and enhancements in 2000.

The Prostate Expression Database (PEDB) is an online resource designed to access and analyze gene expression information derived from the human prostate. PEDB archives >55 000 expressed sequence tags (ESTs) from 43 cDNA libraries in a curated relational database that provides detailed library information including tissue source, library construction methods, sequence diversity and sequence abundance. The differential expression of each EST species can be viewed across all libraries using a Virtual Expression Analysis Tool (VEAT), a graphical user interface written in Java for intra- and inter-library species comparisons. Recent enhancements to PEDB include: (i) the functional categorization of annotated EST assemblies using a classification scheme developed at The Institute for Genome Research; (ii) catalogs of expressed genes in specific prostate tissue sources designated as transcriptomes; and (iii) the addition of prostate proteome information derived from two-dimensional electrophoreses and mass spectrometry of prostate cancer cell lines. PEDB may be accessed via the WWW at http://www.mbt.washington.edu/PEDB/

Databases, Factual↗

The COG database: a tool for genome-scale analysis of protein functions and evolution.

Rational classification of proteins encoded in sequenced genomes is critical for making the genome sequences maximally useful for functional and evolutionary studies. The database of Clusters of Orthologous Groups of proteins (COGs) is an attempt on a phylogenetic classification of the proteins encoded in 21 complete genomes of bacteria, archaea and eukaryotes (http://www. ncbi.nlm. nih.gov/COG). The COGs were constructed by applying the criterion of consistency of genome-specific best hits to the results of an exhaustive comparison of all protein sequences from these genomes. The database comprises 2091 COGs that include 56-83% of the gene products from each of the complete bacterial and archaeal genomes and approximately 35% of those from the yeast Saccharomyces cerevisiae genome. The COG database is accompanied by the COGNITOR program that is used to fit new proteins into the COGs and can be applied to functional and phylogenetic annotation of newly sequenced genomes.

Database Management Systems↗