Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,729 records · Page 96Linked to original sources

GeneViTo: visualizing gene-product functional and structural features in genomic datasets.

BACKGROUND: The availability of increasing amounts of sequence data from completely sequenced genomes boosts the development of new computational methods for automated genome annotation and comparative genomics. Therefore, there is a need for tools that facilitate the visualization of raw data and results produced by bioinformatics analysis, providing new means for interactive genome exploration. Visual inspection can be used as a basis to assess the quality of various analysis algorithms and to aid in-depth genomic studies. RESULTS: GeneViTo is a JAVA-based computer application that serves as a workbench for genome-wide analysis through visual interaction. The application deals with various experimental information concerning both DNA and protein sequences (derived from public sequence databases or proprietary data sources) and meta-data obtained by various prediction algorithms, classification schemes or user-defined features. Interaction with a Graphical User Interface (GUI) allows easy extraction of genomic and proteomic data referring to the sequence itself, sequence features, or general structural and functional features. Emphasis is laid on the potential comparison between annotation and prediction data in order to offer a supplement to the provided information, especially in cases of "poor" annotation, or an evaluation of available predictions. Moreover, desired information can be output in high quality JPEG image files for further elaboration and scientific use. A compilation of properly formatted GeneViTo input data for demonstration is available to interested readers for two completely sequenced prokaryotes, Chlamydia trachomatis and Methanococcus jannaschii. CONCLUSIONS: GeneViTo offers an inspectional view of genomic functional elements, concerning data stemming both from database annotation and analysis tools for an overall analysis of existing genomes. The application is compatible with Linux or Windows ME-2000-XP operating systems, provided that the appropriate Java Runtime Environment is already installed in the system.

Bacterial Proton-Translocating ATPases↗

FlyBase: anatomical data, images and queries.

FlyBase (http://flybase.org/) is a database of genetic and genomic data on the model organism Drosophila melanogaster and the entire insect family Drosophilidae. The FlyBase Consortium curates, annotates, integrates and maintains a wide variety of data within this domain. Access to the data is provided through graphical and textual user interfaces tailored to particular types of data. FlyBase data types include maps at the cytological, genetic and sequence levels, genes and alleles including their products, functions, expression patterns, mutant phenotypes and genetic interactions as well as aberrant chromosomes, annotated genomes, genetic stock collections, transposons, transgene constructs and insertions, anatomy and images, bibliographic data, and community contact information.

Animals↗

PASS2: an automated database of protein alignments organised as structural superfamilies.

BACKGROUND: The functional selection and three-dimensional structural constraints of proteins in nature often relates to the retention of significant sequence similarity between proteins of similar fold and function despite poor sequence identity. Organization of structure-based sequence alignments for distantly related proteins, provides a map of the conserved and critical regions of the protein universe that is useful for the analysis of folding principles, for the evolutionary unification of protein families and for maximizing the information return from experimental structure determination. The Protein Alignment organised as Structural Superfamily (PASS2) database represents continuously updated, structural alignments for evolutionary related, sequentially distant proteins. DESCRIPTION: An automated and updated version of PASS2 is, in direct correspondence with SCOP 1.63, consisting of sequences having identity below 40% among themselves. Protein domains have been grouped into 628 multi-member superfamilies and 566 single member superfamilies. Structure-based sequence alignments for the superfamilies have been obtained using COMPARER, while initial equivalencies have been derived from a preliminary superposition using LSQMAN or STAMP 4.0. The final sequence alignments have been annotated for structural features using JOY4.0. The database is supplemented with sequence relatives belonging to different genomes, conserved spatially interacting and structural motifs, probabilistic hidden markov models of superfamilies based on the alignments and useful links to other databases. Probabilistic models and sensitive position specific profiles obtained from reliable superfamily alignments aid annotation of remote homologues and are useful tools in structural and functional genomics. PASS2 presents the phylogeny of its members both based on sequence and structural dissimilarities. Clustering of members allows us to understand diversification of the family members. The search engine has been improved for simpler browsing of the database. CONCLUSIONS: The database resolves alignments among the structural domains consisting of evolutionarily diverged set of sequences. Availability of reliable sequence alignments of distantly related proteins despite poor sequence identity and single-member superfamilies permit better sampling of structures in libraries for fold recognition of new sequences and for the understanding of protein structure-function relationships of individual superfamilies. PASS2 is accessible at http://www.ncbs.res.in/~faculty/mini/campass/pass2.html

Amino Acid Sequence↗

The role of alternative translation start sites in the generation of human protein diversity.

According to the scanning model, 40S ribosomal subunits initiate translation at the first (5' proximal) AUG codon they encounter. However, if the first AUG is in a suboptimal context, it may not be recognized, and translation can then initiate at downstream AUG(s). In this way, a single RNA can produce several variant products. Earlier experiments suggested that some of these additional protein variants might be functionally important. We have analysed human mRNAs that have AUG triplets in 5' untranslated regions and mRNAs in which the annotated translational start codon is located in a suboptimal context. It was found that 3% of human mRNAs have the potential to encode N-terminally extended variants of the annotated proteins and 12% could code for N-truncated variants. The predicted subcellular localizations of these protein variants were compared: 31% of the N-extended proteins and 30% of the N-truncated proteins were predicted to localize to subcellular compartments that differed from those targeted by the annotated protein forms. These results suggest that additional AUGs may frequently be exploited for the synthesis of proteins that possess novel functional properties.

5' Untranslated Regions↗

Annotation, nomenclature and evolution of four novel homeobox genes expressed in the human germ line.

The homeobox genes comprise a large gene superfamily characterised by a conserved DNA motif encoding the homeodomain. Most homeodomain proteins function as transcription factors, and many have important roles in embryonic development and cell differentiation. Here we describe, annotate and name four novel homeobox genes in the human genome: ARGFX, DPRX, TPRX1 and DUXA. Each has generated multiple retrotransposed (processed) pseudogenes; these are reliable indicators of germ-line expression because only in germ-line cells can retrotransposition result in inheritance to the next generation. The retrotransposed sequences were exploited here as a novel means to deduce exon-intron boundaries. All four novel genes show accelerated rates of protein sequence evolution. This fast rate of sequence change may be connected with roles in human reproductive biology. Deducing the evolutionary origins of these genes is not straightforward, but we propose that TPRX1, DPRX and DUXA are highly divergent derivatives of the CRX gene, itself a member of the Otx homeobox gene family.

Evolution, Molecular↗

ESTviewer: a web interface for visualizing mouse, rat, cattle, pig and chicken conserved ESTs in human genes and human alternatively spliced variants.

ESTviewer is a web application for interactively visualizing human gene structures, with emphasis on mammalian and avian expressed sequence tags (ESTs) that are conserved in the human genome and alternatively spliced (AS) variants. AS variants from the UCSC, Vega and PSEP annotations are presented in this application for comparison. EST data from six species, human, mouse, rat, cattle, pig and chicken, are mapped to the human genome to show cross-species EST conservation in annotated exonic and intronic regions. Cross-species EST conservation is evolutionarily and functionally important because it represents the effects of selection pressure on genic regions and transcriptome over evolutionary time. Emphatically, ESTviewer provides a convenient tool to compare highly conserved non-human ESTs and human AS variants. The application takes human gene accession Ids or coordinates of genomic sequences as inputs and presents annotated gene structures and their AS variants. In addition, the lengths and percentages of human genic regions covered by ESTs are displayed to show the level of EST coverage of different species. The percentages of the UCSC, Vega and PSEP annotated exons covered by ESTs of the six studied species are also displayed in the interface.

Animals↗

Genome-wide transcript profiles in aging and calorically restricted Drosophila melanogaster.

BACKGROUND: We characterized RNA transcript levels for the whole Drosophila genome during normal aging. We compared age-dependent profiles from animals aged under full-nutrient conditions with profiles obtained from animals maintained on a low-calorie medium to determine if caloric restriction slows the aging process. Specific biological functions impacted by caloric restriction were identified using the Gene Ontology annotation. We used the global patterns of expression profiles to test if particular genomic regions contribute differentially to changes in transcript profiles with age and if global disregulation of gene expression occurs during aging. RESULTS: Whole-genome transcript profiles contained a statistically powerful genetic signature of normal aging. Nearly 23% of the genome changed in transcript representation with age. Caloric restriction was accompanied by a slowing of the progression of normal, age-related changes in transcript levels. Many genes, including those associated with stress response and oogenesis, showed age-dependent transcript representation. Caloric restriction resulted in the downregulation of genes primarily involved in cell growth, metabolism, and reproduction. We found no evidence that age-dependent changes in transcription level were confined to genes localized to specific regions of the genome and found no support for widespread disregulation of gene expression with age. CONCLUSIONS: Aging is characterized by highly dynamic changes in the expression of many genes, which provides a powerful molecular description of the normal aging process. Caloric restriction extends life span by slowing down the rate of normal aging. Transcription levels of genes from a wide variety of biological functions and processes are impacted by age and dietary conditions.

Aging↗

XcisClique: analysis of regulatory bicliques.

BACKGROUND: Modeling of cis-elements or regulatory motifs in promoter (upstream) regions of genes is a challenging computational problem. In this work, set of regulatory motifs simultaneously present in the promoters of a set of genes is modeled as a biclique in a suitably defined bipartite graph. A biologically meaningful co-occurrence of multiple cis-elements in a gene promoter is assessed by the combined analysis of genomic and gene expression data. Greater statistical significance is associated with a set of genes that shares a common set of regulatory motifs, while simultaneously exhibiting highly correlated gene expression under given experimental conditions. METHODS: XcisClique, the system developed in this work, is a comprehensive infrastructure that associates annotated genome and gene expression data, models known cis-elements as regular expressions, identifies maximal bicliques in a bipartite gene-motif graph; and ranks bicliques based on their computed statistical significance. Significance is a function of the probability of occurrence of those motifs in a biclique (a hypergeometric distribution), and on the new sum of absolute values statistic (SAV) that uses Spearman correlations of gene expression vectors. SAV is a statistic well-suited for this purpose as described in the discussion. RESULTS: XcisClique identifies new motif and gene combinations that might indicate as yet unidentified involvement of sets of genes in biological functions and processes. It currently supports Arabidopsis thaliana and can be adapted to other organisms, assuming the existence of annotated genomic sequences, suitable gene expression data, and identified regulatory motifs. A subset of Xcis Clique functionalities, including the motif visualization component MotifSee, source code, and supplementary material are available at https://bioinformatics.cs.vt.edu/xcisclique/.

Algorithms↗

Annotation of human chromosome 21 for relevance to Down syndrome: gene structure and expression analysis.

Down syndrome is caused by an extra copy of human chromosome 21 and the resultant dosage-related overexpression of genes contained within it. To efficiently direct experiments to determine specific gene-phenotype correlations, it is necessary to identify all genes within 21q and assess their functional associations and expression patterns. Analysis of the complete finished sequence of 21q resulted in annotated 225 genes and gene models, most of which were incomplete and/or had little or no experimental verification. Here we correct or complete the genomic structures of 16 genes, 4 of which were not reported in the annotation of the complete sequence. Our data include the identification of six genes encoding short or ambiguous open reading frames; the identification of three cases in which alternative splicing produces two structurally unrelated protein sequences; and the identification of six genes encoding proteins with functional motifs, two genes with unusually low similarity to their orthologous mouse proteins, and four genes with significant conservation in Drosophila melanogaster. We further demonstrate that an additional nine gene models represent bona fide transcripts and develop expression patterns for these genes plus nine additional novel chromosome 21 genes and four paralogous genes mapping elsewhere in the human genome. These data have implications for generating complete transcript maps of chromosome 21 and for the entire human genome, and for defining expression abnormalities in Down syndrome and mouse models.

Animals↗

Plasticity of the gene functions for DNA replication in the T4-like phages.

We have completely sequenced and annotated the genomes of several relatives of the bacteriophage T4, including three coliphages (RB43, RB49 and RB69), three Aeromonas salmonicida phages (44RR2.8t, 25 and 31) and one Aeromonas hydrophila phage (Aeh1). In addition, we have partially sequenced and annotated the T4-like genomes of coliphage RB16 (a close relative of RB43), A. salmonicida phage 65, Acinetobacter johnsonii phage 133 and Vibrio natriegens phage nt-1. Each of these phage genomes exhibited a unique sequence that distinguished it from its relatives, although there were examples of genomes that are very similar to each other. As a group the phages compared here diverge from one another by several criteria, including (a) host range, (b) genome size in the range between approximately 160 kb and approximately 250 kb, (c) content and genetic organization of their T4-like genes for DNA metabolism, (d) mutational drift of the predicted T4-like gene products and their regulatory sites and (e) content of open-reading frames that have no counterparts in T4 or other known organisms (novel ORFs). We have observed a number of DNA rearrangements of the T4 genome type, some exhibiting proximity to putative homing endonuclease genes. Also, we cite and discuss examples of sequence divergence in the predicted sites for protein-protein and protein-nucleic acid interactions of homologues of the T4 DNA replication proteins, with emphasis on the diversity in sequence, molecular form and regulation of the phage-encoded DNA polymerase, gp43. Five of the sequenced phage genomes are predicted to encode split forms of this polymerase. Our studies suggest that the modular construction and plasticity of the T4 genome type and several of its replication proteins may offer resilience to mutation, including DNA rearrangements, and facilitate the adaptation of T4-like phages to different bacterial hosts in nature.

Amino Acid Sequence↗

Decay rates of human mRNAs: correlation with functional characteristics and sequence attributes.

Although mRNA decay rates are a key determinant of the steady-state concentration for any given mRNA species, relatively little is known, on a population level, about what factors influence turnover rates and how these rates are integrated into cellular decisions. We decided to measure mRNA decay rates in two human cell lines with high-density oligonucleotide arrays that enable the measurement of decay rates simultaneously for thousands of mRNA species. Using existing annotation and the Gene Ontology hierarchy of biological processes, we assign mRNAs to functional classes at various levels of resolution and compare the decay rate statistics between these classes. The results show statistically significant organizational principles in the variation of decay rates among functional classes. In particular, transcription factor mRNAs have increased average decay rates compared with other transcripts and are enriched in "fast-decaying" mRNAs with half-lives <2 h. In contrast, we find that mRNAs for biosynthetic proteins have decreased average decay rates and are deficient in fast-decaying mRNAs. Our analysis of data from a previously published study of Saccharomyces cerevisiae mRNA decay shows the same functional organization of decay rates, implying that it is a general organizational scheme for eukaryotes. Additionally, we investigated the dependence of decay rates on sequence composition, that is, the presence or absence of short mRNA motifs in various regions of the mRNA transcript. Our analysis recovers the positive correlation of mRNA decay with known AU-rich mRNA motifs, but we also uncover further short mRNA motifs that show statistically significant correlation with decay. However, we also note that none of these motifs are strong predictors of mRNA decay rate, indicating that the regulation of mRNA decay is more complex and may involve the cooperative binding of several RNA-binding proteins at different sites.

Base Composition↗

Identification of functional transcription factor binding sites using closely related Saccharomyces species.

Comparative genomics provides a rapid means of identifying functional DNA elements by their sequence conservation between species. Transcription factor binding sites (TFBSs) may constitute a significant fraction of these conserved sequences, but the annotation of specific TFBSs is complicated by the fact that these short, degenerate sequences may frequently be conserved by chance rather than functional constraint. To identify intergenic sequences that function as TFBSs, we calculated the probability of binding site conservation between Saccharomyces cerevisiae and its two closest relatives under a neutral model of evolution. We found that this probability is <5% for 134 of 163 transcription factor binding motifs, implying that we can reliably annotate binding sites for the majority of these transcription factors by conservation alone. Although our annotation relies on a number of assumptions, mutations in five of five conserved Ume6 binding sites and three of four conserved Ndt80 binding sites show Ume6- and Ndt80-dependent effects on gene expression. We also found that three of five unconserved Ndt80 binding sites show Ndt80-dependent effects on gene expression. Together these data imply that although sequence conservation can be reliably used to predict functional TFBSs, unconserved sequences might also make a significant contribution to a species' biology.

Amino Acid Motifs↗

Annotated genome assemblies of two temperate North American dung beetles, Canthon chalcites and Phanaeus vindex.

Dung beetles serve as cultivators of their natural habitats, improving soil health and functions in both natural and anthropogenic environments. Despite their ecological importance, whole genome sequences for Scarabaeinae are limited. Here, we present the draft annotated genome assemblies for 2 temperate species of North American dung beetles collected from eastern Tennessee: Canthon chalcites and Phanaeus vindex. Both genome assemblies were generated from PacBio long reads and have high completeness, with BUSCO scores of 98.1% and 98.6% for C. chalcites and P. vindex, respectively. For C. chalcites, the BRAKER3 pipeline predicted 12,799 genes, and the gene set was 93.7% complete. For P. vindex, the BRAKER3 predicted 12,252 genes, and the gene set was 94.9% complete. From the annotated gene sets, orthologous protein sequence analyses among C. chalcites, P. vindex, the dung beetle species Onthophagus taurus, and the more evolutionarily distant beetle Tribolium castaneum indicated that there are 260 unique protein clusters for C. chalcites and 210 unique protein clusters for P. vindex. These 2 draft genomes provide valuable data for comparative genomics, evolution, and phylogenic studies for dung beetle species.

Animals↗

Functional and structural genomics using PEDANT.

MOTIVATION: Enormous demand for fast and accurate analysis of biological sequences is fuelled by the pace of genome analysis efforts. There is also an acute need in reliable up-to-date genomic databases integrating both functional and structural information. Here we describe the current status of the PEDANT software system for high-throughput analysis of large biological sequence sets and the genome analysis server associated with it. RESULTS: The principal features of PEDANT are: (i) completely automatic processing of data using a wide range of bioinformatics methods, (ii) manual refinement of annotation, (iii) automatic and manual assignment of gene products to a number of functional and structural categories, (iv) extensive hyperlinked protein reports, and (v) advanced DNA and protein viewers. The system is easily extensible and allows to include custom methods, databases, and categories with minimal or no programming effort. PEDANT is actively used as a collaborative environment to support several on-going genome sequencing projects. The main purpose of the PEDANT genome database is to quickly disseminate well-organized information on completely sequenced and unfinished genomes. It currently includes 80 genomic sequences and in many cases serves as the only source of exhaustive information on a given genome. The database also acts as a vehicle for a number of research projects in bioinformatics. Using SQL queries, it is possible to correlate a large variety of pre-computed properties of gene products encoded in complete genomes with each other and compare them with data sets of special scientific interest. In particular, the availability of structural predictions for over 300 000 genomic proteins makes PEDANT the most extensive structural genomics resource available on the web.

Arabidopsis↗

A post-genomic approach to understanding sphingolipid metabolism in Arabidopsis thaliana.

AIMS: To highlight the importance of sphingolipids and their metabolites in plant biology. SCOPE: The completion of the arabidopsis genome provides a platform for the identification and functional characterization of genes involved in sphingolipid biosynthesis. Using the yeast Saccharomyces cerevisiae as an experimental model, this review annotates arabidopsis open reading frames likely to be involved in sphingolipid metabolism. A number of these open reading frames have already been subject to functional characterization, though the majority still awaits investigation. Plant-specific aspects of sphingolipid biology (such as enhanced long chain base heterogeneity) are considered in the context of the emerging roles for these lipids in plant form and function. CONCLUSIONS: Arabidopsis provides an excellent genetic and post-genomic model for the characterization of the roles of sphingolipids in higher plants.

Arabidopsis↗

An annotation update via cDNA sequence analysis and comprehensive profiling of developmental, hormonal or environmental responsiveness of the Arabidopsis AP2/EREBP transcription factor gene family.

AP2/EREBP transcription factors (TFs) play functionally important roles in plant growth and development, especially in hormonal regulation and in response to environmental stress. Here we reported verification and correction of annotation through an exhaustive cDNA cloning and sequence analysis performed on 145 of 147 gene family members. A RACE analysis performed on genes with potential in-frame up-stream ATG codon resulted in identification of At2g28520 as an authentic AP2/EREBP member and corrected ORF annotations for three other members. A further phylogenetic analysis of this updated and likely complete family divided it into three major subfamilies. The expression patterns of the AP2/EREBP family members among the 11 organ or tissue types were examined using an oligo microarray and their hormonal and environmental responsiveness were further characterized using cDNA custom macroarrays. These detailed expression profile results provide strong support for a role for AP2/EREBP family members in development and in response to environmental stimuli, and a foundation for future functional analysis of this gene family.

Algorithms↗

Biosequence exegesis.

Annotation of large-scale gene sequence data will benefit from comprehensive and consistent application of well-documented, standard analysis methods and from progressive and vigilant efforts to ensure quality and utility and to keep the annotation up to date. However, it is imperative to learn how to apply information derived from functional genomics and proteomics technologies to conceptualize and explain the behaviors of biological systems. Quantitative and dynamical models of systems behaviors will supersede the limited and static forms of single-gene annotation that are now the norm. Molecular biological epistemology will increasingly encompass both teleological and causal explanations.

Animals↗

Subfamily hmms in functional genomics.

The limitations of homology-based methods for prediction of protein molecular function are well known; differences in domain structure, gene duplication events and errors in existing database annotations complicate this process. In this paper we present a method to detect and model protein subfamilies, which can be used in high-throughput, genome-scale phylogenomic inference of protein function. We demonstrate the method on a set of nine PFAM families, and show that subfamily HMMs provide greater separation of homologs and non-homologs than is possible with a single HMM for each family. We also show that subfamily HMMs can be used for functional classification with a very low expected error rate. The BETE method for identifying functional subfamilies is illustrated on a set of serotonin receptors.

Animals↗