Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,297 records · Page 72Linked to original sources

EST-PAC a web package for EST annotation and protein sequence prediction.

With the decreasing cost of DNA sequencing technology and the vast diversity of biological resources, researchers increasingly face the basic challenge of annotating a larger number of expressed sequences tags (EST) from a variety of species. This typically consists of a series of repetitive tasks, which should be automated and easy to use. The results of these annotation tasks need to be stored and organized in a consistent way. All these operations should be self-installing, platform independent, easy to customize and amenable to using distributed bioinformatics resources available on the Internet. In order to address these issues, we present EST-PAC a web oriented multi-platform software package for expressed sequences tag (EST) annotation. EST-PAC provides a solution for the administration of EST and protein sequence annotations accessible through a web interface. Three aspects of EST annotation are automated: 1) searching local or remote biological databases for sequence similarities using Blast services, 2) predicting protein coding sequence from EST data and, 3) annotating predicted protein sequences with functional domain predictions. In practice, EST-PAC integrates the BLASTALL suite, EST-Scan2 and HMMER in a relational database system accessible through a simple web interface. EST-PAC also takes advantage of the relational database to allow consistent storage, powerful queries of results and, management of the annotation process. The system allows users to customize annotation strategies and provides an open-source data-management environment for research and education in bioinformatics.

Journal Article↗

Genome-wide protein interaction maps using two-hybrid systems.

Automated sequence technology has rendered functional biology amenable to genomic scale analysis. Among genome-wide exploratory approaches, the two-hybrid system in yeast (Y2H) has outranked other techniques because it is the system of choice to detect protein-protein interactions. Deciphering the cascade of binding events in a whole cell helps define signal transduction and metabolic pathways or enzymatic complexes. The function of proteins is eventually attributed through whole cell protein interaction maps where totally unknown proteins are partnered with fully annotated proteins belonging to the same functional category. Since its first description in the late 1980's, several versions of the Y2H have been developed in order to overcome the major limitations of the system, namely false positives and false negatives. Optimized versions have been recently applied at multi-molecular and genomic scale. These genome-wide surveys can be methodologically divided into two types of approaches: one either tests combinations of predefined polypeptides (the so-called matrix approach) using various short-cuts to speed up the process, or one screens with a given polypeptide (bait) for potential partners (preys) present in complex libraries of genomic or complementary DNA (library screening). In the former strategy, one tests what one knows, for example pair-wise interactions between full-length open reading frames from recently sequenced and annotated genomes. Although based on a one-by-one scheme, this method is reported to be amenable to large-scale genomics thanks to multicloning strategies and to the use of small robotics workstations. In the latter, highly complex cDNA or genomic libraries of protein domains can be screened to saturation with high-throughput screening systems allowing the discovery of yet unidentified proteins. Both approaches have strengths and drawbacks that will be discussed here. None yields a full proteome-wide screening since certain proteins (e.g. some transcription factors) are not usable in Y2H. Novel two-hybrid assays have been recently described in bacteria. Applications of these time- and cost-effective assays to genomic screening will be discussed and compared to the Y2H technology.

Animals↗

Automated prediction of protein function and detection of functional sites from structure.

Current structural genomics projects are yielding structures for proteins whose functions are unknown. Accordingly, there is a pressing requirement for computational methods for function prediction. Here we present PHUNCTIONER, an automatic method for structure-based function prediction using automatically extracted functional sites (residues associated to functions). The method relates proteins with the same function through structural alignments and extracts 3D profiles of conserved residues. Functional features to train the method are extracted from the Gene Ontology (GO) database. The method extracts these features from the entire GO hierarchy and hence is applicable across the whole range of function specificity. 3D profiles associated with 121 GO annotations were extracted. We tested the power of the method both for the prediction of function and for the extraction of functional sites. The success of function prediction by our method was compared with the standard homology-based method. In the zone of low sequence similarity (approximately 15%), our method assigns the correct GO annotation in 90% of the protein structures considered, approximately 20% higher than inheritance of function from the closest homologue.

Amino Acid Sequence↗

Target SNP selection in complex disease association studies.

BACKGROUND: The massive amount of SNP data stored at public internet sites provides unprecedented access to human genetic variation. Selecting target SNP for disease-gene association studies is currently done more or less randomly as decision rules for the selection of functional relevant SNPs are not available. RESULTS: We implemented a computational pipeline that retrieves the genomic sequence of target genes, collects information about sequence variation and selects functional motifs containing SNPs. Motifs being considered are gene promoter, exon-intron structure, AU-rich mRNA elements, transcription factor binding motifs, cryptic and enhancer splice sites together with expression in target tissue. As a case study, 396 genes on chromosome 6p21 in the extended HLA region were selected that contributed nearly 20,000 SNPs. By computer annotation ~2,500 SNPs in functional motifs could be identified. Most of these SNPs are disrupting transcription factor binding sites but only those introducing new sites had a significant depressing effect on SNP allele frequency. Other decision rules concern position within motifs, the validity of SNP database entries, the unique occurrence in the genome and conserved sequence context in other mammalian genomes. CONCLUSION: Only 10% of all gene-based SNPs have sequence-predicted functional relevance making them a primary target for genotyping in association studies.

Amino Acid Substitution↗

It's all GO for plant scientists.

The Gene Ontology project (http://www.geneontology.org/) produces structured, controlled vocabularies and gene product annotations. Gene products are classified according to the cellular locations and biological process in which they act, and the molecular functions that they carry out. We annotate gene products from a broad range of model species and provide support for those groups that wish to contribute annotation of further model species. The Gene Ontology facilitates the exchange of information between groups of scientists studying similar processes in different model organisms, and so provides a broad range of opportunities for plant scientists.

Genes, Plant↗

Understanding the global properties of functionally-related gene networks using the gene ontology.

The global behavior of interactions between genes can be investigated by forming the network of functionally-related genes using the annotations based on the Gene Ontology. We define two genes to be connected when the pair of genes is involved in the same biological process. There has been other work on the analysis of different kinds of cellular and metabolic networks, such as gene coexpression network, in which genes are paired when they are found to be coexpressed in the microarray experiments. We observe that our functionally-related gene networks among humans, fruit flies, worms and yeast exhibit the small-world property, but all except the network of worms show the existence of the scale-free property.

Animals↗

Whole genome shotgun sequencing of Brassica oleracea and its application to gene discovery and annotation in Arabidopsis.

Through comparative studies of the model organism Arabidopsis thaliana and its close relative Brassica oleracea, we have identified conserved regions that represent potentially functional sequences overlooked by previous Arabidopsis genome annotation methods. A total of 454,274 whole genome shotgun sequences covering 283 Mb (0.44 x) of the estimated 650 Mb Brassica genome were searched against the Arabidopsis genome, and conserved Arabidopsis genome sequences (CAGSs) were identified. Of these 229,735 conserved regions, 167,357 fell within or intersected existing gene models, while 60,378 were located in previously unannotated regions. After removal of sequences matching known proteins, CAGSs that were close to one another were chained together as potentially comprising portions of the same functional unit. This resulted in 27,347 chains of which 15,686 were sufficiently distant from existing gene annotations to be considered a novel conserved unit. Of 192 conserved regions examined, 58 were found to be expressed in our cDNA populations. Rapid amplification of cDNA ends (RACE) was used to obtain potentially full-length transcripts from these 58 regions. The resulting sequences led to the creation of 21 gene models at 17 new Arabidopsis loci and the addition of splice variants or updates to another 19 gene structures. In addition, CAGSs overlapping already annotated genes in Arabidopsis can provide guidance for manual improvement of existing gene models. Published genome-wide expression data based on whole genome tiling arrays and massively parallel signature sequencing were overlaid on the Brassica-Arabidopsis conserved sequences, and 1399 regions of intersection were identified. Collectively our results and these data sets suggest that several thousand new Arabidopsis genes remain to be identified and annotated.

Arabidopsis↗

Complete nucleotide sequence and genome analysis of bacteriophage BFK20--a lytic phage of the industrial producer Brevibacterium flavum.

The entire double-stranded DNA genome of bacteriophage BFK20, a lytic phage of the Brevibacterium flavum CCM 251--industrial producer of L-lysine--was sequenced and analyzed. It consists of 42,968 base pairs with an overall molar G + C content of 56.2%. Fifty-five potential open reading frames were identified and annotated using various bioinformatics tools. Clusters of functionally related putative genes were defined (structural, lytic, replication and regulatory). To verify the annotation of structural proteins, they were resolved by 2D gel electrophoresis and were submitted to N-terminal amino acid sequencing. Structural proteins identified included the portal and major and minor tail proteins. Based on the overall genome sequence comparison, similarities with other known bacteriophage genomes include primarily bacteriophages from Mycobacterium spp. and some regions of Corynebacterium spp. genomes--possible prophages. Our results support the theory that phage genomes are mosaics with respect to each other.

Bacteriophages↗

ECLIPSE: exploring the dark proteome of ESKAPE pathogens through the sequence similarity network of the Protein Universe Atlas.

MOTIVATION: The accelerating crisis of antimicrobial resistance among the critical so-called ESKAPE pathogens demands the urgent identification of novel molecular targets. However, a substantial fraction of ESKAPE proteomes remains functionally uncharacterized, with many genes annotated as encoding hypothetical proteins. These protein sequences often lack significant similarity to known protein families when conventional homology-based annotation methods are used and thus remain "dark". This limits our ability to explore their roles in pathogenicity, and it is thus crucial to bridge this substantial gap in pathogen biology by developing new strategies to illuminate these "dark" regions of the ESKAPE pan-proteome. RESULTS: We introduce ECLIPSE (ESKAPE Connectome Linkage and Inference for Proteome Sequence Exploration), a network-based computational framework that systematically identifies and prioritizes functionally dark protein families in ESKAPE pan-proteomes. ECLIPSE embeds target ESKAPE pathogen proteomes within the global sequence similarity network of the Protein Universe Atlas. It detects connected components composed entirely of unannotated proteins, called the "dark proteome." As a case study, we applied ECLIPSE to a pan-proteome of 3 460 657 protein sequences from 635 strains of Pseudomonas aeruginosa (PA). ECLIPSE identified 120 985 proteins (4%) residing in completely dark connected components. Furthermore, we have performed a taxonomic diversity analysis using normalized Shannon indices to characterize each dark component by its enrichment in ESKAPE pathogens. The analysis utilized the evenness (E) value (see Methods 2.1), which distinguishes Pseudomonas-specific (target-specific) from ESKAPE-enriched dark components. We then developed the Dark Proteome Prioritization Score (DPPS), a composite multidimensional scoring framework (see Methods 2.5). It ranks these dark components by biological relevance across four orthogonal axes: (i) functional darkness, (ii) P. aeruginosa proportion in the Atlas, (iii) AMR-clade taxonomic restriction, and (iv) conservation across the 635 P. aeruginosa strains. This framework outputs a robust four-tier scoring system; the prioritized Tier I components were validated by weight sensitivity analysis and remained stable across 500 Monte Carlo weight perturbations. Structural characterization of one of the top-ranked ESKAPE-enriched dark components revealed that it belongs to the beta-barrel fold DUF1302 (PF06980) family, for which no experimentally solved three-dimensional structure exists in the PDB. The genomic context analysis indicates that it is co-localized with a LuxR-type transcriptional regulator. Collectively, ECLIPSE identifies evolutionarily conserved, structurally defined, and functionally dark proteins enriched across ESKAPE pathogens; these dark proteins can further be utilized as alternative antimicrobial targets for experimental characterization. AVAILABILITY AND IMPLEMENTATION: The source code and dataset are available for free at: Github: https://github.com/surabhilata/ECLIPSE.git, Zenodo: DOI: 10.5281/zenodo.21064323.

Proteome↗

Genome sequences and great expectations.

To assess how automatic function assignment will contribute to genome annotation in the next five years, we have performed an analysis of 31 available genome sequences. An emerging pattern is that function can be predicted for almost two-thirds of the 73,500 genes that were analyzed. Despite progress in computational biology, there will always be a great need for large-scale experimental determination of protein function.

Animals↗

Microarray phenotyping in Dictyostelium reveals a regulon of chemotaxis genes.

MOTIVATION: Coordinate regulation of gene expression can provide information on gene function. To begin a large-scale analysis of Dictyostelium gene function, we clustered genes based on their expression in wild-type and mutant strains and analyzed their functions. RESULTS: We found 17 modes of wild-type gene expression and refined them into 57 submodes considering mutant data. Annotation analyses revealed correlations between co-expression and function and an unexpected correlation between expression and function of genes involved in various aspects of chemotaxis. Co-regulation of chemotaxis genes was also found in published data from neutrophils. To test the predictive power of the analysis, we examined the phenotypes of mutations in seven co-regulated genes that had no published role in chemotaxis. Six mutants exhibited chemotaxis defects, supporting the idea that function can be inferred from co-expression. The clustering and annotation analyses provide a public resource for Dictyostelium functional genomics.

Animals↗

Full-genome RNAi profiling of early embryogenesis in Caenorhabditis elegans.

A key challenge of functional genomics today is to generate well-annotated data sets that can be interpreted across different platforms and technologies. Large-scale functional genomics data often fail to connect to standard experimental approaches of gene characterization in individual laboratories. Furthermore, a lack of universal annotation standards for phenotypic data sets makes it difficult to compare different screening approaches. Here we address this problem in a screen designed to identify all genes required for the first two rounds of cell division in the Caenorhabditis elegans embryo. We used RNA-mediated interference to target 98% of all genes predicted in the C. elegans genome in combination with differential interference contrast time-lapse microscopy. Through systematic annotation of the resulting movies, we developed a phenotypic profiling system, which shows high correlation with cellular processes and biochemical pathways, thus enabling us to predict new functions for previously uncharacterized genes.

Animals↗

Analysis of the chromosome sequence of the legume symbiont Sinorhizobium meliloti strain 1021.

Sinorhizobium meliloti is an alpha-proteobacterium that forms agronomically important N(2)-fixing root nodules in legumes. We report here the complete sequence of the largest constituent of its genome, a 62.7% GC-rich 3,654,135-bp circular chromosome. Annotation allowed assignment of a function to 59% of the 3,341 predicted protein-coding ORFs, the rest exhibiting partial, weak, or no similarity with any known sequence. Unexpectedly, the level of reiteration within this replicon is low, with only two genes duplicated with more than 90% nucleotide sequence identity, transposon elements accounting for 2.2% of the sequence, and a few hundred short repeated palindromic motifs (RIME1, RIME2, and C) widespread over the chromosome. Three regions with a significantly lower GC content are most likely of external origin. Detailed annotation revealed that this replicon contains all housekeeping genes except two essential genes that are located on pSymB. Amino acid/peptide transport and degradation and sugar metabolism appear as two major features of the S. meliloti chromosome. The presence in this replicon of a large number of nucleotide cyclases with a peculiar structure, as well as of genes homologous to virulence determinants of animal and plant pathogens, opens perspectives in the study of this bacterium both as a free-living soil microorganism and as a plant symbiont.

Bacterial Proteins↗

GAMOLA: a new local solution for sequence annotation and analyzing draft and finished prokaryotic genomes.

Laboratories working with draft phase genomes have specific software needs, such as the unattended processing of hundreds of single scaffolds and subsequent sequence annotation. In addition, it is critical to follow the "movement" and the manual annotation of single open reading frames (ORFs) within the successive sequence updates. Even with finished genomes, regular database updates can lead to significant changes in the annotation of single ORFs. In functional genomics it is important to mine data and identify new genetic targets rapidly and easily. Often there is no need for sophisticated relational databases (RDB) that greatly reduce the system-independent access of the results. Another aspect is the internet dependency of most software packages. If users are working with confidential data, this dependency poses a security issue. GAMOLA was designed to handle the numerous scaffolds and changing contents of draft phase genomes in an automated process and stores the results for each predicted ORF in flatfile databases. In addition, annotation transfers, ORF designation tracking, Blast comparisons, and primer design for whole genome microarrays have been implemented. The software is available under the license of North Carolina State University. A website and a downloadable example are accessible under (http://fsweb2.schaub. ncsu.edu/TRKwebsite/index.htm).

Algorithms↗

Genome annotation assessment in Drosophila melanogaster.

Computational methods for automated genome annotation are critical to our community's ability to make full use of the large volume of genomic sequence being generated and released. To explore the accuracy of these automated feature prediction tools in the genomes of higher organisms, we evaluated their performance on a large, well-characterized sequence contig from the Adh region of Drosophila melanogaster. This experiment, known as the Genome Annotation Assessment Project (GASP), was launched in May 1999. Twelve groups, applying state-of-the-art tools, contributed predictions for features including gene structure, protein homologies, promoter sites, and repeat elements. We evaluated these predictions using two standards, one based on previously unreleased high-quality full-length cDNA sequences and a second based on the set of annotations generated as part of an in-depth study of the region by a group of Drosophila experts. Although these standard sets only approximate the unknown distribution of features in this region, we believe that when taken in context the results of an evaluation based on them are meaningful. The results were presented as a tutorial at the conference on Intelligent Systems in Molecular Biology (ISMB-99) in August 1999. Over 95% of the coding nucleotides in the region were correctly identified by the majority of the gene finders, and the correct intron/exon structures were predicted for >40% of the genes. Homology-based annotation techniques recognized and associated functions with almost half of the genes in the region; the remainder were only identified by the ab initio techniques. This experiment also presents the first assessment of promoter prediction techniques for a significant number of genes in a large contiguous region. We discovered that the promoter predictors' high false-positive rates make their predictions difficult to use. Integrating gene finding and cDNA/EST alignments with promoter predictions decreases the number of false-positive classifications but discovers less than one-third of the promoters in the region. We believe that by establishing standards for evaluating genomic annotations and by assessing the performance of existing automated genome annotation tools, this experiment establishes a baseline that contributes to the value of ongoing large-scale annotation projects and should guide further research in genome informatics.

Alcohol Dehydrogenase↗

Visual representation of database search results: the RHIMS Plot.

SUMMARY: An algorithm and software are described that provide a fast method to produce a novel, function-oriented visualization of the results of a sequence database search. Text mining of sequence annotations allows position specific plots of potential functional similarity to be compared in a simple compact representation. AVAILABILITY: The application can be accessed via a web server at http://www.compbio.dundee.ac.uk. The RHIMS software may be obtained by request to the authors.

Algorithms↗

CancerGenes: a gene selection resource for cancer genome projects.

The genome sequence framework provided by the human genome project allows us to precisely map human genetic variations in order to study their association with disease and their direct effects on gene function. Since the description of tumor suppressor genes and oncogenes several decades ago, both germ-line variations and somatic mutations have been established to be important in cancer-in terms of risk, oncogenesis, prognosis and response to therapy. The Cancer Genome Atlas initiative proposed by the NIH is poised to elucidate the contribution of somatic mutations to cancer development and progression through the re-sequencing of a substantial fraction of the total collection of human genes-in hundreds of individual tumors and spanning several tumor types. We have developed the CancerGenes resource to simplify the process of gene selection and prioritization in large collaborative projects. CancerGenes combines gene lists annotated by experts with information from key public databases. Each gene is annotated with gene name(s), functional description, organism, chromosome number, location, Entrez Gene ID, GO terms, InterPro descriptions, gene structure, protein length, transcript count, and experimentally determined transcript control regions, as well as links to Entrez Gene, COSMIC, and iHOP gene pages and the UCSC and Ensembl genome browsers. The user-friendly interface provides for searching, sorting and intersection of gene lists. Users may view tabulated results through a web browser or may dynamically download them as a spreadsheet table. CancerGenes is available at http://cbio.mskcc.org/cancergenes.

Databases, Genetic↗

The prokaryotic selenoproteome.

In the genetic code, the UGA codon has a dual function as it encodes selenocysteine (Sec) and serves as a stop signal. However, only the translation terminator function is used in gene annotation programs, resulting in misannotation of selenoprotein genes. Here, we applied two independent bioinformatics approaches to characterize a selenoprotein set in prokaryotic genomes. One method searched for selenoprotein genes by identifying RNA stem-loop structures, selenocysteine insertion sequence elements; the second approach identified Sec/Cys pairs in homologous sequences. These analyses identified all or almost all selenoproteins in completely sequenced bacterial and archaeal genomes and provided a view on the distribution and composition of prokaryotic selenoproteomes. In addition, lineage-specific and core selenoproteins were detected, which provided insights into the mechanisms of selenoprotein evolution. Characterization of selenoproteomes allows interpretation of other UGA codons in completed genomes of prokaryotes as terminators, addressing the UGA dual-function problem.

Amino Acid Sequence↗