Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 829 records · Page 46Linked to original sources

ASAP: the Alternative Splicing Annotation Project.

Recently, genomics analyses have demonstrated that alternative splicing is widespread in mammalian genomes (30-60% of genes reported to have multiple isoforms), and may be one of their most important mechanisms of functional regulation. However, by comparison with other genomics data such as genome annotation, SNPs, or gene expression, there exists relatively little database infrastructure for the study of alternative splicing. We have constructed an online database ASAP (the Alternative Splicing Annotation Project) for biologists to access and mine the enormous wealth of alternative splicing information coming from genomics and proteomics. ASAP is based on genome-wide analyses of alternative splicing in human (30 793 alternative splice relationships found) from detailed alignment of expressed sequences onto the genomic sequence. ASAP provides precise gene exon-intron structure, alternative splicing, tissue specificity of alternative splice forms, and protein isoform sequences resulting from alternative splicing. Moreover, it can help biologists design probe sequences for distinguishing specific mRNA isoforms. ASAP is intended to be a community resource for collaborative annotation of alternative splice forms, their regulation, and biological functions. The URL for ASAP is http://www.bioinformatics.ucla.edu/ASAP.

Alternative Splicing↗

The InterPro database, an integrated documentation resource for protein families, domains and functional sites.

Signature databases are vital tools for identifying distant relationships in novel sequences and hence for inferring protein function. InterPro is an integrated documentation resource for protein families, domains and functional sites, which amalgamates the efforts of the PROSITE, PRINTS, Pfam and ProDom database projects. Each InterPro entry includes a functional description, annotation, literature references and links back to the relevant member database(s). Release 2.0 of InterPro (October 2000) contains over 3000 entries, representing families, domains, repeats and sites of post-translational modification encoded by a total of 6804 different regular expressions, profiles, fingerprints and Hidden Markov Models. Each InterPro entry lists all the matches against SWISS-PROT and TrEMBL (more than 1,000,000 hits from 462,500 proteins in SWISS-PROT and TrEMBL). The database is accessible for text- and sequence-based searches at http://www.ebi.ac.uk/interpro/. Questions can be emailed to interhelp@ebi.ac.uk.

Databases, Factual↗

Assessment of the total number of human transcription units.

Variation in the estimates of the number of genes encoded by the human genome (28,000-120,000) attests to the difficulty of systematically identifying human genes. Sequencing of human chromosome 22 (Chr22) provided the first comprehensive, unbiased view of an entire human chromosome, and intensive analysis of this sequence identified 545 genes and 134 pseudogenes that had similarity or identity to known proteins and/or ESTs and which were listed in the gene annotation (http://www.sanger.ac.uk/HGP/Chr22). This analysis yielded an estimate of approximately 36,000 functional expressed genes in the human genome (and 9000 pseudogenes). However, a key uncertainty in this estimate was that hundreds of additional genes beyond those annotated in the Chr22 sequence are predicted by the gene prediction program Genscan, an unknown number of which might represent additional expressed genes. To determine what fraction of these "predicted novel genes" (PNGs) represents expressed human genes, we used a sensitive RT-PCR assay to detect predicted transcripts in 17 tissues and one cell line. Our results indicate that at least 5000-9000 additional human genes which lack similarity to known genes or proteins exist in the human genome, increasing baseline gene estimates to approximately 41,000-45,000.

Chromosomes, Human, Pair 22↗

A proteomic view of Desulfovibrio vulgaris metabolism as determined by liquid chromatography coupled with tandem mass spectrometry.

Direct LC-MS/MS was used to examine the proteins extracted from exponential or stationary phase Desulfovibrio vulgaris cells that had been grown on a minimal medium containing either lactate or formate as the primary carbon source. Across all four growth conditions, 976 gene products were identified with high confidence, which is equal to approximately 28% of all predicted proteins in the D. vulgaris genome. Bioinformatic analysis showed that the proteins identified were distributed among almost all functional classes, with the energy metabolism category containing the greatest number of identified proteins. At least 154 ORFs originally annotated as hypothetical proteins were found to encode the expressed proteins, which provided verification for the authenticity of these hypothetical proteins. Proteomic analysis showed that proteins potentially involved in ATP biosynthesis using the proton gradient across membrane, such as ATPase, alcohol dehydrogenases, heterodisulfide reductases, and [NiFe] hydrogenase (HynAB-1) of the hydrogen cycling were highly expressed in all four growth conditions, suggesting they may be the primary pathways for ATP synthesis in D. vulgaris. Most of the enzymes involved in substrate-level phosphorylation were also detected in all tested conditions. However, no enzyme involved in CO cycling or formate cycling was detected, suggesting that they are not the primary ATP-biosynthesis pathways under the tested conditions. This study provides the first proteomic overview of the cellular metabolism of D. vulgaris. The complete list of proteins identified in this study and their abundances (peptide hits) is provided in Supplementary Table 1.

Alcohol Dehydrogenase↗

A genomics approach to crop pest and disease research.

Genome-wide analyses of gene function and gene expression are beginning to yield valuable information in many areas of biological research, and these genomic tools are now being applied to crop pest and disease research. DNA sequencing of cDNA libraries to generate sets of expressed sequence tags (ESTs) are allowing gene compendiums for crop diseases to be compiled. Annotation of such data collections is also providing a wealth of functional information about gene products through similarities to proteins with known function. The next phase of the functional genomics era will be to employ large-scale techniques to knock out or silence genes in order to synthesize gene-specific mutants for phenotypic analysis and to use micro-array methodology to analyze global gene expression, protein turnover and protein processing during the processes of parasitism and colonization. Application of these technologies promises to accelerate the pace that biological information relevant to crop protection accrues. The ability of researchers to assimilate this information into complex models and workable hypotheses is, thus, set to revolutionize the way we study pests and diseases of crop plants.

Crops, Agricultural↗

The anthracnose resistance locus Co-4 of common bean is located on chromosome 3 and contains putative disease resistance-related genes.

The broadest based resistance to anthracnose of common bean ( Phaseolus vulgaris L.) is conferred by the Co-4 locus. We sequenced a bacterial artificial chromosome clone harboring part of the Co-4 locus of the bean genotype Sprite and assembled a single contig of 106.5 kb for functional annotation. This region contained five copies of the COK-4 gene that encodes for a serine threonine kinase protein previously mapped to the Co-4 locus and 19 novel genes with no similarity to any previously identified genes of common bean. Several putative genes of the Co-4 locus seemed to be expressed as they matched perfectly with bean expressed sequence tags. The expression of the COK-4 genes was assessed by reverse transcription (RT)-PCR, and a single 850-bp cDNA fragment was sequenced and compared with the genomic sequences of the COK-4 homologs. Although the COK-4 cDNA was isolated from a different bean cultivar, it showed high similarity (95%) to the exons of genes BA17 and BA21, suggesting that they were expressed. In a phylogenetic tree including all currently available Pto-like sequences from Phaseolus species, the COK-4 homologs formed a single cluster with the Pto gene, whereas two sequences from P. coccineus and all sequences of P. vulgaris formed two closely related clusters. The Co-4 locus was physically mapped to the short arm of bean chromosome 3, which corresponds to linkage group B8. This study represents a first step in gaining an understanding of the genomic organization of an anthracnose resistance locus of common bean and provides molecular data for comparative analysis with other plant species.

Ascomycota↗

Mining metagenomes from extremophiles as a resource for novel glycoside hydrolases for industrial applications.

The exploration of metagenomes from extremophiles has emerged as a promising approach for discovering novel glycoside hydrolases (GHs) with potential industrial applications. Extremophiles, which thrive in harsh conditions such as high salinity, extreme temperatures, and acidic or alkaline environments, produce enzymes naturally adapted to function under these conditions. This unique adaptability makes them highly desirable for industrial processes requiring robust and efficient biocatalysts. These biocatalysts reduce reliance on harsh chemicals and energy-intensive processes, contributing to greener industrial operations. This review underscores the power of metagenomics in bypassing the need to culture large libraries of extremophiles in the lab. High-throughput sequencing and bioinformatics enable the identification of novel GH-encoding genes directly from environmental DNA. While metagenomic mining has yielded promising results, challenges such as the expression of extremophile-derived genes in mesophilic hosts, low activity yields, and scalability remain. Advances in synthetic biology and protein engineering could address these bottlenecks, enabling more efficient utilization of GHs. Additionally, integrating machine learning for predictive functional annotation may accelerate the identification of high-value candidates.

Glycoside Hydrolases↗

Identification of differentially expressed cDNA sequences in ovaries of sexual and apomictic plants of Brachiaria brizantha.

The isolation of genes associated with apomixis would improve understanding of the molecular mechanism of this mode of reproduction in plants as well as open the possibility of transfer of apomixis to sexual plants, enabling cloning of crops through seeds. Brachiaria brizantha is a highly apomictic grass species with 274 tetraploid apomicts accessions and only one diploid sexual. In this study we have compared gene expression in ovaries at megasporogenesis and megagametogenesis of sexual and apomictic accessions of B. brizantha by differential display (DD-PCR), with 60 primer combinations. Specificity of 65 cloned fragments, checked by reverse northern blot analysis, showed that 11 clones were differentially expressed, 6 in apomictic ovaries, 2 in sexual and 3 in apomictic and sexual, but at different stages. Of the 6 sequences isolated that were preferentially expressed in the apomictic accession: one sequence was from ovaries at megasporogenesis stage; three were from megagametogenesis stage; two were from both stages. Of the two sequences isolated from the sexual accessions, one showed expression in ovaries at megagametogenesis, while the other sequence was shown to be specific to both stages. Three sequences were from megasporogenesis stage in apomicts but were also detected at megagametogenesis in sexual plants. Sequence analysis showed that 5 of the 11 clones had no apparent homologues in the protein database. Some of the clones identified as apomictic-specific shared homology with known genes enabling their functional annotation. The relationships of these functions to the generation of the apomictic trait are discussed.

Amino Acid Sequence↗

Origin and molecular evolution of receptor tyrosine kinases with immunoglobulin-like domains.

Receptor tyrosine kinases (RTKs) are involved in the control of fundamental cellular processes in metazoans. In vertebrates, RTK could be grouped in distinct classes based on the nature of their cognate ligand and modular composition of their extracellular domain. RTK with immunoglobulin-like domains (IG-like RTK) encompass several RTK classes and have been found in early metazoans, including sponges. Evolution of IG-like RTK is characterized by extended molecular and functional diversification, which prompted us to study their evolutionary history. For that purpose, a nonredundant data set including annotated protein sequences of IG-like RTK (n = 85) was built, representing 19 species ranging from sponges to humans. Phylogenetic trees were generated from alignment of conserved regions using maximum likelihood approach. Molecular phylogeny strongly suggests that IG-like RTK diversification occurred according to a complex scenario. In particular, we propose that specific cis duplications of a common ancestor to both platelet-derived growth factor receptor (class III) and vascular endothelial growth factor receptor (class V) families preceded two trans duplications. In contrast, other IG-like RTK genes, like Musk and PTK7, apparently did not evolve by duplications, whereas fibroblast growth factor receptors (class IV) evolved through two rounds of trans duplications. The proposed model of IG-like RTK evolution is supported by high bootstrap values and by the clustering of genes encoding class III and class V RTKs at specific chromosomal locations in mouse and human genomes.

Animals↗

HapScope: a software system for automated and visual analysis of functionally annotated haplotypes.

We have developed a software analysis package, HapScope, which includes a comprehensive analysis pipeline and a sophisticated visualization tool for analyzing functionally annotated haplotypes. The HapScope analysis pipeline supports: (i) computational haplotype construction with an expectation-maximization or Bayesian statistical algorithm; (ii) SNP classification by protein coding change, homology to model organisms or putative regulatory regions; and (iii) minimum SNP subset selection by either a Brute Force Algorithm or a Greedy Partition Algorithm. The HapScope viewer displays genomic structure with haplotype information in an integrated environment, providing eight alternative views for assessing genetic and functional correlation. It has a user-friendly interface for: (i) haplotype block visualization; (ii) SNP subset selection; (iii) haplotype consolidation with subset SNP markers; (iv) incorporation of both experimentally determined haplotypes and computational results; and (v) data export for additional analysis. Comparison of haplotypes constructed by the statistical algorithms with those determined experimentally shows variation in haplotype prediction accuracies in genomic regions with different levels of nucleotide diversity. We have applied HapScope in analyzing haplotypes for candidate genes and genomic regions with extensive SNP and genotype data. We envision that the systematic approach of integrating functional genomic analysis with population haplotypes, supported by HapScope, will greatly facilitate current genetic disease research.

Algorithms↗

Sequence, annotation, and analysis of synteny between rice chromosome 3 and diverged grass species.

Rice (Oryza sativa L.) chromosome 3 is evolutionarily conserved across the cultivated cereals and shares large blocks of synteny with maize and sorghum, which diverged from rice more than 50 million years ago. To begin to completely understand this chromosome, we sequenced, finished, and annotated 36.1 Mb ( approximately 97%) from O. sativa subsp. japonica cv Nipponbare. Annotation features of the chromosome include 5915 genes, of which 913 are related to transposable elements. A putative function could be assigned to 3064 genes, with another 757 genes annotated as expressed, leaving 2094 that encode hypothetical proteins. Similarity searches against the proteome of Arabidopsis thaliana revealed putative homologs for 67% of the chromosome 3 proteins. Further searches of a nonredundant amino acid database, the Pfam domain database, plant Expressed Sequence Tags, and genomic assemblies from sorghum and maize revealed only 853 nontransposable element related proteins from chromosome 3 that lacked similarity to other known sequences. Interestingly, 426 of these have a paralog within the rice genome. A comparative physical map of the wild progenitor species, Oryza nivara, with japonica chromosome 3 revealed a high degree of sequence identity and synteny between these two species, which diverged approximately 10,000 years ago. Although no major rearrangements were detected, the deduced size of the O. nivara chromosome 3 was 21% smaller than that of japonica. Synteny between rice and other cereals using an integrated maize physical map and wheat genetic map was strikingly high, further supporting the use of rice and, in particular, chromosome 3, as a model for comparative studies among the cereals.

Arabidopsis↗

From immunogenetics to immunomics: functional prospecting of genes and transcripts.

Human and mouse genome and transcriptome projects have expanded the field of 'immunogenetics' beyond the traditional study of the genetics and evolution of MHC, TCR and Ig loci into the new interdisciplinary area of 'immunomics'. Immunomics is the study of the molecular functions associated with all immune-related coding and non-coding mRNA transcripts. To unravel the function, regulation and diversity of the immunome requires that we identify and correctly categorize all immune-related transcripts. The importance of intercalated genes, antisense transcripts and non-coding RNAs and their potential role in regulation of immune development and function are only just starting to be appreciated. To better understand immune function and regulation, transcriptome projects (e.g. Functional Annotation of the Mouse, FANTOM), that focus on sequencing full-length transcripts from multiple tissue sources, ideally should include specific immune cells (e.g. T cell, B cells, macrophages, dendritic cells) at various states of development, in activated and unactivated states and in different disease contexts. Progress in deciphering immune regulatory networks will require the cooperative efforts of immunologists, immunogeneticists, molecular biologists and bioinformaticians. Although primary sequence analysis remains useful for annotation of new transcripts it is less useful for identifying novel functions of known transcripts in a new context (protein interaction network or pathway). The most efficient approach to mine useful information from the vast a priori knowledge contained in biological databases and the scientific literature, is to use a combination of computational and expert-driven knowledge discovery strategies. This paper will illustrate the challenges posed in attempts to functionally infer transcriptional regulation and interaction of immune-related genes from text and sequence-based data sources.

Alternative Splicing↗

Differences in the evolutionary history of disease genes affected by dominant or recessive mutations.

BACKGROUND: Global analyses of human disease genes by computational methods have yielded important advances in the understanding of human diseases. Generally these studies have treated the group of disease genes uniformly, thus ignoring the type of disease-causing mutations (dominant or recessive). In this report we present a comprehensive study of the evolutionary history of autosomal disease genes separated by mode of inheritance. RESULTS: We examine differences in protein and coding sequence conservation between dominant and recessive human disease genes. Our analysis shows that disease genes affected by dominant mutations are more conserved than those affected by recessive mutations. This could be a consequence of the fact that recessive mutations remain hidden from selection while heterozygous. Furthermore, we employ functional annotation analysis and investigations into disease severity to support this hypothesis. CONCLUSION: This study elucidates important differences between dominantly- and recessively-acting disease genes in terms of protein and DNA sequence conservation, paralogy and essentiality. We propose that the division of disease genes by mode of inheritance will enhance both understanding of the disease process and prediction of candidate disease genes in the future.

Animals↗

Protein domain analysis in the era of complete genomes.

Domains present one of the most useful levels at which to understand protein function, and domain family-based analysis has had a profound impact on the study of individual proteins. Protein domain discovery has been progressing steadily over the past 30 years. What are the realistically achievable goals of sequence-based domain analysis, and how far off are they for the sequences encoded in eukaryotic genomes? Here we address some of the issues involved in better coverage of sequence-based domain annotation, and the integration of these results within the wider context of genomes, structures and function.

Amino Acid Sequence↗

A gene-centered C. elegans protein-DNA interaction network.

Transcription regulatory networks consist of physical and functional interactions between transcription factors (TFs) and their target genes. The systematic mapping of TF-target gene interactions has been pioneered in unicellular systems, using "TF-centered" methods (e.g., chromatin immunoprecipitation). However, metazoan systems are less amenable to such methods. Here, we used "gene-centered" high-throughput yeast one-hybrid (Y1H) assays to identify 283 interactions between 72 C. elegans digestive tract gene promoters and 117 proteins. The resulting protein-DNA interaction (PDI) network is highly connected and enriched for TFs that are expressed in the digestive tract. We provide functional annotations for approximately 10% of all worm TFs, many of which were previously uncharacterized, and find ten novel putative TFs, illustrating the power of a gene-centered approach. We provide additional in vivo evidence for multiple PDIs and illustrate how the PDI network provides insights into metazoan differential gene expression at a systems level.

Animals↗

Identification of divergent functions in homologous proteins by induction over conserved modules.

Homologous proteins do not necessarily exhibit identical biochemical function. Despite this fact, local or global sequence similarity is widely used as an indication of functional identity. Of the 1327 Enzyme Commission defined functional classes with more than one annotated example in the sequence databases, similarity scores alone are inadequate in 251 (19%) of the cases. We test the hypothesis that conserved domains, as defined in the ProDom database, can be used to discriminate between alternative functions for homologous proteins in these cases. Using machine learning methods, we were able to induce correct discriminators for more than half of these 251 challenging functional classes. These results show that the combination of modular representations of proteins with sequence similarity improves the ability to infer function from sequence over similarity scores alone.

Alcohol Dehydrogenase↗

Prediction of rho-independent transcriptional terminators in Escherichia coli.

A new algorithm called RNAMotif containing RNA structure and sequence constraints and a thermodynamic scoring system was used to search for intrinsic rho-independent terminators in the Escherichia coli K-12 genome. We identified all 135 reported terminators and 940 putative terminator sequences beginning no more than 60 nt away from the 3'-end of the annotated transcription units (TU). Putative and reported terminators with the scores above our chosen threshold were found for 37 of the 53 non-coding RNA TU and for almost 50% of the 2592 annotated protein-encoding TU, which correlates well with the number of TU expected to contain rho-independent terminators. We also identified 439 terminators that could function in a bi-directional fashion, servicing one gene on the positive strand and a different gene on the negative strand. Approximately 700 additional termination signals in non-coding regions (NCR) far away from the nearest annotated gene were predicted. This number correlates well with the excess number of predicted 'orphan' promoters in the NCR, and these promoters and terminators may be associated with as yet unidentified TU. The significant number of high scoring hits that occurred within the reading frame of annotated genes suggests that either an additional component of rho-independent terminators exists or that a suppressive mechanism to prevent unwanted termination remains to be discovered.

Algorithms↗

Reverse genetics in the Arabidopsis chloroplast genome identifies rps16 as a transcribed pseudogene.

The plastid (chloroplast) genomes of seed plants contain a conserved set of ribosomal protein genes. The rps16 gene represents an exception: It has been lost from the plastid genomes of gymnosperms and several lineages of angiosperms, and may have undergone pseudogenization in a few other lineages, including members of the Brassicaceae family. Here we report a reverse genetic approach to test the annotated rps16 gene in the Arabidopsis plastid genome for functionality. Employing the recently developed plastid transformation technology for the model plant Arabidopsis, we have deleted the putative rps16 gene from the Arabidopsis plastid genome. We report that the resulting transplastomic plants display wild-type-like growth and photosynthetic performance under a wide range of conditions. Moreover, genome-wide analyses of chloroplast transcript levels and ribosome footprints revealed unaltered plastid translational activity in Δrps16 mutants compared with wild-type plants. We conclude that the annotated rps16 gene in the plastid genome of Arabidopsis is a transcribed pseudogene that has been replaced in evolution by a nuclear gene copy that supplies functional S16 protein to chloroplasts.

Arabidopsis↗