Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 865 records · Page 48Linked to original sources

CancerGenes: a gene selection resource for cancer genome projects.

The genome sequence framework provided by the human genome project allows us to precisely map human genetic variations in order to study their association with disease and their direct effects on gene function. Since the description of tumor suppressor genes and oncogenes several decades ago, both germ-line variations and somatic mutations have been established to be important in cancer-in terms of risk, oncogenesis, prognosis and response to therapy. The Cancer Genome Atlas initiative proposed by the NIH is poised to elucidate the contribution of somatic mutations to cancer development and progression through the re-sequencing of a substantial fraction of the total collection of human genes-in hundreds of individual tumors and spanning several tumor types. We have developed the CancerGenes resource to simplify the process of gene selection and prioritization in large collaborative projects. CancerGenes combines gene lists annotated by experts with information from key public databases. Each gene is annotated with gene name(s), functional description, organism, chromosome number, location, Entrez Gene ID, GO terms, InterPro descriptions, gene structure, protein length, transcript count, and experimentally determined transcript control regions, as well as links to Entrez Gene, COSMIC, and iHOP gene pages and the UCSC and Ensembl genome browsers. The user-friendly interface provides for searching, sorting and intersection of gene lists. Users may view tabulated results through a web browser or may dynamically download them as a spreadsheet table. CancerGenes is available at http://cbio.mskcc.org/cancergenes.

Databases, Genetic↗

Evaluation of methods for determination of a reconstructed history of gene sequence evolution.

With whole-genome sequences being completed at an increasing rate, it is important to develop and assess tools to analyze them. Following annotation of the protein content of a genome, one can compare sequences with previously characterized homologous genes to detect novel functions within specific proteins in the evolution of the newly sequenced genome. One common statistical method to detect such changes is to compare the ratios of nonsynonymous (K(a)) to synonymous (K(s)) nucleotide substitution rates. Here, the effects of several parameters that can influence this calculation (sequence reconstruction method, phylogenetic tree branch length weighting, GC content, and codon bias) are examined. Also, two new alternative measures of adaptive evolution, the point accepted mutations (PAM)/neutral evolutionary distance (NED) ratio and the sequence space assessment (SSA) statistic are presented. All of these methods are compared using two sequence families: the recent divergence of leptin orthologs in primates, and the more ancient divergence of the deoxyribonucleoside kinase family. The examination of these and other measures to detect changes of gene function along branches of a phylogenetic tree will become increasingly important in the postgenomic era.

Algorithms↗

Graph-based iterative Group Analysis enhances microarray interpretation.

BACKGROUND: One of the most time-consuming tasks after performing a gene expression experiment is the biological interpretation of the results by identifying physiologically important associations between the differentially expressed genes. A large part of the relevant functional evidence can be represented in the form of graphs, e.g. metabolic and signaling pathways, protein interaction maps, shared GeneOntology annotations, or literature co-citation relations. Such graphs are easily constructed from available genome annotation data. The problem of biological interpretation can then be described as identifying the subgraphs showing the most significant patterns of gene expression. We applied a graph-based extension of our iterative Group Analysis (iGA) approach to obtain a statistically rigorous identification of the subgraphs of interest in any evidence graph. RESULTS: We validated the Graph-based iterative Group Analysis (GiGA) by applying it to the classic yeast diauxic shift experiment of DeRisi et al., using GeneOntology and metabolic network information. GiGA reliably identified and summarized all the biological processes discussed in the original publication. Visualization of the detected subgraphs allowed the convenient exploration of the results. The method also identified several processes that were not presented in the original paper but are of obvious relevance to the yeast starvation response. CONCLUSIONS: GiGA provides a fast and flexible delimitation of the most interesting areas in a microarray experiment, and leads to a considerable speed-up and improvement of the interpretation process.

Algorithms↗

Identification and expression of odorant-binding proteins of the malaria-carrying mosquitoes Anopheles gambiae and Anopheles arabiensis.

Host preference and blood feeding are restricted to female mosquitoes. Olfaction plays a major role in host-seeking behaviour, which is likely to be associated with a subset of mosquito olfactory genes. Proteins involved in olfaction include the odorant receptors (ORs) and the odorant-binding proteins (OBPs). OBPs are thought to function as a carrier within insect antennae for transporting odours to the olfactory receptors. Here we report the annotation of 32 genes encoding putative OBPs in the malaria mosquito Anopheles gambiae and their tissue-specific expression in two mosquito species of the Anopheles complex; a highly anthropophilic species An. gambiae sensu stricto and an opportunistic, but more zoophilic species, An. arabiensis. RT-PCR shows that some of the genes are expressed mainly in head tissue and a subset of these show highest expression in female heads. One of the genes (agCP1588) which has not been identified as an OBP, has a high similarity (40%) to the Drosophila pheromone-binding protein 4 (PBPRP4) and is only expressed in heads of both An. gambiae and An. arabiensis, and at higher levels in female heads. Two genes (agCP3071 and agCP15554) are expressed only in female heads and agC15554 also shows higher expression levels in An. gambiae. The expression profiles of the genes in the two members of the Anopheles complex provides the first step towards further molecular analysis of the mosquito olfactory apparatus.

Amino Acid Sequence↗

The Yeast Proteome Database (YPD): a model for the organization and presentation of genome-wide functional data.

The Yeast Proteome Database (YPD) is a model for the organization and presentation of comprehensive protein information. Based on the detailed curation of the scientific literature for the yeast Saccharomyces cerevisiae, YPD contains more than 50 000 annotations lines derived from the review of 8500 research publications. The information concerning each of the approximately 6100 yeast proteins is structured around a convenient one-page format, the Yeast Protein Report, with additional information provided as pop-up windows. Protein classification schema have been revised this year, defining each protein's cellular role, function and pathway, and adding a Functional to the Yeast Protein Report. These changes provide the user with a succinct summary of the protein's function and its place in the biology of the cell, and they enhance the power of YPD Search functions. Precalculated sequence alignments have been added, to provide a crossover point for comparative genomics. The first transcript profiling data has been integrated into the YPD Protein Reports, providing the framework for the presentation of genome-wide functional data. The Yeast Proteome Database can be accessed on the Web at http://www.proteome.com/YPDhome.html

Computational Biology↗

Genomes with distinct function composition.

The functional composition of organisms can be analysed for the first time with the appearance of complete or sizeable parts of various genomes. We have reduced the problem of protein function classification to a simple scheme with three classes of protein function: energy-, information- and communication-associated proteins. Finer classification schemes can be easily mapped to the above three classes. To deal with the vast amount of information, a system for automatic function classification using database annotations has been developed. The system is able to classify correctly about 80% of the query sequences with annotations. Using this system, we can analyse samples from the genomes of the most represented species in sequence databases and compare their genomic composition. The similarities and differences for different taxonomic groups are strikingly intuitive. Viruses have the highest proportion of proteins involved in the control and expression of genetic information. Bacteria have the highest proportion of their genes dedicated to the production of proteins associated with small molecule transformations and transport. Animals have a very large proportion of proteins associated with intra- and intercellular communication and other regulatory processes. In general, the proportion of communication-related proteins increases during evolution, indicating trends that led to the emergence of the eukaryotic cell and later the transition from unicellular to multicellular organisms.

Animals↗

Whole-genome sequencing and analysis of the endophytic fungus Alternaria alternata Y-2 from Leymus chinensis.

To explore the genetic basis and functional potential of beneficial symbiosis between the endophytic fungus Alternaria alternata Y-2 and its host Leymus chinensis, we performed Illumina-based draft whole-genome sequencing and systematic bioinformatic analysis. Although this assembly does not reach telomere-to-telomere completeness, it provides high-quality gene-level information for gene prediction, functional annotation, carbohydrate-active enzyme (CAZyme) identification, and secondary metabolite biosynthetic gene cluster analysis. The final genome size of A. alternata Y-2 was 34,383,676 bp with a GC content of 51.0%, containing 12,724 predicted protein-coding genes, 90 tRNAs, and 12 rRNAs. BUSCO assessment showed 98.9% completeness, supporting the high quality of this draft genome. A total of 12,627 genes were successfully annotated in the NCBI NR database, and 17,183 genes were functionally categorized using GO terms. In total, 448 CAZyme genes and 21 secondary metabolite biosynthetic gene clusters were identified, which are potentially involved in lignocellulose degradation, cellular redox homeostasis and biosynthesis of bioactive metabolites. Based on ITS sequence alignment, NR annotation, and phylogenetic analysis of single-copy orthologous genes, the strain was confidently identified as A. alternata. This study firstly reports the draft genome of an endophytic A. alternata strain derived from L. chinensis and provides valuable genetic resources for exploring the endophytic lifestyle, stress tolerance, and bioactive metabolite potential of this fungus.

Alternaria↗

ProtoNet 4.0: a hierarchical classification of one million protein sequences.

ProtoNet is an automatic hierarchical classification of the protein sequence space. In 2004, the ProtoNet (version 4.0) presents the analysis of over one million proteins merged from SwissProt and TrEMBL databases. In addition to rich visualization and analysis tools to navigate the clustering hierarchy, we incorporated several improvements that allow a simplified view of the scaffold of the proteins. An unsupervised, biologically valid method that was developed resulted in a condensation of the ProtoNet hierarchy to only 12% of the clusters. A large portion of these clusters was automatically assigned high confidence biological names according to their correspondence with functional annotations. ProtoNet is available at: http://www.protonet.cs.huji.ac.il.

Animals↗

Mining the Giardia lamblia genome for new cyst wall proteins.

The Giardia lamblia cyst wall (CW), which is required for survival outside the host and infection, is a primitive extracellular matrix. Because of the importance of the CW, we queried the Giardia Genome Project Database with the coding sequences of the only two known CW proteins, which are cysteine-rich and contain leucine-rich repeats (LRRs). We identified five new LRR-containing proteins, of which only one (CWP3) is up-regulated during encystation and incorporated into the cyst wall. Sequence comparison with CWP1 and -2 revealed conservation within the LRRs and the 44-amino-acid N-flanking region, although CWP3 is more divergent. Interestingly, all 14 cysteine residues of CWP3 are positionally conserved with CWP1 and -2. During encystation, C-terminal epitope-tagged CWP3 was transported to the wall of water-resistant cysts via the novel regulated secretory pathway in encystation-secretory vesicles (ESVs). Deletion analysis revealed that the four LRRs are each essential to target CWP3 to the ESVs and cyst wall. In a deletion of the most C-terminal region, fewer ESVs were stained in encysting cells, and there was no staining in cysts. In contrast, deletion of the 44 amino acids between the signal sequence and the LRRs or the region just C-terminal to the LRRs only decreased the number of cells with CWP3 targeting to ESVs and cyst wall by approximately 50%. Our studies indicate that virtually every portion of the CWP3 protein is needed for efficient targeting to the regulated secretory pathway and incorporation into the cyst wall. Further, these data demonstrate the power of genomics in combination with rigorous functional analyses to verify annotation.

Amino Acid Sequence↗

Mnd1p: an evolutionarily conserved protein required for meiotic recombination.

We used a functional genomics approach to identify a gene required for meiotic recombination, YGL183c or MND1. MND1 was spliced in meiotic cells, extending the annotated YGL183c ORF N terminus by 45 aa. Saccharomyces cerevisiae mnd1-1 mutants, in which the majority of the MND1 coding sequence was removed, arrested before the first meiotic division with a phenotype reminiscent of dmc1 mutants. Physical and genetic analysis showed that these cells initiated recombination, but did not form heteroduplex DNA or double Holliday junctions, suggesting that Mnd1p is involved in strand invasion. Orthologs of MND1 were identified in protists, several yeasts, plants, and mammals, suggesting that its function has been conserved throughout evolution.

Amino Acid Sequence↗

Bioinformatic approaches to assigning protein function from novel sequence data.

The current pace of functional genomic initiatives and genome sequencing projects has provided researchers with a bewildering array of sequence and biological data to analyze. The disease system-driven approach to identifying key genes frequently identifies nucleotide and protein sequences for which the gene and protein function are not known in sufficient detail to allow informed follow-up. Using a range of bioinformatic tools and sequence-based clues, most of unassigned sequences can now be annotated. This chapter takes as an example an unannotated expressed sequence tag, describing how to identify its related gene, and how to annotate the encoded protein using sequence, profile, and structure-based annotation methodologies.

Computational Biology↗

Identification of unstable transcripts in Arabidopsis by cDNA microarray analysis: rapid decay is associated with a group of touch- and specific clock-controlled genes.

mRNA degradation provides a powerful means for controlling gene expression during growth, development, and many physiological transitions in plants and other systems. Rates of decay help define the steady state levels to which transcripts accumulate in the cytoplasm and determine the speed with which these levels change in response to the appropriate signals. When fast responses are to be achieved, rapid decay of mRNAs is necessary. Accordingly, genes with unstable transcripts often encode proteins that play important regulatory roles. Although detailed studies have been carried out on individual genes with unstable transcripts, there is limited knowledge regarding their nature and associations from a genomic perspective, or the physiological significance of rapid mRNA turnover in intact organisms. To address these problems, we have applied cDNA microarray analysis to identify and characterize genes with unstable transcripts in Arabidopsis thaliana (AtGUTs). Our studies showed that at least 1% of the 11,521 clones represented on Arabidopsis Functional Genomics Consortium microarrays correspond to transcripts that are rapidly degraded, with estimated half-lives of less than 60 min. AtGUTs encode proteins that are predicted to participate in a broad range of cellular processes, with transcriptional functions being over-represented relative to the whole Arabidopsis genome annotation. Analysis of public microarray expression data for these genes argues that mRNA instability is of high significance during plant responses to mechanical stimulation and is associated with specific genes controlled by the circadian clock.

Arabidopsis↗

Crystal structure of a putative methyltransferase from Mycobacterium tuberculosis: misannotation of a genome clarified by protein structural analysis.

Bioinformatic analyses of whole genome sequences highlight the problem of identifying the biochemical and cellular functions of many gene products that are at present uncharacterized. The open reading frame Rv3853 from Mycobacterium tuberculosis has been annotated as menG and assumed to encode an S-adenosylmethionine (SAM)-dependent methyltransferase that catalyzes the final step in menaquinone biosynthesis. The Rv3853 gene product has been expressed, refolded, purified, and crystallized in the context of a structural genomics program. Its crystal structure has been determined by isomorphous replacement and refined at 1.9 A resolution to an R factor of 19.0% and R(free) of 22.0%. The structure strongly suggests that this protein is not a SAM-dependent methyltransferase and that the gene has been misannotated in this and other genomes that contain homologs. The protein forms a tightly associated, disk-like trimer. The monomer fold is unlike that of any known SAM-dependent methyltransferase, most closely resembling the phosphohistidine domains of several phosphotransfer systems. Attempts to bind cofactor and substrate molecules have been unsuccessful, but two adventitiously bound small-molecule ligands, modeled as tartrate and glyoxalate, are present on each monomer. These may point to biologically relevant binding sites but do not suggest a function. In silico screening indicates a range of ligands that could occupy these and other sites. The nature of these ligands, coupled with the location of binding sites on the trimer, suggests that proteins of the Rv3853 family, which are distributed throughout microbial and plant species, may be part of a larger assembly binding to nucleic acids or proteins.

Amino Acid Sequence↗

Molecular profile and partial functional analysis of novel endothelial cell-derived growth factors that regulate hematopoiesis.

Recent progress has been made in the identification of the osteoblastic cellular niche for hematopoietic stem cells (HSCs) within the bone marrow (BM). Attempts to identify the soluble factors that regulate HSC self-renewal have been less successful. We have demonstrated that primary human brain endothelial cells (HUBECs) support the ex vivo amplification of primitive human BM and cord blood cells capable of repopulating non-obese diabetic/severe combined immunodeficient repopulating (SCID) mice (SCID repopulating cells [SRCs]). In this study, we sought to characterize the soluble hematopoietic activity produced by HUBECs and to identify the growth factors secreted by HUBECs that contribute to this HSC-supportive effect. Extended noncontact HUBEC cultures supported an eight-fold increase in SRCs when combined with thrombopoietin, stem cell factor, and Flt-3 ligand compared with input CD34(+) cells or cytokines alone. Gene expression analysis of HUBEC biological replicates identified 65 differentially expressed, nonredundant transcripts without annotated hematopoietic activity. Gene ontology studies of the HUBEC transcriptome revealed a high concentration of genes encoding extracellular proteins with cell-cell signaling function. Functional analyses demonstrated that adrenomedullin, a vasodilatory hormone, synergized with stem cell factor and Flt-3 ligand to induce the proliferation of primitive human CD34(+)CD38(-)lin(-) cells and promoted the expansion of CD34(+) progenitors in culture. These data demonstrate the potential of primary HUBECs as a reservoir for the discovery of novel secreted proteins that regulate human hematopoiesis.

Adrenomedullin↗

"Plus-C" odorant-binding protein genes in two Drosophila species and the malaria mosquito Anopheles gambiae.

Olfaction plays a crucial role in many aspects of insect behaviour, including host selection by agricultural pests and vectors of human disease. Insect odorant-binding proteins (OBPs) are thought to function as the first step in molecular recognition and the transport of semiochemicals. The whole genome sequence of the fruit fly Drosophila melanogaster has been completed and a large number of genes have been annotated as OBPs, based on the presence of six conserved cysteine residues and a conserved spacing between the cysteines. These proteins can be divided into three distinct subgroups; those with only one six-cysteine motif, those with two such motifs and those with one motif, three extra conserved cysteines and a conserved proline immediately after the sixth cysteine. This study concentrates on the last two subgroups, referred to as 'dimer' OBPs and 'Plus-C' OBPs, respectively. We determined the tissue-specific transcript levels of all of these OBP genes of D. melanogaster using semiquantitative RT-PCR. The results showed that the expression patterns can vary within a subgroup of genes and that this technique is valuable for assessing which of the putative OBP genes are likely to be involved in Drosophila olfaction. The publicly available genomes of another fruit fly Drosophila pseudoobscura, the malaria mosquito Anopheles gambiae and the yellow fever mosquito Aedes aegypti were searched by Blast against each Plus-C OBP and dimer OBP of D. melanogaster. Related genes were found in all of the other species and the relationships of these with the D. melanogaster genes and their possible biological functions are discussed.

Amino Acid Sequence↗

Cause and effect considerations in diagnostic pathology and pathology phenotyping of genetically engineered mice (GEM).

Over the next several decades, biology is embarking on its most ambitious project yet: to annotate the human genome functionally, prioritizing and focusing on those genes relevant to development and disease. Model systems are fundamental prerequisites for this task, and genetically engineered mice (GEM) are by far the most accessible mammalian system because of their anatomical, physiological, and genetic similarity to humans. The scientific utility of GEM has become commonplace since the technology to produce them was established in the early 1980s. Conceptually, however, an efficiently coordinated high-throughput approach that permits correlation between newly discovered genes, functional properties of their protein products, and biological relevance of these products as drug targets has yet to be established. The discipline of veterinary anatomical pathology (hereafter referred to as pathology) is not immune to this requirement for evolution and adaptation, and to address relationships and tissue consequences between tens of thousands of genes and their cognate proteins, novel interdisciplinary technologies and approaches must emerge. Although many of the techniques of pathology are well established, in the context of pathology's contribution to functional annotation of the genome, several conceptually important and unresolved issues remain to be addressed. While an ever-increasing arsenal of genetic and molecular tool-sets are available to evaluate and understand the function of genes and their pathophysiological mechanisms, pathology will continue to play an essential role in confirming cause and effect relationships of gene function in development and disease. This role will continue to be dependent on keen observation, a systematic but disciplined approach, expert knowledge of strain-dependent anatomical differences and incidental lesions, and relevant tissue-based evidence. Miniaturization and high-throughput adaptation of these methods must also continue so that they can complement parallel phenotyping efforts, provide pathology-based data in pace with concurrent phenotyping efforts, and continue to find new utility in the collective effort of functional annotation.

Animals↗

A novel beta-glucanase gene from Bacillus halodurans C-125.

A novel endo-beta-1,3(4)-D-glucanase gene was found in the complete genome sequence of Bacillus halodurans C-125. The gene was previously annotated as an "unknown" protein and assigned an incorrect open reading frame (ORF). However, determining the biochemical characteristics has elucidated the function and correct ORF of the gene. The gene encodes 231 amino acids, and its calculated molecular mass was estimated to be 26743.16 Da. The amino acid sequence alignment showed that the highest sequence identity was only 28% with that of the beta-1,3-1,4-glucanase from Bacillus subtilis. Moreover, the nucleotide sequence did not match any other known Bacillus beta-glucanase gene. The member of the gene cluster that includes this novel gene was apparently different from that of the gene cluster including the putative beta-glucanase genes (bh3231 and bh3232) from B. halodurans C-125. Therefore, the novel gene is not a copy of either of these genes, and in B. halodurans cells, the putative role of the encoded protein may differ from that of bh3231 and bh3232. To examine the activity of the gene product, the gene was cloned as a His-tagged protein and expressed in Escherichia coli. The purified enzyme showed activity against lichenan, barley beta-glucan, laminarin, and carboxymethyl curdlan. Thin-layer chromatography showed that the enzyme hydrolyzes substrates in an endo-type manner. When beta-glucan was used as a substrate, the pH optimum was between 6 and 8, and the temperature optimum was 60 degrees C. After 2 h incubation at 50 and 60 degrees C, the residual activity remained 100% and 50%, respectively. The enzymatic activity was abolished after 30 min incubation at 70 degrees C. Based on the results, the gene encodes an endo-type beta-1,3(4)-D-glucanase (E.C. 3.2.1.6).

Amino Acid Sequence↗

An overview of the genome of Nostoc punctiforme, a multicellular, symbiotic cyanobacterium.

Nostoc punctiforme is a filamentous cyanobacterium with extensive phenotypic characteristics and a relatively large genome, approaching 10 Mb. The phenotypic characteristics include a photoautotrophic, diazotrophic mode of growth, but N. punctiforme is also facultatively heterotrophic; its vegetative cells have multiple developmental alternatives, including terminal differentiation into nitrogen-fixing heterocysts and transient differentiation into spore-like akinetes or motile filaments called hormogonia; and N. punctiforme has broad symbiotic competence with fungi and terrestrial plants, including bryophytes, gymnosperms and an angiosperm. The shotgun-sequencing phase of the N. punctiforme strain ATCC 29133 genome has been completed by the Joint Genome Institute. Annotation of an 8.9 Mb database yielded 7432 open reading frames, 45% of which encode proteins with known or probable known function and 29% of which are unique to N. punctiforme. Comparative analysis of the sequence indicates a genome that is highly plastic and in a state of flux, with numerous insertion sequences and multilocus repeats, as well as genes encoding transposases and DNA modification enzymes. The sequence also reveals the presence of genes encoding putative proteins that collectively define almost all characteristics of cyanobacteria as a group. N. punctiforme has an extensive potential to sense and respond to environmental signals as reflected by the presence of more than 400 genes encoding sensor protein kinases, response regulators and other transcriptional factors. The signal transduction systems and any of the large number of unique genes may play essential roles in the cell differentiation and symbiotic interaction properties of N. punctiforme.

Journal Article↗