Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,333 records · Page 74Linked to original sources

Identification of novel human genes evolutionarily conserved in Caenorhabditis elegans by comparative proteomics.

Modern biomedical research greatly benefits from large-scale genome-sequencing projects ranging from studies of viruses, bacteria, and yeast to multicellular organisms, like Caenorhabditis elegans. Comparative genomic studies offer a vast array of prospects for identification and functional annotation of human ortholog genes. We presented a novel comparative proteomic approach for assembling human gene contigs and assisting gene discovery. The C. elegans proteome was used as an alignment template to assist in novel human gene identification from human EST nucleotide databases. Among the available 18,452 C. elegans protein sequences, our results indicate that at least 83% (15,344 sequences) of C. elegans proteome has human homologous genes, with 7,954 records of C. elegans proteins matching known human gene transcripts. Only 11% or less of C. elegans proteome contains nematode-specific genes. We found that the remaining 7,390 sequences might lead to discoveries of novel human genes, and over 150 putative full-length human gene transcripts were assembled upon further database analyses. [The sequence data described in this paper have been submitted to the

Amino Acid Sequence↗

Control of SXT integration and excision.

The Vibrio cholerae SXT element is a conjugative self-transmissible chromosomally integrating element that encodes resistance to multiple antibiotics. SXT integrates in a site-specific fashion at prfC and excises from the chromosome to form a circular but nonreplicative extrachromosomal form. Both chromosomal integration and excision depend on an SXT-encoded recombinase, Int. Here we found that Int is necessary and sufficient for SXT integration and that int expression in recipient cells requires the SXT activators SetC and SetD. Although no xis-like gene was annotated in the SXT genome, Int was not sufficient to mediate efficient SXT chromosomal excision. We identified a novel SXT Xis that seems to function as a recombination directionality factor (RDF), facilitating SXT excision and inhibiting SXT integration. Although unrelated to any previously characterized RDF, Xis is similar to five hypothetical proteins that together may constitute a new family of RDFs. Using real-time quantitative PCR assays to study SXT excision from the chromosome, we determined that while SXT excision is required for SXT transfer, the percentage of cells containing an excised circular SXT does not appear to be a major factor limiting SXT transfer; i.e., we found that most cells harboring an excised circular SXT molecule do not act as SXT donors. In the absence of prfC, SXT integrated into several secondary attachment sites but preferentially into the 5' end of pntB. SXT excision and transfer from a donor containing pntB::SXT were reduced, suggesting that the SXT integration site may also influence the element's transmissibility.

Amino Acid Sequence↗

Drosophila melanogaster as a model for studying protein-encoding genes that are resident in constitutive heterochromatin.

The organization of chromosomes into euchromatin and heterochromatin is one of the most enigmatic aspects of genome evolution. For a long time, heterochromatin was considered to be a genomic wasteland, incompatible with gene expression. However, recent studies--primarily conducted in Drosophila melanogaster--have shown that this peculiar genomic component performs important cellular functions and carries essential genes. New research on the molecular organization, function and evolution of heterochromatin has been facilitated by the sequencing and annotation of heterochromatic DNA. About 450 predicted genes have been identified in the heterochromatin of D. melanogaster, indicating that the number of active genes is higher than had been suggested by genetic analysis. Most of the essential genes are still unknown at the molecular level, and a detailed functional analysis of the predicted genes is difficult owing to the lack of mutant alleles. Far from being a peculiarity of Drosophila, heterochromatic genes have also been found in Saccharomyces cerevisiae, Schizosaccharomyces pombe, Oryza sativa and Arabidopsis thaliana, as well as in humans. The presence of expressed genes in heterochromatin seems paradoxical because they appear to function in an environment that has been considered incompatible with gene expression. In the future, genetic, functional genomic and proteomic analyses will offer powerful approaches with which to explore the functions of heterochromatic genes and to elucidate the mechanisms driving their expression.

Animals↗

Transcriptomic insights into the coordinated regulation of signaling, apoptosis, immunity, and metabolism during Sinonovacula constricta larval metamorphosis.

Metamorphosis is a critical ontogenetic transition for marine bivalves, marking the shift from planktonic to benthic lifestyles, where successful transformation dictates survival. The razor clam Sinonovacula constricta is economically important; however, low larval metamorphosis rates remain a major bottleneck in seedling production. To elucidate the mechanisms governing this process, we performed a comparative transcriptome analysis of S. constricta larvae at pre- and post-metamorphosis stages using Illumina sequencing. A total of 3701 differentially expressed genes (DEGs) were identified, including 3254 up-regulated and 447 down-regulated genes. Functional annotation of the respective top 20 significantly up-regulated and down-regulated DEGs indicated their potential pivotal roles in signal transduction (e.g., up-regulated: CAV1, CHRNA2; down-regulated: APP, NOTCH1), cellular proliferation and differentiation (e.g., up-regulated: TUBA, EGF1; down-regulated: KIF23, TTC25), transcriptional and epigenetic regulation (e.g., up-regulated: NFIL3; down-regulated: OVO, HMX1), substance transport (e.g., up-regulated: LRP2, LRP1B; down-regulated: SLC51A, Slc33a1), substance metabolism (e.g., up-regulated: CPK3, CYP26A1; down-regulated: RDMT1, ADAC), immunomodulation (e.g., up-regulated: CPN2, CRISP2), and protein homeostasis (e.g., up-regulated: HSP27, NAS-27). Functional enrichment analysis further revealed that DEGs were significantly enriched in pathways related to signal transduction and developmental regulation (e.g., Ras, TNF), cell death and homeostasis (e.g., apoptosis), immune responses (e.g., Toll-like receptor), energy metabolism (e.g., lipid), cardiovascular related (e.g., Fluid shear stress), cell junction and architecture (e.g., Tight junction), and infectious disease (e.g., measles). These results suggest a synergistic interplay between signaling, apoptosis, immunity, and metabolism during S. constricta metamorphosis. This study advances our understanding of marine bivalve metamorphosis and offers candidate genes for further mechanistic studies.

Animals↗

Genome-wide analysis of polymerase III-transcribed Alu elements suggests cell-type-specific enhancer function.

Alu elements are one of the most successful families of transposons in the human genome. A portion of Alu elements is transcribed by RNA Pol III, whereas the remaining ones are part of Pol II transcripts. Because Alu elements are highly repetitive, it has been difficult to identify the Pol III-transcribed elements and quantify their expression levels. In this study, we generated high-resolution, long-genomic-span RAMPAGE data in 155 biosamples all with matching RNA-seq data and built an atlas of 17,249 Pol III-transcribed Alu elements. We further performed an integrative analysis on the ChIP-seq data of 10 histone marks and hundreds of transcription factors, whole-genome bisulfite sequencing data, ChIA-PET data, and functional data in several biosamples, and our results revealed that although the human-specific Alu elements are transcriptionally repressed, the older, expressed Alu elements may be exapted by the human host to function as cell-type-specific enhancers for their nearby protein-coding genes.

Alu Elements↗

Ensemble attribute profile clustering: discovering and characterizing groups of genes with similar patterns of biological features.

BACKGROUND: Ensemble attribute profile clustering is a novel, text-based strategy for analyzing a user-defined list of genes and/or proteins. The strategy exploits annotation data present in gene-centered corpora and utilizes ideas from statistical information retrieval to discover and characterize properties shared by subsets of the list. The practical utility of this method is demonstrated by employing it in a retrospective study of two non-overlapping sets of genes defined by a published investigation as markers for normal human breast luminal epithelial cells and myoepithelial cells. RESULTS: Each genetic locus was characterized using a finite set of biological properties and represented as a vector of features indicating attributes associated with the locus (a gene attribute profile). In this study, the vector space models for a pre-defined list of genes were constructed from the Gene Ontology (GO) terms and the Conserved Domain Database (CDD) protein domain terms assigned to the loci by the gene-centered corpus LocusLink. This data set of GO- and CDD-based gene attribute profiles, vectors of binary random variables, was used to estimate multiple finite mixture models and each ensuing model utilized to partition the profiles into clusters. The resultant partitionings were combined using a unanimous voting scheme to produce consensus clusters, sets of profiles that co-occurred consistently in the same cluster. Attributes that were important in defining the genes assigned to a consensus cluster were identified. The clusters and their attributes were inspected to ascertain the GO and CDD terms most associated with subsets of genes and in conjunction with external knowledge such as chromosomal location, used to gain functional insights into human breast biology. The 52 luminal epithelial cell markers and 89 myoepithelial cell markers are disjoint sets of genes. Ensemble attribute profile clustering-based analysis indicated that both lists contained groups of genes with the functional properties of membrane receptor biology/signal transduction and nucleic acid binding/transcription. A subset of the luminal markers was associated with metabolic and oxidoreductase activities, whereas a subset of myoepithelial markers was associated with protein hydrolase activity. CONCLUSION: Given a set of genes and/or proteins associated with a phenomenon, process or system of interest, ensemble attribute profile clustering provides a simple method for collating and sythesizing the annotation data pertaining to them that are present in text-based, gene-centered corpora. The results provide information about properties common and unique to subsets of the list and hence insights into the biology of the problem under investigation.

Algorithms↗

The repertoire of solute carriers of family 6: identification of new human and rodent genes.

Tremendous amount of primary sequence information has been made available from the genome sequencing projects, although a complete annotation and identification of all genes is still far from being complete. Here, we present the identification of two new human genes from the pharmacologically important family of transporter proteins, solute carriers family 6 (SLC6). These were named SLC6A17 and SLC6A18 by HUGO. The human repertoire of SLC6 proteins now consists of 19 functional members and four pseudogenes. We also identified the corresponding orthologues and additional genes from mouse and rat genomes. Detailed phylogenetic analysis of the entire family of SLC6 proteins in mammals shows that this family can be divided into four subgroups. We used Hidden Markov Models for these subgroups and identified in total 430 unique SLC6 proteins from 10 animal, one plant, two fungi, and 196 bacterial genomes. It is evident that SLC6 proteins are present in both animals and bacteria, and that three of the four subfamilies of mammalian SLC6 proteins are present in Caenorhabditis elegans, showing that these subfamilies are evolutionary very ancient. Moreover, we performed tissue localization studies on the entire family of SLC6 proteins on a panel of 15 rat tissues and further, the expression of three of the new genes was studied using quantitative real-time PCR showing expression in multiple central and peripheral tissues. This paper presents an overall overview of the gene repertoire of the SLC6 gene family and its expression profile in rats.

Animals↗

Http://C. elegans: mining the functional genomic landscape.

Caenorhabditis elegans is a powerful animal model for the study of functional genomics. The completed and well-annotated DNA sequence is available and a systematic study of gene function by RNA-interference-mediated knockdown of every gene is in progress. Full-genome DNA microarrays and DNA chips can be used to determine expression changes at different stages of development and in different mutant backgrounds, and a protein-interaction map based on the yeast two-hybrid approach is in progress. These high-capacity approaches to studying gene function will provide new insights into invertebrate and vertebrate biology.

Animals↗

Identifying secretomes in people, pufferfish and pigs.

The proteins processed by the secretory pathway (secretome) are critical players in the development of multi-cellular eukaryotic organisms but have yet to be comprehensively studied at the genomic level. In this study, we use the Target P algorithm to predict human (13-20% of proteins found in individual datasets) and Fugu (14%) secretomes based on analysis of their nearly complete proteomes. We combine internal processing with prediction software to automate secreted protein identification and overcome one of the major challenges associated with EST data: identification of the minority of clones that encode N-terminally-complete proteins. We discuss the use of these methods to predict secreted proteins in EST-based consensus sequence sets, and we validate these predictions using an assay for cell-free cotranslational translocation. Analysis of TIGR Porcine Gene Index 4.0 as a test dataset resulted in the identification of 352 N-terminally-complete, putative secreted proteins. In functional agreement with our predictions, 34 of 40 (85%) of these cDNAs were verified to be cotranslationally translocated in an in vitro translation system. The methods developed here are specifically designed to accept partial open reading frames and improve secreted protein predictions in eukaryotic transcriptomes, and are valuable for the analysis and annotation of eukaryotic EST databases.

Algorithms↗

Genome annotation errors in pathway databases due to semantic ambiguity in partial EC numbers.

We report on a new type of systematic annotation error in genome and pathway databases that results from the misinterpretation of partial Enzyme Commission (EC) numbers such as '1.1.1.-'. This error results in the assignment of genes annotated with a partial EC number to many or all biochemical reactions that are annotated with the same partial EC number. That inference is faulty because of the ambiguous nature of partial EC numbers. We have observed this type of error in multiple databases, including KEGG, VIMSS and IMG, all of which assign genes to KEGG pathways. The Escherichia coli subset of the KEGG database exhibits this error for 6.8% of its gene-reaction assignments. For example, KEGG contains 17 reactions that are annotated with EC 1.1.1.-. A group of three E.coli genes, b1580 [putative dehydrogenase, NAD(P)-binding, starvation-sensing protein], b3787 (UDP-N-acetyl-D-mannosaminuronic acid dehydrogenase) and b0207 (2,5-diketo-D-gluconate reductase B), is assigned to 15 of those reactions, despite experimental evidence indicating different single functions for two of the three genes. Furthermore, the databases (DBs) are internally inconsistent in that the description of gene functions for genes with partial EC numbers is inconsistent with the activities implied by reactions to which the genes were assigned. We infer that these inconsistencies result from the processing used to match gene products to reactions within KEGG's metabolic pathways. These errors affect scientists who use these DBs as online encyclopedias and they affect bioinformaticists who use these DBs to train and validate newly developed algorithms.

Base Sequence↗

Characterization of adhesion threads of Deinococcus geothermalis as type IV pili.

Deinococcus geothermalis E50051 forms tenuous biofilms on paper machine surfaces. Field emission electron microscopy analysis revealed peritrichous appendages which mediated cell-to-surface and cell-to-cell interactions but were absent in planktonically grown cells. The major protein component of the extracellular extract of D. geothermalis had an N-terminal sequence similar to the fimbrial protein pilin annotated in the D. geothermalis DSM 11300 draft sequence. It also showed similarity to the type IV pilin sequence of D. radiodurans and several gram-negative pathogenic bacteria. Other proteins in the extract had N-terminal sequences identical to D. geothermalis proteins with conservative motifs for serine proteases, metallophosphoesterases, and proteins whose function is unknown. Periodic acid-Schiff staining for carbohydrates indicated that these extracellular proteins may be glycosylated. A further confirmation for the presence of glycoconjugates on the cell surface was obtained by confocal laser scanning imaging of living D. geothermalis cells stained with Amaranthus caudatus lectin, which specifically binds to galactose residues. The results indicate that the thread-like appendages of D. geothermalis E50051 are glycosylated type IV pili, bacterial attachment organelles which have thus far not been described for the genus Deinococcus.

Bacterial Adhesion↗

Phylogenetic and functional classification of ATP-binding cassette (ABC) systems.

ATP binding cassette (ABC) systems constitute one of the most abundant superfamilies of proteins. They are involved in the transport of a wide variety of substances, but also in many cellular processes and in their regulation. In this paper, we made a comparative analysis of the properties of ABC systems and we provide a phylogenetic and functional classification. This analysis will be helpful to accurately annotate ABC systems discovered during the sequencing of the genome of living organisms and to identify the partners of the ABC ATPases.

ATP-Binding Cassette Transporters↗

Genomic organization and molecular characterization of Clostridium difficile bacteriophage PhiCD119.

In this study, we have isolated a temperate phage (PhiCD119) from a pathogenic Clostridium difficile strain and sequenced and annotated its genome. This virus has an icosahedral capsid and a contractile tail covered by a sheath and contains a double-stranded DNA genome. It belongs to the Myoviridae family of the tailed phages and the order Caudovirales. The genome was circularly permuted, with no physical ends detected by sequencing or restriction enzyme digestion analysis, and lacked a cos site. The DNA sequence of this phage consists of 53,325 bp, which carries 79 putative open reading frames (ORFs). A function could be assigned to 23 putative gene products, based upon bioinformatic analyses. The PhiCD119 genome is organized in a modular format, which includes modules for lysogeny, DNA replication, DNA packaging, structural proteins, and host cell lysis. The PhiCD119 attachment site attP lies in a noncoding region close to the putative integrase (int) gene. We have identified the phage integration site on the C. difficile chromosome (attB) located in a noncoding region just upstream of gene gltP, which encodes a carrier protein for glutamate and aspartate. This genetic analysis represents the first complete DNA sequence and annotation of a C. difficile phage.

Bacteriophages↗

Proteome analysis based on motif statistics.

MOTIVATION: Even for the amino acid motifs collected in the Prosite database there may be chance occurences as opposed to those occurences where the motif is involved in fold or function of a protein. With recent mathematical advances in assessing the significance of observing such a motif a particular number of times, we can now study the over- or under-representation of particular motifs in a complete genome and attempt to make functional deductions. RESULTS: We demonstrate that statistical over- or under-representation of motifs in complete proteomes may be an indicator of whether, in that organism, we are looking at chance occurrences of the motif or whether the occurrences are sufficiently numerous to suggest a systematic, and thus functionally important occurrence. This has important implications on databank annotations. AVAILABILITY: The complete dataset comprising the plotted statistics of 266 Prosite motifs on 42 proteomes is available at http://algo.inria.fr/nicodeme/proteomes/proteocomp.html. The software used to compute this data has been described by Nicodème (2000, 2001). They are available either by web access as mentioned in these articles or by direct request from Pierre Nicodème.

Amino Acid Motifs↗

Using the CATH domain database to assign structures and functions to the genome sequences.

The CATH database of protein structures contains approximately 18000 domains organized according to their (C)lass, (A)rchitecture, (T)opology and (H)omologous superfamily. Relationships between evolutionary related structures (homologues) within the database have been used to test the sensitivity of various sequence search methods in order to identify relatives in Genbank and other sequence databases. Subsequent application of the most sensitive and efficient algorithms, gapped blast and the profile based method, Position Specific Iterated Basic Local Alignment Tool (PSI-BLAST), could be used to assign structural data to between 22 and 36 % of microbial genomes in order to improve functional annotation and enhance understanding of biological mechanism. However, on a cautionary note, an analysis of functional conservation within fold groups and homologous superfamilies in the CATH database, revealed that whilst function was conserved in nearly 55% of enzyme families, function had diverged considerably, in some highly populated families. In these families, functional properties should be inherited far more cautiously and the probable effects of substitutions in key functional residues carefully assessed.

Algorithms↗

From masking repeats to identifying functional repeats in the mouse transcriptome.

The back-to-back release of the mouse genome and the functionally annotated RIKEN mouse full-length cDNA collection was an important milestone in mammalian genomics. Yet much of the data remain to be explored in terms of biological effects and mechanisms. For example, interspersed repeats account for 39 per cent of the mouse genome sequence and 11 per cent of representative transcripts. A considerable number of transposable repeat elements are still active and propagating in mouse compared with human. While existing repeat databases and tools assist the classification of repeats or identification of new repeats, there is little bioinformatic support towards exploring the extent and role of repeats in transcriptional variation, modulation of protein function, or gene regulatory events. Since the mouse is used as a model organism to study human genes and their disease associations, this review focuses on information extraction and collation that captures the functional context of repeats in mouse transcripts to facilitate the biological interpretation and extrapolation of findings to the human.

Animals↗

Physical network models.

We develop a new framework for inferring models of transcriptional regulation. The models, which we call physical network models, are annotated molecular interaction graphs. The attributes in the model correspond to verifiable properties of the underlying biological system such as the existence of protein-protein and protein-DNA interactions, the directionality of signal transduction in protein-protein interactions, as well as signs of the immediate effects of these interactions. Possible configurations of these variables are constrained by the available data sources. Some of the data sources, such as factor-binding data, involve measurements that are directly tied to the variables in the model. Other sources, such as gene knock-outs, are functional in nature and provide only indirect evidence about the variables. We associate each observed knock-out effect in the deletion mutant data with a set of causal paths (molecular cascades) that could in principle explain the effect, resulting in aggregate constraints about the physical variables in the model. The most likely settings of all the variables, specifying the most likely graph annotations, are found by a recursive application of the max-product algorithm. By testing our approach on datasets related to the pheromone response pathway in S. cerevisiae, we demonstrate that the resulting model is consistent with previous studies about the pathway. Moreover, we successfully predict gene knock-out effects with a high degree of accuracy in a cross-validation setting. When applying this approach genome-wide, we extract submodels consistent with previous studies. The approach can be readily extended to other data sources or to facilitate automated experimental design.

Computational Biology↗

Tools for comparative protein structure modeling and analysis.

The following resources for comparative protein structure modeling and analysis are described (http://salilab.org): MODELLER, a program for comparative modeling by satisfaction of spatial restraints; MODWEB, a web server for automated comparative modeling that relies on PSI-BLAST, IMPALA and MODELLER; MODLOOP, a web server for automated loop modeling that relies on MODELLER; MOULDER, a CPU intensive protocol of MODWEB for building comparative models based on distant known structures; MODBASE, a comprehensive database of annotated comparative models for all sequences detectably related to a known structure; MODVIEW, a Netscape plugin for Linux that integrates viewing of multiple sequences and structures; and SNPWEB, a web server for structure-based prediction of the functional impact of a single amino acid substitution.

Internet↗