Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,207 records · Page 67Linked to original sources

BacTregulators: a database of transcriptional regulators in bacteria and archaea.

MOTIVATION: The BacTregulators database is intended to collect and to integrate information on proteins belonging to defined families of transcriptional regulators in prokaryotes. RESULTS: The BacTregulators database currently contains data on two families of transcriptional regulators: AraC-XylS and TetR. The proteins included in the BacTregulators database have been identified by screening 123 genomes from archaea and bacteria and the SWISS-PROT and TrEMBL databases with profiles defining each family. As the result of an integration process, we have included 1326 different protein sequences from the AraC-XylS family and 1487 different protein sequences from the TetR family. The definition of an entry in BacTregulators is based on protein sequence, source organism, genome element and position in this genome element. The BacTregulators site allows the user to retrieve protein sequences, functional features and experimental evidence supporting the functions, references and the three-dimensional structure of the regulator when available. BacTregulators supplies an innovative tool that allows the researcher to obtain an integrated report that shows the data corresponding to other entries which are related by sequence similarity to the query entry. BacTregulators detects and classifies the regulators belonging to AraC-XylS and TetR families present in prokaryotic genomes, and thus contributes to a more accurate annotation of regulators in genomes. The information collected on each protein in the family can be useful to characterize a new regulator or compile information on the biological properties of a known regulator. AVAILABILITY: The BacTregulators is available at www.bactregulators.org

Archaea↗

A telomere-to-telomere gap-free genome assembly of the endangered humphead wrasse (Cheilinus undulatus).

Humphead wrasse, Cheilinus undulatus, is an endangered fish species with high economic and ecological value as well as natural sex change from female to male, while sexual selection occurs in breeding aggregations. In our present study, we constructed the first gap-free telomere-to-telomere (T2T) genome assembly for humphead wrasse, by integration of PacBio HiFi, ONT Ultra-long and Hi-C sequencing techniques. With 99% of the entire sequences anchored into 24 chromosomes, this haplotypic genome assembly spans approximately 1.25 Gb and presents a complete set of 48 telomeres and 24 centromeres. In terms of correctness (quality value QV: 53.447) and completeness (BUSCO score: 99.3%), this chromosome-scale assembly is indeed of high quality. We predicted 658.03 Mb of repetitive sequences and annotated 26,609 protein-coding genes in the assembled genome. This high-quality T2T genome assembly not only facilitates the genetic conservation of humphead wrasse, but also offers fundamental genomic data for supporting in-depth investigations on functional genomics, genetic diversity, and selective breeding for this economically important teleost.

Animals↗

Co-evolution analysis on endocrine research: a methodological approach.

The rapid growth of different kinds of biological information allows a good opportunity to analyze the co-evolutionary characteristics in endocrine regulatory pathways. Data ranging from kinds of species' genome, gene sequence, protein structure, and expression profile of different organisms can reveal the inner co-evolutionary relationship of ligands, receptors, and other related molecules. In return, these co-evolutionary characteristics can help us determine uncharacterized ligands and receptors, annotate gene functions, highlight amino acid residues with biochemical significance, and identify regulated genes in the endocrine process. Encouraging examples in this field, although at their starting stage, have emerged. Here we focus on recent progress in endocrine-related co-evolution research from a methodological approach.

Endocrine System↗

Beyond synexpression relationships: local clustering of time-shifted and inverted gene expression profiles identifies new, biologically relevant interactions.

The complexity of biological systems provides for a great diversity of relationships between genes. The current analysis of whole-genome expression data focuses on relationships based on global correlation over a whole time-course, identifying clusters of genes whose expression levels simultaneously rise and fall. There are, of course, other potential relationships between genes, which are missed by such global clustering. These include activation, where one expects a time-delay between related expression profiles, and inhibition, where one expects an inverted relationship. Here, we propose a new method, which we call local clustering, for identifying these time-delayed and inverted relationships. It is related to conventional gene-expression clustering in a fashion analogous to the way local sequence alignment (the Smith-Waterman algorithm) is derived from global alignment (Needleman-Wunsch). An integral part of our method is the use of random score distributions to assess the statistical significance of each cluster. We applied our method to the yeast cell-cycle expression dataset and were able to detect a considerable number of additional biological relationships between genes, beyond those resulting from conventional correlation. We related these new relationships between genes to their similarity in function (as determined from the MIPS scheme) or their having known protein-protein interactions (as determined from the large-scale two-hybrid experiment); we found that genes strongly related by local clustering were considerably more likely than random to have a known interaction or a similar cellular role. This suggests that local clustering may be useful in functional annotation of uncharacterized genes. We examined many of the new relationships in detail. Some of them were already well-documented examples of inhibition or activation, which provide corroboration for our results. For instance, we found an inverted expression profile relationship between genes YME1 and YNT20, where the latter has been experimentally documented as a bypass suppressor of the former. We also found new relationships involving uncharacterized yeast genes and were able to suggest functions for many of them. In particular, we found a time-delayed expression relationship between J0544 (which has not yet been functionally characterized) and four genes associated with the mitochondria. This suggests that J0544 may be involved in the control or activation of mitochondrial genes. We have also looked at other, less extensive datasets than the yeast cell-cycle and found further interesting relationships. Our clustering program and a detailed website of clustering results is available at http://www.bioinfo.mbb.yale.edu/expression/cluster (or http://www.genecensus.org/expression/cluster).

Algorithms↗

Functional characterization of XendoU, the endoribonuclease involved in small nucleolar RNA biosynthesis.

XendoU is the endoribonuclease involved in the biosynthesis of a specific subclass of Xenopus laevis intron-encoded small nucleolar RNAs. XendoU has no homology to any known cellular RNase, although it has sequence similarity with proteins tentatively annotated as serine proteases. It has been recently shown that XendoU represents the cellular counterpart of a nidovirus replicative endoribonuclease (NendoU), which plays a critical role in viral replication and transcription. In this paper, we combined prediction and experimental data to define the amino acid residues directly involved in XendoU catalysis. Specifically, we find that XendoU residues Glu-161, Glu-167, His-162, His-178, and Lys-224 are essential for RNA cleavage, which occurs in the presence of manganese ions. Furthermore, we identified the RNA sequence required for XendoU binding and showed that the formation of XendoU-RNA complex is Mn2+-independent.

Amino Acid Sequence↗

pSTIING: a 'systems' approach towards integrating signalling pathways, interaction and transcriptional regulatory networks in inflammation and cancer.

pSTIING (http://pstiing.licr.org) is a new publicly accessible web-based application and knowledgebase featuring 65 228 distinct molecular associations (comprising protein-protein, protein-lipid, protein-small molecule interactions and transcriptional regulatory associations), ligand-receptor-cell type information and signal transduction modules. It has a particular major focus on regulatory networks relevant to chronic inflammation, cell migration and cancer. The web application and interface provide graphical representations of networks allowing users to combine and extend transcriptional regulatory and signalling modules, infer molecular interactions across species and explore networks via protein domains/motifs, gene ontology annotations and human diseases. pSTIING also supports the direct cross-correlation of experimental results with interaction information in the knowledgebase via the CLADIST tool associated with pSTIING, which currently analyses and clusters gene expression, proteomic and phenotypic datasets. This allows the contextual projection of co-expression patterns onto prior network information, facilitating the identification of functional modules in physiologically relevant systems.

Amino Acid Motifs↗

VKCDB: voltage-gated potassium channel database.

BACKGROUND: The family of voltage-gated potassium channels comprises a functionally diverse group of membrane proteins. They help maintain and regulate the potassium ion-based component of the membrane potential and are thus central to many critical physiological processes. VKCDB (Voltage-gated potassium [K] Channel DataBase) is a database of structural and functional data on these channels. It is designed as a resource for research on the molecular basis of voltage-gated potassium channel function. DESCRIPTION: Voltage-gated potassium channel sequences were identified by using BLASTP to search GENBANK and SWISSPROT. Annotations for all voltage-gated potassium channels were selectively parsed and integrated into VKCDB. Electrophysiological and pharmacological data for the channels were collected from published journal articles. Transmembrane domain predictions by TMHMM and PHD are included for each VKCDB entry. Multiple sequence alignments of conserved domains of channels of the four Kv families and the KCNQ family are also included. Currently VKCDB contains 346 channel entries. It can be browsed and searched using a set of functionally relevant categories. Protein sequences can also be searched using a local BLAST engine. CONCLUSIONS: VKCDB is a resource for comparative studies of voltage-gated potassium channels. The methods used to construct VKCDB are general; they can be used to create specialized databases for other protein families. VKCDB is accessible at http://vkcdb.biology.ualberta.ca.

Animals↗

Systematic discovery of new genes in the Saccharomyces cerevisiae genome.

We used genome-wide comparative analysis of predicted protein sequences to identify many novel small genes, named smORFs for small open reading frames, within the budding yeast genome. Further analysis of 117 of these new genes showed that 84 are transcribed. We extended our analysis of one smORF conserved from yeast to human. This investigation provides an updated and comprehensive annotation of the yeast genome, validates additional concepts in the study of genomes in silico, and increases the expected numbers of coding sequences in a genome with the corresponding impact on future functional genomics and proteomics studies.

Amino Acid Sequence↗

Genome and protein evolution in eukaryotes.

The past year has seen the completion of the genome sequence of the flowering plant Arabidopsis thaliana and the initial sequence reports of the human genome. The availability of completely sequenced eukaryotic genomes from disparate phylogenetic lineages has opened the door to comparative analyses and a better understanding of the evolutionary processes shaping genomes. Complex many-to-many relationships between genes from different species appear to be the norm, suggesting that transfer of detailed functional annotation will not be straightforward. In addition to expansion and contraction of gene families, new genes evolve from recombination of pre-existing domains, although some domain families do appear to have evolved recently and to be specific to restricted phylogenetic lineages. The overall picture is of a huge diversity of gene content within eukaryotic genomes, reflecting different functional demands in different species.

Animals↗

Fundamentals of massive automatic pairwise alignments of protein sequences: theoretical significance of Z-value statistics.

MOTIVATION: Different automatic methods of sequence alignments are routinely used as a starting point for homology searches and function inference. Confidence in an alignment probability is one of the major fundamentals of massive automatic genome-scale pairwise comparisons, for clustering of putative orthologs and paralogs, sequenced genome annotation or multiple-genomic tree constructions. Extreme value distribution based on the Karlin-Altschul model, usually advised for large-scale comparisons are not always valid, particularly in the case of comparisons of non-biased with nucleotide-biased genomes (such that of Plasmodium falciparum). Z-values estimates based on Monte Carlo technics, can be calculated experimentally for any alignment output, whatever the method used. Empirically, a Z-value higher than approximately 8 is supposed reasonable to assess that an alignment score is significant, but this arbitrary figure was never theoretically justified. RESULTS: In this paper, we used the Bienaymé-Chebyshev inequality to demonstrate a theorem of the upper limit of an alignment score probability (or P-value). This theorem implies that a computed Z-value is a statistical test, a single-linkage clustering criterion and that 1/Z-value(2) is an upper limit to the probability of an alignment score whatever the actual probability law is. Therefore, this study provides the missing theoretical link between a Z-value cut-off used for an automatic clustering of putative orthologs and/or paralogs, and the corresponding statistical risk in such genome-scale comparisons (using non-biased or biased genomes).

Algorithms↗

Systematic comparison of catalytic mechanisms of hydrolysis and transfer reactions classified in the EzCatDB database.

Catalytic mechanisms of 270 enzymes from 131 superfamilies, mainly hydrolases and transferases, were analyzed based on their enzyme structures. A method of systematic comparison and classification of the catalytic reactions was developed. Hydrolysis and transfer reactions closely resemble one another, displaying common mechanisms, single displacement, and double displacement. These displacement mechanisms might be further subclassified according to the type of catalytic factors and nucleophilic substitution involved. Several types of catalytic factors exist: nucleophile, acid, base, stabilizer, modulator, cofactors. Nucleophilic substitution might be categorized as S(N)1/S(N)2 (or dissociative/associative) reactions. The classification indicates that some mechanisms favor particular types of catalytic factors. In hydrolyses of amide bonds and phosphoric ester bonds, mechanisms with single displacement tend to use inorganic cofactors such as zinc and magnesium ions as important catalysts, whereas those with double displacement frequently do not use such cofactors. In contrast, hydrolyses of O-glycoside bond rarely use such cofactors, with one exception. The trypsin-like hydrolytic reaction, which is catalyzed by the classic catalytic triad comprising serine/histidine/aspartate, can be considered as a "super-reaction" because it is observed in at least three nonhomologous enzymes, whereas most reactions are singlets without any nonhomologous enzymes. By dividing complex reactions into several reactions, correlations between active site structures and catalytic functions can be suggested. This classification method is applicable to other reactions such as elimination and isomerization. Furthermore, it will facilitate annotation of enzyme functions from 3D patterns of enzyme active sites. The classification is available at http://mbs.cbrc.jp/EzCatDB/RLCP/index.html.

Binding Sites↗

Whole-genome analysis: annotations and updates.

The most important advances in the field of genome annotation over the past two years involve the use of cDNA sequences, protein structures and gene expression data to predict genes. These types of information not only improve gene identification, but they also give insights into variation in gene structure and function.

Alternative Splicing↗

Structuring the universe of proteins.

High-throughput sequencing of human genomes and those of important model organisms (mouse, Drosophila melanogaster, Caenorhabditis elegans, fungi, archaea) and bacterial pathogens has laid the foundation for another "big science" initiative in biology. Together, X-ray crystallographers, nuclear magnetic resonance (NMR) spectroscopists, and computational biologists are pursuing high-throughput structural studies aimed at developing a comprehensive three-dimensional view of the protein structure universe. The new science of structural genomics promises more than 10,000 experimental protein structures and millions of calculated homology models of related proteins. The evolutionary underpinnings and technological challenges of automating target selection, protein expression and purification, sample preparation, NMR and X-ray data measurement/analysis, homology modeling, and structure/function annotation are discussed in detail. An informative case study from one of the structural genomics centers funded by the National Institutes of Health and the National Institute of General Medical Sciences (NIH/NIGMS) demonstrates how this experimental/computational pipeline will reveal important links between form and function in biology and provide new insights into evolution and human health and disease.

Amino Acid Sequence↗

A catalogue of the effector secretome of plant pathogenic oomycetes.

The oomycetes form a phylogenetically distinct group of eukaryotic microorganisms that includes some of the most notorious pathogens of plants. Oomycetes accomplish parasitic colonization of plants by modulating host cell defenses through an array of disease effector proteins. The biology of effectors is poorly understood but tremendous progress has been made in recent years. This review classifies and catalogues the effector secretome of oomycetes. Two classes of effectors target distinct sites in the host plant: Apoplastic effectors are secreted into the plant extracellular space, and cytoplasmic effectors are translocated inside the plant cell, where they target different subcellular compartments. Considering that five species are undergoing genome sequencing and annotation, we are rapidly moving toward genome-wide catalogues of oomycete effectors. Already, it is evident that the effector secretome of pathogenic oomycetes is more complex than expected, with perhaps several hundred proteins dedicated to manipulating host cell structure and function.

Fungal Proteins↗

PATTERNFINDER: combined analysis of DNA regulatory sequences and double-helix stability.

BACKGROUND: Regulatory regions that function in DNA replication and gene transcription contain specific sequences that bind proteins as well as less-specific sequences in which the double helix is often easy to unwind. Progress towards predicting and characterizing regulatory regions could be accelerated by computer programs that perform a combined analysis of specific sequences and DNA unwinding properties. RESULTS: Here we present PATTERNFINDER, a web server that searches DNA sequences for matches to specific or flexible patterns, and analyzes DNA helical stability. A batch mode of the program generates a tabular map of matches to multiple, different patterns. Regions flanking pattern matches can be targeted for helical stability analysis to identify sequences with a minimum free energy for DNA unwinding. As an example application, we analyzed a regulatory region of the human c-myc proto-oncogene consisting of a single-strand-specific protein binding site within a DNA region that unwindsin vivo. The predicted region of minimal helical stability overlapped both the protein binding site and the unwound DNA region identified experimentally. CONCLUSIONS: The PATTERNFINDER web server permits localization of known functional elements or landmarks in DNA sequences as well as prediction of potential new elements. Batch analysis of multiple patterns facilitates the annotation of DNA regulatory regions. Identifying specific pattern matches linked to DNA with low helical stability is useful in characterizing regulatory regions for transcription, replication and other processes and may predict functional DNA unwinding elements.PATTERNFINDER can be accessed freely at: http://wings.buffalo.edu/gsa/dna/dk/PFP/

Computational Biology↗

SMART, a simple modular architecture research tool: identification of signaling domains.

Accurate multiple alignments of 86 domains that occur in signaling proteins have been constructed and used to provide a Web-based tool (SMART: simple modular architecture research tool) that allows rapid identification and annotation of signaling domain sequences. The majority of signaling proteins are multidomain in character with a considerable variety of domain combinations known. Comparison with established databases showed that 25% of our domain set could not be deduced from SwissProt and 41% could not be annotated by Pfam. SMART is able to determine the modular architectures of single sequences or genomes; application to the entire yeast genome revealed that at least 6.7% of its genes contain one or more signaling domains, approximately 350 greater than previously annotated. The process of constructing SMART predicted (i) novel domain homologues in unexpected locations such as band 4.1-homologous domains in focal adhesion kinases; (ii) previously unknown domain families, including a citron-homology domain; (iii) putative functions of domain families after identification of additional family members, for example, a ubiquitin-binding role for ubiquitin-associated domains (UBA); (iv) cellular roles for proteins, such predicted DEATH domains in netrin receptors further implicating these molecules in axonal guidance; (v) signaling domains in known disease genes such as SPRY domains in both marenostrin/pyrin and Midline 1; (vi) domains in unexpected phylogenetic contexts such as diacylglycerol kinase homologues in yeast and bacteria; and (vii) likely protein misclassifications exemplified by a predicted pleckstrin homology domain in a Candida albicans protein, previously described as an integrin.

Amino Acid Sequence↗

VSD: a database for schizophrenia candidate genes focusing on variations.

Schizophrenia is a common mental disease characterized by delusions, hallucinations, and formal thought disorder. It has been demonstrated with genetic evidence that the disease is a polygenic disorder. Pharmacological, neurochemical, and clinical studies have suggested a number of schizophrenia susceptibility loci. In order to systematically search for genes with small effect in the development of schizophrenia, a database called VSD was established to provide variation data for publicly available candidate genes. Most of the genes encode neurotransmitter receptors, neurotransmitter transporters, and the enzymes involved in their metabolism. Other candidate genes extracted from published literature are also included. The variation information has been collected from publicly available mutation and polymorphism databases such as dbSNP, HGVbase, and OMIM, with single nucleotide polymorphism (SNP) being the most abundant form of collected variations. Reference sequences from NCBI's RefSeq database are used as references when positioning variation at transcript and protein levels. The nonsynonymous SNPs (nsSNPs) that lead to amino acid changes in the functional sites or domains of proteins are distinguished since they are more likely to affect protein function and would be target SNPs for association studies. In addition to variation data, gene descriptions, enzyme information, and other biological information for each gene locus are also included. The latest version of VSD contains 23,648 variations assigned to a total of 186 genes. Five-hundred eighty-eight domains and sites annotated in the SWISS-PROT and InterPro databases are found to contain nsSNPs. VSD may be accessed via the World Wide Web (www.chgb.org.cn/vsd.htm) and will be developed as an up-to-date and comprehensive locus-specific resource for identifying susceptibility genes for schizophrenia.

Databases, Nucleic Acid↗

Pseudo-messenger RNA: phantoms of the transcriptome.

The mammalian transcriptome harbours shadowy entities that resist classification and analysis. In analogy with pseudogenes, we define pseudo-messenger RNA to be RNA molecules that resemble protein-coding mRNA, but cannot encode full-length proteins owing to disruptions of the reading frame. Using a rigorous computational pipeline, which rules out sequencing errors, we identify 10,679 pseudo-messenger RNAs (approximately half of which are transposon-associated) among the 102,801 FANTOM3 mouse cDNAs: just over 10% of the FANTOM3 transcriptome. These comprise not only transcribed pseudogenes, but also disrupted splice variants of otherwise protein-coding genes. Some may encode truncated proteins, only a minority of which appear subject to nonsense-mediated decay. The presence of an excess of transcripts whose only disruptions are opal stop codons suggests that there are more selenoproteins than currently estimated. We also describe compensatory frameshifts, where a segment of the gene has changed frame but remains translatable. In summary, we survey a large class of non-standard but potentially functional transcripts that are likely to encode genetic information and effect biological processes in novel ways. Many of these transcripts do not correspond cleanly to any identifiable object in the genome, implying fundamental limits to the goal of annotating all functional elements at the genome sequence level.

Animals↗