Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 739 records · Page 41Linked to original sources

Structural characterization of the human proteome.

This paper reports an analysis of the encoded proteins (the proteome) of the genomes of human, fly, worm, yeast, and representatives of bacteria and archaea in terms of the three-dimensional structures of their globular domains together with a general sequence-based study. We show that 39% of the human proteome can be assigned to known structures. We estimate that for 77% of the proteome, there is some functional annotation, but only 26% of the proteome can be assigned to standard sequence motifs that characterize function. Of the human protein sequences, 13% are transmembrane proteins, but only 3% of the residues in the proteome form membrane-spanning regions. There are substantial differences in the composition of globular domains of transmembrane proteins between the proteomes we have analyzed. Commonly occurring structural superfamilies are identified within the proteome. The frequencies of these superfamilies enable us to estimate that 98% of the human proteome evolved by domain duplication, with four of the 10 most duplicated superfamilies specific for multicellular organisms. The zinc-finger superfamily is massively duplicated in human compared to fly and worm, and occurrence of domains in repeats is more common in metazoa than in single cellular organisms. Structural superfamilies over- and underrepresented in human disease genes have been identified. Data and results can be downloaded and analyzed via web-based applications at http://www.sbg.bio.ic.ac.uk.

Algorithms↗

IdentiCS--identification of coding sequence and in silico reconstruction of the metabolic network directly from unannotated low-coverage bacterial genome sequence.

BACKGROUND: A necessary step for a genome level analysis of the cellular metabolism is the in silico reconstruction of the metabolic network from genome sequences. The available methods are mainly based on the annotation of genome sequences including two successive steps, the prediction of coding sequences (CDS) and their function assignment. The annotation process takes time. The available methods often encounter difficulties when dealing with unfinished error-containing genomic sequence. RESULTS: In this work a fast method is proposed to use unannotated genome sequence for predicting CDSs and for an in silico reconstruction of metabolic networks. Instead of using predicted genes or CDSs to query public databases, entries from public DNA or protein databases are used as queries to search a local database of the unannotated genome sequence to predict CDSs. Functions are assigned to the predicted CDSs simultaneously. The well-annotated genome of Salmonella typhimurium LT2 is used as an example to demonstrate the applicability of the method. 97.7% of the CDSs in the original annotation are correctly identified. The use of SWISS-PROT-TrEMBL databases resulted in an identification of 98.9% of CDSs that have EC-numbers in the published annotation. Furthermore, two versions of sequences of the bacterium Klebsiella pneumoniae with different genome coverage (3.9 and 7.9 fold, respectively) are examined. The results suggest that a 3.9-fold coverage of the bacterial genome could be sufficiently used for the in silico reconstruction of the metabolic network. Compared to other gene finding methods such as CRITICA our method is more suitable for exploiting sequences of low genome coverage. Based on the new method, a program called IdentiCS (Identification of Coding Sequences from Unfinished Genome Sequences) is delivered that combines the identification of CDSs with the reconstruction, comparison and visualization of metabolic networks (free to download at http://genome.gbf.de/bioinformatics/index.html). CONCLUSIONS: The reversed querying process and the program IdentiCS allow a fast and adequate prediction protein coding sequences and reconstruction of the potential metabolic network from low coverage genome sequences of bacteria. The new method can accelerate the use of genomic data for studying cellular metabolism.

Base Sequence↗

Functional grouping based on signatures in protein termini.

The two ends of each protein are known as the amino (N-) and carboxyl (C-) termini. Short signatures in a protein's termini often carry vital cellular function. No systematic research has been conducted to address the importance of short signatures (3 to 10 amino acids) in protein termini at the proteomic level. Specifically, it is unknown whether such signatures are evolutionarily conserved, and if so, whether this conservation confers shared biological functions. Current signature detection methods fail to detect such short signatures due to inadequate statistical scores. The findings presented in this study strongly support the notion that functional significance of protein sets may be captured by short signatures at their termini. A positional search method was applied to over one million proteins from the UniProt database. The result is a collection of about a thousand significant signature groups (SIGs) that include previously identified as well as many novel signatures in protein termini. These SIGs represent protein sets with minimal or no overall sequence similarity excepting the similarity at their termini. The most significant SIGs are assigned by their strong correspondence to functional annotations derived from external databases such as Gene Ontology. Each of the SIGs is associated with the statistical significance of its functional association. These SIGs provide a valuable source for testing previously overlooked signatures in protein termini and allow for the investigation of the role played by such signatures throughout evolution. The SIGs archive and advanced search options are available at http://www.proteus.cs.huji.ac.il.

Amino Acid Sequence↗

The structure and computational analysis of Mycobacterium tuberculosis protein CitE suggest a novel enzymatic function.

Fatty acid biosynthesis is essential for the survival of Mycobacterium tuberculosis and acetyl-coenzyme A (acetyl-CoA) is an essential precursor in this pathway. We have determined the 3-D crystal structure of M. tuberculosis citrate lyase beta-subunit (CitE), which as annotated should cleave protein bound citryl-CoA to oxaloacetate and a protein-bound CoA derivative. The CitE structure has the (beta/alpha)(8) TIM barrel fold with an additional alpha-helix, and is trimeric. We have determined the ternary complex bound with oxaloacetate and magnesium, revealing some of the conserved residues involved in catalysis. While the bacterial citrate lyase is a complex with three subunits, the M. tuberculosis genome does not contain the alpha and gamma subunits of this complex, implying that M. tuberculosis CitE acts differently from other bacterial CitE proteins. The analysis of gene clusters containing the CitE protein from 168 fully sequenced organisms has led us to identify a grouping of functionally related genes preserved in M. tuberculosis, Rattus norvegicus, Homo sapiens, and Mus musculus. We propose a novel enzymatic function for M. tuberculosis CitE in fatty acid biosynthesis that is analogous to bacterial citrate lyase but producing acetyl-CoA rather than a protein-bound CoA derivative.

Amino Acid Sequence↗

Expression of hypothetical proteins in human fetal brain: increased expression of hypothetical protein 28.5kDa in Down syndrome, a clue for its tentative role.

Major advances have been made in annotation of sequences of the human genome, although elucidating the functions of these newly discovered genes remains to be a strong challenge. In an effort to give insight into how triplication of chromosome 21 leads to mental retardation in Down syndrome, we have constructed a two-dimensional protein map from control and Down syndrome fetal brain and identified hypothetical proteins with no known functions. Subsequent quantitative analysis of these proteins revealed no apparent change in expression of hypothetical proteins DKZp564P0562.1 (fragment), 16.6, 21.4, 39.5, and 40kDa as well as putative 55kDa protein between controls and Down syndrome fetuses. By contrast, hypothetical protein 28.5kDa was significantly elevated (P<0.05) in fetal Down syndrome. This finding offers an important clue that a hypothetical protein might be involved in the pathomechanisms of brain abnormality in Down syndrome.

Brain↗

Strategies for the identification, the assembly and the classification of integrated biological systems in completely sequenced genomes.

The proteins involved in a single biological process may form a stable supra-molecular assembly or be transiently in interaction. Although, the first annotation steps of a complete genome may allow the identification of the different partners, their assembly in a functional system, referred to as an integrated system, is a domain where methodological effort has to be done. Indeed, the knowledge required to assemble partners of such systems should be explicitly included in annotation software. The availability of a complete genome, and therefore of all the proteins encoded by that genome, motivated the development of automated approaches through the coordinated combination of different bio-informatic methods allowing the identification of the different partners, their assembly and the classification of the reconstructed systems in functional categories. In this data flux, the identification of the sequence partners represents the principal bottleneck. Here, we describe and compare the results obtained with different classes of methods (BLASTP2, PSI-BLAST, MAST and META-MEME) applied to the identification in complete genomes of a given family of integrated systems: the ABC transporters. PSI-BLAST appears to significantly outperform motif-based methods, and the results are discussed according to the nature of the proteins and the structure of the sub-families.

Binding Sites↗

Exploiting conserved structure for faster annotation of non-coding RNAs without loss of accuracy.

MOTIVATION: Non-coding RNAs (ncRNAs)-functional RNA molecules not coding for proteins-are grouped into hundreds of families of homologs. To find new members of an ncRNA gene family in a large genome database, covariance models (CMs) are a useful statistical tool, as they use both sequence and RNA secondary structure information. Unfortunately, CM searches are slow. Previously, we introduced 'rigorous filters', which provably sacrifice none of CMs' accuracy, although often scanning much faster. A rigorous filter, using a profile hidden Markov model (HMM), is built based on the CM, and filters the genome database, eliminating sequences that provably could not be annotated as homologs. The CM is run only on the remainder. Some biologically important ncRNA families could not be scanned efficiently with this technique, largely due to the significance of conserved secondary structure relative to primary sequence in identifying these families. Current heuristic filters are also expected to perform poorly on such families. RESULTS: By augmenting profile HMMs with limited secondary structure information, we obtain rigorous filters that accelerate CM searches for virtually all known ncRNA families from the Rfam Database and tRNA models in tRNAscan-SE. These filters scan an 8 gigabase database in weeks instead of years, and uncover homologs missed by heuristic techniques to speed CM searches. AVAILABILITY: Software in development; contact the authors.

Algorithms↗

Predicting eukaryotic protein subcellular location by fusing optimized evidence-theoretic K-Nearest Neighbor classifiers.

Facing the explosion of newly generated protein sequences in the post genomic era, we are challenged to develop an automated method for fast and reliably annotating their subcellular locations. Knowledge of subcellular locations of proteins can provide useful hints for revealing their functions and understanding how they interact with each other in cellular networking. Unfortunately, it is both expensive and time-consuming to determine the localization of an uncharacterized protein in a living cell purely based on experiments. To tackle the challenge, a novel hybridization classifier was developed by fusing many basic individual classifiers through a voting system. The "engine" of these basic classifiers was operated by the OET-KNN (Optimized Evidence-Theoretic K-Nearest Neighbor) rule. As a demonstration, predictions were performed with the fusion classifier for proteins among the following 16 localizations: (1) cell wall, (2) centriole, (3) chloroplast, (4) cyanelle, (5) cytoplasm, (6) cytoskeleton, (7) endoplasmic reticulum, (8) extracell, (9) Golgi apparatus, (10) lysosome, (11) mitochondria, (12) nucleus, (13) peroxisome, (14) plasma membrane, (15) plastid, and (16) vacuole. To get rid of redundancy and homology bias, none of the proteins investigated here had >/=25% sequence identity to any other in a same subcellular location. The overall success rates thus obtained via the jack-knife cross-validation test and independent dataset test were 81.6% and 83.7%, respectively, which were 46 approximately 63% higher than those performed by the other existing methods on the same benchmark datasets. Also, it is clearly elucidated that the overwhelmingly high success rates obtained by the fusion classifier is by no means a trivial utilization of the GO annotations as prone to be misinterpreted because there is a huge number of proteins with given accession numbers and the corresponding GO numbers, but their subcellular locations are still unknown, and that the percentage of proteins with GO annotations indicating their subcellular components is even less than the percentage of proteins with known subcellular location annotation in the Swiss-Prot database. It is anticipated that the powerful fusion classifier may also become a very useful high throughput tool in characterizing other attributes of proteins according to their sequences, such as enzyme class, membrane protein type, and nuclear receptor subfamily, among many others. A web server, called "Euk-OET-PLoc", has been designed at http://202.120.37.186/bioinf/euk-oet for public to predict subcellular locations of eukaryotic proteins by the fusion OET-KNN classifier.

Amino Acids↗

Improving spliced alignment by modeling splice sites with deep learning.

MOTIVATION: Spliced alignment refers to the alignment of messenger RNA (mRNA) or protein sequences to eukaryotic genomes. It plays a critical role in gene annotation and the study of gene functions. Accurate spliced alignment demands sophisticated modeling of splice sites, but current aligners use simple models, which may affect their accuracy given dissimilar sequences. RESULTS: We implemented minisplice to learn splice signals with a one-dimensional convolutional neural network (1D-CNN) and trained a model with 7,026 parameters for vertebrate and insect genomes. It captures conserved splice signals across phyla and reveals GC-rich introns specific to mammals and birds. We used this model to estimate the empirical splicing probability for every GT and AG in genomes, and modified minimap2 and miniprot to leverage pre-computed splicing probability during alignment. Evaluation on human long-read RNA-seq data and cross-species protein datasets showed our method greatly improves the junction accuracy especially for noisy long RNA-seq reads and proteins of distant homology. AVAILABILITY AND IMPLEMENTATION: https://github.com/lh3/minisplice.

Journal Article↗

Predicting function: from genes to genomes and back.

Predicting function from sequence using computational tools is a highly complicated procedure that is generally done for each gene individually. This review focuses on the added value that is provided by completely sequenced genomes in function prediction. Various levels of sequence annotation and function prediction are discussed, ranging from genomic sequence to that of complex cellular processes. Protein function is currently best described in the context of molecular interactions. In the near future it will be possible to predict protein function in the context of higher order processes such as the regulation of gene expression, metabolic pathways and signalling cascades. The analysis of such higher levels of function description uses, besides the information from completely sequenced genomes, also the additional information from proteomics and expression data. The final goal will be to elucidate the mapping between genotype and phenotype.

Bacterial Proteins↗

SNAPping up functionally related genes based on context information: a colinearity-free approach.

We describe a computational approach for finding genes that are functionally related but do not possess any noticeable sequence similarity. Our method, which we call SNAP (similarity-neighborhood approach), reveals the conservation of gene order on bacterial chromosomes based on both cross-genome comparison and context information. The novel feature of this method is that it does not rely on detection of conserved colinear gene strings. Instead, we introduce the notion of a similarity-neighborhood graph (SN-graph), which is constructed from the chains of similarity and neighborhood relationships between orthologous genes in different genomes and adjacent genes in the same genome, respectively. An SN-cycle is defined as a closed path on the SN-graph and is postulated to preferentially join functionally related gene products that participate in the same biochemical or regulatory process. We demonstrate the substantial non-randomness and functional significance of SN-cycles derived from real genome data and estimate the prediction accuracy of SNAP in assigning broad function to uncharacterized proteins. Examples of practical application of SNAP for improving the quality of genome annotation are described.

Algorithms↗

FISH analysis of Drosophila melanogaster heterochromatin using BACs and P elements.

The heterochromatin of chromosomes 2 and 3 of Drosophila melanogaster contains about 30 essential genes defined by genetic analysis. In the last decade only a few of these genes have been molecularly characterized and found to correspond to protein-coding genes involved in important cellular functions. Moreover, several predicted genes have been identified by annotation of genomic sequence that are associated with polytene chromosome divisions 40, 41 and 80 but their locations on the cytogenetic map of the heterochromatin are still uncertain. To expand our current knowledge of the genetic functions located in heterochromatin, we have performed fluorescence in situ hybridization (FISH) mapping to mitotic chromosomes of nine bacterial artificial chromosomes (BACs) carrying several predicted genes and of 13 P element insertions assigned to the proximal regions of 2R and 3L. We found that 22 predicted genes map to the h46 region of 2R and eight map to the h47 regions of 3L. This amounts to at least 30 predicted genes located in these heterochromatic regions, whereas previous studies detected only seven vital genes. Finally, another 58 genes localize either in the euchromatin-heterochromatin transition regions or in the proximal euchromatin of 2R and 3L.

Animals↗

The database of epoxide hydrolases and haloalkane dehalogenases: one structure, many functions.

UNLABELLED: The epoxide hydrolases and haloalkane dehalogenases database (EH/HD) integrates sequence and structure of a highly diverse protein family, including mainly the Asp-hydrolases of EHs and HDs but also proteins, such as Ser-hydrolases non-heme peroxidases, prolyl iminopetidases and 2-hydroxymuconic semialdehyde hydrolases. These proteins have a highly conserved structure, but display a remarkable diversity in sequence and function. A total of 305 protein entries were assigned to 14 homologous families, forming two superfamilies. Annotated multisequence alignments and phylogenetic trees are provided for each homologous family and superfamily. Experimentally derived structures of 19 proteins are superposed and consistently annotated. Sequence and structure of all 305 proteins were systematically analysed. Thus, deeper insight is gained into the role of a highly conserved sequence motifs and structural elements. AVAILABILITY: The EH/HD database is available at http://www.led.uni-stuttgart.de

Amino Acid Motifs↗

PEP: Predictions for Entire Proteomes.

PEP is a database of Predictions for Entire Proteomes. The database contains summaries of analyses of protein sequences from a range of organisms representing all three major kingdoms of life: eukaryotes, prokaryotes and archaea. All proteins publicly available for organisms were aligned against SWISS-PROT, TrEMBL and PDB. Additionally, the following annotations are provided: secondary structure, transmembrane helices, coiled coils, regions of low complexity, signal peptides, PROSITE motifs, nuclear localization signals and classes of cellular function. Proteins that contain long regions without regular secondary structure are also identified. We have produced a related database of structural domain-like fragments derived from PEP and clusters based on homology between all fragments. The PEP database, fragments and clusters are distributed freely as a set of flat files and have been integrated into SRS. The PEP group of databases can be accessed from: http://cubic.bioc.columbia.edu/pep.

Animals↗

A classification of disulfide patterns and its relationship to protein structure and function.

We report a detailed classification of disulfide patterns to further understand the role of disulfides in protein structure and function. The classification is applied to a unique searchable database of disulfide patterns derived from the SwissProt and Pfam databases. The disulfide database contains seven times the number of publicly available disulfide annotations. Each disulfide pattern in the database captures the topology and cysteine spacing of a protein domain. We have clustered the domains by their disulfide patterns and visualized the results using a novel representation termed the "classification wheel." The classification is applied to 40,620 protein domains with 2-10 disulfides. The effectiveness of the classification is evaluated by determining the extent to which proteins of similar structure and function are grouped together through comparison with the SCOP and Pfam databases, respectively. In general, proteins with similar disulfide patterns have similar structure and function, even in cases of low sequence similarity, and we illustrate this with specific examples. Using a measure of disulfide topology complexity, we find that there is a predominance of less complex topologies. We also explored the importance of loss or addition of disulfides to protein structure and function by linking classification wheels through disulfide subpattern comparisons. This classification, when coupled with our disulfide database, will serve as a useful resource for searching and comparing disulfide patterns, and understanding their role in protein structure, folding, and stability. Proteins in the disulfide clusters that do not contain structural information are prime candidates for structural genomics initiatives, because they may correspond to novel structures.

Animals↗

The Feasibility of Using Proteome Expression Profile for Genome Annotation.

By investigating into the expression data from ECO2DBASE (Edition 6),the feasibility of using proteome expression profile for genome annotation was tested. Based on our newly developed CRC (cellular role cluster) method,79 proteins extracted from ECO2DBASE were clustered into 4 CRCs. Function related proteins tend to be clustered into same CRC. Total 9 aminoacyl-tRNA synthetases were clustered into CRC2, whereas 4 heat-shock proteins into CRC3. These results indicate with enough proteome expression data and the efficient algorithm, proteome expression profile can provide very important information for genome annotation, while this kind of information is sequence-independent.

Journal Article↗

Functional cloning, sorting, and expression profiling of nucleic acid-binding proteins.

A major challenge in the post-sequencing era is to elucidate the activity and biological function of genes that reside in the human genome. An important subset includes genes that encode proteins that regulate gene expression or maintain the structural integrity of the genome. Using a novel oligonucleotide-binding substrate as bait, we show the feasibility of a modified functional expression-cloning strategy to identify human cDNAs that encode a spectrum of nucleic acid-binding proteins (NBPs). Approximately 170 cDNAs were identified from screening phage libraries derived from a human colorectal adenocarcinoma cell line and from noncancerous fetal lung tissue. Sequence analysis confirmed that virtually every clone contained a known DNA- or RNA-binding motif. We also report on a complementary sorting strategy that, in the absence of subcloning and protein purification, can distinguish different classes of NBPs according to their particular binding properties. To extend our functional annotation of NBPs, we have used GeneChip expression profiling of 14 different breast-derived cell lines to examine the relative transcriptional activity of genes identified in our screen and cluster analysis to discover other genes that have similar expression patterns. Finally, we present strategies to analyze the upstream regulatory region of each gene within a cluster group and select unique combinations of transcription factor binding sites that may be responsible for dictating the observed synexpression.

Adenocarcinoma↗

Slr2013 is a novel protein regulating functional assembly of photosystem II in Synechocystis sp. strain PCC 6803.

The Synechocystis sp. strain PCC 6803, which has a T192H mutation in the D2 protein of photosystem II, is an obligate photoheterotroph due to the lack of assembled photosystem II complexes. A secondary mutant, Rg2, has been selected that retains the T192H mutation but is able to grow photoautotrophically. Restoration of photoautotrophic growth in this mutant was caused by early termination at position 294 in the Slr2013 protein. The T192H mutant with truncated Slr2013 forms fully functional photosystem II reaction centers that differ from wild-type reaction centers only by a 30% higher rate of charge recombination between the primary electron acceptor, QA-, and the donor side and by a reduced stability of the oxidized form of the redox-active Tyr residue, YD, in the D2 protein. This suggests that the T192H mutation itself did not directly affect electron transfer components, but rather affected protein folding and/or stable assembly of photosystem II, and that Slr2013 is involved in the folding of the D2 protein and the assembly of photosystem II. Besides participation in photosystem II assembly, Slr2013 plays a critical role in the cell, because the corresponding gene cannot be deleted completely under conditions in which photosystem II is dispensable. Truncation of Slr2013 by itself does not affect photosynthetic activity of Synechocystis sp. strain PCC 6803. Slr2013 is annotated in CyanoBase as a hypothetical protein and shares a DUF58 family signature with other hypothetical proteins of unknown function. Genes for close homologues of Slr2013 are found in other cyanobacteria (Nostoc punctiforme, Anabaena sp. strain PCC 7120, and Thermosynechococcus elongatus BP-1), and apparent orthologs of this protein are found in Eubacteria and Archaea, but not in eukaryotes. We suggest that Slr2013 regulates functional assembly of photosystem II and has at least one other important function in the cell.

Amino Acid Sequence↗