Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Small envelope protein E of SARS: cloning, expression, purification, CD determination, and bioinformatics analysis.

AIM: To obtain the pure sample of SARS small envelope E protein (SARS E protein), study its properties and analyze its possible functions. METHODS: The plasmid of SARS E protein was constructed by the polymerase chain reaction (PCR), and the protein was expressed in the E coli strain. The secondary structure feature of the protein was determined by circular dichroism (CD) technique. The possible functions of this protein were annotated by bioinformatics methods, and its possible three-dimensional model was constructed by molecular modeling. RESULTS: The pure sample of SARS E protein was obtained. The secondary structure feature derived from CD determination is similar to that from the secondary structure prediction. Bioinformatics analysis indicated that the key residues of SARS E protein were much conserved compared to the E proteins of other coronaviruses. In particular, the primary amino acid sequence of SARS E protein is much more similar to that of murine hepatitis virus (MHV) and other mammal coronaviruses. The transmembrane (TM) segment of the SARS E protein is relatively more conserved in the whole protein than other regions. CONCLUSION: The success of expressing the SARS E protein is a good starting point for investigating the structure and functions of this protein and SARS coronavirus itself as well. The SARS E protein may fold in water solution in a similar way as it in membrane-water mixed environment. It is possible that beta-sheet I of the SARS E protein interacts with the membrane surface via hydrogen bonding, this beta-sheet may uncoil to a random structure in water solution.

Circular Dichroism↗

Transcriptomal profiling of the cellular transformation induced by Rho subfamily GTPases.

We have used microarray technology to identify the transcriptional targets of Rho subfamily guanosine 5'-triphosphate (GTP)ases in NIH3T3 cells. This analysis indicated that murine fibroblasts transformed by these proteins show similar transcriptomal profiles. Functional annotation of the regulated genes indicate that Rho subfamily GTPases target a wide spectrum of functions, although loci encoding proteins linked to proliferation and DNA synthesis/transcription are upregulated preferentially. Rho proteins promote four main networks of interacting proteins nucleated around E2F, c-Jun, c-Myc and p53. Of those, E2F, c-Jun and c-Myc are essential for the maintenance of cell transformation. Inhibition of Rock, one of the main Rho GTPase targets, leads to small changes in the transcriptome of Rho-transformed cells. Rock inhibition decreases c-myc gene expression without affecting the E2F and c-Jun pathways. Loss-of-function studies demonstrate that c-Myc is important for the blockage of cell-contact inhibition rather than for promoting the proliferation of Rho-transformed cells. However, c-Myc overexpression does not bypass the inhibition of cell transformation induced by Rock blockage, indicating that c-Myc is essential, but not sufficient, for Rock-dependent transformation. These results reveal the complexity of the genetic program orchestrated by the Rho subfamily and pinpoint protein networks that mediate different aspects of the malignant phenotype of Rho-transformed cells.

Amino Acid Substitution↗

Struct2net: integrating structure into protein-protein interaction prediction.

UNLABELLED: This paper presents a framework for predicting protein-protein interactions (PPI) that integrates structure-based information with other functional annotations, e.g. GO, co-expression and co-localization, etc., Given two protein sequences, the structure-based interaction prediction technique threads these two sequences to all the protein complexes in the PDB and then chooses the best potential match. Based on this match, structural information is incorporated into logistic regression to evaluate the probability of these two proteins interacting. This paper also describes a random forest classifier which can effectively combine the structure-based prediction results and other functional annotations together to predict protein interactions. Experimental results indicate that the predictive power of the structure-based method is better than many other information sources. Also, combining the structure-based method with other information sources allows us to achieve a better performance than when structure information is not used. We also tested our method on a set of approximately 1000 yeast genes and, interestingly, the predicted interaction network is a scale-free network. Our method predicted some potential interactions involving yeast homologs of human disease-related proteins. SUPPLEMENTARY INFORMATION: http://theory.csail.mit.edu/struct2net

Algorithms↗

Large-scale functional annotation establishes a reference framework for human LRRK2 variants.

Pathogenic variants in leucine-rich repeat kinase 2 (LRRK2)1are among the most frequent monogenic causes of Parkinson's disease (PD)2 and act through a gain-of-function mechanism of increased kinase activity. LRRK2-targeted therapies are in clinical development, but interpretation of the rapidly expanding catalogue of rare LRRK2 variants remains a barrier to translation. Here, we present functionally annotated data on >350 LRRK2 coding variants using a standardized cellular assay with Rab10 phosphorylation as a readout of kinase activity and integrated these data with curated genetic and clinical annotations from the Movement Disorders Society Genetic Mutation Database (MDSGene). Variants differed in activation magnitude, ranging from modest increases (e.g., p.G2019S) to strongly activating substitutions such as p.Y1699C or p.L1795F. Activating variants occurred across the full length of LRRK2, although the largest effects clustered within the ROC-COR regulatory hub, where structural analysis identified subdomains forming an allosteric scaffold controlling kinase output. All known/established pathogenic variants showed increased activity, whereas benign and likely benign variants remained within the wild-type range. Functional effect sizes correlated with pathway activation in patient-derived immune cells, altogether providing a framework for ACMG-based variant interpretation in which kinase activation can support PS3 functional evidence for reclassification of variants.

Protein phosphorylation↗

The SBASE protein domain library, release 9.0: an online resource for protein domain identification.

SBASE (http://www.icgeb.trieste.it/sbase) is an online resource of protein domain sequences designed to facilitate detection of domain homologies based on a simple database search. The ninth release of the SBASE library of protein domain sequences contains 320 000 annotated structural, functional, ligand-binding and topogenic segments of proteins clustered into over 3481 domain groups and 483 protein families. Domain identification and functional prediction are based on a comparison of BLAST search outputs with a knowledge base of within-group ('self') and out-of-group ('non-self') similarities of the known domain groups. This is a memory-based approach wherein class-specific similarity functions are automatically learned from the database [Stanfill,C. and Waltz,D. (1986) COMMUN: ACM, 29, 1213-1228].

Animals↗

Functional annotation of class I lysyl-tRNA synthetase phylogeny indicates a limited role for gene transfer.

Functional and comparative genomic studies have previously shown that the essential protein lysyl-tRNA synthetase (LysRS) exists in two unrelated forms. Most prokaryotes and all eukaryotes contain a class II LysRS, whereas most archaea and a few bacteria contain a less common class I LysRS. In bacteria the class I LysRS is only found in the alpha-proteobacteria and a scattering of other groups, including the spirochetes, while the class I protein is by far the most common form of LysRS in archaea. To investigate this unusual distribution we functionally annotated a representative phylogenetic sampling of LysRS proteins. Class I LysRS proteins from a variety of bacteria and archaea were characterized in vitro by their ability to recognize Escherichia coli tRNA(Lys) anticodon mutants. Class I LysRS proteins were found to fall into two distinct groups, those that preferentially recognize the third anticodon nucleotide of tRNA(Lys) (U36) and those that recognize both the second and third positions (U35 and U36). Strong recognition of U35 and U36 was confined to the pyrococcus-spirochete grouping within the archaeal branch of the class I LysRS phylogenetic tree, while U36 recognition was seen in other archaea and an example from the alpha-proteobacteria. Together with the corresponding phylogenetic relationships, these results suggest that despite its comparative rarity the distribution of class I LysRS conforms to the canonical archaeal-bacterial division. The only exception, suggested from both functional and phylogenetic data, appears to be the horizontal transfer of class I LysRS from a pyrococcal progenitor to a limited number of bacteria.

Acylation↗

The use of edge-betweenness clustering to investigate biological function in protein interaction networks.

BACKGROUND: This paper describes an automated method for finding clusters of interconnected proteins in protein interaction networks and retrieving protein annotations associated with these clusters. RESULTS: Protein interaction graphs were separated into subgraphs of interconnected proteins, using the JUNG implementation of Girvan and Newman's Edge-Betweenness algorithm. Functions were sought for these subgraphs by detecting significant correlations with the distribution of Gene Ontology terms which had been used to annotate the proteins within each cluster. The method was implemented using freely available software (JUNG and the R statistical package). Protein clusters with significant correlations to functional annotations could be identified and included groups of proteins know to cooperate in cell metabolism. The method appears to be resilient against the presence of false positive interactions. CONCLUSION: This method provides a useful tool for rapid screening of small to medium size protein interaction datasets.

Algorithms↗

Systematic Proteome Profiling of Maternal Plasma for Development of Preeclampsia Biomarkers.

Preeclampsia (PE) is a hypertensive disorder of pregnancy with various clinical symptoms. However, traditional markers for the disease including high blood pressure and proteinuria are poor indicators of the related adverse outcomes. Here, we performed systematic proteome profiling of plasma samples obtained from pregnant women with PE to identify clinically effective diagnostic biomarkers. Proteome profiling was performed using TMT-based liquid chromatography-mass spectrometry (LC-MS/MS) followed by subsequent verification by multiple reaction monitoring (MRM) analysis on normal and PE maternal plasma samples. Functional annotations of differentially expressed proteins (DEPs) in PE were predicted using bioinformatic tools. The diagnostic accuracies of the biomarkers for PE were estimated according to the area under the receiver-operating characteristics curve (AUC). A total of 1307 proteins were identified, and 870 proteins of them were quantified from plasma samples. Significant differences were evident in 138 DEPs, including 71 upregulated DEPs and 67 downregulated DEPs in the PE group, compared with those in the control group. Upregulated proteins were significantly associated with biological processes including platelet degranulation, proteolysis, lipoprotein metabolism, and cholesterol efflux. Biological processes including blood coagulation and acute-phase response were enriched for down-regulated proteins. Of these, 40 proteins were subsequently validated in an independent cohort of 26 PE patients and 29 healthy controls. APOM, LCN2, and QSOX1 showed high diagnostic accuracies for PE detection (AUC >0.9 and p&#xa0;<&#xa0;0.001, for all) as validated by MRM and ELISA. Our data demonstrate that three plasma biomarkers, identified by systematic proteomic profiling, present a possibility for the assessment of PE, independent of the clinical characteristics of pregnant women.

Humans↗

Chromosome-level genome assembly with telomeric repeats at scaffold ends for Rhabdosargus sarba.

Rhabdosargus sarba, the goldlined seabream, is a euryhaline marine fish of great aquaculture potential. Genome sequencing and assembly of R. sarba was carried utilizing a multi-platform sequencing strategy that included long-read sequencing (PacBio HiFi), short-read sequencing (Illumina), and chromatin interaction mapping (Hi-C). The final genome assembly size after scaffolding was 764.59&#x2009;Mb in 31 scaffolds with an N50 length of 33.98&#x2009;Mb. Repeat profiling of primary assembly showed that 28.71% of the genome comprises of repeat elements. Gene prediction utilising the evidence from ab initio prediction and transcriptome data revealed 26,913 protein encoding genes and functional annotation and pathway analysis showed their participation in 332 pathways. This genome is an excellent resource for future research on genetic improvement and molecular breeding programmes for R. sarba.

Animals↗

3MATRIX and 3MOTIF: a protein structure visualization system for conserved sequence motifs.

Computational methods such as sequence alignment and motif construction are useful in grouping related proteins into families, as well as helping to annotate new proteins of unknown function. These methods identify conserved amino acids in protein sequences, but cannot determine the specific functional or structural roles of conserved amino acids without additional study. In this work, we present 3MATRIX (http://3matrix.stanford.edu) and 3MOTIF (http://3motif.stanford.edu), a web-based sequence motif visualization system that displays sequence motif information in its appropriate three-dimensional (3D) context. This system is flexible in that users can enter sequences, keywords, structures or sequence motifs to generate visualizations. In 3MOTIF, users can search using discrete sequence motifs such as PROSITE patterns, eMOTIFs, or any other regular expression-like motif. Similarly, 3MATRIX accepts an eMATRIX position-specific scoring matrix, or will convert a multiple sequence alignment block into an eMATRIX for visualization. Each query motif is used to search the protein structure database for matches, in which the motif is then visually highlighted in three dimensions. Important properties of motifs such as sequence conservation and solvent accessible surface area are also displayed in the visualizations, using carefully chosen color shading schemes.

Amino Acid Motifs↗

Nuclease activity of the MutS homologue MutS2 from Thermus thermophilus is confined to the Smr domain.

MutS homologues are highly conserved enzymes engaged in DNA mismatch repair (MMR), meiotic recombination and other DNA modifications. Genome sequencing projects have revealed that bacteria and plants possess a MutS homologue, MutS2. MutS2 lacks the mismatch-recognition domain of MutS, but contains an extra C-terminal region called the small MutS-related (Smr) domain. Sequences homologous to the Smr domain are annotated as 'proteins of unknown function' in various organisms ranging from bacteria to human. Although recent in vivo studies indicate that MutS2 plays an important role in recombinational events, there had been only limited characterization of the biochemical function of MutS2 and the Smr domain. We previously established that Thermus thermophilus MutS2 (ttMutS2) possesses endonuclease activity. In this study, we report that a Smr-deleted ttMutS2 mutant retains the dimerization, ATPase and DNA-binding activities, but has no endonuclease activity. Furthermore, the Smr domain alone was stable and functional in binding and incising DNA. It is noteworthy that an endonuclease activity is associated with a MutS homologue, which is generally thought to recognize specific DNA structures.

Adenosine Triphosphatases↗

Solution structure of Archaeglobus fulgidis peptidyl-tRNA hydrolase (Pth2) provides evidence for an extensive conserved family of Pth2 enzymes in archea, bacteria, and eukaryotes.

The solution structure of protein AF2095 from the thermophilic archaea Archaeglobus fulgidis, a 123-residue (13.6-kDa) protein, has been determined by NMR methods. The structure of AF2095 is comprised of four alpha-helices and a mixed beta-sheet consisting of four parallel and anti-parallel beta-strands, where the alpha-helices sandwich the beta-sheet. Sequence and structural comparison of AF2095 with proteins from Homo sapiens, Methanocaldococcus jannaschii, and Sulfolobus solfataricus reveals that AF2095 is a peptidyl-tRNA hydrolase (Pth2). This structural comparison also identifies putative catalytic residues and a tRNA interaction region for AF2095. The structure of AF2095 is also similar to the structure of protein TA0108 from archaea Thermoplasma acidophilum, which is deposited in the Protein Data Bank but not functionally annotated. The NMR structure of AF2095 has been further leveraged to obtain good-quality structural models for 55 other proteins. Although earlier studies have proposed that the Pth2 protein family is restricted to archeal and eukaryotic organisms, the similarity of the AF2095 structure to human Pth2, the conservation of key active-site residues, and the good quality of the resulting homology models demonstrate a large family of homologous Pth2 proteins that are conserved in eukaryotic, archaeal, and bacterial organisms, providing novel insights in the evolution of the Pth and Pth2 enzyme families.

Archaea↗

I-superfamily conotoxins: sequence and structure analysis.

I-superfamily conotoxins have four-disulfide bonds with cysteine arrangement C-C-CC-CC-C-C, and they inhibit or modify ion channels of nerve cells. They have been characterized only recently and are relatively less well studied compared to other superfamily conotoxins. We have detected selective and sensitive sequence pattern for I-superfamily conotoxins. The availability of sequence pattern should be useful in protein family classification and functional annotation. We have built by homology modeling, a theoretical structural 3D model of ViTx from Conus virgo, a typical member of I-superfamily conotoxins. The modeling was based on the available 3D structure of Janus-atracotoxin-Hv1c of Janus-atracotoxin family whose members have been suggested as possible biopesticides. A study comparing the theoretically modeled structure of ViTx, with experimentally determined structures of other toxins, which share functional similarity with ViTx, reveals the crucial role of C-terminal region of ViTx in blocking therapeutically important voltage-gated potassium channels.

Amino Acid Sequence↗

The cadherin superfamily database.

The cadherin superfamily is a large protein family with diverse structures and functions. Because of this diversity and the growing biological interest in cell adhesion and signaling processes, in which many members of the cadherin superfamily play a crucial role, it is becoming increasingly important to develop tools to manage, distribute and analyze sequences in this protein family. Current profile and motif databases classify protein sequences into a broad spectrum of protein superfamilies, however to provide a more specific functional annotation, the next step should include classification of subfamilies of these protein superfamilies. Here, we present a tool that classified greater than 90% of the proteins belonging to the cadherin superfamily found in the SWISS PROT database. Therefore, for most members of the cadherin superfamily, this tool can assist in adding more specific functional annotations than can be achieved with current profile and motif databases. Finally, the classification tool and the results of our analysis were integrated into a web-accessible database (http://calcium.uhnres. utoronto.ca/cadherin).

Amino Acid Motifs↗

Evaluation of features for catalytic residue prediction in novel folds.

Structural genomics projects are determining the three-dimensional structure of proteins without full characterization of their function. A critical part of the annotation process involves appropriate knowledge representation and prediction of functionally important residue environments. We have developed a method to extract features from sequence, sequence alignments, three-dimensional structure, and structural environment conservation, and used support vector machines to annotate homologous and nonhomologous residue positions based on a specific training set of residue functions. In order to evaluate this pipeline for automated protein annotation, we applied it to the challenging problem of prediction of catalytic residues in enzymes. We also ranked the features based on their ability to discriminate catalytic from noncatalytic residues. When applying our method to a well-annotated set of protein structures, we found that top-ranked features were a measure of sequence conservation, a measure of structural conservation, a degree of uniqueness of a residue's structural environment, solvent accessibility, and residue hydrophobicity. We also found that features based on structural conservation were complementary to those based on sequence conservation and that they were capable of increasing predictor performance. Using a family nonredundant version of the ASTRAL 40 v1.65 data set, we estimated that the true catalytic residues were correctly predicted in 57.0% of the cases, with a precision of 18.5%. When testing on proteins containing novel folds not used in training, the best features were highly correlated with the training on families, thus validating the approach to nonhomologous catalytic residue prediction in general. We then applied the method to 2781 coordinate files from the structural genomics target pipeline and identified both highly ranked and highly clustered groups of predicted catalytic residues.

Algorithms↗

Protein-interaction networks: from experiments to analysis.

Functional proteomics approaches aim to characterize comprehensively the function of gene products, and provide a first-level understanding of cellular mechanisms. Here, we review recent techniques for the construction and prediction of large-scale protein-interaction networks, with a particular emphasis on computational processing steps and comparative assessment of the reliability and completeness of the various approaches. We also discuss the use of protein-interaction network information in functional annotation and in the generation of higher-level biological hypotheses on pathways.

Computational Biology↗

Probabilistic model of the human protein-protein interaction network.

A catalog of all human protein-protein interactions would provide scientists with a framework to study protein deregulation in complex diseases such as cancer. Here we demonstrate that a probabilistic analysis integrating model organism interactome data, protein domain data, genome-wide gene expression data and functional annotation data predicts nearly 40,000 protein-protein interactions in humans-a result comparable to those obtained with experimental and computational approaches in model organisms. We validated the accuracy of the predictive model on an independent test set of known interactions and also experimentally confirmed two predicted interactions relevant to human cancer, implicating uncharacterized proteins into definitive pathways. We also applied the human interactome network to cancer genomics data and identified several interaction subnetworks activated in cancer. This integrative analysis provides a comprehensive framework for exploring the human protein interaction network.

Chromosome Mapping↗

Markovian domain fingerprinting: statistical segmentation of protein sequences.

MOTIVATION: Characterization of a protein family by its distinct sequence domains is crucial for functional annotation and correct classification of newly discovered proteins. Conventional Multiple Sequence Alignment (MSA) based methods find difficulties when faced with heterogeneous groups of proteins. However, even many families of proteins that do share a common domain contain instances of several other domains, without any common underlying linear ordering. Ignoring this modularity may lead to poor or even false classification results. An automated method that can analyze a group of proteins into the sequence domains it contains is therefore highly desirable. RESULTS: We apply a novel method to the problem of protein domain detection. The method takes as input an unaligned group of protein sequences. It segments them and clusters the segments into groups sharing the same underlying statistics. A Variable Memory Markov (VMM) model is built using a Prediction Suffix Tree (PST) data structure for each group of segments. Refinement is achieved by letting the PSTs compete over the segments, and a deterministic annealing framework infers the number of underlying PST models while avoiding many inferior solutions. We show that regions of similar statistics correlate well with protein sequence domains, by matching a unique signature to each domain. This is done in a fully automated manner, and does not require or attempt an MSA. Several representative cases are analyzed. We identify a protein fusion event, refine an HMM superfamily classification into the underlying families the HMM cannot separate, and detect all 12 instances of a short domain in a group of 396 sequences. CONTACT: jill@cs.huji.ac.il; tishby@cs.huji.ac.il.

Algorithms↗