Search PubMedSearch

SEARCH · Search PubMed

Results for “Protein function prediction”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Modification of plant proteins by immobilized proteases.

A potential application of plant proteins could be a replacement of animal proteins now in use in the food industry on the basis of certain specific functional properties plant proteins have. Modification of the chemical structure of selected plant proteins is needed to replace more expensive animal proteins as food ingredients that have specific functional characteristics. Structure modification may be achieved by physical, chemical, or microbiological methods, or by a combination of these. Immobilized enzyme techniques offer significant advantages for protein modification. Knowledge of the molecular properties of plant proteins is essential to understand the basis of protein functionality, to modify proteins so that they acquire desirable functional properties, and to predict potential applications of modified plant proteins. This paper reviews all the above mentioned aspects of plant protein chemistry and potential utilization.

Amino Acids

Cloning and sequencing of the yeast gene for dolichol phosphate mannose synthase, an essential protein.

Dolichol phosphate mannose (Dol-P-Man) synthase (EC 2.4.1.83) catalyzes the formation of Dol-P-Man from Dol-P and GDP-Man. The structural gene for yeast Dol-P-Man synthase (DPM1) was isolated by screening a yeast genomic DNA library for colonies that overexpressed Dol-P-Man synthase activity. This approach relied on a method to screen for Dol-P-Man synthase activity in lysed yeast colonies and used a yeast mutant with very low Dol-P-Man synthase activity in colony lysates. Transformants isolated using this technique expressed Dol-P-Man synthase activity 9-14-fold higher than that of a wild type strain, and all seven plasmids conferring this overproduction had a common region in their yeast genomic DNA insert. DPM1 is the structural gene for yeast Dol-P-Man synthase since Escherichia coli transformants harboring this gene express Dol-P-Man synthase activity in vitro. DNA sequencing of the DPM1 gene revealed an open reading frame of 801 bases. The 30-kDa size of the predicted protein is in excellent agreement with the size of the purified yeast enzyme (Haselbeck, A., and Tanner, W. (1982) Proc. Natl. Acad. Sci. U. S. A. 79, 1520-1524). Analysis of the predicted amino acid sequence reveals the protein has a potential membrane spanning domain of 25 amino acids at its COOH terminus. The protein's NH2 terminus, though not hydrophobic, meets existing criteria for yeast signal sequences, but there is no site for cleavage by signal peptidase. If the NH2 terminus is a functional signal sequence, the protein is predicted to be oriented toward the lumen of the endoplasmic reticulum with both NH2 and COOH termini serving as membrane anchors. If there is no signal sequence, the enzyme is predicted to face the cytoplasm and be anchored only by its COOH terminus. The DPM1 gene is essential for viability in yeast since disruption of the gene is lethal. We suspect Dol-P-Man synthase is not an essential protein due to its role in N-glycosylation since mutations in other genes that affect the late steps in lipid-linked oligosaccharide synthesis do not affect cell growth. Instead, DPM1 may be an essential gene because its product is required for O-glycosylation in yeast or because Dol-P-Man synthase is needed in some unidentified pathway.

Alleles

The PRP31 gene encodes a novel protein required for pre-mRNA splicing in Saccharomyces cerevisiae.

The pre-mRNA splicing factor Prp31p was identified in a screen of temperature-sensitive yeast strains for those exhibiting a splicing defect upon shift to the non- permissive temperature. The wild-type PRP31 gene was cloned and shown to be essential for cell viability. The PRP31 gene is predicted to encode a 60 kDa polypeptide. No similarities with other known splicing factors or motifs indicative of protein-protein or RNA-protein interaction domains are discernible in the predicted amino acid sequence. A PRP31 allele bearing a triple repeat of the hemagglutinin epitope has been generated. The tagged protein is functional in vivo and a single polypeptide species of the predicted size was detected by Western analysis with proteins from yeast cell extracts. Functional Prp31p is required for the processing of pre-mRNA species both in vivo and in vitro, indicating that the protein is directly involved in the splicing pathway.

Cloning, Molecular

MRP-8 and MRP-14, two abundant Ca(2+)-binding proteins of neutrophils and monocytes.

Two calcium-binding proteins, named migration inhibitory factor-related proteins-8 (MRP-8) and MRP-14, are primarily expressed by circulating human neutrophils and monocytes. Evidence accumulating from the investigations of several independent groups is now leading to an improved understanding of the biology of these proteins. Both MRP-8 and MRP-14 display features characteristic of members of the S100 family of calcium-binding proteins. Some of these features predict functions for MRP-8 and MRP-14 but to date an exact and well-defined function remains elusive. Here we review the available information and highlight evidence that suggests the function of MRP-8 and MRP-14 may be associated with both monocyte and neutrophil activation and the accumulation of these cells in inflammatory sites.

Amino Acid Sequence

Translational control mediates the developmental regulation of the Trypanosoma brucei Nrk protein kinase.

The expression and function of eukaryotic protein kinases is highly regulated, primarily through transcriptional and post-translational processes. In this report we demonstrate an unusual mechanism for controlling protein kinase function, translational control. The Trypanosoma brucei Nrk loci encode predicted protein kinases. Here we show that Nrk has protein serine-threonine kinase activity and examine the expression and activity of Nrk during parasite development. While Nrk transcripts were previously found to be constitutively expressed throughout the life cycle, we now find that expression of Nrk protein is highly stage-regulated. Immunoblot analysis revealed that Nrk expression dramatically increased as the parasites differentiated from proliferative slender bloodforms to the non-proliferative stumpy bloodforms. Procyclic form organisms expressed moderate levels of Nrk. Analysis of Nrk activity demonstrated that it too was highest in stumpy bloodforms. Metabolic labeling and pulse-chase analysis demonstrated that Nrk accumulation was highest in stumpy bloodforms and indicated that Nrk abundance is primarily controlled at the level of biosynthesis rather than turnover. All Nrk mRNA was contained in the poly(A)+ fraction, and the 5' ends of the transcript were the same in each developmental stage. Thus, Nrk is under translational control. The strict developmental regulation of the Nrk enzymes within the trypanosome life cycle suggests that the Nrk protein kinase may play a role in parasite differentiation.

Animals

Structural and functional features of Drosophila chorion proteins s36 and s38 from analysis of primary structure and infrared spectroscopy.

Amino acid composition, Fourier transform analysis and secondary structure prediction methods strongly support a tripartite structure for Drosophila chorion proteins s36 and s38. Each protein consists of a central domain and two flanking 'arms'. The central domain contains tandemly repetitive peptides, which apparently generate a secondary structure of beta-sheet strands alternating with beta-turns, most probably, forming a twisted beta-pleated sheet or beta-barrel. The central domains of s36 and s38 share similarities, but they are recognizably different. The flanking 'arms', with different primary and secondary structure features, presumably serve protein-specific functions. The possible roles of the protein domains for the establishment of higher order structure in Drosophila chorion and the possible function of the molecules are discussed. The predicted secondary structure of Drosophila chorion proteins s36 and s38 is supported by experimental information obtained from Fourier transform infrared spectroscopic studies of Drosophila chorions.

Amino Acid Sequence

Identification of IDH3G, encoding the gamma subunit of mitochondrial isocitrate dehydrogenase, as a novel candidate gene for X-linked retinitis pigmentosa.

PURPOSE: Retinitis pigmentosa (RP) is a genetically heterogeneous group of retinal degenerative disorders characterized by the loss of rod and cone photoreceptors, leading to visual impairment and blindness. To date, to our knowledge, X-linked RP has been associated with variants in 3 genes (RPGR, RP2, and OFD1), whereas genetic defects at 3 loci (RP6, RP24, and RP34) are yet unidentified. The aim of this study was to identify a novel candidate gene underlying X-linked RP. METHODS: Participants were identified from cohorts of genetically unsolved male individuals affected by RP, who underwent genome sequencing, exome sequencing, or candidate gene screening via direct Sanger sequencing at 3 referral centers. Specifically, 2 probands were identified at the National Reference Centre for Rare Retinal Diseases (Paris, France), 2 at the Massachusetts Eye and Ear Hospital (Boston, MA), and 1 at the National Reference Centre for Inherited Sensory Diseases (Montpellier, France). The pathogenicity of the identified variants was assessed using bioinformatic predictions, protein expression analyses, and mitochondrial function assays. RESULTS: We identified 4 rare single-nucleotide variants in IDH3G (HGNC:5386), located at the RP34 locus on the X chromosome, and a complete gene deletion, in 5 unrelated male individuals affected with nonsyndromic RP. The variants segregated with the phenotype in all available family members. In all cases, the disease severity was intermediate. None had high myopia. IDH3G encodes the γ subunit of mitochondrial isocitrate dehydrogenase (IDH3), an enzyme involved in the citric acid cycle, which is expressed in the inner segments of photoreceptors. Variants in IDH3A and IDH3B, encoding the other subunits of IDH3, have already been associated with nonsyndromic autosomal recessive RP. Bioinformatic predictions and functional assays support a pathogenic role for the variants identified in this study, possibly through partial loss of enzymatic activity and mitochondrial function. CONCLUSION: Our findings suggest that variants in IDH3G are a novel cause of X-linked RP.

Humans

Protein phosphorylation and neuronal function.

Following the initial demonstration of phosphorylation of endogenous brain proteins (Johnson et al., 1971), two decades of work have shown that this biochemical mechanism represents one of the most important means by which extracellular signals are transduced into changes in neuronal functions. Evidence discussed in this review shows that neural cells contain a plethora of protein kinases, protein phosphatases, and phosphorylated proteins and that many of these systems appear essential for the regulation of cell functions as diverse as membrane excitability, neuronal secretory processes, cytoskeletal organization, neuronal morphology, and cellular metabolism. Moreover, there exists intricate functional relationships between many of the neuronal protein phosphorylation systems, which allow "cross-talk" between distinct signals to take place in various brain cells. The properties of protein phosphorylation systems allow these regulatory systems to influence events taking place on a microsecond scale (e.g., neurotransmitter release) and events lasting for hours and days (e.g., LTP). Our present knowledge concerning neuronal protein phosphorylation has also allowed studies to be initiated regarding the possible involvement of protein phosphorylation in various clinical disorders affecting signal transduction and brain function. It seems safe to predict that continued studies of neuronal protein phosphorylation systems will continue to improve our understanding of the anatomical, physiological, and pharmacological basis for nervous system function in both health and disease.

Animals

The helix-hairpin-helix DNA-binding motif: a structural basis for non-sequence-specific recognition of DNA.

One, two or four copies of the 'helix-hairpin-helix' (HhH) DNA-binding motif are predicted to occur in 14 homologous families of proteins. The predicted DNA-binding function of this motif is shown to be consistent with the crystallographic structure of rat polymerase beta, complexed with DNA template-primer [Pelletier, H., Sawaya, M.R., Kumar, A., Wilson, S.H. and Kraut, J. (1994) Science 264, 1891-1903] and with biochemical data. Five crystal structures of predicted HhH motifs are currently known: two from rat pol beta and one each in endonuclease III, AlkA and the 5' nuclease domain of Taq pol I. These motifs are more structurally similar to each other than to any other structure in current databases, including helix-turn-helix motifs. The clustering of the five HhH structures separately from other bi-helical structures in searches indicates that all members of the 14 families of proteins described herein possess similar HhH structures. By analogy with the rat pol beta structure, it is suggested that each of these HhH motifs bind DNA in a non-sequence-specific manner, via the formation of hydrogen bonds between protein backbone nitrogens and DNA phosphate groups. This type of interaction contrasts with the sequence-specific interactions of other motifs, including helix-turn-helix structures. Additional evidence is provided that alphaherpesvirus virion host shutoff proteins are members of the polymerase I 5'-nuclease and FEN1-like endonuclease gene family, and that a novel HhH-containing DNA-binding domain occurs in the kinesin-like molecule nod, and in other proteins such as cnjB, emb-5 and SPT6.

Amino Acid Sequence

p34cdc2 homologue is located in nucleoli of the nervous and endocrine systems.

p34cdc2 protein kinase is a component of M phase-promoting factor (MPF), which plays an important role in controlling the mitotic and meiotic cell cycle. p34cdc2 contains a unique 16 amino acid sequence (PSTAIR) that is conserved from fission yeast to human. Using polyclonal anti-PSTAIR antibody, we detected the p34cdc2 homologue in the central nervous system of adult mice by western blotting. By immunohistochemical technique, we found that the p34cdc2 homologue was located in the nucleoli of neurons and glia in the central and peripheral nervous systems. In the central nervous system, positive cells were widely distributed from the cerebral cortex to the spinal cord. Immunoreactive cells were also detected in retina and pituitary. The evidence that the p34cdc2 is present in neurons which have lost the ability of cell division predicts another function of p34cdc2 family proteins besides the one that has generally recognized.

Amino Acid Sequence

cDNA sequence, genomic organization, and evolutionary conservation of a novel gene from the WAGR region.

A new gene (239FB) with predominant and differential expression in fetal brain has recently been isolated from a chromosome 11p13-p14 boundary area near FSHB. The corresponding mRNA has an open reading frame of 294 amino acids, a 3' untranslated region of 1247 nucleotides, and a highly GC-rich 5' untranslated region. The coding and 3' UT sequence is specified by 6 exons within nearly 87 kb of isolated genomic locus. The 5' end region of the transcript maps adjacent to the only genomically defined CpG island in a chromosomal subregion that may be associated with part of the mental retardation of some WAGR (Wilms tumor, aniridia, genitourinary anomalies, and mental retardation) syndrome patients. In addition to nucleotide and amino acid similarity to an EST from a normalized infant brain cDNA library, the predicted protein has extensive similarity to two Caenorhabditis elegans polypeptides of, as yet, unknown function. The 239FB locus is, therefore, likely part of a family of genes with two members expressed in human brain. The extensive conservation of the predicted protein suggests a fundamental function of the gene product and will enable evaluation of the role of the 239FB gene in neurogenesis in model organisms.

Amino Acid Sequence

Identification of functional domains in the plasma apolipoproteins by analysis of inter-species sequence variability.

Molecular evolution theory posits that sequence motifs essential for protein function are constrained by selective pressure from changing over long stretches of evolutionary time. Thus, analysis of inter-species amino acid sequence variability, by identifying highly conserved intervals, should predict the location of domains critical for protein function. We have analyzed the amino acid sequences of the mammalian apolipoproteins A-I, A-IV, C-I, C-II, C-III, D, and E with a computer algorithm that calculates numerical residue variability scores. The application of a median sieve filter to the data facilitated identification of the exact boundaries of highly conserved domains, which coincided with the location of known structural features and functional domains in this family of proteins. The analysis also identified highly conserved intervals in every apolipoprotein whose function is unknown at present, but which are candidates for regions with specific functional roles.

Amino Acid Sequence

Regulation of the mitochondrial ATP synthase/ATPase complex: cDNA cloning, sequence, overexpression, and secondary structural characterization of a functional protein inhibitor.

The ATPase inhibitor protein of the rat liver mitochondrial ATP synthase/ATPase complex has been cloned from a rat liver cDNA library, and its nucleotide sequence determined. The sequence is highly homologous to both the bovine heart (approximately 70%) and the yeast inhibitor proteins (approximately 40%). The deduced protein sequence is 107 amino acids in length, and based on homology to the bovine heart protein, the first 25 N-terminal amino acids encode a putative mitochondrial targeting sequence. The "mature" protein (without the targeting sequence) fused to the maltose binding protein has been overexpressed in Escherichia coli. The maltose binding protein was used as a handle for the development of a rapid one-step purification of the fusion protein by affinity chromatography on an amylose resin. The purified fusion protein was cleaved with Factor Xa protease at the fusion junction, and the resulting ATPase inhibitor protein was purified to > 90% purity. The purified, overexpressed inhibitor protein displays normal inhibitor activity. The protein inhibits ATP hydrolysis catalyzed by the ATP synthase/ATPase complex in submitochondrial particles in a manner kinetically indistinguishable from the same protein purified from rat liver mitochondria, and exhibits a specific activity of approximately 10,000 units/mg. The secondary structure of the inhibitor protein was determined by circular dichroism spectropolarimetry. The experimentally determined structure shows a high content of alpha-helix and is in good agreement with sequence-based structural predictions. As the function of the inhibitor protein is known to exhibit a high dependence on pH, a study of the pH dependence of inhibitor secondary structure was performed. It is shown that as pH is lowered, conditions which activate inhibitory capacity, the protein loses significant alpha-helical structure. This is the first report of the overexpression in E. coli of a functional ATPase inhibitor protein. Secondary structural analysis of this protein indicates that conversion from its active to its inactive form involves a significant conformational change.

Adenosine Triphosphatases

Gonococcal transferrin-binding protein 1 is required for transferrin utilization and is homologous to TonB-dependent outer membrane receptors.

The pathogenic Neisseria species are capable of utilizing transferrin as their sole source of iron. A neisserial transferrin receptor has been identified and its characteristics defined; however, the biochemical identities of proteins which are required for transferrin receptor function have not yet been determined. We identified two iron-repressible transferrin-binding proteins in Neisseria gonorrhoeae, TBP1 and TBP2. Two approaches were taken to clone genes required for gonococcal transferrin receptor function. First, polyclonal antiserum raised against TBP1 was used to identify clones expressing TBP1 epitopes. Second, a wild-type gene copy was cloned that repaired the defect in a transferrin receptor function (trf) mutant. The clones obtained by these two approaches were shown to overlap by DNA sequencing. Transposon mutagenesis of both clones and recombination of mutagenized fragments into the gonococcal chromosome generated mutants that showed reduced binding of transferrin to whole cells and that were incapable of growth on transferrin. No TBP1 was produced in these mutants, but TBP2 expression was normal. The DNA sequence of the gene encoding gonococcal TBP1 (tbpA) predicted a protein sequence homologous to the Escherichia coli and Pseudomonas putida TonB-dependent outer membrane receptors. Thus, both the function and the predicted protein sequence of TBP1 were consistent with this protein serving as a transferrin receptor.

Amino Acid Sequence

A yeast mitogen-activated protein kinase homolog (Mpk1p) mediates signalling by protein kinase C.

Mitogen-activated protein (MAP) kinases are activated in response to a variety of stimuli through a protein kinase cascade that results in their phosphorylation on tyrosine and threonine residues. The molecular nature of this cascade is just beginning to emerge. Here we report the isolation of a Saccharomyces cerevisiae gene encoding a functional analog of mammalian MAP kinases, designated MPK1 (for MAP kinase). The MPK1 gene was isolated as a dosage-dependent suppressor of the cell lysis defect associated with deletion of the BCK1 gene. The BCK1 gene is also predicted to encode a protein kinase which has been proposed to function downstream of the protein kinase C isozyme encoded by PKC1. The MPK1 gene possesses a 1.5-kb uninterrupted open reading frame predicted to encode a 53-kDa protein. The predicted Mpk1 protein (Mpk1p) shares 48 to 50% sequence identity with Xenopus MAP kinase and with the yeast mating pheromone response pathway components, Fus3p and Kss1p. Deletion of MPK1 resulted in a temperature-dependent cell lysis defect that was virtually indistinguishable from that resulting from deletion of BCK1, suggesting that the protein kinases encoded by these genes function in a common pathway. Expression of Xenopus MAP kinase suppressed the defect associated with loss of MPK1 but not the mating-related defects associated with loss of FUS3 or KSS1, indicating functional conservation between the former two protein kinases. Mutation of the presumptive phosphorylated tyrosine and threonine residues of Mpk1p individually to phenylalanine and alanine, respectively, severely impaired Mpk1p function. Additional epistasis experiments, and the overall architectural similarity between the PKC1-mediated pathway and the pheromone response pathway, suggest that Pkc1p regulates a protein kinase cascade in which Bck1p activates a pair of protein kinases, designated Mkk1p and Mkk2p (for MAP kinase-kinase), which in turn activate Mpk1p.

Amino Acid Sequence

Resonant recognition model and protein topography. Model studies with myoglobin, hemoglobin and lysozyme.

This study describes the further extension of the resonant recognition model for the analysis and prediction of protein--protein and protein--DNA structure/function dependencies. The model is based on the significant correlation between spectra of numerical presentations of the amino acid or nucleotide sequences of proteins and their coded biological activity. According to this physico-mathematical method, it is possible to define amino acids in the sequence which are predicted to be the most critical for protein function. Using sperm whale myoglobin, human hemoglobin and hen egg white lysozyme as model protein examples, sets of predicted amino acids, or so-called 'hot spots', have been identified within the tertiary structure. It was found for each protein that the predicted 'hot spots', which are distributed along the primary sequence, are spatially grouped in a dome-like arrangement over the active site. The identified amino acids did not correspond to the amino acid residues which are involved in the chemical reaction site of these proteins. It is thus proposed that the resonant recognition model helps to identify amino acid residues which are important for the creation of the molecular structure around the catalytic active site and also the associated physical field conditions required for biorecognition, docking of the specific substrate and full biological activity.

Animals

Genomic organization of the human skeletal muscle sodium channel gene.

Voltage-dependent sodium channels are essential for normal membrane excitability and contractility in adult skeletal muscle. The gene encoding the principal sodium channel alpha-subunit isoform in human skeletal muscle (SCN4A) has recently been shown to harbor point mutations in certain hereditary forms of periodic paralysis. We have carried out an analysis of the detailed structure of this gene including delineation of intron-exon boundaries by genomic DNA cloning and sequence analysis. The complete coding region of SCN4A is found in 32.5 kb of genomic DNA and consists of 24 exons (54 to > 2.2 kb) and 23 introns (97 bp-4.85 kb). The exon organization of the gene shows no relationship to the predicted functional domains of the channel protein and splice junctions interrupt many of the transmembrane segments. The genomic organization of sodium channels may have been partially conserved during evolution as evidenced by the observation that 10 of the 24 splice junctions in SCN4A are positioned in homologous locations in a putative sodium channel gene in Drosophila (para). The information presented here should be extremely useful both for further identifying sodium channel mutations and for gaining a better understanding of sodium channel evolution.

Amino Acid Sequence

Rare variant contribution to the heritability of coronary artery disease.

Whole genome sequences (WGS) enable discovery of rare variants which may contribute to missing heritability of coronary artery disease (CAD). To measure their contribution, we apply the GREML-LDMS-I approach to WGS of 4949 cases and 17,494 controls of European ancestry from the NHLBI TOPMed program. We estimate CAD heritability at 34.3% assuming a prevalence of 8.2%. Ultra-rare (minor allele frequency ≤ 0.1%) variants with low linkage disequilibrium (LD) score contribute ~50% of the heritability. We also investigate CAD heritability enrichment using a diverse set of functional annotations: i) constraint; ii) predicted protein-altering impact; iii) cis-regulatory elements from a cell-specific chromatin atlas of the human coronary; and iv) annotation principal components representing a wide range of functional processes. We observe marked enrichment of CAD heritability for most functional annotations. These results reveal the predominant role of ultra-rare variants in low LD on the heritability of CAD. Moreover, they highlight several functional processes including cell type-specific regulatory mechanisms as key drivers of CAD genetic risk.

Humans