Getting the message: identifying transcribed sequences.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to R J Mural.
Explore the source record for details and available documents.
This paper presents an algorithm for detecting and 'correcting' sequencing errors that occur in DNA coding regions. The types of sequencing errors addressed are insertions and deletions (indels) of DNA bases. The goal is to provide a capability which makes single-pass or low-redundancy sequence data more informative, reducing the need for high-redundancy sequencing for gene identification and characterization purposes. This would permit improved sequencing efficiency and reduce genome sequencing costs. The algorithm detects sequencing errors by discovering changes in the statistically preferred reading frame within a putative coding region and then inserts a number of 'neutral' bases at a perceived reading frame transition point to make the putative exon candidate frame consistent. We have implemented the algorithm as a front-end subsystem of the GRAIL DNA sequence analysis system to construct a version which is very error tolerant and also intend to use this as a testbed for further development of sequencing error-correction technology. Preliminary test results have shown the usefulness of this algorithm and also exhibited some of its weakness, providing possible directions for further improvement. On a test set consisting of 68 human DNA sequences with 1% randomly generated indels in coding regions, the algorithm detected and corrected 76% of the indels. The average distance between the position of an indel and the predicted one was 9.4 bases. With this subsystem in place, GRAIL correctly predicted 89% of the coding messages with 10% false message on the 'corrected' sequences, compared to 69% correctly predicted coding messages and 11% falsely predicted messages on the 'corrupted' sequences using standard GRAIL II method (version 1.2).(ABSTRACT TRUNCATED AT 250 WORDS)
An important open problem in molecular biology is how to use computational methods to understand the structure and function of proteins given only their primary sequences. We describe and evaluate an original machine-learning approach to classifying protein sequences according to their structural folding class. Our work is novel in several respects: we use a set of protein classes that previously have not been used for classifying primary sequences, and we use a unique set of attributes to represent protein sequences to the learners. We evaluate our approach by measuring its ability to correctly classify proteins that were not in its training set. We compare our input representation to a commonly used input representation--amino acid composition--and show that our approach more accurately classifies proteins that have very limited homology to the sequences on which the systems are trained.
Explore the source record for details and available documents.
This paper presents a computationally efficient algorithm, the Gene Assembly Program III (GAP III), for constructing gene models from a set of accurately-predicted 'exons'. The input to the algorithm is a set of clusters of exon candidates, generated by a new version of the GRAIL coding region recognition system. The exon candidates of a cluster differ in their presumed edges and occasionally in their reading frames. Each exon candidate has a numerical score representing its 'probability' of being an actual exon. GAP III uses a dynamic programming algorithm to construct a gene model, complete or partial, by optimizing a predefined objective function. The optimal gene models constructed by GAP III correspond very well with the structures of genes which have been determined experimentally and reported in the Genome Sequence Database (GSDB). On a test set of 137 human and mouse DNA sequences consisting of 954 true exons, GAP III constructed 137 gene models using 892 exons, among which 859 (859/954 = 90%) are true exons and 33 (33/892 = 3%) are false positive. Among the 859 true positives, 635 (74%) match the actual exons exactly, and 838 (98%) have at least one edge correct. GAP III is computationally efficient. If we use E and C to represent the total number of exon candidates in all clusters and the number of clusters, respectively, the running time of GAP III is proportional to (E x C).
A new version of the GRAIL system (Uberbacher and Mural, 1991; Mural et al., 1992; Uberbacher et al., 1993), called GRAIL II, has recently been developed (Xu et al., 1994). GRAIL II is a hybrid AI system that supports a number of DNA sequence analysis tools including protein-coding region recognition, PolyA site and transcription promoter recognition, gene model construction, translation to protein, and DNA/protein database searching capabilities. This paper presents the core of GRAIL II, the coding exon recognition and gene model construction algorithms. The exon recognition algorithm recognizes coding exons by combining coding feature analysis and edge signal (acceptor/donor/translation-start sites) detection. Unlike the original GRAIL system (Uberbacher and Mural, 1991; Mural et al., 1992), this algorithm uses variable-length windows tailored to each potential exon candidate, making its performance almost exon length-independent. In this algorithm, the recognition process is divided into four steps. Initially a large number of possible coding exon candidates are generated. Then a rule-based prescreening algorithm eliminates the majority of the improbable candidates. As the kernel of the recognition algorithm, three neural networks are trained to evaluate the remaining candidates. The outputs of the neural networks are then divided into clusters of candidates, corresponding to presumed exons. The algorithm makes its final prediction by picking the best canadidate from each cluster. The gene construction algorithm (Xu, Mural and Uberbacher, 1994) uses a dynamic programming approach to build gene models by using as input the clusters predicted by the exon recognition algorithm. Extensive testing has been done on these two algorithms.(ABSTRACT TRUNCATED AT 250 WORDS)
Based on selective labeling by ATP analogues, Lys68 of the Calvin Cycle enzyme phosphoribulokinase (PRK) from spinach has been assigned to the active-site region [Miziorko et al. (1990), J. Biol. Chem. 265, 3642-3647]. The equivalent position is occupied by lysyl or arginyl residues in the PRK from both prokaryotic and eukaryotic sources, suggesting a requirement for a basic residue at this location. To examine this possibility, we have replaced Lys68 of the spinach enzyme with arginyl, glutaminyl, alanyl, or glutamyl residues by site-directed mutagenesis. All of the mutant enzymes retain substantial kinase activity; and even in the case of the radical substitution by glutamate, the Km values for ATP and ribulose 5-phosphate are not perturbed significantly. Glutamate at position-68 may destabilize tertiary structure, because the yield of this mutant protein from transformed E. coli is quite low compared to that of the other proteins in this series. Despite the active-site proximity of Lys68, our results show that this residue does not play a key role in catalysis or substrate binding.
Crystallographic studies of ribulose-1,5-bisphosphate carboxylase/oxygenase from Rhodospirillum rubrum suggest that active-site Asn111 interacts with Mg2+ and/or substrate (Lundqvist, T., and Schneider, G. (1991) J. Biol. Chem. 266, 12604-12611). To examine possible catalytic roles of Asn111, we have used site-directed mutagenesis to replace it with a glutaminyl, aspartyl, seryl, or lysyl residue. Although the mutant proteins are devoid of detectable carboxylase activity, their ability to form a quaternary complex comprised of CO2, Mg2+, and a reaction-intermediate analogue is indicative of competence in activation chemistry and substrate binding. The mutant proteins retain enolization activity, as measured by exchange of the C3 proton of ribulose bisphosphate with solvent, thereby demonstrating a preferential role of Asn111 in some later step of overall catalysis. The active sites of this homodimeric enzyme are formed by interactive domains from adjacent subunits (Larimer, F. W., Lee, E. H., Mural, R. J., Soper, T. S., and Hartman, F. C. (1987) J. Biol. Chem. 262, 15327-15329). Crystallography assigns Asn111 to the amino-terminal domain of the active site (Knight, S., Anderson, I., and Brändén, C.-I. (1990) J. Mol. Biol. 215, 113-160). The observed formation of enzymatically active heterodimers by the in vivo hybridization of an inactive position-111 mutant with inactive carboxyl-terminal domain mutants is consistent with this assignment.
Genes in higher eukaryotes may span tens or hundreds of kilobases with the protein-coding regions accounting for only a few percent of the total sequence. Identifying genes within large regions of uncharacterized DNA is a difficult undertaking and is currently the focus of many research efforts. We describe a reliable computational approach for locating protein-coding portions of genes in anonymous DNA sequence. Using a concept suggested by robotic environmental sensing, our method combines a set of sensor algorithms and a neural network to localize the coding regions. Several algorithms that report local characteristics of the DNA sequence, and therefore act as sensors, are also described. In its current configuration the "coding recognition module" identifies 90% of coding exons of length 100 bases or greater with less than one false positive coding exon indicated per five coding exons indicated. This is a significantly lower false positive rate than any method of which we are aware. This module demonstrates a method with general applicability to sequence-pattern recognition problems and is available for current research efforts.
The Calvin Cycle enzyme phosphoribulokinase is activated in higher plants by the reversible reduction of a disulfide bond, which is located at the active site. To determine the possible contribution of the two regulatory residues (Cys16 and Cys55) to catalysis, site-directed mutagenesis has been used to replace each of them in the spinach enzyme with serine or alanine. The only other cysteinyl residues of the kinase, Cys244 and Cys250, were also replaced individually by serine or alanine. A comparison of specific activities of native and mutant enzymes reveals that substitutions at positions 244 or 250 are inconsequential. The position 16 mutants retain 45-90% of the wild-type activity and display normal Km values for both ATP and ribulose 5-phosphate. In contrast, substitution at position 55 results in 85-95% loss of wild-type activity, with less than a 2-fold increase in the Km for ATP and a 4-8-fold increase in the Km for ribulose 5-phosphate. These results are consistent with moderate facilitation of catalysis by Cys55 and demonstrate that the other three cysteinyl residues do not contribute significantly either to structure or catalysis. The enhanced stability, relative to wild-type enzyme, of the Ser16 mutant protein to a sulfhydryl reagent supports earlier suggestions that Cys16 is the initial target of the oxidative deactivation process.
Explore the source record for details and available documents.
The active site of ribulose-bisphosphate carboxylase/oxygenase is constituted from domains of adjacent subunits and includes an intersubunit electrostatic interaction between Lys 168 and Glu48, which has been recently identified by x-ray crystallography (Andersson, I., Knight, S., Schneider, G., Lindqvist, Y., Lundqvist, T., Brändén, C.-I., and Lorimer, G.H. (1989) Nature 337, 229-234; Lundqvist, T., and Schneider, G. (1989) J. Biol. Chem. 264, 7078-7083). To examine the structural and functional requirements for this interaction, we have used site-directed mutagenesis to replace Lys168 of the homodimeric enzyme from Rhodospirillum rubrum with arginine, glutamine, or glutamic acid. All three substitutions result in mutant enzymes with less than or equal to 0.1% of wild-type activity. The nonconservative substitution of Lys168 with a glutamyl residue precludes the formation of a stable dimer, explaining the consequential abolition of enzymic activity. Both the Arg168 and Gln168 mutant proteins are isolated as stable dimers, even though the latter obviously lacks an electrostatic interaction present in the wild-type enzyme. Despite the absence of overall carboxylase activity, these two mutant proteins serve as catalysts for the enolization of ribulose bisphosphate, as measured by exchange of the C3 proton with solvent. These observations, as well as ligand-binding properties of the mutant proteins, are consistent with Lys168 facilitating a catalytic step subsequent to enolization.
Explore the source record for details and available documents.
The two active sites of homodimeric ribulose bisphosphate carboxylase/oxygenase from Rhodospirillum rubrum are constituted by interacting domains of adjacent subunits, in which residues from each are required for catalytic activity. Active-site residues include Lys-166 of one domain and Glu-48 of the interacting domain from the adjacent subunit. Whereas all substitutions for Lys-166, introduced by site-directed mutagenesis, abolished catalytic activity, only a negatively charged residue (e.g., aspartic acid) resulted in the disruption of the subunit interactions (Lee et al., 1987). This disruption could result from improper folding of the individual polypeptide chains or to more localized effects (e.g., charge-charge repulsion due to proximal negative charges of Asp-166 and Glu-48 of adjacent domains or conformational changes restricted to a single domain). To address these questions, we have examined the ability of the Asp-166 mutant subunit to associate with a mutant subunit in which the negatively charged Glu-48 has been replaced by the neutral glutaminyl residue. Coexpression in Escherichia coli of the genes for both mutant subunits results in formation of a catalytically active hybrid, despite the absence of activity when either gene is expressed individually. Isolation and characterization of the hybrid show that it is composed of one Asp-166 subunit and one Gln-48 subunit, presumably with only one functional active site per dimeric molecule. This association of dissimilar subunits shows that introduction of a negative charge at position 166 does not lead to overall distortion of subunit conformation. In contrast to the wild-type enzyme, the hybrid dissociates spontaneously at low protein concentration but is stabilized by elevated ionic strengths or by glycerol.
Oligodeoxynucleotide-mediated mutagenesis of the ada gene of Escherichia coli was used to produce two mutant Ada proteins. In mutant I the methyl acceptor Cys-321 for O6-methylguanine was replaced by histidine; and in mutant II the positions of Cys-321 and His-322 of the wild-type protein were inverted. Neither mutant protein had O6-methylguanine-DNA methyltransferase activity, but both retained the phosphotriester-DNA methyltransferase activity involving methyl group transfer to Cys-69. Under the control of the endogenous promoter, synthesis of mutant I protein was undetectable before or after adaptation treatment with promoter, synthesis of mutant I protein was undetectable before or after adaptation treatment with N-methyl-N'-nitro-N-nitrosoguanidine. This appeared to be due to both inhibition of transcription of the mutant gene and degradation of the synthesized protein. On the other hand, mutant II protein was inducible by N-methyl-N'-nitro-N-nitrosoguanidine, although to a smaller extent than the wild-type protein was, and the phosphotriester-DNA methyltransferase activity appeared to reside in 24- to 30-kilodalton cleavage products. Mutant I protein could be produced under lac promoter control, and its cleavage products, unlike those of mutant II protein, tended to aggregate. These results indicate that (i) Cys-321 cannot be replaced or transposed with the nucleophilic amino acid histidine for O6-methylguanine-DNA methyltransferase function, (ii) single amino acid replacement or transposition at the O6-methylguanine methyl acceptor site can have a profound effect on the in vivo stability and regulatory function of the Ada protein, and (iii) the integrity of the protein may not be absolutely needed for its transcription-activation function.
Phosphoribulokinase (PRK) is a key enzyme in the Calvin cycle of autotrophic organisms. We have constructed a spinach leaf cDNA library in the phage expression vector, lambda gt11, and used a rabbit polyclonal antibody raised against spinach PRK to identify PRK clones. Analyses of the nucleotide sequences of two antibody-positive clones, 1.47 and 1.35 kb in length, showed that they encode a protein which contains the N-terminal amino acid (aa) sequence [Porter et al., Arch. Biochem. Biophys. 245 (1986) 14-23] of mature spinach PRK. The codon for the N-terminal serine of the mature protein occurs 170 bp from the 5' end of the open reading frame (ORF), suggesting that PRK is synthesized with a rather long transit peptide which is removed from the mature enzyme. The ORF, ending with an amber (TAG) codon at position 1054, predicts a mature enzyme of 351 aa with a calculated Mr of 39232.
The unusual chemical properties of active-site Lys-329 of ribulose bisphosphate carboxylase/oxygenase from Rhodospirillum rubrum have suggested that this residue is required for catalysis. To test this postulate Lys-329 was replaced with glycine, serine, alanine, cysteine, arginine, glutamic acid or glutamine by site-directed mutagenesis. These single amino acid substitutions do not appear to induce major conformational changes because (i) intersubunit interactions are unperturbed in that the purified mutant proteins are stable dimers like the wild-type enzyme and (ii) intrasubunit folding is normal in that the mutant proteins bind the competitive inhibitor 6-phosphogluconate with an affinity similar to that of wild-type enzyme. In contrast, all of the mutant proteins are severely deficient in carboxylase activity (less than 0.01% of wild-type) and are unable to form the exchange-inert complex, characteristic of the wild-type enzyme, with the transition-state analogue carboxyarabinitol bisphosphate. These results underscore the stringency of the requirement for a lysyl side-chain at position 329 and imply that Lys-329 is involved in catalysis, perhaps stabilizing a transition state in the overall reaction pathway.
Ribulose bisphosphate carboxylase/oxygenase from Rhodospirillum rubrum is a homodimer of 50.5-kDa subunits with two substrate binding sites per molecule of dimer. To determine whether each subunit contains an independent active site or whether the active sites are created by intersubunit interactions, we have used a novel in vivo approach for producing heterodimers from catalytically inactive, site-directed mutants of the carboxylase. When the alleles encoding these mutant proteins are placed separately into compatible plasmids and coexpressed in the same Escherichia coli host, activity is observed at about 20% of the wild-type level. Analysis of the carboxylase purified from these cells reveals the presence of heterodimers of the two mutant proteins. This interallelic complementation demonstrates that domains from each of the subunits interact to form a shared active site.