Search PubMed⌕ Search

Biomedical subjects

J R Einstein

Publications and source records attributed to J R Einstein.

12 recordsLinked to original sources

A computational method for NMR-constrained protein threading.

Protein threading provides an effective method for fold recognition and backbone structure prediction. But its application is currently limited due to its level of prediction accuracy and scope of applicability. One way to significantly improve its usefulness is through the incorporation of underconstrained (or partial) NMR data. It is well known that the NMR method for protein structure determination applies only to small proteins and that its effectiveness decreases rapidly as the protein mass increases beyond about 30 kD. We present, in this paper, a computational framework for applying underconstrained NMR data (that alone are insufficient for structure determination) as constraints in protein threading and also in all-atom model construction. In this study, we consider both secondary structure assignments from chemical shifts and NOE distance restraints. Our results have shown that both secondary structure assignments and a small number of long-range NOEs can significantly improve the threading quality in both fold recognition and threading-alignment accuracy, and can possibly extend threading's scope of applicability from homologs to analogs. An accurate backbone structure generated by NMR-constrained threading can then provide a great amount of structural information, equivalent to that provided by many NMR data; and hence can help reduce the number of NMR data typically required for an accurate structure determination. This new technique can potentially accelerate current NMR structure determination processes and possibly expand NMR's capability to larger proteins.

Algorithms↗

Detection of RNA polymerase II promoters and polyadenylation sites in human DNA sequence.

Detection of RNA polymerase II promoters and polyadenylation sites helps to locate gene boundaries and can enhance accurate gene recognition and modeling in genomic DNA sequence. We describe a system which can be used to detect polyadenylation sites and thus delineate the 3' boundary of a gene, and discuss improvements to a system first described in Matis et al. (1995) [Matis S., Shah M., Mural R. J. & Uberbacher E.C. (1995) Proc. First Wrld Conf. Computat. Med., Public Hlth, Biotechnol. (Wrld Sci.) (in press).], which predicts a large subset of RNA polymerase II promoters. The promoter system used statistical matrices and distance information as inputs for a neural network which was trained to provide initial promoter recognition. The output of the network was further refined by applying rules which use the gene context information predicted by GRAIL. We have reconstructed the rule-based system which uses gene context information and significantly improved the sensitivity and selectivity of promoter detection.

Algorithms↗

An improved system for exon recognition and gene modeling in human DNA sequences.

A new version of the GRAIL system (Uberbacher and Mural, 1991; Mural et al., 1992; Uberbacher et al., 1993), called GRAIL II, has recently been developed (Xu et al., 1994). GRAIL II is a hybrid AI system that supports a number of DNA sequence analysis tools including protein-coding region recognition, PolyA site and transcription promoter recognition, gene model construction, translation to protein, and DNA/protein database searching capabilities. This paper presents the core of GRAIL II, the coding exon recognition and gene model construction algorithms. The exon recognition algorithm recognizes coding exons by combining coding feature analysis and edge signal (acceptor/donor/translation-start sites) detection. Unlike the original GRAIL system (Uberbacher and Mural, 1991; Mural et al., 1992), this algorithm uses variable-length windows tailored to each potential exon candidate, making its performance almost exon length-independent. In this algorithm, the recognition process is divided into four steps. Initially a large number of possible coding exon candidates are generated. Then a rule-based prescreening algorithm eliminates the majority of the improbable candidates. As the kernel of the recognition algorithm, three neural networks are trained to evaluate the remaining candidates. The outputs of the neural networks are then divided into clusters of candidates, corresponding to presumed exons. The algorithm makes its final prediction by picking the best canadidate from each cluster. The gene construction algorithm (Xu, Mural and Uberbacher, 1994) uses a dynamic programming approach to build gene models by using as input the clusters predicted by the exon recognition algorithm. Extensive testing has been done on these two algorithms.(ABSTRACT TRUNCATED AT 250 WORDS)

Algorithms↗

Preliminary crystallographic data for Bowman-Birk inhibitor from soybean seeds.

A well characterized soybean protease inhibitor, the Bowman-Birk inhibitor, has been crystallized at room temperature in the presence of polyethylene glycol 4000 by vapor diffusion against an ammonium sulfate solution containing 2-methyl-2,4-pentanediol. An x-ray diffraction study reveals that the inhibitor crystallizes in a hexagonal unit cell of symmetry P6122 (or P6522) and dimensions a = b = 91.36(2) A and c = 63.92(2) A. Each of the 12 asymmetric units contains 2 molecules of molecular weight 8000. The crystal, which diffracts barely to 3-A spacings, is fairly stable to x-irradiation and has a solvent content of approximately 52% by volume.

Crystallization↗

Purification, properties, and crystallographic data for a principal nontoxic lectin from seeds of Abrus precatorius.

Nontoxic abrus lectin has been prepared by a new purification procedure. The method is accomplished by 45% saturation ammonium sulfate fractionation from a 5% acetic acid extract of the seeds of Abrus precatorius followed by diethylaminoethyl-Sephadex A-50 and Sepharose 4B affinity chromatography. The abrus lectin appeared homogeneous as judged by electrophoresis, analytical ultracentrifugation, and isoelectric focusing. The lectin molecule has a weight of 126,000 as determined by sedimentation equilibrium. It is composed of four subunits, of which two pairs have either identical or closely similar molecular sizes (33,800 and 32,200 daltons) as revealed by polyacrylamide gel electrophoresis in the presence of sodium dodecyl sulfate. The results of amino acid analyses are given; none of the cysteic acid appears to arise from cysteine. An electrofocusing experiment indicated the isoelectric point to be 5.0. Crystals large enough for x-ray investigation were obtained by a vapor diffusion technique. X-ray precession photographs revealed that abrus lectin crystallizes in a tetragonal unit cell of symmetry P41212 and dimensions a = 140 and c = 210 A. The asymmetric unit contains 2 protein molecules of molecular weight 126,000 and has a solvent content of approximately 41% by volume.

Amino Acids↗

An artificial intelligence approach to DNA sequence feature recognition.

The ultimate goal of the Human Genome project is to extract the biologically relevant information recorded in the estimated 100,000 genes encoded by the 3 x 10(9) bases of the human genome. This necessitates development of reliable computer-based methods capable of analysing and correctly identifying genes in the vast amounts of DNA-sequence data generated. Such tools may save time and labour by simplifying, for example, screening of cDNA libraries. They may also facilitate the localization of human disease genes by identifying candidate genes in promising regions of anonymous DNA sequence.

Artificial Intelligence↗