Search PubMed⌕ Search

Biomedical subjects

Rajan Munshi

Publications and source records attributed to Rajan Munshi.

3 recordsLinked to original sources

A sequence alignment-independent method for protein classification.

Annotation of the rapidly accumulating body of sequence data relies heavily on the detection of remote homologues and functional motifs in protein families. The most popular methods rely on sequence alignment. These include programs that use a scoring matrix to compare the probability of a potential alignment with random chance and programs that use curated multiple alignments to train profile hidden Markov models (HMMs). Related approaches depend on bootstrapping multiple alignments from a single sequence. However, alignment-based programs have limitations. They make the assumption that contiguity is conserved between homologous segments, which may not be true in genetic recombination or horizontal transfer. Alignments also become ambiguous when sequence similarity drops below 40%. This has kindled interest in classification methods that do not rely on alignment. An approach to classification without alignment based on the distribution of contiguous sequences of four amino acids (4-grams) was developed. Interest in 4-grams stemmed from the observation that almost all theoretically possible 4-grams (20(4)) occur in natural sequences and the majority of 4-grams are uniformly distributed. This implies that the probability of finding identical 4-grams by random chance in unrelated sequences is low. A Bayesian probabilistic model was developed to test this hypothesis. For each protein family in Pfam-A and PIR-PSD, a feature vector called a probe was constructed from the set of 4-grams that best characterised the family. In rigorous jackknife tests, unknown sequences from Pfam-A and PIR-PSD were compared with the probes for each family. A classification result was deemed a true positive if the probe match with the highest probability was in first place in a rank-ordered list. This was achieved in 70% of cases. Analysis of false positives suggested that the precision might approach 85% if selected families were clustered into subsets. Case studies indicated that the 4-grams in common between an unknown and the best matching probe correlated with functional motifs from PRINTS. The results showed that remote homologues and functional motifs could be identified from an analysis of 4-gram patterns.

Algorithms↗

Biochemical characterization of the Staphylococcus aureus PcrA helicase and its role in plasmid rolling circle replication.

Previous genetic studies have suggested that a putative chromosome-encoded helicase, PcrA, is required for the rolling circle replication of plasmid pT181 in Staphylococcus aureus. We have overexpressed and purified the staphylococcal PcrA protein and studied its biochemical properties in vitro. Purified PcrA helicase supported the in vitro replication of plasmid pT181. It had ATPase activity that was stimulated in the presence of single-stranded DNA. Unlike many replicative helicases, PcrA was highly active as a 5' --> 3' helicase and had a weaker 3' --> 5' helicase activity. The RepC initiator protein encoded by pT181 nicks at the origin of replication and becomes covalently attached to the 5' end of the DNA. The 3' OH end at the nick then serves as a primer for displacement synthesis. PcrA helicase showed an origin-specific unwinding activity with supercoiled plasmid pT181 DNA that had been nicked at the origin by RepC. We also provide direct evidence for a protein-protein interaction between PcrA and RepC proteins. Our results are consistent with a model in which the PcrA helicase is targeted to the pT181 origin through a protein-protein interaction with RepC and facilitates the movement of the replisome by initiating unwinding from the RepC-generated nick.

Bacterial Proteins↗

An introduction to simulation and visualization of biological systems at multiple scales: a summer training program for interdisciplinary research.

Advances in biomedical research require a new generation of researchers having a strong background in both the life and physical sciences and a knowledge of computational, mathematical, and engineering tools for tackling biological problems. The NIH-NSF Bioengineering and Bioinformatics Summer Institute at the University of Pittsburgh (BBSI @ Pitt; www.ccbb.pitt.edu/bbsi) is a multi-institutional 10-week summer program hosted by the University of Pittsburgh, Duquesne University, the Pittsburgh Supercomputing Center, and Carnegie Mellon University, and is one of nine Institutes throughout the nation currently participating in the NIH-NSF program. Each BBSI focuses on a different area; the BBSI @ Pitt, entitled "Simulation and Computer Visualization of Biological Systems at Multiple Scales", focuses on computational and mathematical approaches to understanding the complex machinery of molecular-to-cellular systems at three levels, namely, molecular, subcellular (microphysiological), and cellular. We present here an overview of the BBSI @ Pitt, the objectives and focus of the program, and a description of the didactic training activities that distinguish it from other traditional summer research programs. Furthermore, we also report several challenges that have been identified in implementing such an interdisciplinary program that brings together students from diverse academic programs for a limited period of time. These challenges notwithstanding, presenting an integrative view of molecular-to-system analytical models has introduced these students to the field of computational biology and has allowed them to make an informed decision regarding their future career prospects.

Computational Biology↗