Search PubMed⌕ Search

Biomedical subjects

A S Kolaskar

Publications and source records attributed to A S Kolaskar.

At least 19 recordsLinked to original sources

SEGE: A database on 'intron less/single exonic' genes from eukaryotes.

UNLABELLED: Eukaryotes have both 'intron containing' and 'intron less' genes. Several databases are available for 'intron containing' genes in eukaryotes. In this note, we describe a database for 'intron less' genes from eukaryotes. 'Intron less' eukaryotic genes having prokaryotic architecture will help to understand gene evolution in a much simpler way unlike 'intron containing' genes. AVAILABILITY: SEGE is available at http://intron.bic.nus.edu.sg/seg/ CONTACT: mmeena@ntu.edu.sg

Animals↗

Online identification of viruses.

A computerized animal virus information system is developed in the Sequence Retrieval System (SRS) format. This database is available on the Word Wide Web (WWW) at the site http://bioinfo.ernet.in/www/avis/avis++ +.html. The database has been used to generate large number of identification matrices for each family. The software is developed in C. Unix shell scripts and Hypertext Marked-up Language (HTML) to assign the family to an unknown virus deterministically and to identify the virus probabilistically. It has been shown that such web based virus identification approach provides results with high confidence in those cases where identification matrix uses large number of independent characters. Protein sequence data for animal viruses have been analyzed and oligopeptides specific to each virus family and also specific to each virus species are identified for several viruses. These peptides thus could be used to identify the virus and to assign the virus family with high confidence showing the usefulness of sequence data in virus identification.

Databases as Topic↗

Prediction of three-dimensional structure and mapping of conformational epitopes of envelope glycoprotein of Japanese encephalitis virus.

Japanese encephalitis virus (JEV), a mosquito-borne flavivirus, is an important human pathogen. The envelope glycoprotein (Egp), a major structural antigen, is responsible for viral haemagglutination and eliciting neutralising antibodies. The three-dimensional structure of the Egp of JEV was predicted using the knowledge-based homology modeling approach and X-ray structure data of the Egp of tick-borne encephalitis virus as a template (Rey et al., 1995). In the initial stages of optimisation, a distance-dependent dielectric constant of 4r(ij) was used to simulate the solvent effect. The predicted structure was refined by solvating the protein in a 10-A layer of water by explicitly considering 4867 water molecules. Four independent structure evaluation methods report this structure to be acceptable stereochemically and geometrically. The Egp of JEV has an extended structure with seven beta-sheets, two alpha-helices, and three domains. The water-solvated structure was used to delineate conformational and sequential epitopes. These results document the importance of tertiary structure in understanding the antigenic properties of flaviviruses in general and JEV in particular. The conformational epitope prediction method could be used to identify conformational epitopes on any protein antigen with known three-dimensional structure. This is one of the largest proteins whose three-dimensional structure has been predicted using an homology modeling approach and water as a solvent.

Amino Acid Sequence↗

Determination of primary amino acid sequence and unique three-dimensional structure of WGH1, a monoclonal human IgM antibody with anti-PR3 specificity.

Transformed B cells making monoclonal IgM-lambda anti-PR3 antibody WGH1 from a patient with Wegener's granulomatosis were used to prepare mRNA and synthesize cDNA. PCR primers for human micro and lambda chains were then employed to amplify heavy- and light-chain V-regions followed by cloning into pCR2-1 vector and sequencing. Molecular modeling of VH regions employed knowledge-based homology modeling to obtain minimum energy conformation. The VH sequence was subgroup III with marked overall homology to VH1.9III. The VHCDR3 region of WGH1 was unique, consisting of 21 amino acid residues which included seven tyrosines as well as three negatively charged aspartic acid residues. The VL region was subgroup II with a negatively charged glutamic acid at position 100 in CDR3. Molecular modeling of VH revealed a major conformational difference in the shape of CDR3 compared with other antibodies for which three-dimensional structures have been determined. Monoclonal antibody WGH1 reacting with PR3 (a highly positively charged molecule) shows a unique reactive cassette within VHCDR3 with a number of negatively charged aspartic acid residues. WGH1 VHCDR3 contains a loop which shows a major projection not usually recorded in other previously studied antibody molecules.

Amino Acid Sequence↗

Computer-aided virus identification on the World Wide Web.

An attempt has been made to devise computer software that will aid virologists to identify unknown virus isolates using the World Wide Web. Computerized information from the Animal Virus Information System was used to obtain data on various characters of a virus species. Sequence data banks are used to obtain the molecular data. A probabilistic method of virus identification based on Willcox's implementation of Bayes' theorem is implemented. The program provides hints to the users to carry out additional tests required to obtain higher confidence in identification of virus species. Signature peptides of the virus can also be used to confirm identification. The software is implemented on a UNIX machine and is written in C, UNIX shell scripts and HTML to run on the World Wide Web. This is the first species identification software that allows the user to carry out identification online through Internet.

Computer Communication Networks↗

Molecular dynamics simulation of a 13-mer duplex DNA: a PvuII substrate.

Parallel version of AMBER 4.1 was ported and optimised on the Indian parallel supercomputer PARAM OpenFrame built around Sun Ultra Sparc processors. This version of AMBER program was then used to carry out molecular dynamics (MD) simulations on 5'-TGACCAGCTGGTC-3', a substrate for PvuII enzyme. MD simulations in water are carried out under following conditions: (i) unconstrained at 300 K (230 ps); (ii) unconstrained at 283 K (500 ps); (iii) Watson-Crick basepair constrained at 283 K (1 ns); and (iv) Watson-Crick basepair constrained with ions at 283 K (1.2 ns). In all these simulation studies, the molecule was observed to be bending and maximum distortions in the double helix around was seen around the G7:C7' basepair, which is the phosphodiester bond that is cleaved by PvuII. Analysis of MD simulation with ions carried out for 1.2 ns also pointed out that the conformation of double helix alternates between a conformation close to B-form and close to A-form. It is argued that a bent non-standard conformation is recognised by the PvuII enzyme. The maximum bend occurs at the G7:C7' region, weakening the phosphodiester bond and allows His48 to get placed in such a fashion to permit the scission through a general base mechanism. The bending and distortion observed is a property of the sequence which acts as a substrate for PvuII enzyme. This is confirmed by carrying out MD studies on the Dickerson's sequence d(CGCGAATTCGCG)2 as a reference molecule, which practically does not bend or get deformed.

Base Composition↗

Antigenic determinants reacting with rheumatoid factor: epitopes with different primary sequences share similar conformation.

Polyclonal or monoclonal human IgM rheumatoid factors (RF) react with eight antigenic sites on the CH3 IgG domain, four sites on CH2 and two on human beta 2-microglobulin. All 14 of these RF-reactive epitopes are linear 7-11 amino acid peptides with different primary sequence. We questioned whether RF reactivity with such a variety of epitopes showing no obvious sequence homology might result from conformational similarities shared by various RF-reactive regions. Strong support for this concept was obtained using rabbit antisera as well as mouse mAbs to individual CH3, CH2 or beta 2m RF-reactive peptides. Major cross-reactivity was demonstrated between most of the 14 different CH3, CH2, or beta 2m RF-reactive peptides using individual anti-epitope antibodies. Molecular modelling studies of these peptides showed striking similarities in three-dimensional shape among many RF-reactive peptides. Main-chain atoms rather than side chains seemed to contribute most directly to conformational similarity. Molecular simulation studies on control peptides showed no conformational similarities with RF-reactive peptides. Our studies indicate that autoantibodies such as RF recognize main-chain conformations of reactive epitopes and react with a number of antigenic determinants of quite different primary sequence but similar main chain conformations.

Amino Acid Sequence↗

Contextual constraints in the choice of synonymous codons.

From EMBL Nucleotide Sequence Database, protein coding sequences of all E. coli and its DNA phages, were extracted using our computer programme. Same programme has been used to form a database of sequence of oligonucleotides of length 18 nucleotides on both sides of each of the 61 codons. From analysis of this database and study of variations in twist parameter (Tw) values, as an indicator of sequence dependent variations in B-DNA helix, a method is developed to fix the codon among the set of synonymous codons. The accuracy of the method was checked on enlarged data set by adding data from more prokaryotes. Our method assign the codon 85-90% times correctly if the selection has to be made between codons having different sequence in terms of R and Y. The accuracy of the method is somewhat lower when choice of the codon has to be made between codons having same codes in terms of R and Y. This study points out that the major factors which decide the choice of a codon from a set of synonymous codons are contextual constraints arising from flanking regions.

Base Sequence↗

Multiple alignment of sequences on parallel computers.

A software package that allows one to carry out multiple alignment of protein and nucleic acid sequences of almost unlimited length and number of sequences is developed on C-DAC parallel computer--a transputer-based machine. The farming approach is used for data parallelization. The speed gains are almost linear when the number of transputers is increased from 4 to 64. The software is used to carry out multiple alignment of 100 sequences each of alpha-chain and beta-chain of hemoglobin and 83 cytochrome c sequences. The signature sequence of cytochrome c was found to be PGTKMXF. The single parameter, multiple alignment score, S, has been used to categorize proteins in different subfamilies and groups.

Algorithms↗

Analysis of computer-predicted antibody inducing epitope on Japanese encephalitis virus.

Theoretical methods to delineate antibody inducing epitopes have been employed to predict antigenic determinants on envelope glycoprotein (gpE) of Japanese encephalitis (JE), West Nile (WN) and Dengue (DEN) I-IV viruses. A predicted region on JE virus gpE 74CPTTGEAHNEKRAD87 was synthesized, conjugated to KLH (KLH-peptide) and used in immunization of mice. A mouse monoclonal antibody (MoAb IVB4) reactive to the peptide was also found to react with native JE virus gpE. Characterization of the idiotypic (ID) determinants with the help of polyclonal domain-specific anti-ID antibodies revealed that polyclonal anti-KLH-peptide antibodies and MoAb IVB4 are flavivirus-cross-reactive to Hx and NHx domains, respectively. The region 74-87 in JE virus gpE has been mapped as a linking area between Hx and NHx domains. Reactivity of the peptide with sera from JE patients and vaccinees also indicated the feasibility of using predicted peptides for diagnostic and prophylastic purposes.

Amino Acid Sequence↗

Sequence alignment approach to pick up conformationally similar protein fragments.

Crystal structure data of globular proteins were used to prepare (phi, psi) probability maps of 20 proteinous amino acids. These maps were compared grid-wise with each other and a conformational similarity index was calculated for each pair of amino acids. A weight matrix, called Conformational Similarity Weight (CSW) matrix, was prepared using the conformational similarity index. This weight matrix was used to align sequences of 21 pairs of proteins whose crystal structures are known. The aligned regions with more than seven contiguous amino acids were further analysed by plotting average weight (W) values of overlapping hepatapeptides in these regions and carrying out curve fitting by Fourier series having TEN harmonics. The protein fragments corresponding to the half-linewidth of peaks were predicted as fragments having similar conformation in the protein pair under consideration. Such an approach allows us to pick up conformationally similar protein fragments with more than 67% accuracy.

Amino Acids↗

Computerization of virus data and its usefulness in virus classification.

Data on 537 Arboviruses and 180 other viruses have been collected and coded in two different formats. These data include information not only regarding the taxonomy and history of isolation, but also regarding the properties of biomacromolecules, proteins and nucleic acids. Information on antigenic relationships, histopathology and experimental viremia is also included. This information is stored in formats which allow the manipulation and analysis of data by dBASE III PLUS and MICRO-IS. A set of programs was written for interconversion and editing purposes. Transmission electron micrographs are scanned and stored. This stored information can be used in viral classification as shown by carrying out analysis of data on the Bunyaviridae family.

Arboviruses↗

Analysis of inverted repeats in primary structure of proteins.

A computer program has been developed to locate exact inverted repeating subsequences present anywhere in the given primary structure of proteins or nucleic acids. The output is amenable to protein sequence/nucleic acid query (PSQ/NAQ) packages. Our analysis has shown that there is a large number of proteins which have inverted repeats of more than four amino acid residues in length. However, the number is small when conditions such as the existence of more than 20 inverted repeats in given sequence or the existence of inverted repeats having more than five different types of amino acids are applied.

Amino Acid Sequence↗

A protein secondary structure database (PSS).

A protein secondary structure database (PSS) has been designed to correlate the Protein Sequence Database of the PIR-International with the atomic coordinates and bond connectivities database of the Protein Data Bank in the Brookhaven National Laboratory. The present database includes secondary structures determined by X-ray diffraction analysis, but not predicted structures. The database currently contains data from both the Protein Sequence Database and the Protein Data Bank Database, and will encompass the NMR database in the future. The main characteristics of the database are as follows: (1) the secondary structures, sites, regions and domains of structural interest are displayed together with protein primary structures; and (2) the secondary structure of a desired length of peptide fragment is displayed upon request, as are the peptide fragment(s) that correspond to a defined secondary structure. This database also has software to indicate amino acid pairs having hydrogen bonds and to count the occurrence frequency of each pair as well as the conformational parameters widely used in semi-empirical methods of secondary structure prediction.

Amino Acid Sequence↗

A semi-empirical method for prediction of antigenic determinants on protein antigens.

Analysis of data from experimentally determined antigenic sites on proteins has revealed that the hydrophobic residues Cys, Leu and Val, if they occur on the surface of a protein, are more likely to be a part of antigenic sites. A semi-empirical method which makes use of physicochemical properties of amino acid residues and their frequencies of occurrence in experimentally known segmental epitopes was developed to predict antigenic determinants on proteins. Application of this method to a large number of proteins has shown that our method can predict antigenic determinants with about 75% accuracy which is better than most of the known methods. This method is based on a single parameter and thus very simple to use.

Algorithms↗

An extension of the graph theoretical approach to predict the secondary structure of large RNAs: the complex of 16S and 23S rRNAs from E. coli as a case study.

An algorithm using the graph theoretical approach to predict secondary structures of large nucleic acids is discussed. Reliability of prediction can be improved by incorporating available experimental data and sequence homology information. As a case study, this algorithm is applied to predict the secondary structure of the 16S-23S rRNA complex from E. coli. It was found that several structures of the complex can coexist. The computer program developed to predict the secondary structure of large RNAs can be run on IBM PC/AT compatible systems.

Algorithms↗

Prediction of the recognition sites on 16S and 23S rRNAs from E. coli for the formation of 16S-23S rRNA complex.

Interactions between RNA molecules have been postulated to play an important role in the assembly of ribosomes. Using the sequence analysis and the search of continuous complementary regions on 16S rRNA and 23S rRNA, the recognition sites involved in the formation of ribosome of E. coli are postulated. The number of postulated sites was narrowed down by taking available experimental data. The suggestive evidence for correct postulation is obtained from sequence comparison studies of 16S and 23S rRNAs from various species. The sites 891-899 and 1195-1203 on 16S rRNA along with the corresponding complementary sites 1904-1912 and 760-768 on 23S rRNA are predicted to be the most probable candidates for the sites of recognition between 16S and 23S rRNAs. The possibility of the involvement of the additional site 630-638 on 16S rRNA with its complementary site 2031-2039 on 23S rRNA cannot be ruled out.

Computer Simulation↗

Contextual constraints on codon pair usage: structural and biological implications.

Complementary DNA sequence data of 278 protein coding genes from prokaryotic systems have been analysed at the level of near neighbour codon pairs. Our analysis points out that constraints exist even at the level of near neighbour codon pairs. These constraints are in addition to those which arise due to relative levels of tRNA. Codon pairs, which in the data base have different occurrence values from their expected values, neither have common secondary structure nor do have better stabilization due to high base stacking. Our study points out that there are strong interaction between constituent codons in these codon pairs. These strongly interacting codon pairs, we suggest, are involved in the formation of three dimensional structural elements of cDNA/mRNA and interact with ribosome and thus modulate translation.

Base Sequence↗