Search PubMed⌕ Search

Biomedical subjects

Chandrika B-Rao

Publications and source records attributed to Chandrika B-Rao.

4 recordsLinked to original sources

A novel complexity measure for comparative analysis of protein sequences from complete genomes.

Analysis of sequence complexities of proteins is an important step in the characterization and classification of new genomes. A new measure has been proposed to compute sequence complexity in protein sequences based on linguistic complexity. The algorithm requires a single parameter, is computationally simple and provides a framework for comparative genomic analysis. Protein sequences were classified into groups of high or low complexity based on a quantitative measure termed F(c), which is proportional to the fraction of low complexity sequence present in the protein. The algorithm was tested on sequences of 196 non-homologous proteins whose crystal structures are available at </=2.0 A resolution. Protein sequences of high complexity had 'globular' structures (95% agreement), whereas those of low complexity had non-globular structures (80% agreement). Application of this measure to proteins of unknown structure/function from different genomes revealed that the sequences of high complexity constitute the majority in all genomes (about 90% in Archaea, about 93% in Eubacteria, 89% in Saccharomyces cerevisiae and 90% in Caenorhabditis elegans). Aeropyrum pernix among Archaeae and Deinococcus radiodurans among Eubacteria have the lowest fraction of high complexity proteins (75% and 80% respectively). Further, it was observed that a few bacterial pathogens (Mycobacterium tuberculosis, Pseudomonas aeruginosa) have high fraction of low complexity proteins. The program ScanCom is available from the authors as a PERL script (UNIX system).

Algorithms↗

(TG:CA)(n) repeats in human housekeeping genes.

The unravelling of human genome sequence gives a new opportunity to investigate the role of repetitive sequences in gene regulation. Among the various types of repetitive sequences, the dinucleotide (TG:CA)(n) repeats are one of the most abundant in human genome and exhibit polymorphism. Early on, it was observed that the (TG:CA)(n) repeats could modulate gene expression and has the propensity to undergo conformational transitions in in vivo conditions. Recent reports describe the role of polymorphic (TG:CA)(n) repeats in gene regulation in several genes. In this work, we have analysed the distribution of (TG:CA)(n) (n >or= 6) repeats in human 'housekeeping genes' on which recently released Gene Chip data is available. Our results indicate that (i). The number of short intragenic (TG:CA)(n) repeats is significantly higher than the number of long repeats (ii). the proportion of genes with (TG:CA)(n) repeats (n >or= 12 units) had lower mean expression levels compared to those without these repeats, (iii). the genes belonging to the functional class of 'signalling and communication' had a positive association with repeats in contrast to the genes belonging to the 'information' class that were negatively associated with repeats.

Base Composition↗

Study of the single nucleotide polymorphism (SNP) at the palindromic sequence of hypersensitive site (HS)4 of the human beta-globin locus control region (LCR) in Indian population.

LCR, a genetic regulatory element, was examined in beta-thalassemia patients who do not show any mutation in the beta-globin genes. We sequenced LCR-HS2, HS3, and HS4 in samples from 16 such patients from the Indian population and found only one SNP A-G in the inverted repeat in HS4. A significant association was observed between the G allele and occurrence of beta-thalassemia by Fisher's exact test. The AG and GG genotypes showed higher relative risk as compared to the AA genotype. We also observed linkage disequilibrium between the A/G polymorphism and the AT-rich motif of the LCR HS2 region, suggesting that the G allele could be an evolutionarily new mutation in the study population.

Globins↗

Comparative genomics using data mining tools.

We have analysed the genomes of representatives of three kingdoms of life, namely, archaea, eubacteria and eukaryota using data mining tools based on compositional analyses of the protein sequences. The representatives chosen in this analysis were Methanococcus jannaschii, Haemophilus influenzae and Saccharomyces cerevisiae. We have identified the common and different features between the three genomes in the protein evolution patterns. M. jannaschii has been seen to have a greater number of proteins with more charged amino acids whereas S. cerevisiae has been observed to have a greater number of hydrophilic proteins. Despite the differences in intrinsic compositional characteristics between the proteins from the different genomes we have also identified certain common characteristics. We have carried out exploratory Principal Component Analysis of the multivariate data on the proteins of each organism in an effort to classify the proteins into clusters. Interestingly, we found that most of the proteins in each organism cluster closely together, but there are a few 'outliers'. We focus on the outliers for the functional investigations, which may aid in revealing any unique features of the biology of the respective organisms

Archaeal Proteins↗