Search PubMed⌕ Search

Biomedical subjects

I B Rogozin

Publications and source records attributed to I B Rogozin.

At least 19 recordsLinked to original sources

Mutagenesis by transient misalignment in the human mitochondrial DNA control region.

To study spontaneous base substitutions in human mitochondrial DNA (mtDNA), we reconstructed the mutation spectra of the hypervariable segments I and II (HVS I and II) using published data on polymorphisms from various human populations. Classification analysis revealed numerous mutation hotspots in HVS I and II mutation spectra. Statistical analysis suggested that strand dislocation mutagenesis, operating in monotonous runs of nucleotides, plays an important role in generating base substitutions in the mtDNA control region. The frequency of mutations compatible with the primer strand dislocation in the HVS I region was almost twice as high as that for template strand dislocation. Frequencies of mutations compatible with the primer and template strand dislocation models are almost equal in the HVS II region. Further analysis of strand dislocation models suggested that an excess of pyrimidine transitions in mutation spectra, reconstructed on the basis of the L-strand sequence, is caused by an excess of both L-strand pyrimidine transitions and H-strand purine transitions. In general, no significant bias toward parent H-strand-specific dislocation mutagenesis was found in the HVS I and II regions.

Base Sequence↗

Cloning of human centromeres by transformation-associated recombination in yeast and generation of functional human artificial chromosomes.

Human centromeres remain poorly characterized regions of the human genome despite their importance for the maintenance of chromosomes. In part this is due to the difficulty of cloning of highly repetitive DNA fragments and distinguishing chromosome-specific clones in a genomic library. In this work we report the highly selective isolation of human centromeric DNA using transformation-associated recombination (TAR) cloning. A TAR vector with alphoid DNA monomers as targeting sequences was used to isolate large centromeric regions of human chromosomes 2, 5, 8, 11, 15, 19, 21 and 22 from human cells as well as monochromosomal hybrid cells. The alphoid DNA array was also isolated from the 12 Mb human mini-chromosome DeltaYq74 that contained the minimum amount of alphoid DNA required for proper chromosome segregation. Preliminary results of the structural analyses of different centromeres are reported in this paper. The ability of the cloned human centromeric regions to support human artificial chromosome (HAC) formation was assessed by transfection into human HT1080 cells. Centromeric clones from DeltaYq74 did not support the formation of HACs, indicating that the requirements for the existence of a functional centromere on an endogenous chromosome and those for forming a de novo centromere may be distinct. A construct with an alphoid DNA array from chromosome 22 with no detectable CENP-B motifs formed mitotically stable HACs in the absence of drug selection without detectable acquisition of host DNAs. In summary, our results demonstrated that TAR cloning is a useful tool for investigating human centromere organization and the structural requirements for formation of HAC vectors that might have a potential for therapeutic applications.

Base Sequence↗

Constant relative rate of protein evolution and detection of functional diversification among bacterial, archaeal and eukaryotic proteins.

BACKGROUND: Detection of changes in a protein's evolutionary rate may reveal cases of change in that protein's function. We developed and implemented a simple relative rates test in an attempt to assess the rate constancy of protein evolution and to detect cases of functional diversification between orthologous proteins. The test was performed on clusters of orthologous protein sequences from complete bacterial genomes (Chlamydia trachomatis, C. muridarum and Chlamydophila pneumoniae), complete archaeal genomes (Pyrococcus horikoshii, P. abyssi and P. furiosus) and partially sequenced mammalian genomes (human, mouse and rat). RESULTS: Amino-acid sequence evolution rates are significantly correlated on different branches of phylogenetic trees representing the great majority of analyzed orthologous protein sets from all three domains of life. However, approximately 1% of the proteins from each group of species deviates from this pattern and instead shows variation that is consistent with an acceleration of the rate of amino-acid substitution, which may be due to functional diversification. Most of the putative functionally diversified proteins from all three species groups are predicted to function at the periphery of the cells and mediate their interaction with the environment. CONCLUSIONS: Relative rates of protein evolution are remarkably constant for the three species groups analyzed here. Deviations from this rate constancy are probably due to changes in selective constraints associated with diversification between orthologs. Functional diversification between orthologs is thought to be a relatively rare event. However, the resolution afforded by the test designed specifically for genomic-scale datasets allowed us to identify numerous cases of possible functional diversification between orthologous proteins.

Animals↗

Genome trees constructed using five different approaches suggest new major bacterial clades.

BACKGROUND: The availability of multiple complete genome sequences from diverse taxa prompts the development of new phylogenetic approaches, which attempt to incorporate information derived from comparative analysis of complete gene sets or large subsets thereof. Such attempts are particularly relevant because of the major role of horizontal gene transfer and lineage-specific gene loss, at least in the evolution of prokaryotes. RESULTS: Five largely independent approaches were employed to construct trees for completely sequenced bacterial and archaeal genomes: i) presence-absence of genomes in clusters of orthologous genes; ii) conservation of local gene order (gene pairs) among prokaryotic genomes; iii) parameters of identity distribution for probable orthologs; iv) analysis of concatenated alignments of ribosomal proteins; v) comparison of trees constructed for multiple protein families. All constructed trees support the separation of the two primary prokaryotic domains, bacteria and archaea, as well as some terminal bifurcations within the bacterial and archaeal domains. Beyond these obvious groupings, the trees made with different methods appeared to differ substantially in terms of the relative contributions of phylogenetic relationships and similarities in gene repertoires caused by similar life styles and horizontal gene transfer to the tree topology. The trees based on presence-absence of genomes in orthologous clusters and the trees based on conserved gene pairs appear to be strongly affected by gene loss and horizontal gene transfer. The trees based on identity distributions for orthologs and particularly the tree made of concatenated ribosomal protein sequences seemed to carry a stronger phylogenetic signal. The latter tree supported three potential high-level bacterial clades,: i) Chlamydia-Spirochetes, ii) Thermotogales-Aquificales (bacterial hyperthermophiles), and ii) Actinomycetes-Deinococcales-Cyanobacteria. The latter group also appeared to join the low-GC Gram-positive bacteria at a deeper tree node. These new groupings of bacteria were supported by the analysis of alternative topologies in the concatenated ribosomal protein tree using the Kishino-Hasegawa test and by a census of the topologies of 132 individual groups of orthologous proteins. Additionally, the results of this analysis put into question the sister-group relationship between the two major archaeal groups, Euryarchaeota and Crenarchaeota, and suggest instead that Euryarchaeota might be a paraphyletic group with respect to Crenarchaeota. CONCLUSIONS: We conclude that, the extensive horizontal gene flow and lineage-specific gene loss notwithstanding, extension of phylogenetic analysis to the genome scale has the potential of uncovering deep evolutionary relationships between prokaryotic lineages.

Bacteria↗

Error rate and specificity of human and murine DNA polymerase eta.

We describe here the error specificity of mammalian DNA polymerase eta (pol eta), an enzyme that performs translesion DNA synthesis and may participate in somatic hypermutation of immunoglobulin genes. Both mouse and human pol eta lack intrinsic proofreading exonuclease activity and both copy undamaged DNA inaccurately. Analysis of more than 1500 single-base substitutions by human pol eta indicates that error rates for all 12 mismatches are high and variable depending on the composition and symmetry of the mismatch and its location. pol eta also generates tandem base substitutions at an unprecedented rate, and kinetic analysis indicates that it extends a tandem double mismatch about as efficiently as other replicative enzymes extend single-base mismatches. This ability to use an aberrant primer terminus and the high rate of single and double-base substitutions support the idea that pol eta may forego strict shape complementarity in order to facilitate highly efficient lesion bypass. Relaxed discrimination is further indicated by pol eta infidelity for a wide variety of nucleotide deletion and addition errors. The nature and location of these errors suggest that some may be initiated by strand slippage, while others result from additional mechanisms.

Animals↗

Mutagenic specificity of the base analog 6-N-hydroxylaminopurine in the LYS2 gene of yeast Saccharomyces cerevisiae.

We used the LYS2 gene mutational system to study mutation specificity of the base analog 6-N-hydroxylaminopurine (HAP) in yeast. We characterized phenotypes of mutations using codon-specific nonsense suppressors and the test employing inactivation of the release factor Sup35 due to overexpression and formation of prion-like derivative [PSI]. We have shown that HAP induces predominantly nonsense mutations. While the tests using codon-specific nonsense-suppressors allowed to identify only about 50% of nonsense-mutations, all the nonsense-mutations were identified in the test with defective Sup35. We determined and analyzed the spectrum of HAP-induced nucleotide changes in two regions of the gene. HAP induces predominantly GC-->AT transitions in a hotspots of a central position of trinucleotide GGA or AGG. Directionality of these transitions is consistent with the idea that initial dHAPMP incorporation in the leading strand is more genetically dangerous than in lagging DNA strand. We revealed a specific context inhibitory for HAP mutagenesis, a "T" in -1 position to mutation site.

Adenine↗

Comparative study and prediction of DNA fragments associated with various elements of the nuclear matrix.

Scaffold/matrix-associated region (S/MAR) sequences are DNA regions that are attached to the nuclear matrix, and participate in many cellular processes. The nuclear matrix is a complex structure consisting of various elements. In this paper we compared frequencies of simple nucleotide motifs in S/MAR sequences and in sequences extracted directly from various nuclear matrix elements, such as nuclear lamina, cores of rosette-like structures, synaptonemal complex. Multivariate linear discriminant analysis revealed significant differences between these sequences. Based on this result we have developed a program, ChrClass (Win/NT version, ftp.bionet.nsc.ru/pub/biology/chrclass/chrclass.zip), for the prediction of the regions associated with various elements of the nuclear matrix in a query sequence. Subsequently, several test samples were analyzed by using two S/MAR prediction programs (a ChrClass and MAR-Finder) and a simple MRS criterion (S/MAR recognition signature) indicating the presence of S/MARs. Some overlap between the predictions of all MAR prediction tools has been found. Simultaneous use of the ChrClass, MRS criterion and MAR-Finder programs may help to obtain a more clearcut picture of S/MAR distribution in a query sequence. In general, our results suggest that the proportion of missed S/MARs is lower for ChrClass, whereas the proportion of wrong S/MARs is lower for MAR-Finder and MRS.

Animals↗

Somatic mutation hotspots correlate with DNA polymerase eta error spectrum.

Mutational spectra analysis of 15 immunoglobulin genes suggested that consensus motifs RGYW and WA were universal descriptors of somatic hypermutation. Highly mutable sites, "hotspots", that matched WA were preferentially found in one DNA strand and RGYW hotspots were found in both strands. Analysis of base-substitution hotspots in DNA polymerase error spectra showed that 33 of 36 hotspots in the human polymerase eta spectrum conformed to the WA consensus. This and four other characteristics of polymerase eta substitution specificity suggest that errors introduced by this enzyme during synthesis of the nontranscribed DNA strand in variable regions may contribute to strand-specific somatic hypermutagenesis of immunoglobulin genes at A-T base pairs.

Amino Acid Motifs↗

Cloning and functional analysis of SEL1L promoter region, a pancreas-specific gene.

We examined the promoter activity of SEL1L, the human ortholog of the C. elegans gene sel-1, a negative regulator of LIN-12/NOTCH receptor proteins. To understand the relation in SEL1L transcription pattern observed in different epithelial cells, we determined the transcription start site and sequenced the 5' flanking region. Sequence analysis revealed the presence of consensus promoter elements--GC boxes and a CAAT box--but the absence of a TATA motif. Potential binding sites for transcription factors that are involved in tissue-specific gene expression were identified, including: activator protein-2 (AP-2), hepatocyte nuclear factor-3 (HNF3 beta), homeobox Nkx2-5 and GATA-1. Transcription activity of the TATA-less SEL1L promoter was analyzed by transient transfection using luciferase reporter gene constructs. A core basal promoter of 302 bp was sufficient for constitutive promoter activity in all the cell types studied. This genomic fragment contains a CAAT and several GC boxes. The activity of the SEL1L promoter was considerably higher in mouse pancreatic beta cells (beta TC3) than in several human pancreatic neoplastic cell lines; an even greater reduction of its activity was observed in cells of nonpancreatic origin. These results suggest that SEL1L promoter may be a useful tool in gene therapy applications for pancreatic pathologies.

Animals↗

Presence of ATG triplets in 5' untranslated regions of eukaryotic cDNAs correlates with a 'weak' context of the start codon.

MOTIVATION: The context of the start codon (typically, AUG) and the features of the 5' Untranslated Regions (5' UTRs) are important for understanding translation regulation in eukaryotic mRNAs and for accurate prediction of the coding region in genomic and cDNA sequences. The presence of AUG triplets in 5' UTRs (upstream AUGs) might effect the initiation rate and, in the context of gene prediction, could reduce the accuracy of the identification of the authentic start. To reveal potential connections between the presence of upstream AUGs and other features of 5' UTRs, such as their length and the start codon context, we undertook a systematic analysis of the available eukaryotic 5' UTR sequences. RESULTS: We show that a large fraction of 5' UTRs in the available cDNA sequences, 15-53% depending on the organism, contain upstream ATGs. A negative correlation was observed between the information content of the translation start signal and the length of the 5' UTR. Similarly, a negative correlation exists between the 'strength' of the start context and the number of upstream ATGs. Typically, cDNAs containing long 5' UTRs with multiple upstream ATGs have a 'weak' start context, and in contrast, cDNAs containing short 5' UTRs without ATGs have 'strong' starts. These counter-intuitive results may be interpreted in terms of upstream AUGs having an important role in the regulation of translation efficiency by ensuring low basal translation level via double negative control and creating the potential for additional regulatory mechanisms. One of such mechanisms, supported by experimental studies of some mRNAs, includes removal of the AUG-containing portion of the 5' UTR by alternative splicing. AVAILABILITY: An ATG_ EVALUATOR program is available upon request or at www.itba.mi.cnr.it/webgene. CONTACT: rogozin@ncbi.nlm.nih.gov, milanesi@itba.mi.cnr.it.

5' Untranslated Regions↗

Rapid evolution of a cyclin A inhibitor gene, roughex, in Drosophila.

The recent sequencing of the complete genome of the fruit fly Drosophila melanogaster has yielded about 30% of the predicted genes with no obvious counterparts in other organisms. These rapidly evolving genes remain largely unexplored. Here, we present evidence for a striking variability in an important Drosophila cell cycle regulator encoded by the gene roughex (rux) in closely related fly species. The unusual level of Rux protein variability indicates that there are very low overall constraints on amino acid substitutions. Despite the lack of sequence similarity, certain common features, including the presence of a C-terminal nuclear localization signal and a functionally important N-terminal RXL cyclin-binding motif, exist between Rux and cyclin-dependent kinase inhibitors of the Cip/Kip family. These results indicate that even some genes involved in key regulatory processes in eukaryotes evolve at extremely high rates.

Amino Acid Sequence↗

Characterization of the genomic Xist locus in rodents reveals conservation of overall gene structure and tandem repeats but rapid evolution of unique sequence.

The Xist locus plays a central role in the regulation of X chromosome inactivation in mammals, although its exact mode of action remains to be elucidated. Evolutionary studies are important in identifying conserved genomic regions and defining their possible function. Here we report cloning, sequence analysis, and detailed characterization of the Xist gene from four closely related species of common vole (field mouse), Microtus arvalis. Our analysis reveals that there is overall conservation of Xist gene structure both between different vole species and relative to mouse and human Xist/XIST. Within transcribed sequence, there is significant conservation over five short regions of unique sequence and also over Xist-specific tandem repeats. The majority of unique sequences, however, are evolving at an unexpectedly high rate. This is also evident from analysis of flanking sequences, which reveals a very high rate of rearrangement and invasion of dispersed repeats. We discuss these results in the context of Xist gene function and evolution.

3' Untranslated Regions↗

Genome alignment, evolution of prokaryotic genome organization, and prediction of gene function using genomic context.

Gene order in prokaryotes is conserved to a much lesser extent than protein sequences. Only several operons, primarily those that code for physically interacting proteins, are conserved in all or most of the bacterial and archaeal genomes. Nevertheless, even the limited conservation of operon organization that is observed can provide valuable evolutionary and functional clues through multiple genome comparisons. A program for constructing gapped local alignments of conserved gene strings in two genomes was developed. The statistical significance of the local alignments was assessed using Monte Carlo simulations. Sets of local alignments were generated for all pairs of completely sequenced bacterial and archaeal genomes, and for each genome a template-anchored multiple alignment was constructed. In most pairwise genome comparisons, <10% of the genes in each genome belonged to conserved gene strings. When closely related pairs of species (i.e., two mycoplasmas) are excluded, the total coverage of genomes by conserved gene strings ranged from <5% for the cyanobacterium Synechocystis sp to 24% for the minimal genome of Mycoplasma genitalium, and 23% in Thermotoga maritima. The coverage of the archaeal genomes was only slightly lower than that of bacterial genomes. The majority of the conserved gene strings are known operons, with the ribosomal superoperon being the top-scoring string in most genome comparisons. However, in some of the bacterial-archaeal pairs, the superoperon is rearranged to the extent that other operons, primarily those subject to horizontal transfer, show the greatest level of conservation, such as the archaeal-type H+-ATPase operon or ABC-type transport cassettes. The level of gene order conservation among prokaryotic genomes was compared to the cooccurrence of genomes in clusters of orthologous genes (COGs) and to the conservation of protein sequences themselves. Only limited correlation was observed between these evolutionary variables. Gene order conservation shows a much lower variance than the cooccurrence of genomes in COGs, which indicates that intragenome homogenization via recombination occurs in evolution much faster than intergenome homogenization via horizontal gene transfer and lineage-specific gene loss. The potential of using template-anchored multiple-genome alignments for predicting functions of uncharacterized genes was quantitatively assessed. Functions were predicted or significantly clarified for approximately 90 COGs (approximately 4% of the total of 2414 analyzed COGs). The most significant predictions were obtained for the poorly characterized archaeal genomes; these include a previously uncharacterized restriction-modification system, a nuclease-helicase combination implicated in DNA repair, and the probable archaeal counterpart of the eukaryotic exosome. Multiple genome alignments are a resource for studies on operon rearrangement and disruption, which is central to our understanding of the evolution of prokaryotic genomes. Because of the rapid evolution of the gene order, the potential of genome alignment for prediction of gene functions is limited, but nevertheless, such predictions information significantly complements the results obtained through protein sequence and structure analysis.

Computational Biology↗

[Study of the DNA primary structure effect on induction of mutations by alkylating agents].

Analysis and comparison of mutation spectra is one of the major tasks of molecular biology, since mutation spectra often reveal important properties of various mutagens and proteins involved in the repair/replication systems. Mutability is known to vary significantly along the nucleotide sequence. Mutations are abundant at certain positions (mutation hotspots). In this work, we applied regression analysis based on the basic logic patterns to understand the role of the nucleotide sequence context in mutation induction. The spectra of mutations induced by various alkylating agents were studied. The nucleotide bases at positions -2, -1, +1 and +2 were shown to have the most significant effect in G:C-->A:T replacements.

Alkylating Agents↗

Purifying selection and birth-and-death evolution in the ubiquitin gene family.

Ubiquitin is a highly conserved protein that is encoded by a multigene family. It is generally believed that this gene family is subject to concerted evolution, which homogenizes the member genes of the family. However, protein homogeneity can be attained also by strong purifying selection. We therefore studied the proportion (p(S)) of synonymous nucleotide differences between members of the ubiquitin gene family from 28 species of fungi, plants, and animals. The results have shown that p(S) is generally very high and is often close to the saturation level, although the protein sequence is virtually identical for all ubiquitins from fungi, plants, and animals. A small proportion of species showed a low level of p(S) values, but these values appeared to be caused by recent gene duplication. It was also found that the number of repeat copies of the gene family varies considerably with species, and some species harbor pseudogenes. These observations suggest that the members of this gene family evolve almost independently by silent nucleotide substitution and are subjected to birth-and-death evolution at the DNA level.

Animals↗