Search PubMed⌕ Search

Biomedical subjects

Igor B Rogozin

Publications and source records attributed to Igor B Rogozin.

At least 19 recordsLinked to original sources

Remarkable interkingdom conservation of intron positions and massive, lineage-specific intron loss and gain in eukaryotic evolution.

Sequencing of eukaryotic genomes allows one to address major evolutionary problems, such as the evolution of gene structure. We compared the intron positions in 684 orthologous gene sets from 8 complete genomes of animals, plants, fungi, and protists and constructed parsimonious scenarios of evolution of the exon-intron structure for the respective genes. Approximately one-third of the introns in the malaria parasite Plasmodium falciparum are shared with at least one crown group eukaryote; this number indicates that these introns have been conserved through >1.5 billion years of evolution that separate Plasmodium from the crown group. Paradoxically, humans share many more introns with the plant Arabidopsis thaliana than with the fly or nematode. The inferred evolutionary scenario holds that the common ancestor of Plasmodium and the crown group and, especially, the common ancestor of animals, plants, and fungi had numerous introns. Most of these ancestral introns, which are retained in the genomes of vertebrates and plants, have been lost in fungi, nematodes, arthropods, and probably Plasmodium. In addition, numerous introns have been inserted into vertebrate and plant genes, whereas, in other lineages, intron gain was much less prominent.

Amino Acid Sequence↗

Evolution of mosaic operons by horizontal gene transfer and gene displacement in situ.

BACKGROUND: Shuffling and disruption of operons and horizontal gene transfer are major contributions to the new, dynamic view of prokaryotic evolution. Under the 'selfish operon' hypothesis, operons are viewed as mobile genetic entities that are constantly disseminated via horizontal gene transfer, although their retention could be favored by the advantage of coregulation of functionally linked genes. Here we apply comparative genomics and phylogenetic analysis to examine horizontal transfer of entire operons versus displacement of individual genes within operons by horizontally acquired orthologs and independent assembly of the same or similar operons from genes with different phylogenetic affinities. RESULTS: Since a substantial number of operons have been identified experimentally in only a few model bacteria, evolutionarily conserved gene strings were analyzed as surrogates of operons. The phylogenetic affinities within these predicted operons were assessed first by sequence similarity analysis and then by phylogenetic analysis, including statistical tests of tree topology. Numerous cases of apparent horizontal transfer of entire operons were detected. However, it was shown that apparent horizontal transfer of individual genes or arrays of genes within operons is not uncommon either and results in xenologous gene displacement in situ, that is, displacement of an ancestral gene by a horizontally transferred ortholog from a taxonomically distant organism without change of the local gene organization. On rarer occasions, operons might have evolved via independent assembly, in part from horizontally acquired genes. CONCLUSIONS: The discovery of in situ gene displacement shows that combination of rampant horizontal gene transfer with selection for preservation of operon structure provides for events in prokaryotic evolution that, a priori, seem improbable. These findings also emphasize that not all aspects of operon evolution are selfish, with operon integrity maintained by purifying selection at the organism level.

Alkyl and Aryl Transferases↗

129-derived strains of mice are deficient in DNA polymerase iota and have normal immunoglobulin hypermutation.

Recent studies suggest that DNA polymerase eta (poleta) and DNA polymerase iota (poliota) are involved in somatic hypermutation of immunoglobulin variable genes. To test the role of poliota in generating mutations in an animal model, we first characterized the biochemical properties of murine poliota. Like its human counterpart, murine poliota is extremely error-prone when catalyzing synthesis on a variety of DNA templates in vitro. Interestingly, when filling in a 1 base-pair gap, DNA synthesis and subsequent strand displacement was greatest in the presence of both pols iota and eta. Genomic sequence analysis of Poli led to the serendipitous discovery that 129-derived strains of mice have a nonsense codon mutation in exon 2 that abrogates production of poliota. Analysis of hypermutation in variable genes from 129/SvJ (Poli-/-) and C57BL/6J (Poli+/+) mice revealed that the overall frequency and spectrum of mutation were normal in poliota-deficient mice. Thus, either poliota does not participate in hypermutation, or its role is nonessential and can be readily assumed by another low-fidelity polymerase.

Animals↗

Genome sequence of the cyanobacterium Prochlorococcus marinus SS120, a nearly minimal oxyphototrophic genome.

Prochlorococcus marinus, the dominant photosynthetic organism in the ocean, is found in two main ecological forms: high-light-adapted genotypes in the upper part of the water column and low-light-adapted genotypes at the bottom of the illuminated layer. P. marinus SS120, the complete genome sequence reported here, is an extremely low-light-adapted form. The genome of P. marinus SS120 is composed of a single circular chromosome of 1,751,080 bp with an average G+C content of 36.4%. It contains 1,884 predicted protein-coding genes with an average size of 825 bp, a single rRNA operon, and 40 tRNA genes. Together with the 1.66-Mbp genome of P. marinus MED4, the genome of P. marinus SS120 is one of the two smallest genomes of a photosynthetic organism known to date. It lacks many genes that are involved in photosynthesis, DNA repair, solute uptake, intermediary metabolism, motility, phototaxis, and other functions that are conserved among other cyanobacteria. Systems of signal transduction and environmental stress response show a particularly drastic reduction in the number of components, even taking into account the small size of the SS120 genome. In contrast, housekeeping genes, which encode enzymes of amino acid, nucleotide, cofactor, and cell wall biosynthesis, are all present. Because of its remarkable compactness, the genome of P. marinus SS120 might approximate the minimal gene complement of a photosynthetic organism.

Adaptation, Physiological↗

Unique error signature of the four-subunit yeast DNA polymerase epsilon.

We have purified wild type and exonuclease-deficient four-subunit DNA polymerase epsilon (Pol epsilon) complex from Saccharomyces cerevisiae and analyzed the fidelity of DNA synthesis by the two enzymes. Wild type Pol epsilon synthesizes DNA accurately, generating single-base substitutions and deletions at average error rates of </=2 x 10-5 and </=5 x 10-7, respectively. Pol epsilon lacking 3' --> 5' exonuclease activity is less accurate to a degree suggesting that wild type Pol epsilon proofreads at least 92% of base substitution errors and at least 99% of frameshift errors made by the polymerase. Surprisingly the base substitution fidelity of exonuclease-deficient Pol epsilon is severalfold lower than that of proofreading-deficient forms of other replicative polymerases. Moreover the spectrum of errors shows a feature not seen with other A, B, C, or X family polymerases: a high proportion of transversions resulting from T.dTTP, T.dCTP, and C.dTTP mispairs. This unique error specificity and amino acid sequence alignments suggest that the structure of the polymerase active site of Pol epsilon differs from those of other B family members. We observed both similarities and differences between the spectrum of substitutions generated by proofreading-deficient Pol epsilon in vitro and substitutions occurring in vivo in a yeast strain defective in Pol epsilon proofreading and DNA mismatch repair. We discuss the implications of these findings for the role of Pol epsilon polymerase activity in DNA replication.

Amino Acid Sequence↗

Getting positive about selection.

A report on the 68th Symposium on Quantitative Biology, The Genome of Homo Sapiens', Cold Spring Harbor, USA, 28 May-2 June 2003.

Animals↗

Transcriptome dynamics of Deinococcus radiodurans recovering from ionizing radiation.

Deinococcus radiodurans R1 (DEIRA) is a bacterium best known for its extreme resistance to the lethal effects of ionizing radiation, but the molecular mechanisms underlying this phenotype remain poorly understood. To define the repertoire of DEIRA genes responding to acute irradiation (15 kGy), transcriptome dynamics were examined in cells representing early, middle, and late phases of recovery by using DNA microarrays covering approximately 94% of its predicted genes. At least at one time point during DEIRA recovery, 832 genes (28% of the genome) were induced and 451 genes (15%) were repressed 2-fold or more. The expression patterns of the majority of the induced genes resemble the previously characterized expression profile of recA after irradiation. DEIRA recA, which is central to genomic restoration after irradiation, is substantially up-regulated on DNA damage (early phase) and down-regulated before the onset of exponential growth (late phase). Many other genes were expressed later in recovery, displaying a growth-related pattern of induction. Genes induced in the early phase of recovery included those involved in DNA replication, repair, and recombination, cell wall metabolism, cellular transport, and many encoding uncharacterized proteins. Collectively, the microarray data suggest that DEIRA cells efficiently coordinate their recovery by a complex network, within which both DNA repair and metabolic functions play critical roles. Components of this network include a predicted distinct ATP-dependent DNA ligase and metabolic pathway switching that could prevent additional genomic damage elicited by metabolism-induced free radicals.

Amino Acid Sequence↗

Differential action of natural selection on the N and C-terminal domains of 2'-5' oligoadenylate synthetases and the potential nuclease function of the C-terminal domain.

2'-5' Oligoadenylate synthetases (OAS) are a family of enzymes, which are best known for their important role in interferon-dependent antiviral mechanisms, but are also involved in the regulation of apoptosis, cell growth and differentiation in vertebrates. These enzymes bind double-stranded RNA and catalyze the synthesis of 2'-5' oligoadenylates from ATP. Several 2'-5' oligoadenylate synthetase-like proteins, which lack the ability to synthesize 2'-5' A, have been recently identified in humans and mice; the functions of these inactivated OAS derivatives remain unknown. Examination of phylogenetic trees shows that OAS inactivation in mammals occurred on several independent occasions. Comparative sequence analysis of OAS, poly(A)-polymerases, TRF4/sigma-family polymerases, archaeal CCA-adding enzymes and uridilyltransferases from trypanosomes resulted in the identification of a C-terminal domain, which is conserved in all these enzymes and is distinct from the nucleotidyltransferase domain. Secondary structure prediction shows that this domain has a four-helix core, which is most closely related to the ATP-cone domain, a regulatory nucleotide-binding domain present in ribonucleotide reductases and several other enzymes and transcription regulators. These observations, taken together with the experimental evidence of nuclease activity in the TRF4/sigma-family of polymerases, suggest that the C-terminal domain of OAS and their homologs might have nuclease activity. The putative nuclease domain is preferentially conserved in OAS derivatives that lack an active nucleotidyltransferase domain and, as indicated by the analysis of the ratio of synonymous to non-synonymous substitutions, appears to be subject to purifying selection in these proteins. In contrast, phylogenetic analysis provided evidence of episodic positive selection in the mouse OAS-like proteins with inactivated nucleotidyltransferase domains, which suggests that some of these proteins might have distinct antiviral functions.

2',5'-Oligoadenylate Synthetase↗

The rhomboids: a nearly ubiquitous family of intramembrane serine proteases that probably evolved by multiple ancient horizontal gene transfers.

BACKGROUND: The rhomboid family of polytopic membrane proteins shows a level of evolutionary conservation unique among membrane proteins. They are present in nearly all the sequenced genomes of archaea, bacteria and eukaryotes, with the exception of several species with small genomes. On the basis of experimental studies with the developmental regulator rhomboid from Drosophila and the AarA protein from the bacterium Providencia stuartii, the rhomboids are thought to be intramembrane serine proteases whose signaling function is conserved in eukaryotes and prokaryotes. RESULTS: Phylogenetic tree analysis carried out using several independent methods for tree constructions and the corresponding statistical tests suggests that, despite its broad distribution in all three superkingdoms, the rhomboid family was not present in the last universal common ancestor of extant life forms. Instead, we propose that rhomboids evolved in bacteria and have been acquired by archaea and eukaryotes through several independent horizontal gene transfers. In eukaryotes, two distinct, ancient acquisitions apparently gave rise to the two major subfamilies, typified by rhomboid and PARL (presenilins-associated rhomboid-like protein), respectively. Subsequent evolution of the rhomboid family in eukaryotes proceeded by multiple duplications and functional diversification through the addition of extra transmembrane helices and other domains in different orientations relative to the conserved core that harbors the protease activity. CONCLUSIONS: Although the near-universal presence of the rhomboid family in bacteria, archaea and eukaryotes appears to suggest that this protein is part of the heritage of the last universal common ancestor, phylogenetic tree analysis indicates a likely bacterial origin with subsequent dissemination by horizontal gene transfer. This emphasizes the importance of explicit phylogenetic analysis for the reconstruction of ancestral life forms. A hypothetical scenario for the origin of intracellular membrane proteases from membrane transporters is proposed.

Amino Acid Sequence↗

Origin of a substantial fraction of human regulatory sequences from transposable elements.

Transposable elements (TEs) are abundant in mammalian genomes and have potentially contributed to their hosts' evolution by providing novel regulatory or coding sequences. We surveyed different classes of regulatory region in the human genome to assess systematically the potential contribution of TEs to gene regulation. Almost 25% of the analyzed promoter regions contain TE-derived sequences, including many experimentally characterized cis-regulatory elements. Scaffold/matrix attachment regions (S/MARs) and locus control regions (LCRs) that are involved in the simultaneous regulation of multiple genes also contain numerous TE-derived sequences. Thus, TEs have probably contributed substantially to the evolution of both gene-specific and global patterns of human gene regulation.

5' Untranslated Regions↗

A significant fraction of conserved noncoding DNA in human and mouse consists of predicted matrix attachment regions.

Noncoding DNA in the human-mouse orthologous intergenic regions contains "islands" of conserved sequences, the functions of which remain largely unknown. We hypothesized that some of these regions might be matrix-scaffold attachment regions, MARs (or S/MARs). MARs comprise one of the few classes of eukaryotic noncoding DNA with an experimentally characterized function, being involved in the attachment of chromatin to the nuclear matrix, chromatin remodeling and transcription regulation. To test our hypothesis, we analyzed the co-occurrence of predicted MARs with highly conserved noncoding DNA regions in human-mouse genomic alignments. We found that 11% of the conserved noncoding DNA consists of predicted MARs. Conversely, more than half of the predicted MARs co-occur with one or more independently identified conserved sequence blocks. An excess of conserved predicted MARs is seen in intergenic regions preceding 5' ends of genes, suggesting that these MARs are primarily involved in transcriptional control.

Animals↗

Theoretical analysis of mutation hotspots and their DNA sequence context specificity.

Mutation frequencies vary significantly along nucleotide sequences such that mutations often concentrate at certain positions called hotspots. Mutation hotspots in DNA reflect intrinsic properties of the mutation process, such as sequence specificity, that manifests itself at the level of interaction between mutagens, DNA, and the action of the repair and replication machineries. The hotspots might also reflect structural and functional features of the respective DNA sequences. When mutations in a gene are identified using a particular experimental system, resulting hotspots could reflect the properties of the gene product and the mutant selection scheme. Analysis of the nucleotide sequence context of hotspots can provide information on the molecular mechanisms of mutagenesis. However, the determinants of mutation frequency and specificity are complex, and there are many analytical methods for their study. Here we review computational approaches for analyzing mutation spectra (distribution of mutations along the target genes) that include many mutable (detectable) positions. The following methods are reviewed: derivation of a consensus sequence, application of regression approaches to correlate nucleotide sequence features with mutation frequency, mutation hotspot prediction, analysis of oligonucleotide composition of regions containing mutations, pairwise comparison of mutation spectra, analysis of multiple spectra, and analysis of "context-free" characteristics. The advantages and pitfalls of these methods are discussed and illustrated by examples from the literature. The most reliable analyses were obtained when several methods were combined and information from theoretical analysis and experimental observations was considered simultaneously. Simple, robust approaches should be used with small samples of mutations, whereas combinations of simple and complex approaches may be required for large samples. We discuss several well-documented studies where analysis of mutation spectra has substantially contributed to the current understanding of molecular mechanisms of mutagenesis. The nucleotide sequence context of mutational hotspots is a fingerprint of interactions between DNA and DNA repair, replication, and modification enzymes, and the analysis of hotspot context provides evidence of such interactions.

Animals↗

Congruent evolution of different classes of non-coding DNA in prokaryotic genomes.

Prokaryotic genomes are considered to be 'wall-to-wall' genomes, which consist largely of genes for proteins and structural RNAs, with only a small fraction of the genomic DNA allotted to intergenic regions, which are thought to typically contain regulatory signals. The majority of bacterial and archaeal genomes contain 6-14% non-coding DNA. Significant positive correlations were detected between the fraction of non-coding DNA and inter- and intra-operonic distances, suggesting that different classes of non-coding DNA evolve congruently. In contrast, no correlation was found between any of these characteristics of non-coding sequences and the number of genes or genome size. Thus, the non-coding regions and the gene sets in prokaryotes seem to evolve in different regimes. The evolution of non-coding regions appears to be determined primarily by the selective pressure to minimize the amount of non-functional DNA, while maintaining essential regulatory signals, because of which the content of non-coding DNA in different genomes is relatively uniform and intra- and inter-operonic non-coding regions evolve congruently. In contrast, the gene set is optimized for the particular environmental niche of the given microbe, which results in the lack of correlation between the gene number and the characteristics of non-coding regions.

DNA, Intergenic↗

Correlation of somatic hypermutation specificity and A-T base pair substitution errors by DNA polymerase eta during copying of a mouse immunoglobulin kappa light chain transgene.

To test the hypothesis that inaccurate DNA synthesis by mammalian DNA polymerase eta (pol eta) contributes to somatic hypermutation (SHM) of Ig genes, we measured the error specificity of mouse pol eta during synthesis of each strand of a mouse Ig kappa light chain transgene. We then compared the results to the base substitution specificity of SHM of this same gene in the mouse. The in vitro and in vivo base substitution spectra shared a number of common features. A highly significant correlation was observed for overall substitutions at A-T pairs but not for substitutions at G-C pairs. Sixteen mutational hotspots at A-T pairs observed in vivo were also found in spectra generated by mouse pol eta in vitro. The correlation was strongest for errors made by pol eta during synthesis of the non-transcribed strand, but it was also observed for synthesis of the transcribed strand. These facts, and the distribution of substitutions generated in vivo, support the hypothesis that pol eta contributes to SHM of Ig genes at A-T pairs via short patches of low fidelity DNA synthesis of both strands, but with a preference for the non-transcribed strand.

Adenine↗

Analysis of phylogenetically reconstructed mutational spectra in human mitochondrial DNA control region.

Analysis of mutations in mitochondrial DNA is an important issue in population and evolutionary genetics. To study spontaneous base substitutions in human mitochondrial DNA we reconstructed the mutational spectra of the hypervariable segments I and II (HVS I and II) using published data on polymorphisms from various human populations. An excess of pyrimidine transitions was found both in HVS I and II regions. By means of classification analysis numerous mutational hotspots were revealed in these spectra. Context analysis of hotspots revealed a complex influence of neighboring bases on mutagenesis in the HVS I region. Further statistical analysis suggested that a transient misalignment dislocation mutagenesis operating in monotonous runs of nucleotides play an important role for generating base substitutions in mitochondrial DNA and define context properties of mtDNA. Our results suggest that dislocation mutagenesis in HVS I and II is a fingerprint of errors produced by DNA polymerase gamma in the course of human mitochondrial DNA replication

Base Sequence↗

Connected gene neighborhoods in prokaryotic genomes.

A computational method was developed for delineating connected gene neighborhoods in bacterial and archaeal genomes. These gene neighborhoods are not typically present, in their entirety, in any single genome, but are held together by overlapping, partially conserved gene arrays. The procedure was applied to comparing the orders of orthologous genes, which were extracted from the database of Clusters of Orthologous Groups of proteins (COGs), in 31 prokaryotic genomes and resulted in the identification of 188 clusters of gene arrays, which included 1001 of 2890 COGs. These clusters were projected onto actual genomes to produce extended neighborhoods including additional genes, which are adjacent to the genes from the clusters and are transcribed in the same direction, which resulted in a total of 2387 COGs being included in the neighborhoods. Most of the neighborhoods consist predominantly of genes united by a coherent functional theme, but also include a minority of genes without an obvious functional connection to the main theme. We hypothesize that although some of the latter genes might have unsuspected roles, others are maintained within gene arrays because of the advantage of expression at a level that is typical of the given neighborhood. We designate this phenomenon 'genomic hitchhiking'. The largest neighborhood includes 79 genes (COGs) and consists of overlapping, rearranged ribosomal protein superoperons; apparent genome hitchhiking is particularly typical of this neighborhood and other neighborhoods that consist of genes coding for translation machinery components. Several neighborhoods involve previously undetected connections between genes, allowing new functional predictions. Gene neighborhoods appear to evolve via complex rearrangement, with different combinations of genes from a neighborhood fixed in different lineages.

Algorithms↗

The complete genome of hyperthermophile Methanopyrus kandleri AV19 and monophyly of archaeal methanogens.

We have determined the complete 1,694,969-nt sequence of the GC-rich genome of Methanopyrus kandleri by using a whole direct genome sequencing approach. This approach is based on unlinking of genomic DNA with the ThermoFidelase version of M. kandleri topoisomerase V and cycle sequencing directed by 2'-modified oligonucleotides (Fimers). Sequencing redundancy (3.3x) was sufficient to assemble the genome with less than one error per 40 kb. Using a combination of sequence database searches and coding potential prediction, 1,692 protein-coding genes and 39 genes for structural RNAs were identified. M. kandleri proteins show an unusually high content of negatively charged amino acids, which might be an adaptation to the high intracellular salinity. Previous phylogenetic analysis of 16S RNA suggested that M. kandleri belonged to a very deep branch, close to the root of the archaeal tree. However, genome comparisons indicate that, in both trees constructed using concatenated alignments of ribosomal proteins and trees based on gene content, M. kandleri consistently groups with other archaeal methanogens. M. kandleri shares the set of genes implicated in methanogenesis and, in part, its operon organization with Methanococcus jannaschii and Methanothermobacter thermoautotrophicum. These findings indicate that archaeal methanogens are monophyletic. A distinctive feature of M. kandleri is the paucity of proteins involved in signaling and regulation of gene expression. Also, M. kandleri appears to have fewer genes acquired via lateral transfer than other archaea. These features might reflect the extreme habitat of this organism.

Base Sequence↗

A DNA repair system specific for thermophilic Archaea and bacteria predicted by genomic context analysis.

During a systematic analysis of conserved gene context in prokaryotic genomes, a previously undetected, complex, partially conserved neighborhood consisting of more than 20 genes was discovered in most Archaea (with the exception of Thermoplasma acidophilum and Halobacterium NRC-1) and some bacteria, including the hyperthermophiles Thermotoga maritima and Aquifex aeolicus. The gene composition and gene order in this neighborhood vary greatly between species, but all versions have a stable, conserved core that consists of five genes. One of the core genes encodes a predicted DNA helicase, often fused to a predicted HD-superfamily hydrolase, and another encodes a RecB family exonuclease; three core genes remain uncharacterized, but one of these might encode a nuclease of a new family. Two more genes that belong to this neighborhood and are present in most of the genomes in which the neighborhood was detected encode, respectively, a predicted HD-superfamily hydrolase (possibly a nuclease) of a distinct family and a predicted, novel DNA polymerase. Another characteristic feature of this neighborhood is the expansion of a superfamily of paralogous, uncharacterized proteins, which are encoded by at least 20-30% of the genes in the neighborhood. The functional features of the proteins encoded in this neighborhood suggest that they comprise a previously undetected DNA repair system, which, to our knowledge, is the first repair system largely specific for thermophiles to be identified. This hypothetical repair system might be functionally analogous to the bacterial-eukaryotic system of translesion, mutagenic repair whose central components are DNA polymerases of the UmuC-DinB-Rad30-Rev1 superfamily, which typically are missing in thermophiles.

Amino Acid Sequence↗