Search PubMed⌕ Search

Biomedical subjects

Alexey S Kondrashov

Publications and source records attributed to Alexey S Kondrashov.

At least 19 recordsLinked to original sources

Bursts of nonsynonymous substitutions in HIV-1 evolution reveal instances of positive selection at conservative protein sites.

The fixation of a new allele can be driven by Darwinian positive selection or can be due to random genetic drift. Identifying instances of positive selection is a difficult task, because its impact is routinely obscured by the action of negative selection. The nature of the genetic code dictates that positive selection in favor of an amino acid replacement should often cause a burst of two or three nucleotide substitutions at a single codon site, because a large fraction of amino acid replacements cannot be achieved after just one nucleotide substitution. Here, we study pairs of successive nonsynonymous substitutions at one codon in the course of evolution of HIV-1 genes within HIV-1 populations inhabiting infected individuals. Such pairs are more numerous and more clumped than expected if different substitutions were independent and than what is observed for pairs of successive synonymous substitutions. Bursts of nonsynonymous substitutions in HIV-1 evolution cannot be explained by mutational biases and must, therefore, be due to positive selection. Both reversals, exact or imprecise, of fixed deleterious mutations and acquisitions of amino acids with new properties are responsible for the bursts. Temporal clumping is strongest at codon sites with a low overall rate of nonsynonymous evolution, implying that a substantial fraction of replacements of conservative amino acids are driven by positive selection. We identified many conservative sites of HIV-1 proteins that occasionally experience positive selection.

Amino Acid Substitution↗

Sympatric speciation under incompatibility selection.

The existing theory of sympatric speciation assumes that a local population splits into two species under one-dimensional disruptive selection, which favors both of the opposite extreme values of a quantitative trait. Here we model sympatric speciation under selection that favors high values of either of the two independently inherited traits, each required to efficiently consume one of the two available resources, but acts, because of a tradeoff, against those possessing high values of both traits. Such two-dimensional incompatibility selection is similar to that involved in allopatric speciation. Using a hypergeometric phenotypic model, we show that incompatibility selection readily leads to sympatric speciation. In contrast to disruptive selection, two distinct modes of sympatric speciation exist under incompatibility selection: under strong tradeoffs both of the new species are specialists, each consuming its own resource, but under moderate tradeoffs speciation may be asymmetric and involve the origin of a specialist and a generalist species. Also, incompatibility selection may lead to irreversible specialization: under strong tradeoffs, the population speciates if it consists mostly of unspecialized individuals, but remains undivided if most of the individuals are specialized to consume one of the resources. Incompatibility selection appears to be more realistic than disruptive selection, implying that incompatibility between individually adaptive alleles or trait states drives both allopatric and sympatric speciation.

Animals↗

Distant conserved sequences flanking endothelial-specific promoters contain tissue-specific DNase-hypersensitive sites and over-represented motifs.

The transcriptional regulation of genes is a complex process, particularly for genes exhibiting a tissue-specific pattern of expression. We studied 28 genes that are expressed primarily in endothelial cells, another 28 genes that are expressed highly, but not exclusively, in cultured endothelial cells, and three control sets, consisting of genes not expressed in endothelium, genes expressed in neural tissues and housekeeping genes. For each gene, we identified conserved non-coding sequences (CNSs) of lengths 50 to >1000 nucleotides, located within the upstream intergenic region (from 500 to as far as 200 000 nucleotides upstream from the transcription start) or within the first intron. As a functional test, we assayed the CNSs from the set of endothelial cell-specific genes (EC-CNSs) for DNase hypersensitivity. Among 262 distant EC-CNSs, 33% are hypersensitive (HS) in endothelial cells, whereas only 16% are HS in control fibroblasts. A search for short sequence patterns revealed a number of motifs which are over-represented in EC-CNSs relative to CNSs from the control gene sets. In particular, the motif SAGGAAR is strongly and consistently over-represented among EC-CNSs, and is more over-represented in HS CNSs than in non-HS CNSs. CNSs which contain this motif are no closer to the promoter than an average CNS. This motif contains the core element of binding sites from the Ets family of transcription factors. Thus, one or several factors from this family may play a key role in the regulation of endothelial gene expression.

5' Flanking Region↗

Analysis of internal loops within the RNA secondary structure in almost quadratic time.

MOTIVATION: Evaluating all possible internal loops is one of the key steps in predicting the optimal secondary structure of an RNA molecule. The best algorithm available runs in time O(L(3)), L is the length of the RNA. RESULTS: We propose a new algorithm for evaluating internal loops, its run-time is O(M(*)log(2)L), M < L(2) is a number of possible nucleotide pairings. We created a software tool Afold which predicts the optimal secondary structure of RNA molecules of lengths up to 28 000 nt, using a computer with 2 Gb RAM. We also propose algorithms constructing sets of conditionally optimal multi-branch loop free (MLF) structures, e.g. the set that for every possible pairing (x, y) contains an optimal MLF structure in which nucleotides x and y form a pair. All the algorithms have run-time O(M(*)log(2)L).

Algorithms↗

Rate of promoter class turn-over in yeast evolution.

BACKGROUND: Phylogenetic conservation at the DNA level is routinely used as evidence of molecular function, under the assumption that locations and sequences of functional DNA segments remain invariant in evolution. In particular, short DNA segments participating in initiation and regulation of transcription are often conserved between related species. However, transcription of a gene can evolve, and this evolution may involve changes of even such conservative DNA segments. Genes of yeast Saccharomyces have promoters of two classes, class 1 (TATA-containing) and class 2 (non-TATA-containing). RESULTS: Comparison of upstream non-coding regions of orthologous genes from the five species of Saccharomyces sensu stricto group shows that among 212 genes which very likely have class 1 promoters in S. cerevisiae, 17 probably have class 2 promoters in one or more other species. Conversely, among 322 genes which very likely have class 2 promoters in S. cerevisiae, 44 probably have class 1 promoters in one or more other species. Also, for at least 2 genes from the set of 212 S. cerevisiae genes with class 1 promoters, the locations of the TATA consensus sequences are substantially different between the species. CONCLUSION: Our results indicate that, in the course of yeast evolution, a promoter switches its class with the probability at least approximately 0.1 per time required for the accumulation of one nucleotide substitution at a non-coding site. Thus, key sequences involved in initiation of transcription evolve with substantial rates in yeast.

Base Sequence↗

Selection in favor of nucleotides G and C diversifies evolution rates and levels of polymorphism at mammalian synonymous sites.

The impact of synonymous nucleotide substitutions on fitness in mammals remains controversial. Despite some indications of selective constraint, synonymous sites are often assumed to be neutral, and the rate of their evolution is used as a proxy for mutation rate. We subdivide all sites into four classes in terms of the mutable CpG context, nonCpG, postC, preG, and postCpreG, and compare four-fold synonymous sites and intron sites residing outside transposable elements. The distribution of the rate of evolution across all synonymous sites is trimodal. Rate of evolution at nonCpG synonymous sites, not preceded by C and not followed by G, is approximately 10% below that at such intron sites. In contrast, rate of evolution at postCpreG synonymous sites is approximately 30% above that at such intron sites. Finally, synonymous and intron postC and preG sites evolve at similar rates. The relationship between the levels of polymorphism at the corresponding synonymous and intron sites is very similar to that between their rates of evolution. Within every class, synonymous sites are occupied by G or C much more often than intron sites, whose nucleotide composition is consistent with neutral mutation-drift equilibrium. These patterns suggest that synonymous sites are under weak selection in favor of G and C, with the average coefficient s approximately 0.25/Ne approximately 10(-5), where Ne is the effective population size. Such selection decelerates evolution and reduces variability at sites with symmetric mutation, but has the opposite effects at sites where the favored nucleotides are more mutable. The amino-acid composition of proteins dictates that many synonymous sites are CpGprone, which causes them, on average, to evolve faster and to be more polymorphic than intron sites. An average genotype carries approximately 10(7) suboptimal nucleotides at synonymous sites, implying synergistic epistasis in selection against them.

Animals↗

Role of selection in fixation of gene duplications.

New genes commonly appear through complete or partial duplications of pre-existing genes. Duplications of long DNA segments are constantly produced by rare mutations, may become fixed in a population by selection or random drift, and are subject to divergent evolution of the paralogous sequences after fixation, although gene conversion can impede this process. New data shed some light on each of these processes. Mutations which involve duplications can occur through at least two different mechanisms, backward strand slippage during DNA replication and unequal crossing-over. The background rate of duplication of a complete gene in humans is 10(-9)-10(-10) per generation, although many genes located within hot-spots of large-scale mutation are duplicated much more often. Many gene duplications affect fitness strongly, and are responsible, through gene dosage effects, for a number of genetic diseases. However, high levels of intrapopulation polymorphism caused by presence or absence of long, gene-containing DNA segments imply that some duplications are not under strong selection. The polymorphism to fixation ratios appear to be approximately the same for gene duplications and for presumably selectively neutral nucleotide substitutions, which, according to the McDonald-Kreitman test, is consistent with selective neutrality of duplications. However, this pattern can also be due to negative selection against most of segregating duplications and positive selection for at least some duplications which become fixed. Patterns in post-fixation evolution of duplicated genes do not easily reveal the causes of fixations. Many gene duplications which became fixed recently in a variety of organisms were positively selected because the increased expression of the corresponding genes was beneficial. The effects of gene dosage provide a unified framework for studying all phases of the life history of a gene duplication. Application of well-known methods of evolutionary genetics to accumulating data on new, polymorphic, and fixed duplication will enhance our understanding of the role of natural selection in the evolution by gene duplication.

Animals↗

Distribution of the strength of selection against amino acid replacements in human proteins.

The impact of an amino acid replacement on the organism's fitness can vary from lethal to selectively neutral and even, in rare cases, beneficial. Substantial data are available on either pathogenic or acceptable replacements. However, the whole distribution of coefficients of selection against individual replacements is not known for any organism. To ascertain this distribution for human proteins, we combined data on pathogenic missense mutations, on human non-synonymous SNPs and on human-chimpanzee divergence of orthologous proteins. Fractions of amino acid replacements which reduce fitness by >10(-2), 10(-2)-10(-4), 10(-4)-10(-5) and <10(-5) are 25, 49, 14 and 12%, respectively. On average, the strength of selection against a replacement is substantially higher when chemically dissimilar amino acids are involved, and the Grantham's index of a replacement explains 35% of variance in the average logarithm of selection coefficients associated with different replacements. Still, the impact of a replacement depends on its context within the protein more than on its own nature. Reciprocal replacements are often associated with rather different selection coefficients, in particular, replacements of non-polar amino acids with polar ones are typically much more deleterious than replacements in the opposite direction. However, differences between evolutionary fluxes of reciprocal replacements are only weakly correlated with the differences between the corresponding selection coefficients.

Amino Acid Substitution↗

Persistence time of loss-of-function mutations at nonessential loci affecting eye color in Drosophila melanogaster.

Persistence time of a mutant allele, the expected number of generations before its elimination from the population, can be estimated as the ratio of the number of segregating mutations per individual over the mutation rate per generation. We screened two natural populations of Drosophila melanogaster for mutations causing clear-cut eye phenotypes and detected 25 mutant alleles, falling into 19 complementation groups, in 1164 haploid genomes, which implies 0.021 eye mutations/genome. The de novo haploid mutation rate for the same set of loci was estimated as 2 x 10(-4) in a 10-generation mutation-accumulation experiment. Thus, the average persistence time of all mutations causing clear-cut eye phenotypes is approximately 100 generations (95% confidence interval: 61-219). This estimate shows that the strength of selection against phenotypically drastic alleles of nonessential loci is close to that against recessive lethals. In both cases, deleterious alleles are apparently eliminated by selection against heterozygous individuals, which show no visible phenotypic differences from wild type.

Animals↗

A universal trend of amino acid gain and loss in protein evolution.

Amino acid composition of proteins varies substantially between taxa and, thus, can evolve. For example, proteins from organisms with (G + C)-rich (or (A + T)-rich) genomes contain more (or fewer) amino acids encoded by (G + C)-rich codons. However, no universal trends in ongoing changes of amino acid frequencies have been reported. We compared sets of orthologous proteins encoded by triplets of closely related genomes from 15 taxa representing all three domains of life (Bacteria, Archaea and Eukaryota), and used phylogenies to polarize amino acid substitutions. Cys, Met, His, Ser and Phe accrue in at least 14 taxa, whereas Pro, Ala, Glu and Gly are consistently lost. The same nine amino acids are currently accrued or lost in human proteins, as shown by analysis of non-synonymous single-nucleotide polymorphisms. All amino acids with declining frequencies are thought to be among the first incorporated into the genetic code; conversely, all amino acids with increasing frequencies, except Ser, were probably recruited late. Thus, expansion of initially under-represented amino acids, which began over 3,400 million years ago, apparently continues to this day.

AT Rich Sequence↗

Two classes of deleterious recessive alleles in a natural population of zebrafish, Danio rerio.

Natural populations carry deleterious recessive alleles which cause inbreeding depression. We compared mortality and growth of inbred and outbred zebrafish, Danio rerio, between 6 and 48 days of age. Grandparents of the studied fish were caught in the wild. Inbred fish were generated by brother-sister mating. Mortality was 9% in outbred fish, and 42% in inbred fish, which implies at least 3.6 lethal equivalents of deleterious recessive alleles per zygote. There was no significant inbreeding depression in the growth, perhaps because the surviving inbred fish lived under less crowded conditions. In contrast to alleles that cause embryonic and early larval mortality in the same population, alleles responsible for late larval and early juvenile mortality did not result in any gross morphological abnormalities. Thus, deleterious recessive alleles that segregate in a wild zebrafish population belong to two sharply distinct classes: early-acting, morphologically overt, unconditional lethals; and later-acting, morphologically cryptic, and presumably milder alleles.

Alleles↗

Positive selection at sites of multiple amino acid replacements since rat-mouse divergence.

New alleles become fixed owing to random drift of nearly neutral mutations or to positive selection of substantially advantageous mutations. After decades of debate, the fraction of fixations driven by selection remains uncertain. Within 9,390 genes, we analysed 28,196 codons at which rat and mouse differ from each other at two nucleotide sites and 1,982 codons with three differences. At codons where rat-mouse divergence involved two non-synonymous substitutions, both of them occurred in the same lineage, either rat or mouse, in 64% of cases; however, independent substitutions would occur in the same lineage with a probability of only 50%. All three non-synonymous substitutions occurred in the same lineage for 46% of codons, instead of the 25% expected. Furthermore, comparison of 12 pairs of prokaryotic genomes also shows clumping of multiple non-synonymous substitutions in the same lineage. This pattern cannot be explained by correlated mutation or episodes of relaxed negative selection, but instead indicates that positive selection acts at many sites of rapid, successive amino acid replacement.

Alleles↗

Bioinformatical assay of human gene morbidity.

Only a fraction of eukaryotic genes affect the phenotype drastically. We compared 18 parameters in 1273 human morbid genes, known to cause diseases, and in the remaining 16 580 unambiguous human genes. Morbid genes evolve more slowly, have wider phylogenetic distributions, are more similar to essential genes of Drosophila melanogaster, code for longer proteins containing more alanine and glycine and less histidine, lysine and methionine, possess larger numbers of longer introns with more accurate splicing signals and have higher and broader expressions. These differences make it possible to classify as non-morbid 34% of human genes with unknown morbidity, when only 5% of known morbid genes are incorrectly classified as non-morbid. This classification can help to identify disease-causing genes among multiple candidates.

Computational Biology↗

Context of deletions and insertions in human coding sequences.

We studied the dependence of the rate of short deletions and insertions on their contexts using the data on mutations within coding exons at 19 human loci that cause mendelian diseases. We confirm that periodic sequences consisting of three to five or more nucleotides are mutagenic. Mutability of sequences with strongly biased nucleotide composition is also elevated, even when mutations within homonucleotide runs longer than three nucleotides are ignored. In contrast, no elevated mutation rates have been detected for imperfect direct or inverted repeats. Among known candidate contexts, the indel context GTAAGT and regions with purine-pyrimidine imbalance between the two DNA strands are mutagenic in our sample, and many others are not mutagenic. Data on mutation hot spots suggest two novel contexts that increase the deletion rate. Comprehensive analysis of mutability of all possible contexts of lengths four, six, and eight indicates a substantially elevated deletion rate within YYYTG and similar sequences, which is one of the two contexts revealed by the hot spots. Possible contexts that increase the insertion rate (AT(A/C)(A/C)GCC and TACCRC) and decrease deletion (TATCGC) or insertion (GCGG) rates have also been identified. Two-thirds of deletions remove a repeat, and over 80% of insertions create a repeat, i.e., they are duplications.

Chromosome Deletion↗

Indel-based evolutionary distance and mouse-human divergence.

We propose a method for estimating the evolutionary distance between DNA sequences in terms of insertions and deletions (indels), defined as the per site number of indels accumulated in the course of divergence of the two sequences. We derive a maximal likelihood estimate of this distance from differences between lengths of orthologous introns or other segments of sequences delimited by conservative markers. When indels accumulate, lengths of orthologous introns diverge only slightly slower than linearly, because long indels occur with substantial frequencies. Thus, saturation is not a major obstacle for estimating indel-based evolutionary distance. For introns of medium lengths, our method recovers the known evolutionary distance between rat and mouse, 0.014 indels per site, with good precision. We estimate that mouse-human divergence exceeds rat-mouse divergence by a factor of 4, so that mouse-human evolutionary distance in terms of selectively neutral indels is 0.056. Because in mammals, indels are approximately 14 times less frequent than nucleotide substitutions, mouse-human evolutionary distance in terms of selectively neutral substitutions is approximately 0.8.

Animals↗

Sympatric speciation by sexual selection alone is unlikely.

According to Darwin, sympatric speciation is driven by disruptive, frequency-dependent natural selection caused by competition for diverse resources. Recently, several authors have argued that disruptive sexual selection can also cause sympatric speciation. Here, we use hypergeometric phenotypic and individual-based genotypic models to explore sympatric speciation by sexual selection under a broad range of conditions. If variabilities of preference and display traits are each caused by more than one or two polymorphic loci, sympatric speciation requires rather strong sexual selection when females exert preferences for extreme male phenotypes. Under this kind of mate choice, speciation can occur only if initial distributions of preference and display are close to symmetric. Otherwise, the population rapidly loses variability. Thus, unless allele replacements at very few loci are enough for reproductive isolation, female preferences for extreme male displays are unlikely to drive sympatric speciation. By contrast, similarity-based female preferences that do not cause sexual selection are less destabilizing to the maintenance of genetic variability and may result in sympatric speciation across a broader range of initial conditions. Certain groups of African cichlids have served as the exclusive motivation for the hypothesis of sympatric speciation by sexual selection. Mate choice in these fishes appears to be driven by female preferences for extreme male phenotypes rather than similarity-based preferences, and the evolution of premating reproductive isolation commonly involves at least several genes. Therefore, differences in female preferences and male display in cichlids and other species of sympatric origin are more likely to have evolved as isolating mechanisms under disruptive natural selection.

Animals↗

Patterns in interspecies similarity correlate with nucleotide composition in mammalian 3'UTRs.

Post-transcriptional regulation and the formation of mRNA 3' ends are crucial for gene expression in eukaryotes. Interspecies conservation of many sequences within 3'UTRs reveals selective constraint due to similar function. To study the pattern of conservation within 3'UTRs, we compiled and aligned 50 sets of complete orthologous 3'UTRs from four orders of mammals. We observed a mosaic pattern of conservation, with alternating regions of high (phylogenetic footprints) and low similarity. Conservation in 3'UTRs correlates with their base composition and also with the synonymous substitution rate in corresponding coding regions. The non-uniform distribution of conservation is more pronounced for 3'UTRs with a moderate or low level of overall conservation, where invariant nucleotides are more numerous, and their runs of lengths 4-7 occur more frequently than if conservation were random. Many runs of invariant nucleotides are AU-rich or pyrimidine-rich. Some of these runs coincide with known functional cis- elements of eukaryotic mRNAs, such as the U-rich upstream element, polyadenylation signal and DICE regulatory signal. More divergent regions of multiple alignments of 3'UTRs are often more G- and/or C-rich. Our results provide evidence on the importance of moderately conserved regions in 3'UTRs and suggest that regulatory functions of 3'UTRs might utilize gene-specific information in these regions.

3' Untranslated Regions↗