Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome evolution”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,369 records · Page 76Linked to original sources

Seeing chordate evolution through the Ciona genome sequence.

A draft sequence of the compact genome of the sea squirt Ciona intestinalis, a non-vertebrate chordate that diverged very early from other chordates, including vertebrates, illuminates how chordates originated and how vertebrate developmental innovations evolved.

Animals↗

Comparative genomics of the Archaea (Euryarchaeota): evolution of conserved protein families, the stable core, and the variable shell.

Comparative analysis of the protein sequences encoded in the four euryarchaeal species whose genomes have been sequenced completely (Methanococcus jannaschii, Methanobacterium thermoautotrophicum, Archaeoglobus fulgidus, and Pyrococcus horikoshii) revealed 1326 orthologous sets, of which 543 are represented in all four species. The proteins that belong to these conserved euryarchaeal families comprise 31%-35% of the gene complement and may be considered the evolutionarily stable core of the archaeal genomes. The core gene set includes the great majority of genes coding for proteins involved in genome replication and expression, but only a relatively small subset of metabolic functions. For many gene families that are conserved in all euryarchaea, previously undetected orthologs in bacteria and eukaryotes were identified. A number of euryarchaeal synapomorphies (unique shared characters) were identified; these are protein families that possess sequence signatures or domain architectures that are conserved in all euryarchaea but are not found in bacteria or eukaryotes. In addition, euryarchaea-specific expansions of several protein and domain families were detected. In terms of their apparent phylogenetic affinities, the archaeal protein families split into bacterial and eukaryotic families. The majority of the proteins that have only eukaryotic orthologs or show the greatest similarity to their eukaryotic counterparts belong to the core set. The families of euryarchaeal genes that are conserved in only two or three species constitute a relatively mobile component of the genomes whose evolution should have involved multiple events of lineage-specific gene loss and horizontal gene transfer. Frequently these proteins have detectable orthologs only in bacteria or show the greatest similarity to the bacterial homologs, which might suggest a significant role of horizontal gene transfer from bacteria in the evolution of the euryarchaeota.

Amino Acid Sequence↗

Retrotransposon evolution in diverse plant genomes.

Retrotransposon or retrotransposon-like sequences have been reported to be conserved components of cereal centromeres. Here we show that the published sequences are derived from a single conventional Ty3-gypsy family or a nonautonomous derivative. Both autonomous and nonautonomous elements are likely to have colonized Poaceae centromeres at the time of a common ancestor but have been maintained since by active retrotransposition. The retrotransposon family is also present at a lower copy number in the Arabidopsis genome, where it shows less pronounced localization. The history of the family in the two types of genome provides an interesting contrast between "boom and bust" and persistent evolutionary patterns.

Amino Acid Sequence↗

Intervening sequences in paralogous genes: a comparative genomic approach to study the evolution of X chromosome introns.

The enlargement of the genome size and the decrease in genome compactness with increase in the number and size of introns is a general pattern during the evolution of eukaryotes. Among the possible mechanisms for modifying intron size, it has been suggested that the insertion of transposable elements might have an important role in driving intron evolution. The analysis of large portions of the human genome demonstrated that a relatively recent (50 to 100 MYA) accumulation of transposable elements appears to be biased, favoring a preferential insertion of LINE1 transposons into sex chromosomes rather than into autosomes. In the present work, the effect of chromosomal location on the increase in size of introns was evaluated with a comparative analysis performed on pairs of human paralogous genes, one located on the X chromosome and the second on an autosome. A phylogenetic analysis was also performed on the X-encoded proteins and their paralogs to confirm orthology-paralogy and to approximately estimate the time of gene duplication. Statistical analysis of total intron length for each pair of paralogous genes provided no evidence for a larger size of introns in the gene copies located on the X chromosome. On the opposite, introns of autosomal genes were found to be significantly longer than introns of their X-linked paralogs. Likewise, LINE1 elements were not significantly more frequent in X-chromosome introns, whereas the frequency of SINE elements showed a marginally significant bias toward autosomal introns.

Animals↗

Contrasted modes of evolution in the same genome: allozymes and adaptive change in Heliconius.

Butterflies in the South American genus Heliconius have undergone a spectacular adaptive radiation (with convergent evolution between some lines) in their color patterns; this has been produced by natural selection for muellerian mimicry. The genetic basis of this radiation, shown by crossing highly differentiated races within two of the species, is homozygosity for alternative alleles at some half dozen loci. In complete contrast, allozyme loci in these butterflies are strongly heterozygous and show only frequency differences (never amounting to homozygosity of alternative alleles) between races; the amount of allozyme divergence is the same between races of H. erato and H. sara, although in color pattern the first forms marked races and the other does not. For the allozymes, there is a strong correlation over loci for rate of divergence between species and average heterozygosity. This is not true of the genes controlling color pattern. Heterozygosity of the enzymes is correlated with subunit molecular weight. Thus, different parts of the genome can evolve in different ways simultaneously; genes controlling color pattern in the "classical" mode, and allozymes in a different mode in which the rate of evolution is related to their heterozygosity (a "balance" or "neutral" mode).

Animals↗

Clustering of genes coding for DNA binding proteins in a region of atypical evolution of the human genome.

Comparison of the human and mouse genomes has revealed that significant variations in evolutionary rates exist among genomic regions and that a large part of this variation is interchromosomal. We confirm in this work, using a large collection of introns, that human chromosome 19 is the one that shows the highest divergence with respect to mouse. To search for other differences among chromosomes, we examine the distribution of gene functions in human and mouse chromosomes using the Gene Ontology definitions. We found by correspondence analysis that among the strongest clusterings of gene functions in human chromosomes is a group of genes coding for DNA binding proteins in chromosome 19. Interestingly, chromosome 19 also has a very high GC content, a feature that has been proposed to promote an opening of the chromatin, thereby facilitating binding of proteins to the DNA helix. In the mouse genome, however, a similar aggregation of genes coding for DNA binding proteins and high GC content cannot be found. This suggests that the distribution of genes coding for DNA binding proteins and the variations of the chromatin accessibility to these proteins are different in the human and mouse genomes. It is likely that the overall high synonymous and intron rates in chromosome 19 are a by-product of the high GC content of this chromosome.

Base Composition↗

Interphase cell flow cytometry as a means of monitoring genomic size in normal and neoplastoid cell cultures.

The DNA specific fluorescence of mass cultures and clones derived from human skin and bladder tumor tissue was assayed by flow cytometry. In order to detect and quantitate small fluorescence intensity changes, cytogenetically defined triploid or diploid human fibroblast strains were cocultivated, harvested, and stained with the cell strain of unknown karyotype. The triploid standard (derived from human abortus tissue) proved chromosomally unstable at high passage level. Fifteen male, female, and 45,X strains displayed target-to-standard cell fluorescence ratios commensurate with their respective chromosome constitutions. Interstrain variation was highest among the 45,X strains, although mosaicism could not be detected by conventional cytogenetics. Interclonal fluorescence variation was two- to ten-fold higher among the tumor-derived clones tested. Chromosome counts and subcloning experiments indicate that this increased fluorescence variation is due to genome size variation. The clonal evolution of genome size differences was observed in subclones of chromosomally divergent parental clones. These observations suggest that well controlled flow cytometry can adequately resolve subtle degrees of genome size variation in cultivated human cells. The technique is especially suited for monitoring genome size changes in cultivated tumor cells.

Cell Separation↗

OrthoMCL: identification of ortholog groups for eukaryotic genomes.

The identification of orthologous groups is useful for genome annotation, studies on gene/protein evolution, comparative genomics, and the identification of taxonomically restricted sequences. Methods successfully exploited for prokaryotic genome analysis have proved difficult to apply to eukaryotes, however, as larger genomes may contain multiple paralogous genes, and sequence information is often incomplete. OrthoMCL provides a scalable method for constructing orthologous groups across multiple eukaryotic taxa, using a Markov Cluster algorithm to group (putative) orthologs and paralogs. This method performs similarly to the INPARANOID algorithm when applied to two genomes, but can be extended to cluster orthologs from multiple species. OrthoMCL clusters are coherent with groups identified by EGO, but improved recognition of "recent" paralogs permits overlapping EGO groups representing the same gene to be merged. Comparison with previously assigned EC annotations suggests a high degree of reliability, implying utility for automated eukaryotic genome annotation. OrthoMCL has been applied to the proteome data set from seven publicly available genomes (human, fly, worm, yeast, Arabidopsis, the malaria parasite Plasmodium falciparum, and Escherichia coli). A Web interface allows queries based on individual genes or user-defined phylogenetic patterns (http://www.cbil.upenn.edu/gene-family). Analysis of clusters incorporating P. falciparum genes identifies numerous enzymes that were incompletely annotated in first-pass annotation of the parasite genome.

Animals↗

Toward a "new" paradigm of therapeutic action: neuro-psychoanalysis and downward causation.

Freud's metapsychological assumption, which split the mind from the brain, is increasingly recognized as limiting the growth of psychoanalysis and its integration with other fields, including psychiatry. The dual-aspect monist position, sometimes used to rationalize the mind/body split, is seen to contain a mereological (category) error, which can be avoided by the introduction of the paradigm of psychodynamic science, and by the use of a nondualistic, symbiotic mind/brain formulation, well justified by the study of cultural and organic evolution. Therapeutic action may now be seen as a special case of development that occurs within the interactional context of cultural evolution, personal history, and the genome. Human evolution is increasingly dominated by culture operating on the mind/brain through downward causation. This concept refers to the influence of higher organizational entities (e.g., mind) upon lower ones (e.g., brain). Although little recognized, downward causation is tacitly assumed in psychotherapeutic interventions, and is illustrated in recent fMRI studies. The clinical integration of the downward causation concept links therapeutic action to the power of cultural evolution, and facilitates reunion with traditional science.

Biological Evolution↗

Evidence of interaction network evolution by whole-genome duplications: a case study in MADS-box proteins.

Recent investigations on metazoan transcription factors (TFs) indicate that single-gene duplication events and the gain and loss of protein domains are 2 crucial factors in shaping their protein-protein interaction networks. Plant genomes, on the other hand, have a history of polyploidy and whole-genome duplications (WGDs), and thus, their study helps to understand whether WGDs have also had a significant influence on protein network evolution. Here we investigate the evolution of the interaction network in the well-studied MADS domain MIKC-type proteins, a TF family which plays an important role in both the vegetative and the reproductive phases of plant life. We combine phylogenetic reconstruction, protein domain analysis, and interaction data from different species. We show that, unlike previously analyzed interaction networks, the MIKC-type protein network displays a characteristic topology, with overall high inter-subfamily connectivity, shared interactors between paralogs, and conservation of interaction patterns across species. The evaluation of the number of MIKC-type proteins at key time points throughout the evolution of land plants in the lineage leading to Arabidopsis suggested that most duplicates were retained after each round of WGD. We provide evidence that an initial network, formed by 9-11 homodimerizing proteins interacting with each other, existed in the common ancestor of all seed plants. This basic structure has been conserved after each round of WGD, adding layers of paralogs with similar interaction patterns. We thus present the first model where we can show that a network of eukaryotic TFs has evolved via rounds of WGD. Furthermore, we found that in subfamilies in which the K domain is most diverged, the interactions with other subfamilies have been largely lost. We discuss the possibility that such a high proportion of genes were retained after each WGD because of their capacity to form higher order complexes involving proteins from different subfamilies. The simultaneous duplications allowed for the conservation of the quantitative balance between the constituents and facilitated sub- and neofunctionalization through differential expression of whole units.

Computational Biology↗

Origins and evolution of the Europeans' genome: evidence from multiple microsatellite loci.

There is general agreement that the current European gene pool is mainly derived from Palaeolithic hunting-gathering and Neolithic farming ancestors, but different studies disagree on the relative weight of these contributions. We estimated admixture rates in European populations from data on 377 autosomal microsatellite loci in 235 individuals, using five different numerical methods. On average, the Near Eastern (and presumably Neolithic) contribution was between 46 and 66%, and admixture estimates showed, with all methods, a strong and significant negative correlation with distance from the Near East. If the assumptions of the model are approximately correct, i.e. if the Basques' and Near Easterners' genomes represent a good approximation to the Palaeolithic and Neolithic settlers of Europe, respectively, these results imply that half or more of the Europeans' genes are descended from Near Eastern ancestors who immigrated in Europe 10000 years ago. If these assumptions are incorrect, our results show anyway that clinal variation is the rule in the Europeans' genomes and that lower estimates of Near Eastern admixture obtained from the analysis of single markers do not reflect the patterns observed at the genomic level.

Europe↗

Evolution of the chloroplast genome.

We discuss the suggestion that differences in the nucleotide composition between plastid and nuclear genomes may provide a selective advantage in the transposition of genes from plastid to nucleus. We show that in the adenine, thymine (AT)-rich genome of Borrelia burgdorferi several genes have an AT-content lower than the average for the genome as a whole. However, genes whose plant homologues have moved from plastid to nucleus are no less AT-rich than genes whose plant homologues have remained in the plastid, indicating that both classes of gene are able to support a high AT-content. We describe the anomalous organization of dinoflagellate plastid genes. These are located on small circles of 2-3 kbp, in contrast to the usual plastid genome organization of a single large circle of 100-200 kbp. Most circles contain a single gene. Some circles contain two genes and some contain none. Dinoflagellate plastids have retained far fewer genes than other plastids. We discuss a similarity between the dinoflagellate minicircles and the bacterial integron system.

Amino Acid Sequence↗

Evolution and variation of the SARS-CoV genome.

Knowledge of the evolution of pathogens is of great medical and biological significance to the prevention, diagnosis, and therapy of infectious diseases. In order to understand the origin and evolution of the SARS-CoV (severe acute respiratory syndrome-associated coronavirus), we collected complete genome sequences of all viruses available in GenBank, and made comparative analyses with the SARS-CoV. Genomic signature analysis demonstrates that the coronaviruses all take the TGTT as their richest tetranucleotide except the SARS-CoV. A detailed analysis of the forty-two complete SARS-CoV genome sequences revealed the existence of two distinct genotypes, and showed that these isolates could be classified into four groups. Our manual analysis of the BLASTN results demonstrates that the HE (hemagglutinin-esterase) gene exists in the SARS-CoV, and many mutations made it unfamiliar to us.

Amino Acid Motifs↗

Using hidden Markov models and observed evolution to annotate viral genomes.

MOTIVATION: ssRNA (single stranded) viral genomes are generally constrained in length and utilize overlapping reading frames to maximally exploit the coding potential within the genome length restrictions. This overlapping coding phenomenon leads to complex evolutionary constraints operating on the genome. In regions which code for more than one protein, silent mutations in one reading frame generally have a protein coding effect in another. To maximize coding flexibility in all reading frames, overlapping regions are often compositionally biased towards amino acids which are 6-fold degenerate with respect to the 64 codon alphabet. Previous methodologies have used this fact in an ad hoc manner to look for overlapping genes by motif matching. In this paper differentiated nucleotide compositional patterns in overlapping regions are incorporated into a probabilistic hidden Markov model (HMM) framework which is used to annotate ssRNA viral genomes. This work focuses on single sequence annotation and applies an HMM framework to ssRNA viral annotation. A description of how the HMM is parameterized, whilst annotating within a missing data framework is given. A Phylogenetic HMM (Phylo-HMM) extension, as applied to 14 aligned HIV2 sequences is also presented. This evolutionary extension serves as an illustration of the potential of the Phylo-HMM framework for ssRNA viral genomic annotation. RESULTS: The single sequence annotation procedure (SSA) is applied to 14 different strains of the HIV2 virus. Further results on alternative ssRNA viral genomes are presented to illustrate more generally the performance of the method. The results of the SSA method are encouraging however there is still room for improvement, and since there is overwhelming evidence to indicate that comparative methods can improve coding sequence (CDS) annotation, the SSA method is extended to a Phylo-HMM to incorporate evolutionary information. The Phylo-HMM extension is applied to the same set of 14 HIV2 sequences which are pre-aligned. The performance improvement that results from including the evolutionary information in the analysis is illustrated.

Algorithms↗

A 39-kb sequence around a blackbird Mhc class II gene: ghost of selection past and songbird genome architecture.

To gain an understanding of the evolution and genomic context of avian major histocompatibility complex (Mhc) genes, we sequenced a 38.8-kb Mhc-bearing cosmid insert from a red-winged blackbird (Agelaius phoeniceus). The DNA sequence, the longest yet retrieved from a bird other than a chicken, provides a detailed view of the process of gene duplication, divergence, and degeneration ("birth and death") in the avian Mhc, as well as a glimpse into major noncoding features of a songbird genome. The peptide-binding region (PBR) of the single Mhc class II B gene in this region, Agph-DAB2, is almost devoid of polymorphism, and a still-segregating single-base-pair deletion and other features suggest that it is nonfunctional. Agph-DAB2 is estimated to have diverged about 40 MYA from a previously characterized and highly polymorphic blackbird Mhc gene, Aph-DAB1, and is therefore younger than most mammalian Mhc paralogs and arose relatively late in avian evolution. Despite its nonfunctionality, Agph-DAB2 shows very high levels of nonsynonymous divergence from Agph-DAB1 and from reconstructed ancestral sequences in antigen-binding PBR codons-a strong indication of a period of adaptive divergence preceding loss of function. We also found that the region sequenced contains very few other unambiguous genes, a partial Mhc- class II gene fragment, and a paucity of simple-sequence and other repeats. Thus, this sequence exhibits some of the genomic streamlining expected for avian as compared with mammalian genomes, but is not as densely packed with functional genes as is the chicken Mhc.

Amino Acid Sequence↗