Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome evolution”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

Transposable elements and the evolution of genome organization in mammals.

All mammalian transposable elements characterized to date appear to be nonrandomly distributed in the mammalian genome. While no element has been found to be exclusively restricted in its chromosomal location, LINE elements and some retrovirus-like elements are preferentially accumulated in G-banding regions of the chromosomes, and in some cases in the sex chromosomes, while SINE elements occur preferentially in R-banding regions. Four mechanisms are presented which may explain the nonrandom genomic distribution of mammalian transposons: i) sequence-specific insertion, ii) S-phase insertion, iii) ectopic excision, and iv) recombinational editing. Some of the available data are consistent with each of these four models, but no single model is sufficient to explain all of the existing data.

Animals↗

Mycodnaviridae is a clade of giant viruses that persistently infect zoosporic fungi.

Giant viruses of the phylum Nucleocytoviricota have emerged as particularly notable due to their increasingly recognized impacts on eukaryotic genome evolution. Their origins are hypothesized to predate or coincide with the diversification of eukaryotes, and they have been detected in hosts that span the eukaryotic tree of life. But surprisingly, such viruses have not been definitively found in Kingdom Fungi, though earlier genomic and metagenomic work suggests putative associations. Here we report both "viral fossils" and active infection by giant viruses in fungi, particularly in the zoosporic phyla Blastocladiomycota and Chytridiomycota. The recovered viral assemblies span up to 350 kb, encode over 300 genes, and form a monophyletic family-level clade within the Nucleocytoviricota related to orders Imitervirales and Algavirales, which we name Mycodnaviridae. We observed variation in infection status among the isolates including apparent active infection and transcriptionally suppressed states, suggesting that viral activation may be constrained to certain life stages of the host. Our experimental findings add to the limited natural virus-host systems available in culture for the study of giant viruses and expand the known host range of Nucleocytoviricota into a new kingdom that contains many model species. Mycodnaviridae have a global distribution, which invites inquiry into the implications of these infections for host traits, host genome evolution, and the metabolic impacts on ecosystems.

Giant Viruses↗

Organization and sequence of five tRNA genes and of an unidentified reading frame in the wheat chloroplast genome: evidence for gene rearrangements during the evolution of chloroplast genomes.

The genes for the initiator tRNA(Met)CAU, tRNA(Gly)UCC, tRNA(Thr)GGU, tRNA(Glu)UUC and tRNA(Tyr)GUA and an open reading frame of 62 codons have been identified by sequencing a 2,358 bp BamHI and a 1,378 bp BamHI-Sst2 DNA fragments from wheat chloroplasts. A comparison of the organization of these five tRNA genes and of the open reading frame on the wheat, tobacco and spinach chloroplast genomes suggests that at least three genomic inversions must have occurred during the evolution of the wheat chloroplast genome from a spinach-like ancestor genome. Furthermore, it seems that in wheat the 91 bp intergenic region between the genes for the initiator tRNA(Met) and the gene for tRNA(Gly)UCC is one end-point of the 20 kbp genomic inversion proposed by Palmer and Thompson in the case of maize (Palmer and Thompson 1982). A 119 bp duplication is located at this junction: the first copy comprises the 91 bp of the intergenic region and the first 28 bp of the tRNA(Met) gene, the second copy is found downstream of the tRNA(Met) gene.

Base Sequence↗

Evolution of genome organizations of squirrels (Sciuridae) revealed by cross-species chromosome painting.

With complete sets of chromosome-specific painting probes derived from flow-sorted chromosomes of human and grey squirrel (Sciurus carolinensis), the whole genome homologies between human and representatives of tree squirrels (Sciurus carolinensis, Callosciurus erythraeus), flying squirrels (Petaurista albiventer) and chipmunks (Tamias sibiricus) have been defined by cross-species chromosome painting. The results show that, unlike the highly rearranged karyotypes of mouse and rat, the karyotypes of squirrels are highly conserved. Two methods have been used to reconstruct the genome phylogeny of squirrels with the laboratory rabbit (Oryctolagus cuniculus) as the out-group: (1) phylogenetic analysis by parsimony using chromosomal characters identified by comparative cytogenetic approaches; (2) mapping the genome rearrangements onto recently published sequence-based molecular trees. Our chromosome painting results, in combination with molecular data, show that flying squirrels are phylogenetically close to New World tree squirrels. Chromosome painting and G-banding comparisons place chipmunks (Tamias sibiricus ), with a derived karyotype, outside the clade comprising tree and flying squirrels. The superorder Glires (orde Rodentia + order Lagomorpha) is firmly supported by two conserved syntenic associations between human chromosomes 1 and 10p homologues, and between 9 and 11 homologues.

Animals↗

Rapid genome divergence at orthologous low molecular weight glutenin loci of the A and Am genomes of wheat.

To study genome evolution in wheat, we have sequenced and compared two large physical contigs of 285 and 142 kb covering orthologous low molecular weight (LMW) glutenin loci on chromosome 1AS of a diploid wheat species (Triticum monococcum subsp monococcum) and a tetraploid wheat species (Triticum turgidum subsp durum). Sequence conservation between the two species was restricted to small regions containing the orthologous LMW glutenin genes, whereas >90% of the compared sequences were not conserved. Dramatic sequence rearrangements occurred in the regions rich in repetitive elements. Dating of long terminal repeat retrotransposon insertions revealed different insertion events occurring during the last 5.5 million years in both species. These insertions are partially responsible for the lack of homology between the intergenic regions. In addition, the gene space was conserved only partially, because different predicted genes were identified on both contigs. Duplications and deletions of large fragments that might be attributable to illegitimate recombination also have contributed to the differentiation of this region in both species. The striking differences in the intergenic landscape between the A and A(m) genomes that diverged 1 to 3 million years ago provide evidence for a dynamic and rapid genome evolution in wheat species.

Base Sequence↗

Shallow genomics, phylogenetics, and evolution in the family Drosophilidae.

The effects of the genomic revolution are beginning to be felt in all disciplines of the biological sciences. Evolutionary biology in general, and phylogenetic systematics in particular, are being revolutionized by these advances. The advent of rapid nucleotide sequencing techniques have provided phylogenetic biologists with the tools required to quickly and efficiently generate large amounts of character information. We use family Drosophilidae as a model system to study phylogenetics and genome evolution by combining high throughput sequencing methods from the field genomics and standard phylogenetic methodology. This paper presents preliminary results from this work. Separate data partitions, based on either gene function or linkage group, are compared to a combined analysis of all the data to assess support on phylogenetic trees.

Animals↗

The Irx gene family in zebrafish: genomic structure, evolution and initial characterization of irx5b.

Genes of the iroquois ( Iro/Irx) family are highly conserved from Drosophila to mammals and they have been implicated in a number of developmental processes. In flies, the Iro genes participate in patterning events in the early larva and in imaginal disk specification. In vertebrates, the Irx genes regulate developmental events during gastrulation, nervous system regionalization, activation of proneural genes and organ patterning. The Iro genes in Drosophila and the Irx genes of mammals show a clustered organization in the genome. Flies have a single cluster comprising three genes while mammals have two clusters also having three genes each. Moreover, experimental evidence in flies shows that transcriptional regulatory elements are shared among genes within the Iro cluster, suggesting that the same may be true in vertebrates. To date, the genomic organization of the Irx genes in non-mammalian species has not been studied. In this work, we have isolated the irx5b gene from zebrafish, Danio rerio, and have characterized its expression pattern. Furthermore, we have identified the complete set of Irx genes in two fish species, the zebrafish and pufferfish, Takifugu rubripes, and have determined the genomic organization of these genes. Our analysis indicates that early in fish evolutionary history, the Irx gene clusters have been duplicated and that subsequent events have maintained the clustered organization for some of the genes, while others have been lost. In total there are 11 existing Irx genes in zebrafish and 10 in pufferfish. We propose a new nomenclature for the zebrafish Irx genes based on the analysis of their sequences and their genomic relationships.

Animals↗

Correlational selection and the evolution of genomic architecture.

We review and discuss the importance of correlational selection (selection for optimal character combinations) in natural populations. If two or more traits subject to multivariate selection are heritable, correlational selection builds favourable genetic correlations through the formation of linkage disequilibrium at underlying loci governing the traits. However, linkage disequilibria built up by correlational selection are expected to decay rapidly (ie, within a few generations), unless correlational selection is strong and chronic. We argue that frequency-dependent biotic interactions that have 'Red Queen dynamics' (eg, host-parasite interactions, predator-prey relationships or intraspecific arms races) often fuel chronic correlational selection, which is strong enough to maintain adaptive genetic correlations of the kind we describe. We illustrate these processes and phenomena using empirical examples from various plant and animal systems, including our own recent work on the evolutionary dynamics of a heritable throat colour polymorphism in the side-blotched lizard Uta stansburiana. In particular, male and female colour morphs of side-blotched lizards cycle on five- and two-generation (year) timescales under the force of strong frequency-dependent selection. Each morph refines the other morph in a Red Queen dynamic. Strong correlational selection gradients among life history, immunological and morphological traits shape the genetic correlations of the side-blotched lizard polymorphism. We discuss the broader evolutionary consequences of the buildup of co-adapted trait complexes within species, such as the implications for speciation processes.

Animals↗

Birth and death of protein domains: a simple model of evolution explains power law behavior.

BACKGROUND: Power distributions appear in numerous biological, physical and other contexts, which appear to be fundamentally different. In biology, power laws have been claimed to describe the distributions of the connections of enzymes and metabolites in metabolic networks, the number of interactions partners of a given protein, the number of members in paralogous families, and other quantities. In network analysis, power laws imply evolution of the network with preferential attachment, i.e. a greater likelihood of nodes being added to pre-existing hubs. Exploration of different types of evolutionary models in an attempt to determine which of them lead to power law distributions has the potential of revealing non-trivial aspects of genome evolution. RESULTS: A simple model of evolution of the domain composition of proteomes was developed, with the following elementary processes: i) domain birth (duplication with divergence), ii) death (inactivation and/or deletion), and iii) innovation (emergence from non-coding or non-globular sequences or acquisition via horizontal gene transfer). This formalism can be described as a birth, death and innovation model (BDIM). The formulas for equilibrium frequencies of domain families of different size and the total number of families at equilibrium are derived for a general BDIM. All asymptotics of equilibrium frequencies of domain families possible for the given type of models are found and their appearance depending on model parameters is investigated. It is proved that the power law asymptotics appears if, and only if, the model is balanced, i.e. domain duplication and deletion rates are asymptotically equal up to the second order. It is further proved that any power asymptotic with the degree not equal to -1 can appear only if the hypothesis of independence of the duplication/deletion rates on the size of a domain family is rejected. Specific cases of BDIMs, namely simple, linear, polynomial and rational models, are considered in details and the distributions of the equilibrium frequencies of domain families of different size are determined for each case. We apply the BDIM formalism to the analysis of the domain family size distributions in prokaryotic and eukaryotic proteomes and show an excellent fit between these empirical data and a particular form of the model, the second-order balanced linear BDIM. Calculation of the parameters of these models suggests surprisingly high innovation rates, comparable to the total domain birth (duplication) and elimination rates, particularly for prokaryotic genomes. CONCLUSIONS: We show that a straightforward model of genome evolution, which does not explicitly include selection, is sufficient to explain the observed distributions of domain family sizes, in which power laws appear as asymptotic. However, for the model to be compatible with the data, there has to be a precise balance between domain birth, death and innovation rates, and this is likely to be maintained by selection. The developed approach is oriented at a mathematical description of evolution of domain composition of proteomes, but a simple reformulation could be applied to models of other evolving networks with preferential attachment.

Animals↗

A novel allelic variant of serum amyloid A, SAA1 gamma: genomic evidence, evolution, frequency, and implication as a risk factor for reactive systemic AA-amyloidosis.

Reactive systemic amyloidosis, also called AA-amyloidosis is a rare fatal complication of common chronic inflammatory diseases such as rheumatoid arthritis. It has been proposed that as yet undefined factors other than persistent elevation of serum level of the precursor protein, serum amyloid A (SAA), are also important for the development of AA-amyloidosis. In this work we show genomic evidence for a novel allelic variant of human SAA, SAA1 gamma, which we have recently identified at the protein level. The SAA1 gamma [Ala52(GCC), Ala57(GCG)] differed from SAA1 alpha [Val52(GTC), Ala57(GCG)] only at one base, indicating a single point mutation. On the other hand, SAA1 beta [Ala52(GCC), Val57(GTG)] had not only one, but additional differences in a nearby intron and this portion was identical to the SAA2 gene, suggesting a crossing-over between the SAA1 and SAA2 genes. Furthermore, we report that there was a significant difference in the observed numbers of SAA1 alleles between rheumatoid arthritis patients with AA-amyloidosis and the control population (chi 2(2) = 11.59, p = 0.003) with a higher frequency of gamma-allele in the AA-amyloid group (0.70 vs. 0.37). There was also a notable difference in the distribution of SAA1 genotypes (chi 5(2) = 14.63, p = 0.012) with an increased frequency of gamma/gamma-homozygotes in the AA-amyloid group (0.60 vs. 0.18). Thus our findings indicate that this novel allelic variant may be an important risk factor for the development of AA-amyloidosis.

Adult↗

New approaches for reconstructing phylogenies from gene order data.

We report on new techniques we have developed for reconstructing phylogenies on whole genomes. Our mathematical techniques include new polynomial-time methods for bounding the inversion length of a candidate tree and new polynomial-time methods for estimating genomic distances which greatly improve the accuracy of neighbor-joining analyses. We demonstrate the power of these techniques through an extensive performance study based on simulating genome evolution under a wide range of model conditions. Combining these new tools with standard approaches (fast reconstruction with neighbor-joining, exploration of all possible refinements of strict consensus trees, etc.) has allowed us to analyze datasets that were previously considered computationally impractical. In particular, we have conducted a complete phylogenetic analysis of a subset of the Campanulaceae family, confirming various conjectures about the relationships among members of the subset and about the principal mechanism of evolution for their chloroplast genome. We give representative results of the extensive experimentation we conducted on both real and simulated datasets in order to validate and characterize our approaches. We find that our techniques provide very accurate reconstructions of the true tree topology even when the data are generated by processes that include a significant fraction of transpositions and when the data are close to saturation.

Biological Evolution↗

Comparative genomics and evolution of proteins associated with RNA polymerase II C-terminal domain.

The C-terminal domain (CTD) of the largest subunit of RNA polymerase II provides an anchoring point for a wide variety of proteins involved in mRNA synthesis and processing. Most of what is known about CTD-protein interactions comes from animal and yeast models. The consensus sequence and repetitive structure of the CTD is conserved strongly across a wide range of organisms, implying that the same is true of many of its known functions. In some eukaryotic groups, however, the CTD has been allowed to degenerate, suggesting a comparable lack of essential protein interactions. To date, there has been no comprehensive examination of CTD-related proteins across the eukaryotic domain to determine which of its identified functions are correlated with strong stabilizing selection on CTD structure. Here we report a comparative investigation of genes encoding 50 CTD-associated proteins, identifying putative homologs from 12 completed or nearly completed eukaryotic genomes. The presence of a canonical CTD generally is correlated with the apparent presence and conservation of its known protein partners; however, no clear set of interactions emerges that is invariably linked to conservation of the CTD. General rates of evolution, phylogenetic patterns, and the conservation of modeled tertiary structure of capping enzyme guanylyltransferase (Cgt1) indicate a pattern of coevolution of components of a transcription factory organized around the CTD, presumably driven by common functional constraints. These constraints complicate efforts to determine orthologous gene relationships and can mislead phylogenetic and informatic algorithms.

Amino Acid Sequence↗

Complete telomere-to-telomere genome assembly of Guazuma ulmifolia uncovers evolutionary mechanisms, drought adaptation, and flavonoid biosynthesis.

The first T2T reference genome of Guazuma ulmifolia is reported, which serves as a core genomic resource for stress adaptation research and stress-tolerant breeding in cacao wild relatives. Climate change, particularly increased incidence of drought, poses a major threat to food security. Understanding the genomic basis of environmental adaptation in crop wild relatives can provide valuable resources for improving stress resilience. Guazuma ulmifolia, a wild relative of Theobroma cacao with important ecological and medicinal value, lacks high-quality reference genomic resources. Here, we report the first telomere-to-telomere (T2T) chromosome-level genome assembly of G. ulmifolia, with a genome size of 311.31 Mb, contig N50 of 35.19 Mb, and 98.70% BUSCO completeness. Repetitive sequences constitute 27.43% of the G. ulmifolia genome, with LTR retrotransposons as the predominant class. Comparative genomic analyses revealed that genome-size variation among Malvaceae species is associated with differences in polyploidization history and TE dynamics. Ancestral karyotype reconstruction identified five lineage-specific chromosome fusion events distinguishing G. ulmifolia from T. cacao. Comparative analyses further identified tandem duplication-associated expansion of stress-related LEA and GST gene families, suggesting potential genomic features associated with stress responses. Flavonoid biosynthesis genes were largely conserved in copy number but showed tissue-specific expression patterns, providing candidate genes for investigating secondary metabolism. Together, this study establishes a high-quality T2T genome resource for exploring genome evolution, chromosome organization, and stress-related genomic features in Malvaceae.

Genome, Plant↗

Complete mitochondrial DNA sequences of six snakes: phylogenetic relationships and molecular evolution of genomic features.

Complete mitochondrial DNA (mtDNA) sequences were determined for representative species from six snake families: the acrochordid little file snake, the bold boa constrictor, the cylindrophiid red pipe snake, the viperid himehabu, the pythonid ball python, and the xenopeltid sunbeam snake. Thirteen protein-coding genes, 22 tRNA genes, 2 rRNA genes, and 2 control regions were identified in these mtDNAs. Duplication of the control region and translocation of the tRNALeu gene were two notable features of the snake mtDNAs. The duplicate control regions had nearly identical nucleotide sequences within species but they were divergent among species, suggesting concerted sequence evolution of the two control regions. In addition, the duplicate control regions appear to have facilitated an interchange of some flanking tRNA genes in the viperid lineage. Phylogenetic analyses were conducted using a large number of sites (9570 sites in total) derived from the complete mtDNA sequences. Our data strongly suggested a new phylogenetic relationship among the major families of snakes: ((((Viperidae, Colubridae), Acrochordidae), (((Pythonidae, Xenopeltidae), Cylindrophiidae), Boidae)), Leptotyphlopidae). This conclusion was distinct from a widely accepted view based on morphological characters in denying the sister-group relationship of boids and pythonids, as well as the basal divergence of nonmacrostomatan cylindrophiids. These results imply the significance to reconstruct the snake phylogeny with ample molecular data, such as those from complete mtDNA sequences.

Animals↗

Recombination in evolutionary genomics.

Recombination can be a dominant force in shaping genomes and associated phenotypes. To better understand the impact of recombination on genomic evolution, we need to be able to identify recombination in aligned sequences. We review bioinformatic approaches for detecting recombination and measuring recombination rates. We also examine the impact of recombination on the reconstruction of evolutionary histories and the estimation of population genetic parameters. Finally, we review the role of recombination in the evolutionary history of bacteria, viruses, and human mitochondria. We conclude by highlighting a number of areas for future development of tools to help quantify the role of recombination in genomic evolution.

Base Sequence↗

Domain organization, genomic structure, evolution, and regulation of expression of the aggrecan gene family.

Proteoglycans are complex macromolecules, consisting of a polypeptide backbone to which are covalently attached one or more glycosaminoglycan chains. Molecular cloning has allowed identification of the genes encoding the core proteins of various proteoglycans, leading to a better understanding of the diversity of proteoglycan structure and function, as well as to the evolution of a classification of proteoglycans on the basis of emerging gene families that encode the different core proteins. One such family includes several proteoglycans that have been grouped with aggrecan, the large aggregating chondroitin sulfate proteoglycan of cartilage, based on a high number of sequence similarities within the N- and C-terminal domains. Thus far these proteoglycans include versican, neurocan, and brevican. It is now apparent that these proteins, as a group, are truly a gene family with shared structural motifs on the protein and nucleotide (mRNA) levels, and with nearly identical genomic organizations. Clearly a common ancestral origin is indicated for the members of the aggrecan family of proteoglycans. However, differing patterns of amplification and divergence have also occurred within certain exons across species and family members, leading to the class-characteristic protein motifs in the central carbohydrate-rich region exclusively. Thus the overall domain organization strongly suggests that sequence conservation in the terminal globular domains underlies common functions, whereas differences in the central portions of the genes account for functional specialization among the members of this gene family.

Aggrecans↗

Comparative genomic characterization and antimicrobial resistance of bacteremia-causing Enterococcus faecium and Enterococcus faecalis in a Chinese hospital.

Enterococci are common commensals of the human gut and important opportunistic pathogens, with Enterococcus faecium and Enterococcus faecalis being the most clinically prevalent species. A significant epidemiological shift has emerged with an increasing clinical burden of E. faecium. To compare genomic evolution of E. faecium and E. faecalis, we performed whole-genome sequencing on 93 E. faecium and 32 E. faecalis isolates causing bloodstream infections at a single hospital (2022-2024). Analysis of patient demographics revealed that E. faecium infections originated from fewer sources than E. faecalis, with a higher proportion deriving from intra-abdominal infections. Multilocus sequence typing identified ST78 and ST789 as the predominant sequence types for E. faecium, whereas ST16 and ST179 were most common for E. faecalis. E. faecium carried more antimicrobial resistance genes and putative virulence marker (PVM)-type virulence genes than E. faecalis, with vancomycin resistance predominantly mediated by vanHAX (33/93, 35.5%) and a single E. faecalis isolate also carrying vanHAX (1/32, 3.1%); the structurally incomplete vanHMX gene cluster was detected in 11 E. faecium isolates. Pan-genome analysis indicated a larger core genome in E. faecalis compared to E. faecium, consistent with greater plasmid replicon diversity in the latter. Intra-host comparisons showed that two E. faecalis pairs from the same patient were clonally related, with one isolate acquiring a vanHAX plasmid conferring vancomycin resistance. In contrast, E. faecium isolates exhibited marked genomic diversity even among clonally related pairs. These findings suggest that E. faecium possesses greater genomic plasticity and adaptive potential to the clinical environment.IMPORTANCEThis study provides a detailed comparison of clinical and genomic features between Enterococcus faecium and Enterococcus faecalis from the same hospital setting. We show that E. faecium isolates, mainly ST78/ST789, carry more antimicrobial resistance genes and a higher number of putative virulence marker (PVM) genes than E. faecalis, reflecting their hospital-adapted nature. E. faecium also exhibits a smaller core genome and greater diversity of plasmid replicon types, indicating higher genomic plasticity and capacity for horizontal gene transfer. By contrast, E. faecalis retains a larger core genome and a set of classical virulence factors, and its within-host isolates are clonally related. These distinct genomic profiles help to understand how the two species adapt to clinical environments and may inform more targeted infection control strategies and resistance surveillance.

Enterococcus faecium↗

Helicobacter acinonychis: genetic and rodent infection studies of a Helicobacter pylori-like gastric pathogen of cheetahs and other big cats.

Insights into bacterium-host interactions and genome evolution can emerge from comparisons among related species. Here we studied Helicobacter acinonychis (formerly H. acinonyx), a species closely related to the human gastric pathogen Helicobacter pylori. Two groups of strains were identified by randomly amplified polymorphic DNA fingerprinting and gene sequencing: one group from six cheetahs in a U.S. zoo and two lions in a European circus, and the other group from a tiger and a lion-tiger hybrid in the same circus. PCR and DNA sequencing showed that each strain lacked the cag pathogenicity island and contained a degenerate vacuolating cytotoxin (vacA) gene. Analyses of nine other genes (glmM, recA, hp519, glr, cysS, ppa, flaB, flaA, and atpA) revealed a approximately 2% base substitution difference, on average, between the two H. acinonychis groups and a approximately 8% difference between these genes and their homologs in H. pylori reference strains such as 26695. H. acinonychis derivatives that could chronically infect mice were selected and were found to be capable of persistent mixed infection with certain H. pylori strains. Several variants, due variously to recombination or new mutation, were found after 2 months of mixed infection. H. acinonychis ' modest genetic distance from H. pylori, its ability to infect mice, and its ability to coexist and recombine with certain H. pylori strains in vivo should be useful in studies of Helicobacter infection and virulence mechanisms and studies of genome evolution.

Acinonyx↗