Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Evolutionary analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,333 records · Page 74Linked to original sources

Reconstructing contiguous regions of an ancestral genome.

This article analyzes mammalian genome rearrangements at higher resolution than has been published to date. We identify 3171 intervals, covering approximately 92% of the human genome, within which we find no rearrangements larger than 50 kilobases (kb) in the lineages leading to human, mouse, rat, and dog from their most recent common ancestor. Combining intervals that are adjacent in all contemporary species produces 1338 segments that may contain large insertions or deletions but that are free of chromosome fissions or fusions as well as inversions or translocations >50 kb in length. We describe a new method for predicting the ancestral order and orientation of those intervals from their observed adjacencies in modern species. We combine the results from this method with data from chromosome painting experiments to produce a map of an early mammalian genome that accounts for 96.8% of the available human genome sequence data. The precision is further increased by mapping inversions as small as 31 bp. Analysis of the predicted evolutionary breakpoints in the human lineage confirms certain published observations but disagrees with others. Although only a few mammalian genomes are currently sequenced to high precision, our theoretical analyses and computer simulations indicate that our results are reasonably accurate and that they will become highly accurate in the foreseeable future. Our methods were developed as part of a project to reconstruct the genome sequence of the last ancestor of human, dogs, and most other placental mammals.

Algorithms↗

Spectrum and significance of variants and mutations in the Fanconi anaemia group G gene in children with sporadic acute myeloid leukaemia.

Childhood acute myeloid leukaemia (AML) is uncommon. Children with Fanconi anaemia (FA), however, have a very high risk of developing AML. FA is a rare inherited disease caused by mutations in at least 12 genes, of which Fanconi anaemia group G gene (FANCG) is one of the commonest. To address to what extent FANCG variants contribute to sporadic childhood AML, we determined the spectrum of FANCG sequence variants in 107 children diagnosed with sporadic AML, using polymerase chain reaction (PCR), fluorescent single-strand conformational polymorphism (SSCP) and sequencing methodologies. The significance of variants was determined by frequency analysis and assessment of evolutionary conservation. Seven children (6.5%) carried variants in FANCG. Two of these carried two variants, including the known IVS2 + 1G>A mutation with the novel missense mutation S588F, and R513Q with the intronic deletion IVS12-38 (-28)_del11, implying that these patients might have been undiagnosed FA patients. R513Q, which affects a semi-conserved amino acid, was carried in two additional children with AML. Although not significant, the frequency of R513Q was higher in children with AML than unselected cord bloods. While FANCG mutation carrier status does not predispose to sporadic AML, the identification of unrecognised FA patients implies that FA presenting with primary AML in childhood is more common than suspected.

Acute Disease↗

The insect cytochrome oxidase I gene: evolutionary patterns and conserved primers for phylogenetic studies.

Insect mitochondrial cytochrome oxidase I (COI) genes are used as a model to examine the within-gene heterogeneity of evolutionary rate and its implications for evolutionary analyses. The complete sequence (1537 bp) of the meadow grasshopper (Chorthippus parallelus) COI gene has been determined, and compared with eight other insect COI genes at both the DNA and amino acid sequence levels. This reveals that different regions evolve at different rates, and the patterns of sequence variability seems associated with functional constraints on the protein. The COOH-terminal was found to be significantly more variable than internal loops (I), external loops (E), transmembrane helices (M) or the NH2 terminal. The central region of COI (M5-M8) has lower levels of sequence variability, which is related to several important functional domains in this region. Highly conserved primers which amplify regions of different variabilities have been designed to cover the entire insect COI gene. These primers have been shown to amplify COI in a wide range of species, representing all the major insect groups; some even in an arachnid. Implications of the observed evolutionary pattern for phylogenetic analysis are discussed, with particular regard to the choice of regions of suitable variability for specific phylogenetic projects.

Amino Acid Sequence↗

Genetic complexity and serum reactivity of HVR1 quasispecies of hepatitis C virus in patients with cirrhosis.

OBJECTIVE: The RNA genome of hepatitis C virus varies considerably, especially within the hypervariable region 1 (HVR1), a domain located on the 5' end of the E2/NS1 envelope region. Our previous study has suggested there were greater numbers of quasispecies in the liver than in matched serum, independent of the viral load. However, the significance of this finding has not been examined extensively at genetic and serological levels. METHODS: By large scale cloning and sequencing, we studied the genetic complexity of HVR1 quasispecies in two selected patients with cirrhosis. The serum reactivity of peptides representing different HVR1 quasispecies isolated from these cases was also estimated by standard ELISA format. RESULTS: We found the same major (dominant and/or subdominant) viral quasispecies variants in serum and in the cirrhotic liver. Genetic analysis suggested that the evolutionary pressure on HVR1 was higher than on its flanking region in quasispecies derived from the liver, whereas this trend is attenuated in quasispecies from serum. The immunoreactivity to peptides representing different HVR1 quasispecies variants showed considerable cross-reactivity with heterologous sera, whereas the reactivity was strongest against the dominant HVR1 peptide over time in homologous sera. CONCLUSIONS: These findings indicate that the formation and selection of HVR1 quasispecies may not be driven solely by humoral immune pressure, at least in these two cases.

Amino Acid Sequence↗

Evolutionary relationships of pathogenic clones of Vibrio cholerae by sequence analysis of four housekeeping genes.

Studies of the Vibrio cholerae population, using molecular typing techniques, have shown the existence of several pathogenic clones, mainly sixth-pandemic, seventh-pandemic, and U.S. Gulf Coast clones. However, the relationship of the pathogenic clones to environmental V. cholerae isolates remains unclear. A previous study to determine the phylogeny of V. cholerae by sequencing the asd (aspartate semialdehyde dehydrogenase) gene of V. cholerae showed that the sixth-pandemic, seventh-pandemic, and U.S. Gulf Coast clones had very different asd sequences which fell into separate lineages in the V. cholerae population. As gene trees drawn from a single gene may not reflect the true topology of the population, we sequenced the mdh (malate dehydrogenase) and hlyA (hemolysin A) genes from representatives of environmental and clinical isolates of V. cholerae and found that the mdh and hlyA sequences from the three pathogenic clones were identical, except for the previously reported 11-bp deletion in hlyA in the sixth-pandemic clone. Identical sequences were obtained, despite average nucleotide differences in the mdh and hlyA genes of 1.52 and 3.25%, respectively, among all the isolates, suggesting that the three pathogenic clones are closely related. To extend these observations, segments of the recA and dnaE genes were sequenced from a selection of the pathogenic isolates, where the sequences were either identical or substantially different between the clones. The results show that the three pathogenic clones are very closely related and that there has been a high level of recombination in their evolution.

Bacterial Proteins↗

DNA adenine methylation of GATC sequences appeared recently in the Escherichia coli lineage.

We have examined the presence of methylated adenine at GATC sequences (Dam phenotype) in the DNA of 23 eubacteria and 13 archaebacteria by using isoshizomer restriction enzymes. We have found a completely Dam+ phenotype in bacteria of nine genera related to the families Enterobacteriaceae, Parvobacteriaceae, and Vibrionaceae, and in the five cyanobacteria tested. We have found a partial Dam+ phenotype in the two archaebacteria Halobacterium saccharovorum and Methanobacterium sp. strain Ivanov. All of the other archaebacteria (three genera) and eubacteria (nine genera) tested were Dam-. Phylogenetic analysis, based on the evolutionary tree of Fox et al. (Science 209:457-463, 1980), indicates that dam methylation in the Escherichia coli lineage appeared recently in bacterial evolution and is restricted to a small range of closely related bacteria.

Adenine↗

Genetic diversity of the attachment protein of subgroup B respiratory syncytial viruses.

Respiratory syncytial (RS) virus causes repeated infections throughout life. Between the two main antigenic subgroups of RS virus, there is antigenic variation in the attachment protein G. The antigenic differences between the subgroups appear to play a role in allowing repeated infections to occur. Antigenic differences also occur within subgroups; however, neither the extent of these differences nor their contributions to repeat infections are known. We report a molecular analysis of the extent of diversity within the subgroup B RS virus attachment protein genes of viruses isolated from children over a 30-year period. Amino acid sequence differences as high as 12% were observed in the ectodomains of the G proteins among the isolates, whereas the cytoplasmic and transmembrane domains were highly conserved. The changes in the G-protein ectodomain were localized to two areas on either side of a highly conserved region surrounding four cysteine residues. Strikingly, single-amino-acid coding changes generated by substitution mutations were not the only means by which change occurred. Changes also occurred by (i) substitutions that changed the available termination codons, resulting in proteins of various lengths, and (ii) a mutation introduced by a single nucleotide deletion and subsequent nucleotide insertion, which caused a shift in the open reading frame of the protein in comparison to the other G genes analyzed. Fifty-one percent of the G-gene nucleotide changes observed among the isolates resulted in amino acid coding changes in the G protein, indicating a selective pressure for change. Maximum-parsimony analysis demonstrated that distinct evolutionary lineages existed. These data show that sequence diversity exists among the G proteins within the subgroup B RS viruses, and this diversity may be important in the immunobiology of the RS viruses.

Amino Acid Sequence↗

Evolutionary trace residues in noroviruses: importance in receptor binding, antigenicity, virion assembly, and strain diversity.

Noroviruses cause major epidemic gastroenteritis in humans. A large number of strains of these single-stranded RNA viruses have been reported. Due to the absence of infectious clones of noroviruses and the high sequence variability in their capsids, it has not been possible to identify functionally important residues in these capsids. Consequently, norovirus strain diversity is not understood on the basis of capsid functions, and the development of therapeutic compounds has been hampered. To determine functionally important residues in noroviruses, we have analyzed a number of norovirus capsid sequences in the context of the Norwalk virus capsid crystal structure by using the evolutionary trace method. This analysis has identified capsid protein residues that uniquely characterize different norovirus strains and provide new insights into capsid assembly and disassembly pathways and the strain diversity of these viruses. Such residues form specific three-dimensional clusters that may be of functional importance in noroviruses. One of these clusters includes residues known to participate in the proteolytic cleavage of these viruses at high pH. Other clusters are formed in capsid regions known to be important in the binding of antibodies to noroviruses, thereby indicating residues that may be important in the antigenicity of these viruses. The highly variable region of the capsid shows a distinct cluster whose residues may participate in norovirus-receptor interactions.

Amino Acid Sequence↗

The use of functional analysis of the ribosome as a tool to determine archaebacterial phylogeny.

Forty different antibiotics with diverse kingdom and functional specificities were used to measure the functional characteristics of the archaebacterial translation apparatus. The resulting inhibitory curves, which are characteristic of the cell-free system analyzed, were transformed into quantitative values that were used to cluster the different archaebacteria analyzed. This cluster resembles the phylogenetic tree generated by 16S rRNA sequence comparisons. These results strongly suggest that functional analysis of an appropriate evolutionary clock, such as the ribosome, is of intrinsic phylogenetic value. More importantly, they indicate that the study of the nexus between genotypic and phenotypic (functional) information may shed considerable light on the evolution of the protein synthetic machinery.

Archaea↗

On the choice of the offspring population size in evolutionary algorithms.

Evolutionary algorithms (EAs) generally come with a large number of parameters that have to be set before the algorithm can be used. Finding appropriate settings is a difficult task. The influence of these parameters on the efficiency of the search performed by an evolutionary algorithm can be very high. But there is still a lack of theoretically justified guidelines to help the practitioner find good values for these parameters. One such parameter is the offspring population size. Using a simplified but still realistic evolutionary algorithm, a thorough analysis of the effects of the offspring population size is presented. The result is a much better understanding of the role of offspring population size in an EA and suggests a simple way to dynamically adapt this parameter when necessary.

Algorithms↗

Section-level relationships of North American Agalinis (Orobanchaceae) based on DNA sequence analysis of three chloroplast gene regions.

BACKGROUND: The North American Agalinis are representatives of a taxonomically difficult group that has been subject to extensive taxonomic revision from species level through higher sub-generic designations (e.g., subsections and sections). Previous presentations of relationships have been ambiguous and have not conformed to modern phylogenetic standards (e.g., were not presented as phylogenetic trees). Agalinis contains a large number of putatively rare taxa that have some degree of taxonomic uncertainty. We used DNA sequence data from three chloroplast genes to examine phylogenetic relationships among sections within the genus Agalinis Raf. (=Gerardia), and between Agalinis and closely related genera within Orobanchaceae. RESULTS: Maximum likelihood analysis of sequences data from rbcL, ndhF, and matK gene regions (total aligned length 7323 bp) yielded a phylogenetic tree with high bootstrap values for most branches. Likelihood ratio tests showed that all but a few branch lengths were significantly greater than zero, and an additional likelihood ratio test rejected the molecular clock hypothesis. Comparisons of substitution rates between gene regions based on linear models of pairwise distance estimates between taxa show both ndhF and matK evolve more rapidly than rbcL, although the there is substantial rate heterogeneity within gene regions due in part to rate differences among codon positions. CONCLUSIONS: Phylogenetic analysis supports the monophyly of Agalinis, including species formerly in Tomanthera, and this group is sister to a group formed by the genera Aureolaria, Brachystigma, Dasistoma, and Seymeria. Many of the previously described sections within Agalinis are polyphyletic, although many of the subsections appear to form natural groups. The analysis reveals a single evolutionary event leading to a reduction in chromosome number from n = 14 to n = 13 based on the sister group relationship of section Erectae and section Purpureae subsection Pedunculares. Our results establish the evolutionary distinctiveness of A. tenella from the more widespread and common A. obtusifolia. However, further data are required to clearly resolve the relationship between A. acuta and A. tenella.

Chloroplasts↗

Molecular cloning of the cDNAs encoding a novel insulin-like growth factor-binding protein from rat and human.

cDNA clones encoding a novel insulin-like growth factor-binding protein (IGFBP) purified from rat serum and human bone cell-conditioned medium have been isolated from rat liver and human placenta, liver, and ovary cDNA libraries. The deduced amino acid sequences of the cDNAs revealed a mature polypeptide consisting of 233 amino acids for the rat, while the human structure contains an additional four-amino acid sequence in the middle region of the molecule. This protein, now proposed to be named IGFBP-4, contains two extra cysteines compared with the previously characterized IGFBP-1, -2, and -3, but the alignment of the remaining 18 cysteines is conserved across the four IGFBPs. Amino acid sequence comparison among the four binding proteins within the rat species demonstrated that both the amino- and carboxy-terminal one thirds of the molecules are highly conserved, while the middle one third region, where no cysteines are present except for the two that exist in IGFBP-4, is the most divergent. The overall sequence homology among the four rat IGFBPs is very similar (53-59%), suggesting that their individual genes diverged from a single ancestral gene at about the same evolutionary time point. Northern analysis of the IGFBP-4 mRNA in rat tissue demonstrated that transcription of the IGFBP-4 gene is highly active in the liver, although a single 2.6-kilobase IGFBP-4 mRNA band was detectable in all tissues examined, including adrenal, testis, spleen, heart, lung, kidney, liver, stomach, hypothalamus, and brain cortex.

Amino Acid Sequence↗

The relationship of severe acute respiratory syndrome coronavirus with avian and other coronaviruses.

In February 2003, a severe acute respiratory syndrome coronavirus (SARS-CoV) emerged in humans in Guangdong Province, China, and caused an epidemic that had severe impact on public health, travel, and economic trade. Coronaviruses are worldwide in distribution, highly infectious, and extremely difficult to control because they have extensive genetic diversity, a short generation time, and a high mutation rate. They can cause respiratory, enteric, and in some cases hepatic and neurological diseases in a wide variety of animals and humans. An enormous, previously unrecognized reservoir of coronaviruses exists among animals. Because coronaviruses have been shown, both experimentally and in nature, to undergo genetic mutations and recombination at a rate similar to that of influenza viruses, it is not surprising that zoonosis and host switching that leads to epidemic diseases have occurred among coronaviruses. Analysis of coronavirus genomic sequence data indicates that SARS-CoV emerged from an animal reservoir. Scientists examining coronavirus isolates from a variety of animals in and around Guangdong Province reported that SARS-CoV has similarities with many different coronaviruses including avian coronaviruses and SARS-CoV-like viruses from a variety of mammals found in live-animal markets. Although a SARS-like coronavirus isolated from a bat is thought to be the progenitor of SARS-CoV, a lack of genomic sequences for the animal coronaviruses has prevented elucidation of the true origin of SARS-CoV. Sequence analysis of SARS-CoV shows that the 5' polymerase gene has a mammalian ancestry; whereas the 3' end structural genes (excluding the spike glycoprotein) have an avian origin. Spike glycoprotein, the host cell attachment viral surface protein, was shown to be a mosaic of feline coronavirus and avian coronavirus sequences resulting from a recombination event. Based on phylogenetic analysis designed to elucidate evolutionary links among viruses, SARS-CoV is believed to have branched from the modern Group 2 coronaviruses, suggesting that it evolved relatively rapidly. This is significant because SARS-CoV is likely still circulating in an animal reservoir (or reservoirs) and has the potential to quickly emerge and cause a new epidemic.

Animals↗

Bioinformatics and the discovery of novel anti-microbial targets.

Genomic research is playing a critical role in the discovery of new anti-microbial drugs. The rapid increase in bacterial and eukaryotic genome sequences allows for new and innovative ways for obtaining antimicrobial protein targets. Here, we describe a two level strategy for target identification and validation using computers (in silico). First, large scale comparative analyses of genome sequences were used to identify highly conserved genes which might be essential for in vitro and/or in vivo survival of bacterial pathogens. Lab-based experiments provided confirmation or validation of the hypothesis of in silico essentiality for over 350 individual genes. Over 200 validated, broad spectrum; yet highly specific gene targets, were identified in community infection pathogens. The second part of the target discovery strategy is an in-depth evolutionary, structural and cellular analysis of key drug targets. As an example, phylogenetic and structural analyses suggest that sequence and binding-pocket conservation in FabH (beta-ketoacyl-ACP synthase III) would allow for the development of small molecule inhibitors not only effective against a broad species spectrum of community bacterial pathogens but also as potential new therapies for tuberculosis and malaria.

Animals↗

Proteomic Analysis of Biomineralization Proteins in the Shell Plates and Spicules of Chiton Acanthochitona rubrolineata.

Chitons, ancient polyplacophoran mollusks, are ideal models for studying biomineralization evolution due to their conserved morphology since the Cambrian. This study investigates the matrix proteins in shell plates and spicules of Acanthochitona rubrolineata using liquid chromatography-tandem mass spectrometry. By extracting proteins from 30 individuals and using proteomic method, we identified 26 soluble proteins and 22 insoluble proteins in the shell plates and 25 insoluble proteins, and found domains such as von Willebrand factor type A, chitin-binding, ferritin, and cadherin. These domains, prevalent in molluscan biominerals, suggest conserved roles in organic matrix formation. Despite genomic dynamism, the conservation of key domains across species highlights a core biomineralization mechanism. Notably, eight of the shell proteins and eight of the spicule proteins were homologous between A. rubrolineata and chiton Acanthopleura loochooana, indicating functional conservation. Phylogenetic analysis further supported the evolutionary significance of these domains in chitons. The study advances understanding of biomineralization in Polyplacophora, emphasizing the interplay between morphological stasis and molecular evolution.

matrix proteins↗

NF-kappa B regulates BCL3 transcription in T lymphocytes through an intronic enhancer.

Exposure to soluble protein Ags in vivo leads to abortive proliferation of responding T cells. In the absence of a danger signal, artificially provided by adjuvants, most responding cells die, and the remainder typically become anergic. The adjuvant-derived signals provided to T cells are poorly understood, but recent work has identified BCL3 as the gene, of those tested, with the greatest differential transcriptional response to adjuvant administration in vivo. As an initial step in analyzing transcriptional responses of BCL3 in T cells, we have identified candidate regulatory regions within the locus through their evolutionary conservation and by analysis of DNase hypersensitivity. An evolutionarily conserved DNase hypersensitive site (HS3) within intron 2 was found to act as a transcriptional enhancer in response to stimuli that mimic TCR activation, namely, PHA and PMA. In luciferase reporter gene constructs transiently transfected into the Jurkat T cell line, the HS3 enhancer can cooperate not only with the BCL3 promoter, but also with an exogenous promoter from herpes simplex thymidine kinase. Deletional analysis revealed that a minimal sequence of approximately 81 bp is required for full enhancer activity. At the 5' end of this minimal sequence is a kappaB site, as confirmed by EMSAs. Mutation of this site in the context of the full-length HS3 abolished enhancer activity. Cotransfection with NF-kappaB p65 expression constructs dramatically increased luciferase activity, even without stimulation. Conversely, cotransfection with the NF-kappaB inhibitor IkappaBalpha reduced activation. Together, these results demonstrate a critical role for NF-kappaB in BCL3 transcriptional up-regulation by TCR-mimetic signals.

B-Cell Lymphoma 3 Protein↗

A highly prevalent lupus risk haplotype increases IRF7-dependent induction of IFN-α, enhancing antiviral defense and exacerbating autoimmunity.

UNLABELLED: Genome-wide association studies have identified genetic polymorphisms at 11p15 associated with Systemic Lupus Erythematosus (lupus). Statistical fine mapping prioritizes a highly prevalent coding haplotype within the IRF7 gene. Analysis of ancient DNA confirms that this haplotype has persisted at high frequencies in the global population for millennia. The IRF7 risk haplotype is sufficient to increase nuclear localization of IRF7 and transcriptional activity downstream of pattern recognition receptor pathways. This risk haplotype increases IRF7 DNA binding strength and alters IRF7 DNA sequence specificity, resulting in genotype-dependent increases in IFN-α production in numerous biological systems, including monocytes and airway epithelial cells. CRISPR engineering of a homologous risk variant in mouse Irf7 results in both enhanced innate control of virus infection and increased autoantibody titers in a model of autoimmunity. Altogether, we establish a persistent and prominent genetic IRF7 haplotype that amplifies IRF7 activity in a manner that has immunological risks and benefits. HIGHLIGHTS: Genetic analysis using modern and evolutionary datasets identifies a persistent and highly prevalent lupus-associated coding haplotype in IRF7 at 11p15 The IRF7 lupus risk haplotype increases IFN-α production by monocytes and airway epithelial cells The IRF7 lupus risk haplotype increases IRF7 DNA binding strength and alters DNA sequence specificity A homologous lupus risk variant in mouse Irf7 enhances control of vesicular stomatitis virus and exacerbates autoantibody production.

Journal Article↗

Phylogenetic structure in the grass family (Poaceae): evidence from the nuclear gene phytochrome B.

Phylogenetic analyses of partial phytochrome B (PHYB) nuclear DNA sequences provide unambiguous resolution of evolutionary relationships within Poaceae. Analysis of PHYB nucleotides from 51 taxa representing seven traditionally recognized subfamilies clearly distinguishes three early-diverging herbaceous "bambusoid" lineages. First and most basal are Anomochloa and Streptochaeta, second is Pharus, and third is Puelia. The remaining grasses occur in two principal, highly supported clades. The first comprises bambusoid, oryzoid, and pooid genera (the BOP clade); the second comprises panicoid, arundinoid, chloridoid, and centothecoid genera (the PACC clade). The PHYB phylogeny is the first nuclear gene tree to address comprehensively phylogenetic relationships among grasses. It corroborates several inferences made from chloroplast gene trees, including the PACC clade, and the basal position of the herbaceous bamboos Anomochloa, Streptochaeta, and Pharus. However, the clear resolution of the sister group relationship among bambusoids, oryzoids, and pooids in the PHYB tree is novel; the relationship is only weakly supported in ndhF trees and is nonexistent in rbcL and plastid restriction site trees. Nuclear PHYB data support Anomochlooideae, Pharoideae, Pooideae sensu lato, Oryzoideae, Panicoideae, and Chloridoideae, and concur in the polyphyly of both Arundinoideae and Bambusoideae.

Journal Article↗