Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Evolutionary analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

Nonribosomal peptide synthesis and toxigenicity of cyanobacteria.

Nonribosomal peptide synthesis is achieved in prokaryotes and lower eukaryotes by the thiotemplate function of large, modular enzyme complexes known collectively as peptide synthetases. These and other multifunctional enzyme complexes, such as polyketide synthases, are of interest due to their use in unnatural-product or combinatorial biosynthesis (R. McDaniel, S. Ebert-Khosla, D. A. Hopwood, and C. Khosla, Science 262:1546-1557, 1993; T. Stachelhaus, A. Schneider, and M. A. Marahiel, Science 269:69-72, 1995). Most nonribosomal peptides from microorganisms are classified as secondary metabolites; that is, they rarely have a role in primary metabolism, growth, or reproduction but have evolved to somehow benefit the producing organisms. Cyanobacteria produce a myriad array of secondary metabolites, including alkaloids, polyketides, and nonribosomal peptides, some of which are potent toxins. This paper addresses the molecular genetic basis of nonribosomal peptide synthesis in diverse species of cyanobacteria. Amplification of peptide synthetase genes was achieved by use of degenerate primers directed to conserved functional motifs of these modular enzyme complexes. Specific detection of the gene cluster encoding the biosynthetic pathway of the cyanobacterial toxin microcystin was shown for both cultured and uncultured samples. Blot hybridizations, DNA amplifications, sequencing, and evolutionary analysis revealed a broad distribution of peptide synthetase gene orthologues in cyanobacteria. The results demonstrate a molecular approach to assessing preexpression microbial functional diversity in uncultured cyanobacteria. The nonribosomal peptide biosynthetic pathways detected may lead to the discovery and engineering of novel antibiotics, immunosuppressants, or antiviral agents.

Bacterial Toxins↗

Evolution of genomic content in the stepwise emergence of Escherichia coli O157:H7.

Genome comparisons have demonstrated that dramatic genetic change often underlies the emergence of new bacterial pathogens. Evolutionary analysis of Escherichia coli O157:H7, a pathogen that has emerged as a worldwide public health threat in the past two decades, has posited that this toxin-producing pathogen evolved in a series of steps from O55:H7, a recent ancestor of a nontoxigenic pathogenic clone associated with infantile diarrhea. We used comparative genomic hybridization with 50-mer oligonucleotide microarrays containing probes from both pathogenic and nonpathogenic genomes to infer when genes were acquired and lost. Many ancillary virulence genes identified in the O157 genome were already present in an O55:H7-like progenitor, with 27 of 33 genomic islands of >5 kb and specific for O157:H7 (O islands) that were acquired intact before the split from this immediate ancestor. Most (85%) of variably absent or present genes are part of prophages or phage-like elements. Divergence in gene content among these closely related strains was approximately 140 times greater than divergence at the nucleotide sequence level. A >100-kb region around the O-antigen gene cluster contained highly divergent sequences and also appears to be duplicated in its entirety in one lineage, suggesting that the whole region was cotransferred in the antigenic shift from O55 to O157. The beta-glucuronidase-positive O157 variants, although phylogenetically closest to the Sakai strain, were divergent for multiple adherence factors. These observations suggest that, in addition to gains and losses of phage elements, O157:H7 genomes are rapidly diverging and radiating into new niches as the pathogen disseminates.

Chromosome Mapping↗

Two distinct human parainfluenza virus type 1 genotypes detected during the 1991 Milwaukee epidemic.

The extent of genetic and antigenic variation found in a population of human parainfluenza virus type 1 (HPIV-1) during a single local epidemic was investigated. Fifteen HPIV-1 strains isolated from children in 1991 were analyzed. Nucleotide sequence variation in the hemagglutinin-neuraminidase protein (HN) gene demonstrated two distinct genotypes (genotypes C and D). Unique patterns were identified involving 62 nucleotide and 10 amino acid positions. These patterns represented 40% of all mutations within the HN gene. The remaining mutations were randomly distributed, and 74% involved only one (55%) or two isolates. Genotypes were statistically different from each other at both the nucleotide (P = 0.001) and amino acid (P = 0.001) levels and demonstrated unique potential N-linked glycosylation patterns. Thirty-eight monoclonal antibodies (MAbs) made to four different viral proteins (22 HN, 2 fusion [F], 1 phosphoprotein, and 13 nucleoprotein) (originating from two different genotypes [genotypes A and D]) were compared for their ability to bind to the clinical isolates in enzyme-linked immunosorbent assays (ELISAs) and hemagglutinin-inhibition (HI) assays. Twenty-one MAbs bound well to all clinical isolates in ELISAs and HI assays. The remaining 17 MAbs showed variation in all four structural proteins. Three HN MAbs demonstrated genotype C- and D-specific antigenic and neutralization differences. Evolutionary analysis using parsimony methods confirmed the differences between the two genotypes. No differences in either clinical presentation or disease severity between the two genotypes were found. Geographically localized HPIV-1 epidemics can be caused by at least two distinct genotypes with minor but specific antigenic changes. The clinical and immunologic roles of HPIV-1 genotypes have not been determined.

Amino Acid Sequence↗

New hepatitis C virus (HCV) genotyping system that allows for identification of HCV genotypes 1a, 1b, 2a, 2b, 3a, 3b, 4, 5a, and 6a.

Recent studies have focused on whether different hepatitis C virus (HCV) genotypes are associated with different profiles of pathogenicity, infectivity, and response to antiviral therapy. The establishment of a simple and precise genotyping system for HCV is essential to address these issues. A new genotyping system based on PCR of the core region with genotype-specific PCR primers for the determination of HCV genotypes 1a, 1b, 2a, 2b, 3a, 3b, 4, 5a, and 6a was developed. A total of 607 samples (379 from Japan, 63 from the United States, 53 from Korea, 35 from Taiwan, 32 from China, 20 from Hong Kong, 15 from Australia, 6 from Egypt, 3 from Bangladesh, and 1 from South Africa) were tested by both the assay of Okamoto et al. (H. Okamoto, Y. Sugiyama, S. Okada, K. Kurai, Y. Akahane, Y. Sugai, T. Tanaka, K. Sato, F. Tsuda, Y. Miyamura, and M. Mayumi, J. Gen. Virol. 73:673-679, 1992) and this new genotyping system. Comparison of the results showed concordant results for 539 samples (88.8%). Of the 68 samples with discordant results, the nucleotide sequences of the HCV isolates were determined in 23, and their genotypes were determined by molecular evolutionary analysis. In all 23 samples, the assignment of genotype by our new genotyping system was correct. This genotyping system may be useful for large-scale determination of HCV genotypes in clinical studies.

DNA Primers↗

A second gene for the African green monkey poliovirus receptor that has no putative N-glycosylation site in the functional N-terminal immunoglobulin-like domain.

Using cDNA of the human poliovirus receptor (PVR) as a probe, two types of cDNA clones of the monkey homologs were isolated from a cDNA library prepared from an African green monkey kidney cell line. Either type of cDNA clone rendered mouse L cells permissive for poliovirus infection. Homologies of the amino acid sequences deduced from these cDNA sequences with that of human PVR were 90.2 and 86.4%, respectively. These two monkey PVRs were found to be encoded in two different loci of the genome. Evolutionary analysis suggested that duplication of the PVR gene in the monkey genome had occurred after the species differentiation between humans and monkeys. The NH2-terminal immunoglobulin-like domain, domain 1, of the second monkey PVR, which lacks a putative N-glycosylation site, mediated poliovirus infection. In addition, a human PVR mutant without N-glycosylation sites in domain 1 also promoted viral infection. These results suggest that domain 1 of the monkey receptor also harbors the binding site for poliovirus and that sugar moieties possibly attached to this domain of human PVR are dispensable for the virus-receptor interaction.

Amino Acid Sequence↗

Evolution of hypervariable region 1 of hepatitis C virus in primary infection.

The hypervariable region 1 (HVR-1) of the putative envelope encoding E2 region of hepatitis C virus (HCV) RNA was analyzed in sequential samples from three patients with acute type C hepatitis infected from different sources to address (i) the dynamics of intrahost HCV variability during the primary infection and (ii) the role of host selective pressure in driving viral genetic evolution. HVR-1 sequences from 20 clones per each point in time were analyzed after amplification, cloning, and purification of plasmid DNA from single colonies of transformed cells. The intrasample evolutionary analysis (nonsynonymous mutations per nonsynonymous site [Ka], synonymous mutations per synonymous site [Ks], Ka/Ks ratio, and genetic distances [gd]) documented low gd in early samples (ranging from 2. 11 to 7.79%) and a further decrease after seroconversion (from 0 to 4.80%), suggesting that primary HCV infection is an oligoclonal event, and found different levels and dynamics of host pressure in the three cases. The intersample analysis (pairwise comparisons of intrapatient sequences; rKa, rKs, rKa/rKs ratio, and gd) confirmed the individual features of HCV genetic evolution in the three subjects and pointed to the relative contribution of either neutral evolution or selective forces in driving viral variability, documenting that adaptation of HCV for persistence in vivo follows different routes, probably representing the molecular counterpart of the viral fitness for individual environments.

Adolescent↗

Four distinct and unusual linker proteins in a mitotically dividing nucleus are derived from a 71-kilodalton polyprotein, lack p34cdc2 sites, and contain protein kinase A sites.

Tetrahymena thermophila micronuclei contain four linker-associated proteins, alpha, beta, gamma, and delta. Synthetic oligonucleotides based on N-terminal protein sequences of beta and gamma were used to clone the micronuclear linker histone (MLH) gene. The MLH gene is single copy and is transcribed into a 2.4-kb message encoding all four linker-associated proteins. The message is translated into a polypeptide (Mic LH) that is processed at the sequence decreases RTK to give proteins whose amino acid sequences differ markedly from each other, from the sequence of macronuclear H1, and from sequences of typical H1s of other organisms. This represents the first example of multiple chromatin proteins derived from a single polyprotein. The delta protein consists largely of two high-mobility-group (HMG) boxes. An evolutionary analysis of HMG boxes indicates that the delta HMG boxes are similar to the HMG boxes of tsHMG, a protein that appears in elongating mouse spermatids when they condense and cease transcription, suggesting that delta could play a similar role in the micronucleus. The micronucleus divides mitotically, while the macronucleus divides amitotically. Surprisingly, macronuclear H1 but not Mic LH contains sequences resembling p34cdc2 kinase phosphorylation sites, while each of the Mic LH-derived proteins contains a typical protein kinase A phosphorylation site in its carboxy terminus.

Amino Acid Sequence↗

The genomics of long tandem arrays of satellite DNA in the human genome.

At least 10% of DNA in the human genome consists of long arrays of repeated sequences, arranged in tandem head-to-tail arrays in a number of discrete, highly localized chromosomal regions. Different families of these so-called "satellite DNA" sequences have been defined, organized in diverged subsets on different chromosomes. The molecular, cytogenetic, and evolutionary analysis of the hierarchical organization of such sequences in the human and other complex genomes encompasses a variety of approaches, including chromosomal mapping, in situ hybridization, genetic linkage analysis, long-range restriction mapping, and DNA sequencing. Investigation of the organization of satellite arrays constitutes a necessary first step towards eventual elucidation of the origin, evolution, and maintenance of these sequences and their contribution to the structure and behavior of human chromosomes.

Chromosome Mapping↗

Molecular evolution of the human chromosome 15 pericentromeric region.

We present a detailed molecular evolutionary analysis of 1.2 Mb from the pericentromeric region of human 15q11. Sequence analysis indicates the region has been subject to extensive interchromosomal and intrachromosomal duplications during primate evolution. Comparative FISH analyses among non-human primates show remarkable quantitative and qualitative differences in the organization and duplication history of this region - including lineage-specific deletions and duplication expansions. Phylogenetic and comparative analyses reveal that the region is composed of at least 24 distinct segmental duplications or duplicons that have populated the pericentromeric regions of the human genome over the last 40 million years of human evolution. The value of combining both cytogenetic and experimental data in understanding the complex forces which have shaped these regions is discussed.

Animals↗

Accelerated probabilistic inference of RNA structure evolution.

BACKGROUND: Pairwise stochastic context-free grammars (Pair SCFGs) are powerful tools for evolutionary analysis of RNA, including simultaneous RNA sequence alignment and secondary structure prediction, but the associated algorithms are intensive in both CPU and memory usage. The same problem is faced by other RNA alignment-and-folding algorithms based on Sankoff's 1985 algorithm. It is therefore desirable to constrain such algorithms, by pre-processing the sequences and using this first pass to limit the range of structures and/or alignments that can be considered. RESULTS: We demonstrate how flexible classes of constraint can be imposed, greatly reducing the computational costs while maintaining a high quality of structural homology prediction. Any score-attributed context-free grammar (e.g. energy-based scoring schemes, or conditionally normalized Pair SCFGs) is amenable to this treatment. It is now possible to combine independent structural and alignment constraints of unprecedented general flexibility in Pair SCFG alignment algorithms. We outline several applications to the bioinformatics of RNA sequence and structure, including Waterman-Eggert N-best alignments and progressive multiple alignment. We evaluate the performance of the algorithm on test examples from the RFAM database. CONCLUSION: A program, Stemloc, that implements these algorithms for efficient RNA sequence alignment and structure prediction is available under the GNU General Public License.

Algorithms↗

Evolutionary relationships of Aurora kinases: implications for model organism studies and the development of anti-cancer drugs.

BACKGROUND: As key regulators of mitotic chromosome segregation, the Aurora family of serine/threonine kinases play an important role in cell division. Abnormalities in Aurora kinases have been strongly linked with cancer, which has lead to the recent development of new classes of anti-cancer drugs that specifically target the ATP-binding domain of these kinases. From an evolutionary perspective, the species distribution of the Aurora kinase family is complex. Mammals uniquely have three Aurora kinases, Aurora-A, Aurora-B, and Aurora-C, while for other metazoans, including the frog, fruitfly and nematode, only Aurora-A and Aurora-B kinases are known. The fungi have a single Aurora-like homolog. Based on the tacit assumption of orthology to human counterparts, model organism studies have been central to the functional characterization of Aurora kinases. However, the ortholog and paralog relationships of these kinases across various species have not been rigorously examined. Here, we present comprehensive evolutionary analyses of the Aurora kinase family. RESULTS: Phylogenetic trees suggest that all three vertebrate Auroras evolved from a single urochordate ancestor. Specifically, Aurora-A is an orthologous lineage in cold-blooded vertebrates and mammals, while structurally similar Aurora-B and Aurora-C evolved more recently in mammals from a duplication of an ancestral Aurora-B/C gene found in cold-blooded vertebrates. All so-called Aurora-A and Aurora-B kinases of non-chordates are ancestral to the clade of chordate Auroras and, therefore, are not strictly orthologous to vertebrate counterparts. Comparisons of human Aurora-B and Aurora-C sequences to the resolved 3D structure of human Aurora-A lends further support to the evolutionary scenario that vertebrate Aurora-B and Aurora-C are closely related paralogs. Of the 26 residues lining the ATP-binding active site, only three were variant and all were specific to Aurora-A. CONCLUSIONS: In this study, we found that invertebrate Aurora-A and Aurora-B kinases are highly divergent protein families from their chordate counterparts. Furthermore, while the Aurora-A family is ubiquitous among all vertebrates, the Aurora-B and Aurora-C families in humans arose from a gene duplication event in mammals. These findings show the importance of understanding evolutionary relationships in the interpretation and transference of knowledge from studies of model organism systems to human cellular biology. In addition, given the important role of Aurora kinases in cancer, evolutionary analysis and comparisons of ATP-binding domains suggest a rationale for designing dual action anti-tumor drugs that inhibit both Aurora-B and Aurora-C kinases.

Amino Acid Sequence↗

Computing Ka and Ks with a consideration of unequal transitional substitutions.

BACKGROUND: Approximate methods for estimating nonsynonymous and synonymous substitution rates (Ka and Ks) among protein-coding sequences have adopted different mutation (substitution) models. In the past two decades, several methods have been proposed but they have not considered unequal transitional substitutions (between the two purines, A and G, or the two pyrimidines, T and C) that become apparent when sequences data to be compared are vast and significantly diverged. RESULTS: We propose a new method (MYN), a modified version of the Yang-Nielsen algorithm (YN), for evolutionary analysis of protein-coding sequences in general. MYN adopts the Tamura-Nei Model that considers the difference among rates of transitional and transversional substitutions as well as factors in codon frequency bias. We evaluate the performance of MYN by comparing to other methods, especially to YN, and to show that MYN has minimal deviations when parameters vary within normal ranges defined by empirical data. CONCLUSION: Our comparative results deriving from consistency analysis, computer simulations and authentic datasets, indicate that ignoring unequal transitional rates may lead to serious biases and that MYN performs well in most of the tested cases. These results also suggest that acquisitions of reliable synonymous and nonsynonymous substitution rates primarily depend on less biased estimates of transition/transversion rate ratio.

Algorithms↗

i-Genome: a database to summarize oligonucleotide data in genomes.

BACKGROUND: Information on the occurrence of sequence features in genomes is crucial to comparative genomics, evolutionary analysis, the analyses of regulatory sequences and the quantitative evaluation of sequences. Computing the frequencies and the occurrences of a pattern in complete genomes is time-consuming. RESULTS: The proposed database provides information about sequence features generated by exhaustively computing the sequences of the complete genome. The repetitive elements in the eukaryotic genomes, such as LINEs, SINEs, Alu and LTR, are obtained from Repbase. The database supports various complete genomes including human, yeast, worm, and 128 microbial genomes. CONCLUSIONS: This investigation presents and implements an efficiently computational approach to accumulate the occurrences of the oligonucleotides or patterns in complete genomes. A database is established to maintain the information of the sequence features, including the distributions of oligonucleotide, the gene distribution, the distribution of repetitive elements in genomes and the occurrences of the oligonucleotides. The database can provide more effective and efficient way to access the repetitive features in genomes.

Alu Elements↗

Duplication and positive selection among hominin-specific PRAME genes.

BACKGROUND: The physiological and phenotypic differences between human and chimpanzee are largely specified by our genomic differences. We have been particularly interested in recent duplications in the human genome as examples of relatively large-scale changes to our genome. We performed an in-depth evolutionary analysis of a region of chromosome 1, which is copy number polymorphic among humans, and that contains at least 32 PRAME (Preferentially expressed antigen of melanoma) genes and pseudogenes. PRAME-like genes are expressed in the testis and in a large number of tumours, and are thought to possess roles in spermatogenesis and oogenesis. RESULTS: Using nucleotide substitution rate estimates for exons and introns, we show that two large segmental duplications, of six and seven human PRAME genes respectively, occurred in the last 3 million years. These duplicated genes are thus hominin-specific, having arisen in our genome since the divergence from chimpanzee. This cluster of PRAME genes appears to have arisen initially from a translocation approximately 95-85 million years ago. We identified multiple sites within human or mouse PRAME sequences which exhibit strong evidence of positive selection. These form a pronounced cluster on one face of the predicted PRAME protein structure. CONCLUSION: We predict that PRAME genes evolved adaptively due to strong competition between rapidly-dividing cells during spermatogenesis and oogenesis. We suggest that as PRAME gene copy number is polymorphic among individuals, positive selection of PRAME alleles may still prevail within the human population.

Alleles↗

Comparative genomics and evolution of the HSP90 family of genes across all kingdoms of organisms.

BACKGROUND: HSP90 proteins are essential molecular chaperones involved in signal transduction, cell cycle control, stress management, and folding, degradation, and transport of proteins. HSP90 proteins have been found in a variety of organisms suggesting that they are ancient and conserved. In this study we investigate the nuclear genomes of 32 species across all kingdoms of organisms, and all sequences available in GenBank, and address the diversity, evolution, gene structure, conservation and nomenclature of the HSP90 family of genes across all organisms. RESULTS: Twelve new genes and a new type HSP90C2 were identified. The chromosomal location, exon splicing, and prediction of whether they are functional copies were documented, as well as the amino acid length and molecular mass of their polypeptides. The conserved regions across all protein sequences, and signature sequences in each subfamily were determined, and a standardized nomenclature system for this gene family is presented. The proeukaryote HSP90 homologue, HTPG, exists in most Bacteria species but not in Archaea, and it evolved into three lineages (Groups A, B and C) via two gene duplication events. None of the organellar-localized HSP90s were derived from endosymbionts of early eukaryotes. Mitochondrial TRAP and endoplasmic reticulum HSP90B separately originated from the ancestors of HTPG Group A in Firmicutes-like organisms very early in the formation of the eukaryotic cell. TRAP is monophyletic and present in all Animalia and some Protista species, while HSP90B is paraphyletic and present in all eukaryotes with the exception of some Fungi species, which appear to have lost it. Both HSP90C (chloroplast HSP90C1 and location-undetermined SP90C2) and cytosolic HSP90A are monophyletic, and originated from HSP90B by independent gene duplications. HSP90C exists only in Plantae, and was duplicated into HSP90C1 and HSP90C2 isoforms in higher plants. HSP90A occurs across all eukaryotes, and duplicated into HSP90AA and HSP90AB in vertebrates. Diplomonadida was identified as the most basal organism in the eukaryote lineage. CONCLUSION: The present study presents the first comparative genomic study and evolutionary analysis of the HSP90 family of genes across all kingdoms of organisms. HSP90 family members underwent multiple duplications and also subsequent losses during their evolution. This study established an overall framework of information for the family of genes, which may facilitate and stimulate the study of this gene family across all organisms.

Alternative Splicing↗

A murine specific expansion of the Rhox cluster involved in embryonic stem cell biology is under natural selection.

BACKGROUND: The rodent specific reproductive homeobox (Rhox) gene cluster on the X chromosome has been reported to contain twelve homeobox-containing genes, Rhox1-12. RESULTS: We have identified a 40 kb genomic region within the Rhox cluster that is duplicated eight times in tandem resulting in the presence of eight paralogues of Rhox2 and Rhox3 and seven paralogues of Rhox4. Transcripts have been identified for the majority of these paralogues and all but three are predicted to produce full-length proteins with functional potential. We predict that there are a total of thirty-two Rhox genes at this genomic location, making it the most gene-rich homoeobox cluster identified in any species. From the 95% sequence similarity between the eight duplicated genomic regions and the synonymous substitution rate of the Rhox2, 3 and 4 paralogues we predict that the duplications occurred after divergence of mouse and rat and represent the youngest homoeobox cluster identified to date. Molecular evolutionary analysis reveals that this cluster is an actively evolving region with Rhox2 and 4 paralogues under diversifying selection and Rhox3 evolving neutrally. The biological importance of this duplication is emphasised by the identification of an important role for Rhox2 and Rhox4 in regulating the initial stages of embryonic stem (ES) cell differentiation. CONCLUSION: The gene rich Rhox cluster provides the mouse with significant biological novelty that we predict could provide a substrate for speciation. Moreover, this unique cluster may explain species differences in ES cell derivation and maintenance between mouse, rat and human.

Amino Acid Sequence↗

The extent and importance of intragenic recombination.

We have studied the recombination rate behaviour of a set of 140 genes which were investigated for their potential importance in inflammatory disease. Each gene was extensively sequenced in 24 individuals of African descent and 23 individuals of European descent, and the recombination process was studied separately in the two population samples. The results obtained from the two populations were highly correlated, suggesting that demographic bias does not affect our population genetic estimation procedure. We found evidence that levels of recombination correlate with levels of nucleotide diversity. High marker density allowed us to study recombination rate variation on a very fine spatial scale. We found that about 40 per cent of genes showed evidence of uniform recombination, while approximately 12 per cent of genes carried distinct signatures of recombination hotspots. On studying the locations of these hotspots, we found that they are not always confined to introns but can also stretch across exons. An investigation of the protein products of these genes suggested that recombination hotspots can sometimes separate exons belonging to different protein domains; however, this occurs much less frequently than might be expected based on evolutionary studies into the origins of recombination. This suggests that evolutionary analysis of the recombination process is greatly aided by considering nucleotide sequences and protein products jointly.

Africa↗

Identifying transcribed sequences, and beyond.

A report on the 12th International Workshop 'Beyond the Identification of Transcribed Sequences (BITS): Functional, Expression and Evolutionary Analysis', Washington DC, USA, 25-28 October 2002.

Alternative Splicing↗