Search PubMed⌕ Search

Biomedical subjects

Claude W dePamphilis

Publications and source records attributed to Claude W dePamphilis.

15 recordsLinked to original sources

Complete plastid genome sequences of Drimys, Liriodendron, and Piper: implications for the phylogenetic relationships of magnoliids.

BACKGROUND: The magnoliids with four orders, 19 families, and 8,500 species represent one of the largest clades of early diverging angiosperms. Although several recent angiosperm phylogenetic analyses supported the monophyly of magnoliids and suggested relationships among the orders, the limited number of genes examined resulted in only weak support, and these issues remain controversial. Furthermore, considerable incongruence resulted in phylogenetic reconstructions supporting three different sets of relationships among magnoliids and the two large angiosperm clades, monocots and eudicots. We sequenced the plastid genomes of three magnoliids, Drimys (Canellales), Liriodendron (Magnoliales), and Piper (Piperales), and used these data in combination with 32 other angiosperm plastid genomes to assess phylogenetic relationships among magnoliids and to examine patterns of variation of GC content. RESULTS: The Drimys, Liriodendron, and Piper plastid genomes are very similar in size at 160,604, 159,886 bp, and 160,624 bp, respectively. Gene content and order are nearly identical to many other unrearranged angiosperm plastid genomes, including Calycanthus, the other published magnoliid genome. Overall GC content ranges from 34-39%, and coding regions have a substantially higher GC content than non-coding regions. Among protein-coding genes, GC content varies by codon position with 1st codon > 2nd codon > 3rd codon, and it varies by functional group with photosynthetic genes having the highest percentage and NADH genes the lowest. Phylogenetic analyses using parsimony and likelihood methods and sequences of 61 protein-coding genes provided strong support for the monophyly of magnoliids and two strongly supported groups were identified, the Canellales/Piperales and the Laurales/Magnoliales. Strong support is reported for monocots and eudicots as sister clades with magnoliids diverging before the monocot-eudicot split. The trees also provided moderate or strong support for the position of Amborella as sister to a clade including all other angiosperms. CONCLUSION: Evolutionary comparisons of three new magnoliid plastid genome sequences, combined with other published angiosperm genomes, confirm that GC content is unevenly distributed across the genome by location, codon position, and functional group. Furthermore, phylogenetic analyses provide the strongest support so far for the hypothesis that the magnoliids are sister to a large clade that includes both monocots and eudicots.

Base Composition↗

EST database for early flower development in California poppy (Eschscholzia californica Cham., Papaveraceae) tags over 6,000 genes from a basal eudicot.

The Floral Genome Project (FGP) selected California poppy (Eschscholzia californica Cham. ssp. Californica) to help identify new florally-expressed genes related to floral diversity in basal eudicots. A large, non-normalized cDNA library was constructed from premeiotic and meiotic floral buds and sequenced to generate a database of 9,079 high quality Expressed Sequence Tags (ESTs). These sequences clustered into 5,713 unigenes, including 1,414 contigs and 4,299 singletons. Homologs of genes regulating many aspects of flower development were identified, including those for organ identity and development, cell and tissue differentiation, cell cycle control, and secondary metabolism. Over 5% of the transcriptome consisted of homologs to known floral gene families. Most are the first representatives of their respective gene families in basal eudicots and their conservation suggests they are important for floral development and/or function. App. 10% of the transcripts encoded transcription factors and other regulatory genes, including nine genes from the seven major lineages of the important MADS-box family of developmental regulators. Homologs of alkaloid pathway genes were also recovered, providing opportunities to explore adaptive evolution in secondary products. Furthermore, comparison of the poppy ESTs with the Arabidopsis genome provided support for putative Arabidopsis genes that previously lacked annotation. Finally, over 1,800 unique sequences had no observable homology in the public databases. The California poppy EST database and library will help bridge our understanding of flower initiation and development among higher eudicot and monocot model plants and provide new opportunities for comparative analysis of gene families across angiosperm species.

DNA, Complementary↗

Widespread genome duplications throughout the history of flowering plants.

Genomic comparisons provide evidence for ancient genome-wide duplications in a diverse array of animals and plants. We developed a birth-death model to identify evidence for genome duplication in EST data, and applied a mixture model to estimate the age distribution of paralogous pairs identified in EST sets for species representing the basal-most extant flowering plant lineages. We found evidence for episodes of ancient genome-wide duplications in the basal angiosperm lineages including Nuphar advena (yellow water lily: Nymphaeaceae) and the magnoliids Persea americana (avocado: Lauraceae), Liriodendron tulipifera (tulip poplar: Magnoliaceae), and Saruma henryi (Aristolochiaceae). In addition, we detected independent genome duplications in the basal eudicot Eschscholzia californica (California poppy: Papaveraceae) and the basal monocot Acorus americanus (Acoraceae), both of which were distinct from duplications documented for ancestral grass (Poaceae) and core eudicot lineages. Among gymnosperms, we found equivocal evidence for ancient polyploidy in Welwitschia mirabilis (Gnetales) and no evidence for polyploidy in pine, although gymnosperms generally have much larger genomes than the angiosperms investigated. Cross-species sequence divergence estimates suggest that synonymous substitution rates in the basal angiosperms are less than half those previously reported for core eudicots and members of Poaceae. These lower substitution rates permit inference of older duplication events. We hypothesize that evidence of an ancient duplication observed in the Nuphar data may represent a genome duplication in the common ancestor of all or most extant angiosperms, except Amborella.

Cycadopsida↗

Adaptive evolution of chloroplast genome structure inferred using a parametric bootstrap approach.

BACKGROUND: Genome rearrangements influence gene order and configuration of gene clusters in all genomes. Most land plant chloroplast DNAs (cpDNAs) share a highly conserved gene content and with notable exceptions, a largely co-linear gene order. Conserved gene orders may reflect a slow intrinsic rate of neutral chromosomal rearrangements, or selective constraint. It is unknown to what extent observed changes in gene order are random or adaptive. We investigate the influence of natural selection on gene order in association with increased rate of chromosomal rearrangement. We use a novel parametric bootstrap approach to test if directional selection is responsible for the clustering of functionally related genes observed in the highly rearranged chloroplast genome of the unicellular green alga Chlamydomonas reinhardtii, relative to ancestral chloroplast genomes. RESULTS: Ancestral gene orders were inferred and then subjected to simulated rearrangement events under the random breakage model with varying ratios of inversions and transpositions. We found that adjacent chloroplast genes in C. reinhardtii were located on the same strand much more frequently than in simulated genomes that were generated under a random rearrangement processes (increased sidedness; p < 0.0001). In addition, functionally related genes were found to be more clustered than those evolved under random rearrangements (p < 0.0001). We report evidence of co-transcription of neighboring genes, which may be responsible for the observed gene clusters in C. reinhardtii cpDNA. CONCLUSION: Simulations and experimental evidence suggest that both selective maintenance and directional selection for gene clusters are determinants of chloroplast gene order.

Adaptation, Physiological↗

ChloroplastDB: the Chloroplast Genome Database.

The Chloroplast Genome Database (ChloroplastDB) is an interactive, web-based database for fully sequenced plastid genomes, containing genomic, protein, DNA and RNA sequences, gene locations, RNA-editing sites, putative protein families and alignments (http://chloroplast.cbio.psu.edu/). With recent technical advances, the rate of generating new organelle genomes has increased dramatically. However, the established ontology for chloroplast genes and gene features has not been uniformly applied to all chloroplast genomes available in the sequence databases. For example, annotations for some published genome sequences have not evolved with gene naming conventions. ChloroplastDB provides unified annotations, gene name search, BLAST and download functions for chloroplast encoded genes and genomic sequences. A user can retrieve all orthologous sequences with one search regardless of gene names in GenBank. This feature alone greatly facilitates comparative research on sequence evolution including changes in gene content, codon usage, gene structure and post-transcriptional modifications such as RNA editing. Orthologous protein sets are classified by TribeMCL and each set is assigned a standard gene name. Over the next few years, as the number of sequenced chloroplast genomes increases rapidly, the tools available in ChloroplastDB will allow researchers to easily identify and compile target data for comparative analysis of chloroplast genes and genomes.

Chloroplasts↗

Gene capture prediction and overlap estimation in EST sequencing from one or multiple libraries.

BACKGROUND: In expressed sequence tag (EST) sequencing, we are often interested in how many genes we can capture in an EST sample of a targeted size. This information provides insights to sequencing efficiency in experimental design, as well as clues to the diversity of expressed genes in the tissue from which the library was constructed. RESULTS: We propose a compound Poisson process model that can accurately predict the gene capture in a future EST sample based on an initial EST sample. It also allows estimation of the number of expressed genes in one cDNA library or co-expressed in two cDNA libraries. The superior performance of the new prediction method over an existing approach is established by a simulation study. Our analysis of four Arabidopsis thaliana EST sets suggests that the number of expressed genes present in four different cDNA libraries of Arabidopsis thaliana varies from 9155 (root) to 12005 (silique). An observed fraction of co-expressed genes in two different EST sets as low as 25% can correspond to an actual overlap fraction greater than 65%. CONCLUSION: The proposed method provides a convenient tool for gene capture prediction and cDNA library property diagnosis in EST sequencing.

Algorithms↗

Expression pattern shifts following duplication indicative of subfunctionalization and neofunctionalization in regulatory genes of Arabidopsis.

Gene duplication plays an important role in the evolution of diversity and novel function and is especially prevalent in the nuclear genomes of flowering plants. Duplicate genes may be maintained through subfunctionalization and neofunctionalization at the level of expression or coding sequence. In order to test the hypothesis that duplicated regulatory genes will be differentially expressed in a specific manner indicative of regulatory subfunctionalization and/or neofunctionalization, we examined expression pattern shifts in duplicated regulatory genes in Arabidopsis. A two-way analysis of variance was performed on expression data for 280 phylogenetically identified paralogous pairs. Expression data were extracted from global expression profiles for wild-type root, stem, leaf, developing inflorescence, nearly mature flower buds, and seedpod. Gene, organ, and gene by organ interaction (G x O) effects were examined. Results indicate that 85% of the paralogous pairs exhibited a significant G x O effect indicative of regulatory subfunctionalization and/or neofunctionalization. A significant G x O effect was associated with complementary expression patterns in 45% of pairwise comparisons. No association was detected between a G x O effect and a relaxed evolutionary constraint as detected by the ratio of nonsynonymous to synonymous substitutions. Ancestral gene expression patterns inferred across a Type II MADS-box gene phylogeny suggest several cases of regulatory neofunctionalization and organ-specific nonfunctionalization. Complete linkage clustering of gene expression levels across organs suggests that regulatory modules for each organ are independent or ancestral genes had limited expression. We propose a new classification, regulatory hypofunctionalization, for an overall decrease in expression level in one member of a paralogous pair while still having a significant G x O effect. We conclude that expression divergence specifically indicative of subfunctionalization and/or neofunctionalization contributes to the maintenance of most if not all duplicated regulatory genes in Arabidopsis and hypothesize that this results in increasing expression diversity or specificity of regulatory genes after each round of duplication.

Arabidopsis↗

Rate variation in parasitic plants: correlated and uncorrelated patterns among plastid genes of different function.

BACKGROUND: The analysis of synonymous and nonsynonymous rates of DNA change can help in the choice among competing explanations for rate variation, such as differences in constraint, mutation rate, or the strength of genetic drift. Nonphotosynthetic plants of the Orobanchaceae have increased rates of DNA change. In this study 38 taxa of Orobanchaceae and relatives were used and 3 plastid genes were sequenced for each taxon. RESULTS: Phylogenetic reconstructions of relative rates of sequence evolution for three plastid genes (rbcL, matK and rps2) show significant rate heterogeneity among lineages and among genes. Many of the non-photosynthetic plants have increases in both synonymous and nonsynonymous rates, indicating that both (1) selection is relaxed, and (2) there has been a change in the rate at which mutations are entering the population in these species. However, rate increases are not always immediate upon loss of photosynthesis. Overall there is a poor correlation of synonymous and nonsynonymous rates. There is, however, a strong correlation of synonymous rates across the 3 genes studied and the lineage-speccific pattern for each gene is strikingly similar. This indicates that the causes of synonymous rate variation are affecting the whole plastid genome in a similar way. There is a weaker correlation across genes for nonsynonymous rates. Here the picture is more complex, as could be expected if there are many causes of variation, differing from taxon to taxon and gene to gene. CONCLUSIONS: The distinctive pattern of rate increases in Orobanchaceae has at least two causes. It is clear that there is a relaxation of constraint in many (though not all) non-photosynthetic lineages. However, there is also some force affecting synonymous sites as well. At this point, it is not possible to tell whether it is generation time, speciation rate, mutation rate, DNA repair efficiency or some combination of these factors.

Cell Nucleus↗

Methods for obtaining and analyzing whole chloroplast genome sequences.

During the past decade, there has been a rapid increase in our understanding of plastid genome organization and evolution due to the availability of many new completely sequenced genomes. There are 45 complete genomes published and ongoing projects are likely to increase this sampling to nearly 200 genomes during the next 5 years. Several groups of researchers including ours have been developing new techniques for gathering and analyzing entire plastid genome sequences and details of these developments are summarized in this chapter. The most important developments that enhance our ability to generate whole chloroplast genome sequences involve the generation of pure fractions of chloroplast genomes by whole genome amplification using rolling circle amplification, cloning genomes into Fosmid or bacterial artificial chromosome (BAC) vectors, and the development of an organellar annotation program (Dual Organellar GenoMe Annotator [DOGMA]). In addition to providing details of these methods, we provide an overview of methods for analyzing complete plastid genome sequences for repeats and gene content, as well as approaches for using gene order and sequence data for phylogeny reconstruction. This explosive increase in the number of sequenced plastid genomes and improved computational tools will provide many insights into the evolution of these genomes and much new data for assessing relationships at deep nodes in plants and other photosynthetic organisms.

Amino Acid Sequence↗

EST clustering error evaluation and correction.

MOTIVATION: The gene expression intensity information conveyed by (EST) Expressed Sequence Tag data can be used to infer important cDNA library properties, such as gene number and expression patterns. However, EST clustering errors, which often lead to greatly inflated estimates of obtained unique genes, have become a major obstacle in the analyses. The EST clustering error structure, the relationship between clustering error and clustering criteria, and possible error correction methods need to be systematically investigated. RESULTS: We identify and quantify two types of EST clustering error, namely, Type I and II in EST clustering using CAP3 assembling program. A Type I error occurs when ESTs from the same gene do not form a cluster whereas a Type II error occurs when ESTs from distinct genes are falsely clustered together. While the Type II error rate is <1.5% for both 5' and 3' EST clustering, the Type I error in the 5' EST case is approximately 10 times higher than the 3' EST case (30% versus 3%). An over-stringent identity rule, e.g., P >/= 95%, may even inflate the Type I error in both cases. We demonstrate that approximately 80% of the Type I error is due to insufficient overlap among sibling ESTs (ISO error) in 5' EST clustering. A novel statistical approach is proposed to correct ISO error to provide more accurate estimates of the true gene cluster profile.

Algorithms↗

Highly heterogeneous rates of evolution in the SKP1 gene family in plants and animals: functional and evolutionary implications.

Skp1 (S-phase kinase-associated protein 1) is a core component of SCF ubiquitin ligases and mediates protein degradation, thereby regulating eukaryotic fundamental processes such as cell cycle progression, transcriptional regulation, and signal transduction. Among the four components of the SCF complexes, Rbx1 and Cullin form a core catalytic complex, an F-box protein acts as a receptor for target proteins, and Skp1 is an adaptor between one of the variable F-box proteins and Cullin. Whereas protists, fungi, and some vertebrates have a single SKP1 gene, many animal and plant species possess multiple SKP1 homologs. It has been shown that the same Skp1 homolog can interact with two or more F-box proteins, and different Skp1 homologs from the same species sometimes can interact with the same F-box protein. In this paper, we demonstrate that multiple Skp1 homologs from the same species have evolved at highly heterogeneous rates. Parametric bootstrap analyses suggested that the differences in evolutionary rate are so large that true phylogenies were not recoverable from the full data set. Only when the original data set were partitioned into sets of genes with slow, medium, and rapid rates of evolution and analyzed separately, better-resolved relationships were observed. The slowly evolving Skp1 homologs, which are relatively highly conserved in sequence and expressed widely and/or at high levels, usually have very low d(N)/d(S) values, suggesting that they have evolved under functional constraint and serve the most fundamental function(s). On the other hand, the rapidly evolving members are structurally more diverse and usually have limited expression patterns and higher d(N)/d(S) values, suggesting that they may have evolved under relaxed or altered constraint, or even under positive selection. Some rapidly evolving members may have lost their original function(s) and/or acquired new function(s) or become pseudogenes, as suggested by their expression patterns, d(N)/d(S) values, and amino acid changes at key positions. In addition, our analyses revealed several monophyletic groups within the SKP1 gene family, one for each of protists, fungi, animals, and plants, as well as nematodes, arthropods, and angiosperms, suggesting that the extant SKP1 genes within each of these eukaryote groups shared only one common ancestor.

Amino Acid Sequence↗

Antiquity and evolution of the MADS-box gene family controlling flower development in plants.

MADS-box genes in plants control various aspects of development and reproductive processes including flower formation. To obtain some insight into the roles of these genes in morphological evolution, we investigated the origin and diversification of floral MADS-box genes by conducting molecular evolutionary genetics analyses. Our results suggest that the most recent common ancestor of today's floral MADS-box genes evolved roughly 650 MYA, much earlier than the Cambrian explosion. They also suggest that the functional classes T (SVP), B (and Bs), C, F (AGL20 or TM3), A, and G (AGL6) of floral MADS-box genes diverged sequentially in this order from the class E gene lineage. The divergence between the class G and E genes apparently occurred around the time of the angiosperm/gymnosperm split. Furthermore, the ancestors of three classes of genes (class T genes, class B/Bs genes, and the common ancestor of the other classes of genes) might have existed at the time of the Cambrian explosion. We also conducted a phylogenetic analysis of MADS-domain sequences from various species of plants and animals and presented a hypothetical scenario of the evolution of MADS-box genes in plants and animals, taking into account paleontological information. Our study supports the idea that there are two main evolutionary lineages (type I and type II) of MADS-box genes in plants and animals.

Animals↗

Missing links: the genetic architecture of flowers [correction of flower] and floral diversification.

To understand the genetic architecture of floral development, including the origin and subsequent diversification of the flower, data are needed not only for a few model organisms but also for gymnosperms, basal angiosperm lineages and early-diverging eudicots. We must link what is known about derived model plants such as Arabidopsis, snapdragon and maize with other angiosperms. To this end, we suggest a massive evolutionary genomics effort focused on the identification and expression patterns of floral genes and elucidation of their expression patterns in 'missing-link' taxa differing in the arrangement, number and organization of floral parts.

Cycadopsida↗

The Chlamydomonas reinhardtii plastid chromosome: islands of genes in a sea of repeats.

Chlamydomonas reinhardtii is a unicellular eukaryotic alga possessing a single chloroplast that is widely used as a model system for the study of photosynthetic processes. This report analyzes the surprising structural and evolutionary features of the completely sequenced 203,395-bp plastid chromosome. The genome is divided by 21.2-kb inverted repeats into two single-copy regions of approximately 80 kb and contains only 99 genes, including a full complement of tRNAs and atypical genes encoding the RNA polymerase. A remarkable feature is that >20% of the genome is repetitive DNA: the majority of intergenic regions consist of numerous classes of short dispersed repeats (SDRs), which may have structural or evolutionary significance. Among other sequenced chlorophyte plastid genomes, only that of the green alga Chlorella vulgaris appears to share this feature. The program MultiPipMaker was used to compare the genic complement of Chlamydomonas with those of other chloroplast genomes and to scan the genomes for sequence similarities and repetitive DNAs. Among the results was evidence that the SDRs were not derived from extant coding sequences, although some SDRs may have arisen from other genomic fragments. Phylogenetic reconstruction of changes in plastid genome content revealed that an accelerated rate of gene loss also characterized the Chlamydomonas/Chlorella lineage, a phenomenon that might be independent of the proliferation of SDRs. Together, our results reveal a dynamic and unusual plastid genome whose existence in a model organism will allow its features to be tested functionally.

Amino Acid Sequence↗

Conservation and divergence in the AGAMOUS subfamily of MADS-box genes: evidence of independent sub- and neofunctionalization events.

The MADS-box gene AGAMOUS (AG) plays a key role in determining floral meristem and organ identities. We identified three AG homologs, EScaAG1, EScaAG2, and EScaAGL11 from the basal eudicot Eschscholzia californica (California poppy). Phylogenetic analyses indicate that EScaAG1 and EScaAG2 are recent paralogs within the AG clade, independent of the duplication in ancestral core eudicots that gave rise to the euAG and PLENA (PLE) orthologs. EScaAGL11 is basal to core eudicot AGL11 orthologs in a clade representing an older duplication event after the divergence of the angiosperm and gymnosperm lineages. Detailed in situ hybridization experiments show that expression of EScaAG1 and EScaAG2 is similar to AG; however, both genes appear to be expressed earlier in floral development than described in the core eudicots. A thorough examination of available expression and functional data in a phylogenetic context for members of the AG and AGL11 clades reveals that gene expression has been quite variable throughout the evolutionary history of the AG subfamily and that ovule-specific expression might have evolved more than twice. Although sub- and neofunctionalization are inferred to have occurred following gene duplication, functional divergence among orthologs is evident, as is convergence, among paralogs sampled from different species. We propose that retention of multiple AG homologs in several paralogous lineages can be explained by the conservation of ancestral protein activity combined with evolutionarily labile regulation of expression in the AG and AGL11 clades such that the collective functions of the AG subfamily in stamen and carpel development are maintained following gene duplication.

Amino Acid Sequence↗