Search PubMed⌕ Search

Biomedical subjects

Christopher D Town

Publications and source records attributed to Christopher D Town.

At least 19 recordsLinked to original sources

Experimental validation of novel genes predicted in the un-annotated regions of the Arabidopsis genome.

BACKGROUND: Several lines of evidence support the existence of novel genes and other transcribed units which have not yet been annotated in the Arabidopsis genome. Two gene prediction programs which make use of comparative genomic analysis, Twinscan and EuGene, have recently been deployed on the Arabidopsis genome. The ability of these programs to make use of sequence data from other species has allowed both Twinscan and EuGene to predict over 1000 genes that are intergenic with respect to the most recent annotation release. A high throughput RACE pipeline was utilized in an attempt to verify the structure and expression of these novel genes. RESULTS: 1,071 un-annotated loci were targeted by RACE, and full length sequence coverage was obtained for 35% of the targeted genes. We have verified the structure and expression of 378 genes that were not present within the most recent release of the Arabidopsis genome annotation. These 378 genes represent a structurally diverse set of transcripts and encode a functionally diverse set of proteins. CONCLUSION: We have investigated the accuracy of the Twinscan and EuGene gene prediction programs and found them to be reliable predictors of gene structure in Arabidopsis. Several hundred previously un-annotated genes were validated by this work. Based upon this information derived from these efforts it is likely that the Arabidopsis genome annotation continues to overlook several hundred protein coding genes.

Arabidopsis↗

Sequencing Medicago truncatula expressed sequenced tags using 454 Life Sciences technology.

BACKGROUND: In this study, we addressed whether a single 454 Life Science GS20 sequencing run provides new gene discovery from a normalized cDNA library, and whether the short reads produced via this technology are of value in gene structure annotation. RESULTS: A single 454 GS20 sequencing run on adapter-ligated cDNA, from a normalized cDNA library, generated 292,465 reads that were reduced to 252,384 reads with an average read length of 92 nucleotides after cleaning. After clustering and assembly, a total of 184,599 unique sequences were generated containing over 400 SSRs. The 454 sequences generated hits to more genes than a comparable amount of sequence from MtGI. Although short, the 454 reads are of sufficient length to map to a unique genome location as effectively as longer ESTs produced by conventional sequencing. Functional interpretation of the sequences was carried out by Gene Ontology assignments from matches to Arabidopsis and was shown to cover a broad range of GO categories. 53,796 assemblies and singletons (29%) had no match in the existing MtGI. Within the previously unobserved Medicago transcripts, thousands had matches in a comprehensive protein database and one or more of the TIGR Plant Gene Indices. Approximately 20% of these novel sequences could be found in the Medicago genome sequence. A total of 70,026 reads generated by the 454 technology were mapped to 785 Medicago finished BACs using PASA and over 1,000 gene models required modification. In parallel to 454 sequencing, 4,445 5'-prime reads were generated by conventional sequencing using the same library and from the assembled sequences it was shown to contain about 52% full length cDNAs encoding proteins from 50 to over 500 amino acids in length. CONCLUSION: Due to the large number of reads afforded by the 454 DNA sequencing technology, it is effective in revealing the expression of transcripts from a broad range of GO categories and contains many rare transcripts in normalized cDNA libraries, although only a limited portion of their sequence is uncovered. As with longer ESTs, 454 reads can be mapped uniquely onto genomic sequence to provide support for, and modifications of, gene predictions.

Base Sequence↗

Legume genome evolution viewed through the Medicago truncatula and Lotus japonicus genomes.

Genome sequencing of the model legumes, Medicago truncatula and Lotus japonicus, provides an opportunity for large-scale sequence-based comparison of two genomes in the same plant family. Here we report synteny comparisons between these species, including details about chromosome relationships, large-scale synteny blocks, microsynteny within blocks, and genome regions lacking clear correspondence. The Lotus and Medicago genomes share a minimum of 10 large-scale synteny blocks, each with substantial collinearity and frequently extending the length of whole chromosome arms. The proportion of genes syntenic and collinear within each synteny block is relatively homogeneous. Medicago-Lotus comparisons also indicate similar and largely homogeneous gene densities, although gene-containing regions in Mt occupy 20-30% more space than Lj counterparts, primarily because of larger numbers of Mt retrotransposons. Because the interpretation of genome comparisons is complicated by large-scale genome duplications, we describe synteny, synonymous substitutions and phylogenetic analyses to identify and date a probable whole-genome duplication event. There is no direct evidence for any recent large-scale genome duplication in either Medicago or Lotus but instead a duplication predating speciation. Phylogenetic comparisons place this duplication within the Rosid I clade, clearly after the split between legumes and Salicaceae (poplar).

Chromosomes, Plant↗

Comparative sequence and genetic analyses of asparagus BACs reveal no microsynteny with onion or rice.

The Poales (includes the grasses) and Asparagales [includes onion (Allium cepa L.) and asparagus (Asparagus officinalis L.)] are the two most economically important monocot orders. The Poales are a member of the commelinoid monocots, a group of orders sister to the Asparagales. Comparative genomic analyses have revealed a high degree of synteny among the grasses; however, it is not known if this synteny extends to other major monocot groups such as the Asparagales. Although we previously reported no evidence for synteny at the recombinational level between onion and rice, microsynteny may exist across shorter genomic regions in the grasses and Asparagales. We sequenced nine asparagus BACs to reveal physically linked genic-like sequences and determined their most similar positions in the onion and rice genomes. Four of the asparagus BACs were selected using molecular markers tightly linked to the sex-determining M locus on chromosome 5 of asparagus. These BACs possessed only two putative coding regions and had long tracts of degenerated retroviral elements and transposons. Five asparagus BACs were selected after hybridization of three onion cDNAs that mapped to three different onion chromosomes. Genic-like sequences that were physically linked on the cDNA-selected BACs or genetically linked on the M-linked BACs showed significant similarities (e < -20) to expressed sequences on different rice chromosomes, revealing no evidence for microsynteny between asparagus and rice across these regions. Genic-like sequences that were linked in asparagus were used to identify highly similar (e < -20) expressed sequence tags (ESTs) of onion. These onion ESTs mapped to different onion chromosomes and no relationship was observed between physical or genetic linkages in asparagus and genetic linkages in onion. These results further indicate that synteny among grass genomes does not extend to a sister order in the monocots and that asparagus may not be an appropriate smaller genome model for plants in the Asparagales with enormous nuclear genomes.

Asparagus Plant↗

Complete plastid genome sequence of Daucus carota: implications for biotechnology and phylogeny of angiosperms.

BACKGROUND: Carrot (Daucus carota) is a major food crop in the US and worldwide. Its capacity for storage and its lifecycle as a biennial make it an attractive species for the introduction of foreign genes, especially for oral delivery of vaccines and other therapeutic proteins. Until recently efforts to express recombinant proteins in carrot have had limited success in terms of protein accumulation in the edible tap roots. Plastid genetic engineering offers the potential to overcome this limitation, as demonstrated by the accumulation of BADH in chromoplasts of carrot taproots to confer exceedingly high levels of salt resistance. The complete plastid genome of carrot provides essential information required for genetic engineering. Additionally, the sequence data add to the rapidly growing database of plastid genomes for assessing phylogenetic relationships among angiosperms. RESULTS: The complete carrot plastid genome is 155,911 bp in length, with 115 unique genes and 21 duplicated genes within the IR. There are four ribosomal RNAs, 30 distinct tRNA genes and 18 intron-containing genes. Repeat analysis reveals 12 direct and 2 inverted repeats > or = 30 bp with a sequence identity > or = 90%. Phylogenetic analysis of nucleotide sequences for 61 protein-coding genes using both maximum parsimony (MP) and maximum likelihood (ML) were performed for 29 angiosperms. Phylogenies from both methods provide strong support for the monophyly of several major angiosperm clades, including monocots, eudicots, rosids, asterids, eurosids II, euasterids I, and euasterids II. CONCLUSION: The carrot plastid genome contains a number of dispersed direct and inverted repeats scattered throughout coding and non-coding regions. This is the first sequenced plastid genome of the family Apiaceae and only the second published genome sequence of the species-rich euasterid II clade. Both MP and ML trees provide very strong support (100% bootstrap) for the sister relationship of Daucus with Panax in the euasterid II clade. These results provide the best taxon sampling of complete chloroplast genomes and the strongest support yet for the sister relationship of Caryophyllales to the asterids. The availability of the complete plastid genome sequence should facilitate improved transformation efficiency and foreign gene expression in carrot through utilization of endogenous flanking sequences and regulatory elements.

Amino Acid Sequence↗

Accumulation of genome-specific transcripts, transcription factors and phytohormonal regulators during early stages of fiber cell development in allotetraploid cotton.

Gene expression during the early stages of fiber cell development and in allopolyploid crops is poorly understood. Here we report computational and expression analyses of 32 789 high-quality ESTs derived from Gossypium hirsutum L. Texas Marker-1 (TM-1) immature ovules (GH_TMO). The ESTs were assembled into 8540 unique sequences including 4036 tentative consensus sequences (TCs) and 4504 singletons, representing approximately 15% of the unique sequences in the cotton EST collection. Compared with approximately 178 000 existing ESTs derived from elongating fibers and non-fiber tissues, GH_TMO ESTs showed a significant increase in the percentage of genes encoding putative transcription factors such as MYB and WRKY and genes encoding predicted proteins involved in auxin, brassinosteroid (BR), gibberellic acid (GA), abscisic acid (ABA) and ethylene signaling pathways. Cotton homologs related to MIXTA, MYB5, GL2 and eight genes in the auxin, BR, GA and ethylene pathways were induced during fiber cell initiation but repressed in the naked seed mutant (N1N1) that is impaired in fiber formation. The data agree with the known roles of MYB and WRKY transcription factors in Arabidopsis leaf trichome development and the well-documented phytohormonal effects on fiber cell development in immature cotton ovules cultured in vitro. Moreover, the phytohormonal pathway-related genes were induced prior to the activation of MYB-like genes, suggesting an important role of phytohormones in cell fate determination. Significantly, AA sub-genome ESTs of all functional classifications including cell-cycle control and transcription factor activity were selectively enriched in G. hirsutum L., an allotetraploid derived from polyploidization between AA and DD genome species, a result consistent with the production of long lint fibers in AA genome species. These results suggest general roles for genome-specific, phytohormonal and transcriptional gene regulation during the early stages of fiber cell development in cotton allopolyploids.

Arabidopsis Proteins↗

Comparative genomics of Brassica oleracea and Arabidopsis thaliana reveal gene loss, fragmentation, and dispersal after polyploidy.

We sequenced 2.2 Mb representing triplicated genome segments of Brassica oleracea, which are each paralogous with one another and homologous with a segmentally duplicated region of the Arabidopsis thaliana genome. Sequence annotation identified 177 conserved collinear genes in the B. oleracea genome segments. Analysis of synonymous base substitution rates indicated that the triplicated Brassica genome segments diverged from a common ancestor soon after divergence of the Arabidopsis and Brassica lineages. This conclusion was corroborated by phylogenetic analysis of protein families. Using A. thaliana as an outgroup, 35% of the genes inferred to be present when genome triplication occurred in the Brassica lineage have been lost, most likely via a deletion mechanism, in an interspersed pattern. Genes encoding proteins involved in signal transduction or transcription were not found to be significantly more extensively retained than those encoding proteins classified with other functions, but putative proteins predicted in the A. thaliana genome were underrepresented in B. oleracea. We identified one example of gene loss from the Arabidopsis lineage. We found evidence for the frequent insertion of gene fragments of nuclear genomic origin and identified four apparently intact genes in noncollinear positions in the B. oleracea and A. thaliana genomes.

Arabidopsis↗

The complete chloroplast genome sequence of Gossypium hirsutum: organization and phylogenetic relationships to other angiosperms.

BACKGROUND: Cotton (Gossypium hirsutum) is the most important fiber crop grown in 90 countries. In 2004-2005, US farmers planted 79% of the 5.7-million hectares of nuclear transgenic cotton. Unfortunately, genetically modified cotton has the potential to hybridize with other cultivated and wild relatives, resulting in geographical restrictions to cultivation. However, chloroplast genetic engineering offers the possibility of containment because of maternal inheritance of transgenes. The complete chloroplast genome of cotton provides essential information required for genetic engineering. In addition, the sequence data were used to assess phylogenetic relationships among the major clades of rosids using cotton and 25 other completely sequenced angiosperm chloroplast genomes. RESULTS: The complete cotton chloroplast genome is 160,301 bp in length, with 112 unique genes and 19 duplicated genes within the IR, containing a total of 131 genes. There are four ribosomal RNAs, 30 distinct tRNA genes and 17 intron-containing genes. The gene order in cotton is identical to that of tobacco but lacks rpl22 and infA. There are 30 direct and 24 inverted repeats 30 bp or longer with a sequence identity > or = 90%. Most of the direct repeats are within intergenic spacer regions, introns and a 72 bp-long direct repeat is within the psaA and psaB genes. Comparison of protein coding sequences with expressed sequence tags (ESTs) revealed nucleotide substitutions resulting in amino acid changes in ndhC, rpl23, rpl20, rps3 and clpP. Phylogenetic analysis of a data set including 61 protein-coding genes using both maximum likelihood and maximum parsimony were performed for 28 taxa, including cotton and five other angiosperm chloroplast genomes that were not included in any previous phylogenies. CONCLUSION: Cotton chloroplast genome lacks rpl22 and infA and contains a number of dispersed direct and inverted repeats. RNA editing resulted in amino acid changes with significant impact on their hydropathy. Phylogenetic analysis provides strong support for the position of cotton in the Malvales in the eurosids II clade sister to Arabidopsis in the Brassicales. Furthermore, there is strong support for the placement of the Myrtales sister to the eurosid I clade, although expanded taxon sampling is needed to further test this relationship.

Chloroplasts↗

Annotating the genome of Medicago truncatula.

Medicago truncatula will be among the first plant species to benefit from the completion of a whole-genome sequencing project. For each of these species, Arabidopsis, rice and now poplar and Medicago, annotation, the process of identifying gene structures and defining their functions, is essential for the research community to benefit from the sequence data generated. Annotation of the Arabidopsis genome involved gene-by-gene curation of the entire genome, but the larger genomes of rice, Medicago and other species necessitate the automation of the annotation process. Profiting from the experience gained from previous whole-genome efforts, a uniform set of Medicago gene annotations has been generated by coordinated international effort and, along with other views of the genome data, has been provided to the research community at several websites.

Automation↗

The Soybean Genome Database (SoyGD): a browser for display of duplicated, polyploid, regions and sequence tagged sites on the integrated physical and genetic maps of Glycine max.

Genomes that have been highly conserved following increases in ploidy (by duplication or hybridization) like Glycine max (soybean) present challenges during genome analysis. At http://soybeangenome.siu.edu the Soybean Genome Database (SoyGD) genome browser has, since 2002, integrated and served the publicly available soybean physical map, bacterial artificial chromosome (BAC) fingerprint database and genetic map associated genomic data. The browser shows both build 3 and build 4 contiguous sets of clones (contigs) of the soybean physical map. Build 4 consisted of 2854 contigs that encompassed 1.05 Gb and 404 high-quality DNA markers that anchored 742 contigs. Many DNA markers anchored sets of 2-8 different contigs. Each contig in the set represented a homologous region of related sequences. GBrowse was adapted to show sets of homologous contigs at all potential anchor points, spread laterally and prevented from overlapping. About 8064 minimum tiling path (MTP2) clones provided 13,473 BAC end sequences (BES) to decorate the physical map. Analyses of BES placed 2111 gene models, 40 marker anchors and 1053 new microsatellite markers on the map. Estimated sequence tag probes from 201 low-copy gene families located 613 paralogs. The genome browser portal showed each data type as a separate track. Tetraploid, octoploid, diploid and homologous regions are shown clearly in relation to an integrated genetic and physical map.

Chromosome Mapping↗

Development of Arabidopsis whole-genome microarrays and their application to the discovery of binding sites for the TGA2 transcription factor in salicylic acid-treated plants.

We have developed two long-oligonucleotide microarrays for the analysis of genome features in Arabidopsis thaliana, in particular for the high-throughput identification of transcription factor-binding sites. The first platform contains 190,000 probes representing the 2-kb regions upstream of all annotated genes at a density of seven probes per promoter. The second platform is divided into three chips, each of over 390,000 features, and represents the entire Arabidopsis genome at a density of one probe per 90 bases. Protein-DNA complexes resulting from the formaldehyde fixation of leaves of plants 2 h after exposure to 1 mm salicylic acid (SA) were immunoprecipitated using antibodies against the TGA2 transcription factor. After reversal of the cross-links and amplification, the resulting ChIP sample was hybridized to both platforms. High signal ratios of the ChIP sample versus raw chromatin for clusters of neighboring probes provided evidence for 51 putative binding sites for TGA2, including the only previously confirmed site in the promoter of PR-1 (At2g14610). Enrichment of several regions was confirmed by quantitative real-time PCR. Motif search revealed that the palindromic octamer TGACGTCA was found in 55% of the enriched regions. Interestingly, 15 of the putative binding sites for TGA2 lie outside the presumptive promoter regions. The effect of the 2-h SA treatment on gene expression was measured using Affymetrix ATH1 arrays, and SA-induced genes were found to be significantly over-represented among genes neighboring putative TGA2-binding sites.

Arabidopsis↗

Simultaneous high-throughput recombinational cloning of open reading frames in closed and open configurations.

Comprehensive open reading frame (ORF) clone collections, ORFeomes, are key components of functional genomics projects. When recombinational cloning systems are used to capture ORFs in master clones, these DNA sequences can be easily transferred into a variety of expression plasmids, each designed for a specific assay. Depending on downstream applications, an ORF is cloned either with or without a stop codon at its original position, referred to as closed or open configuration, respectively. The former is preferred when the encoded protein is produced in its native form or with an amino-terminal tag; the latter is obligatory when the protein is produced as a fusion with a carboxyl-terminal tag. We developed a streamlined protocol for high-throughput, simultaneous cloning of both open and closed ORF entry clones with the Gateway recombinational cloning system. The protocol is straightforward to set up in large-scale ORF cloning projects, and is cost-effective, because the initial ORF amplification and the cloning in a pDONR vector are performed only once to obtain the two ORF configurations. We illustrated its implementation for the isolation and validation of 346 Arabidopsis ORF entry clones.

Arabidopsis↗

Analysis of the cDNAs of hypothetical genes on Arabidopsis chromosome 2 reveals numerous transcript variants.

In the fully sequenced Arabidopsis (Arabidopsis thaliana) genome, many gene models are annotated as "hypothetical protein," whose gene structures are predicted solely by computer algorithms with no support from either expressed sequence matches from Arabidopsis, or nucleic acid or protein homologs from other species. In order to confirm their existence and predicted gene structures, a high-throughput method of rapid amplification of cDNA ends (RACE) was used to obtain their cDNA sequences from 11 cDNA populations. Primers from all of the 797 hypothetical genes on chromosome 2 were designed, and, through 5' and 3' RACE, clones from 506 genes were sequenced and cDNA sequences from 399 target genes were recovered. The cDNA sequences were obtained by assembling their 5' and 3' RACE polymerase chain reaction products. These sequences revealed that (1) the structures of 151 hypothetical genes were different from their predictions; (2) 116 hypothetical genes had alternatively spliced transcripts and 187 genes displayed polyadenylation sites; and (3) there were transcripts arising from both strands, from the strand opposite to that of the prediction and possible dicistronic transcripts. Promoters from five randomly chosen hypothetical genes (At2g02540, At2g31270, At2g33640, At2g35550, and At2g36340) were cloned into report constructs, and their expressions are tissue or development stage specific. Our results indicate at least 50% of hypothetical genes on chromosome 2 are expressed in the cDNA populations with about 38% of the gene structures differing from their predictions. Thus, by using this targeted approach, high-throughput RACE, we revealed numerous transcripts including many uncharacterized variants from these hypothetical genes.

Alternative Splicing↗

Genetic mapping of expressed sequences in onion and in silico comparisons with rice show scant colinearity.

The Poales (which include the grasses) and Asparagales [which include onion (Allium cepa L.) and other Allium species] are the two most economically important monocot orders. Enormous genomic resources have been developed for the grasses; however, their applicability to other major monocot groups, such as the Asparagales, is unclear. Expressed sequence tags (ESTs) from onion that showed significant similarities (80% similarity over at least 70% of the sequence) to single positions in the rice genome were selected. One hundred new genetic markers developed from these ESTs were added to the intraspecific map derived from the BYG15-23xAC43 segregating family, producing 14 linkage groups encompassing 1,907 cM at LOD 4. Onion linkage groups were assigned to chromosomes using alien addition lines of Allium fistulosum L. carrying single onion chromosomes. Visual comparisons of genetic linkage in onion with physical linkage in rice revealed scant colinearity; however, short regions of colinearity could be identified. Our results demonstrate that the grasses may not be appropriate genomic models for other major monocot groups such as the Asparagales; this will make it necessary to develop genomic resources for these important plants.

Base Sequence↗

Highly syntenic regions in the genomes of soybean, Medicago truncatula, and Arabidopsis thaliana.

BACKGROUND: Recent genome sequencing enables mega-base scale comparisons between related genomes. Comparisons between animals, plants, fungi, and bacteria demonstrate extensive synteny tempered by rearrangements. Within the legume plant family, glimpses of synteny have also been observed. Characterizing syntenic relationships in legumes is important in transferring knowledge from model legumes to crops that are important sources of protein, fixed nitrogen, and health-promoting compounds. RESULTS: We have uncovered two large soybean regions exhibiting synteny with M. truncatula and with a network of segmentally duplicated regions in Arabidopsis. In all, syntenic regions comprise over 500 predicted genes spanning 3 Mb. Up to 75% of soybean genes are colinear with M. truncatula, including one region in which 33 of 35 soybean predicted genes with database support are colinear to M. truncatula. In some regions, 60% of soybean genes share colinearity with a network of A. thaliana duplications. One region is especially interesting because this 500 kbp segment of soybean is syntenic to two paralogous regions in M. truncatula on different chromosomes. Phylogenetic analysis of individual genes within these regions demonstrates that one is orthologous to the soybean region, with which it also shows substantially denser synteny and significantly lower levels of synonymous nucleotide substitutions. The other M. truncatula region is inferred to be paralogous, presumably resulting from a duplication event preceding speciation. CONCLUSION: The presence of well-defined M. truncatula segments showing orthologous and paralogous relationships with soybean allows us to explore the evolution of contiguous genomic regions in the context of ancient genome duplication and speciation events.

Arabidopsis↗

Analysis of indole-3-butyric acid-induced adventitious root formation on Arabidopsis stem segments.

Root induction by auxins is still not well understood at the molecular level. In this study a system has been devised which distinguishes between the two active auxins indole-3-butyric acid (IBA) and indole-3-acetic acid (IAA). IBA, but not IAA, efficiently induced adventitious rooting in Arabidopsis stem segments at a concentration of 10 microM. In wild-type plants, roots formed exclusively out of calli at the basal end of the segments. Root formation was inhibited by 10 microM 3,4,5-triiodobenzoic acid (TIBA), an inhibitor of polar auxin transport. At intermediate IBA concentrations (3-10 microM), root induction was less efficient in trp1, a tryptophan auxotroph of Arabidopsis with a bushy phenotype but no demonstrable reduction in IAA levels. By contrast, two mutants of Arabidopsis with measurably higher levels of IAA (trp2, amt1) show root induction characteristics very similar to the wild type. Using differential display, transcripts specific to the rooting process were identified by devising a protocol that distinguished between callus production only and callus production followed by root initiation. One fragment was identical to the sequence of a putative regulatory subunit B of protein phosphatase 2A. It is suggested that adventitious rooting in Arabidopsis stem segments is due to an interaction between endogenous IAA and exogenous IBA. In stem explants, residual endogenous IAA is transported to the basal end of each segment, thereby inducing root formation. In stem segments in which the polar auxin transport is inhibited by TIBA, root formation does not occur.

Arabidopsis↗

Complete reannotation of the Arabidopsis genome: methods, tools, protocols and the final release.

BACKGROUND: Since the initial publication of its complete genome sequence, Arabidopsis thaliana has become more important than ever as a model for plant research. However, the initial genome annotation was submitted by multiple centers using inconsistent methods, making the data difficult to use for many applications. RESULTS: Over the course of three years, TIGR has completed its effort to standardize the structural and functional annotation of the Arabidopsis genome. Using both manual and automated methods, Arabidopsis gene structures were refined and gene products were renamed and assigned to Gene Ontology categories. We present an overview of the methods employed, tools developed, and protocols followed, summarizing the contents of each data release with special emphasis on our final annotation release (version 5). CONCLUSION: Over the entire period, several thousand new genes and pseudogenes were added to the annotation. Approximately one third of the originally annotated gene models were significantly refined yielding improved gene structure annotations, and every protein-coding gene was manually inspected and classified using Gene Ontology terms.

Alternative Splicing↗

Comparative analysis of 87,000 expressed sequence tags from the fumonisin-producing fungus Fusarium verticillioides.

Fusarium verticillioides (teleomorph Gibberella moniliformis) is a pathogen of maize worldwide and produces fumonisins, a family of mycotoxins that have been associated with several animal diseases as well as cancer in humans. In this study, we sought to identify fungal genes that affect fumonisin production and/or the plant-fungal interaction. We generated over 87,000 expressed sequence tags from nine different cDNA libraries that correspond to 11,119 unique sequences and are estimated to represent 80% of the genomic complement of genes. A comparative analysis of the libraries showed that all 15 genes in the fumonisin gene cluster were differentially expressed. In addition, nine candidate fumonisin regulatory genes and a number of genes that may play a role in plant-fungal interaction were identified. Analysis of over 700 FUM gene transcripts from five different libraries provided evidence for transcripts with unspliced introns and spliced introns with alternative 3' splice sites. The abundance of the alternative splice forms and the frequency with which they were found for genes involved in the biosynthesis of a single family of metabolites as well as their differential expression suggest they may have a biological function. Finally, analysis of an EST that aligns to genomic sequence between FUM12 and FUM13 provided evidence for a previously unidentified gene (FUM20) in the FUM gene cluster.

Amino Acid Sequence↗