Evolution of the Spiroplasma P58 multigene family.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Polyploidy (genome duplication) is thought to have contributed to the evolution of the eukaryotic genome, but complex genome structures and massive gene loss during evolution has complicated detection of these ancestral duplication events. The major factors determining the fate of duplicated genes are currently unclear, as are the processes by which duplicated genes evolve after polyploidy. Fine-scale analysis between homologous regions may allow us to better understand post-polyploidy evolution. Here, using gene-by-gene and gene-by-genome strategies, we identified the S5 region and four homologous regions within the japonica genome. Additional phylogenomic analyses of the comparable duplicated blocks indicate that four successive duplication events gave rise to these five regions, allowing us to propose a model for this local chromosomal evolution. According to this model, gene loss may play a major role in post-duplication genetic evolution at the segmental level. Moreover, we found molecular evidence that one of the sister duplicated blocks experienced more gene loss and a more rapid evolution subsequent to two recent duplication events. Given that these two recent duplication events were likely involved in polyploidy, this asymmetric evolution (gene loss and gene divergence) may be one possible mechanism accounting for the diploidization at the segmental level.
We have constructed a high-resolution cytogenetic map with 168 DNA markers, including 90 RFLP markers for human chromosome 11. The cosmid clones were mapped by fluorescence in situ suppression hybridization, in which discrete fluorescent signals can be detected directly on prometaphase R-banded chromosomes. Although these cosmid clones were distributed throughout the chromosome, they had some tendency to localize in the regions of R-positive band, such as 11p15, 11p11.2, 11q13, 11q23, and 11q25. Since these regions of chromosome 11 are considered to contain genes responsible for certain genetic diseases, cancer breakpoints involved in chromosome rearrangements, and tumor-suppressor genes, this high-resolution cytogenetic map will contribute to the molecular characterization of such genes. This map will also provide many landmarks essential for construction of the complete physical map with contigs of cosmid and YAC clones.
Explore the source record for details and available documents.
We analyzed the expressed sequence tags (ESTs) obtained from a cDNA library of the eyestalk of the kuruma prawn, Marsupenaeus japonicus, to examine gene expression profile with special focus on female reproduction. The assembly of 1988 ESTs created 136 contigs from 738 ESTs; however 1250 ESTs remained singletons. Significant similarities (blast score > or = 50 bits) to the DNA sequences in the databank were found for only 16.7% of the 1386 sequences (136 contigs plus 1250 singletons), suggesting that the eyestalk library contains many unknown genes. Ribosomal RNA and mitochondrial respiration enzymes with significant similarities were found abundantly in the ESTs, whereas genes related to maturation or endocrine systems were scarce. Three ESTs were assumed to encode novel eyestalk hormones with marked similarities to pigment-dispersing hormone, molt-inhibiting hormone and crustacean hyperglycemic hormone. Sequences encoding a product highly homologous to farnesoic acid O-methyltransferase, an enzyme that produces methyl farnesoate, were also found.
Gene duplication, silencing and translocation have all been implicated in shaping the unique genomic architecture of the teleost MH regions. Previously, we demonstrated that trout possess five unlinked regions encoding MH genes. One of these regions harbors ABCB2 which in all other vertebrate classes is found in the MHC class II region. In this study, we sequenced a BAC contig for the trout ABCB2 region. Analysis of this region revealed the presence of genes homologous to those located in the human class II (ABCB2, BRD2, psiDAA), extended class II (RGL2, PHF1, SYGP1) and class III (PBX2, Notch-L) regions. The organization and syntenic relationships of this region were then compared to similar regions in humans, Tetraodon and zebrafish to learn more about the evolutionary history of this region. Our analysis indicates that this region was generated during the teleost-specific duplication event while also providing insight about potential MH paralogous regions in teleosts.
We constructed and characterized a bacterial artificial chromosome (BAC) library for Epichloë festucae, a genetically tractable fungal plant mutualist. The 6144 clone library with an average insert size of 87kb represents at least 18-fold coverage of the 29 Mb genome. We used the library to assemble a 110kb contig spanning the putative ornithine decarboxylase (odc) ortholog and subsequently expanded it to 228kb with a single walking step in each direction. Furthermore, we evaluated conservation of microsynteny between E. festucae and some model filamentous fungi by comparing sequence available from a 43kb region at the end of one BAC to publicly available fungal genome sequences. Orthologs to the 13 contiguous open reading frames (ORFs) identified in E. festucae are syntenic in Neurospora crassa and Magnaporthe grisea occurring in small sets of two, three or four colinear ORFs. This library is a valuable resource for research into traits important for the development and maintenance of a plant-fungus mutualistic symbiosis.
Mycosphaerella graminicola is a major fungal pathogen of wheat as the causal agent of Septoria leaf blotch disease. As a first step toward a greater understanding of the mechanism of host infection we have generated, sequenced, and analyzed three M. graminicola EST libraries from conditions predicted to resemble independent phases of the host infection process, including one library generated from the fungus during interaction with its host. A total of 5180 ESTs were sequenced and clustered into 886 contigs and 2039 singletons to give a set of 2925 unique sequences (unisequences). BLASTX analysis revealed 33% of the unknown M. graminicola unisequences to be orphans. Very limited inter-library overlap of expression was seen with the majority of unisequences (contigs and singletons) being library-specific. Analysis of EST redundancy between libraries demonstrated a significant difference in gene expression in the three conditions. Comparisons made against fully sequenced genomes revealed most M. graminicola sequences to be homologous to genes present in both pathogenic and non-pathogenic Ascomycete filamentous fungi. A range of sequences having significant homology to verified pathogenicity/virulence genes (HvPV-genes) of either plant or mammalian fungal and Oomycete pathogens were also identified (<1e-20). The generation of, and the diversity present within, this EST collection will facilitate future efforts aimed at a more detailed study of the transcriptome of the fungus during host infection.
Rust fungi are plant parasites which colonise host tissue with an intercellular mycelium that forms haustoria within living plant cells. To identify genes expressed during biotrophic growth, EST sequencing was performed with a haustorium-specific cDNA library from Uromyces fabae. One thousand seventeen ESTs were generated, which assembled into 530 contigs. Several of the most frequently represented sequences in the EST database were identical to the in planta induced genes (PIGs) identified previously (Hahn, M., Mendgen, K., 1997. Characterisation of in planta-induced rust genes isolated from a haustorium-specific cDNA library, Mol. Plant-Microbe Interact. 10, 427-437). Virus-encoded sequences were identified, providing evidence for two novel RNA mycoviruses in U. fabae. Microarray hybridisation revealed many cDNAs that were significantly activated in rust-infected leaves compared to germinated uredospores. Very strong in planta expression was found for two PIGs encoding putative metallothioneins. Furthermore, several genes involved in ribosome biogenesis and translation, glycolysis, amino acid metabolism, stress response, and detoxification showed an increased expression in the parasitic mycelium. These data indicate a strong shift in gene expression in rust fungi between germination and the biotrophic stage of development.
The organization and expression of a putative serine/threonine kinase gene (designated hcstk), proposed to relate to a conserved eukaryotic signal transduction pathway, was characterized for the socio-economically important pathogen Haemonchus contortus (Nematoda). The entire hcstk gene is approximately 26.7 kb in size, has 26 exons and is inferred to produce multiple isoforms via alternative splicing in its N-terminal header and spacer domains. Comparison of hcstk with its Caenorhabditis elegans homologue, par-1, revealed major differences in genomic organization, exon number and inferred mRNA processing. The expression of hcstk transcripts was highest in the first- and late-fourth-stage larvae of the parasite compared with other developmental stages, somewhat distinct from par-1 in C. elegans. In spite of a substantial amino acid sequence identity in the functional domains between the predicted proteins HcSTK and PAR-1, overall, the findings suggest a unique functional role for each molecule.
We completely sequenced a 516,013-bp portion of the porcine genome that encompassed a cluster of genes for chemokine (C-C motif) receptors (CC chemokine receptors). We identified genes for six CC chemokine receptors (CCR1, CCR2, CCR3, CCR5, CCR9, and CCRL2) and two other chemokine receptors (CXCR6 and XCR1) in this region. Clarification of the entire structure of the region and the respective genes revealed their high conservation among human, mouse, and pig. Interestingly, much of the 5'UTR of porcine XCR1 shared an identical sequence with CCR1; this sharing does not occur in humans or mice. This finding suggests a mechanism for posttranscriptional switching of tandem-located genes in mammals that depends on alternative splicing. Furthermore, our findings contribute to analyses of lymphocyte trafficking and the functions of immune cells in pigs and other artiodactyls.
We have adopted a method of telomere-mediated chromosome fragmentation in order to demonstrate the alignment of contigs and determination of gaps. We established the order and orientation of four contigs of Candida albicans chromosome 5 and determined the sizes of three gaps between these contigs. We confirmed this proposed alignment of contigs, as well as gap sizes, by sequencing one gap and analyzing three mega deletions of approximately 41 kbp, 58 kbp, and 77 kbp, which covered two other gaps. These gaps could be also conveniently sequenced, which is an important step in establishing a complete sequence. The combined length of contigs and gaps covered approximately 422 kbp, which is one third of chromosome 5. Telomere-mediated chromosome fragmentation, used here for the first time to align the contigs of C. albicans and determine the gaps, proved to be a reliable method. The method could be helpful in sequencing projects of other diploid organisms, in particular those in which centromeres have not been identified. In addition, our approach can be used to assign any contig to a chromosome, or to induce the loss of a specific chromosome.
We made use of 81,635 expressed sequence tags (ESTs) derived from 12 different cDNA libraries of the silkworm, Bombyx mori, inbred strain Dazao (P50), to identify high-quality candidate single nucleotide polymorphisms (SNPs). By PHRAP assembling, 12,980 contigs containing 11,537 contigs assembled by more than one read were obtained, and 101 candidate SNPs and 27 single base insertions/deletions were identified from 117 contigs assembled from 1576 high-quality reads base-called with PHRED and screened on the basis of the neighborhood quality standard (NQS). Simultaneously, we also predicted 40 SNPs in coding regions (cSNPs), of which 26 were predicted to lead to amino acid non-synonymous variations and 14 synonymous substitutions. Also, the 1.66:1 ratio of transition/transversion is different from that of other insects. As the first SNP analysis of a Lepidoptera, B. mori, the single nucleotide polymorphic density is estimated to be 1.3 x 10(-3) by sequence diversity. This analysis shows that expressed sequences from multiple libraries may provide an abundant source of comparative reads to mine for cSNPs from the silkworm genome.
Gregarines are protozoan parasites of invertebrates in the phylum Apicomplexa. We employed an expressed sequence tag strategy in order to dissect the molecular processes of sexual or gametocyst development of gregarines. Expressed sequence tags provide a rapid way to identify genes, particularly in organisms for which we have very little molecular information. Analysis of approximately 1800 expressed sequence tags from the gametocyst stage revealed highly expressed genes related to cell division and differentiation. Evidence was found for the role of degradation and recycling in gametocyst development. Numerous additional genes uncovered by expressed sequence tag sequencing should provide valuable tools to investigate gametocyst development as well as for molecular phylogenetics, and comparative genomics in this important group of parasites.
Lactobacillus paracasei NFBC338 is a probiotic strain that was isolated from the human gastrointestinal tract (GIT) and contains a plasmid genome of 80kb. Using a shotgun sequencing approach, two of the plasmids, pCD01 (19,882bp) and pCD02 (8554bp) have been completely sequenced, and four contiguous sequences (Contigs) have been assembled. Bioinformatic analysis of pCD01 revealed that it contains 23 putative open reading frames (ORFs) and that it contains regions characterised by potential replication functions and multidrug resistance (MDR). In contrast, the content of pCD02 is mainly cryptic, although, it does contain two insertion sequence (IS) elements. Indeed, up to 17% of the entire plasmid genome encodes putative transposable elements. In addition, there are a number of interesting ORFs distributed over the four Contigs that show significant homology to genes such as those involved in adherence and biotin metabolism, which may prove beneficial to Lb. paracasei NFBC338 under certain environmental conditions. This study provides a novel insight into the rich plasmid complement of this probiotic Lactobacillus strain, which may potentially be exploited as the basis for development of improved genetic tools for probiotic lactobacilli.
One of the major challenges in genome research is the identification of the complete set of genes in a genome. Alignments of expressed sequences (RNA and EST) with genomic sequences have been used to characterize genes. However, the number of alignments far exceeds the likely number of genes in a genome, suggesting that, for many genes, two or more alignments can be joined through overlapping sequences to yield accurate gene structures. High-throughput EST sequencing becomes less efficient in closing those alignment gaps due to its nonselective nature. We sought to bridge these alignments through a novel approach: targeted cDNA sequencing. Human expressed sequences from GenBank version 124 were aligned with the genomic sequence from NCBI build 24 using LEADS, Compugen's EST and RNA clustering and assembly software system. Nine hundred forty-eight pairs of alignments were selected based on EST clone information and/or their homology to the same known proteins. Reverse transcriptase PCR and sequencing yielded sequences for 363 of those pairs. These sequences helped characterize over 60 novel or otherwise incomplete genes in the recent UniGene build 153, which included over 1 million additional ESTs. These results indicate that this integrated and targeted strategy, combining computational prediction and experimental cDNA sequencing, can efficiently generate the overlapping sequences and enable the full characterization of genomes. Additional information about the contig pairs, the resultant overlapping sequences, tissue sources, and tissue profiles are available in a supplemental file.
Our previous study described the amplification of a genomic sequence containing exon 9 of CFTR in the human genome. Here we report that this CFTR sequence is part of a large duplicated sequence unit, provisionally named LCR7-20. Through successive screening of two human chromosome 7-specific cosmid libraries to construct a cosmid contig, we assembled two sequenced BAC clones into a single contig containing a prototypic LCR7-20 unit. Subsequent searches of existing human genome sequences identified additional six copies of LCR7-20-like sequences with more than 90% sequence homology. Additional genomic clones containing LCR7-20-like sequences were then isolated from total genomic BAC and PAC libraries. Restriction fragment analysis and limited sequencing data indicated that there could be around 30 copies of LCR7-20-like sequences in the human genome and that the average region of homology could extend over 120 kb. As indicated by fluorescence in situ hybridization analysis, LCR7-20-like sequences are dispersed on different chromosomes, mainly in the centromeric and pericentromeric regions, and some may exist in tandem copies. Our study also indicates that many genomic regions containing LCR7-20's either have been misassembled or are missing in current versions of the human genome sequence.
Sixteen CC chemokine genes localize to a 2.06-Mb interval at 17q11.2-q12 on genomic contig NT_010799.13. Four of these genes comprise two closely related paralogous pairs: CCL3-CCL3L1 and CCL4-CCL4L1. Members within each pair share 95% sequence identity at both the genomic and the amino acid levels. One BAC clone (AC131056.5) on the contig with substantial internal sequence duplication contains two complete copies of CCL3L1 and CCL4L1 and one truncated copy of CCL3L1, while a partially overlapping clone (AC003976.1) contains one copy each of CCL3 and CCL4. Dot-matrix comparison of the regions of AC131056.5 with those of AC003976.1 containing the four genes reveals 90% sequence similarity over 37 kb. These observations support the idea that the multiple copies of CCL3L1 and CCL4L1 present in a single diploid genome are the result of segmental duplication.