Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Contig Mapping”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 739 records · Page 41Linked to original sources

The phusion assembler.

The Phusion assembler has assembled the mouse genome from the whole-genome shotgun (WGS) dataset collected by the Mouse Genome Sequencing Consortium, at ~7.5x sequence coverage, producing a high-quality draft assembly 2.6 gigabases in size, of which 90% of these bases are in 479 scaffolds. For the mouse genome, which is a large and repeat-rich genome, the input dataset was designed to include a high proportion of paired end sequences of various size selected inserts, from 2-200 kbp lengths, into various host vector templates. Phusion uses sequence data, called reads, and information about reads that share common templates, called read pairs, to drive the assembly of this large genome to highly accurate results. The preassembly stage, which clusters the reads into sensible groups, is a key element of the entire assembler, because it permits a simple approach to parallelization of the assembly stage, as each cluster can be treated independent of the others. In addition to the application of Phusion to the mouse genome, we will also present results from the WGS assembly of Caenorhabditis briggsae sequenced to about 11x coverage. The C. briggsae assembly was accessioned through EMBL, http://www.ebi.ac.uk/services/index.html, using the series CAAC01000001-CAAC01000578, however, the Phusion mouse assembly described here was not accessioned. The mouse data was generated by the Mouse Genome Sequencing Consortium. The C. briggsae sequence was generated at The Wellcome Trust Sanger Institute and the Genome Sequencing Center, Washington University School of Medicine.

Animals↗

Analysis of the quality and utility of random shotgun sequencing at low redundancies.

The currently favored approach for sequencing the human genome involves selecting representative large-insert clones (100-200 kb), randomly shearing this DNA to construct shotgun libraries, and then sequencing many different isolates from the library. This method, entitled directed random shotgun sequencing, requires highly redundant sequencing to obtain a complete and accurate finished consensus sequence. Recently it has been suggested that a rapidly generated lower redundancy sequence might be of use to the scientific community. Low-redundancy sequencing has been examined previously using simulated data sets. Here we utilize trace data from a number of projects submitted to GenBank to perform reconstruction experiments that mimic low-redundancy sequencing. These low-redundancy sequences have been examined for the completeness and quality of the consensus product, information content, and usefulness for interspecies comparisons. The data presented here suggest three different sequencing strategies, each with different utilities. (1) Nearly complete sequence data can be obtained by sequencing a random shotgun library at sixfold redundancy. This may therefore represent a good point to switch from a random to directed approach. (2) Sequencing can be performed with as little as twofold redundancy to find most of the information about exons, EST hits, and putative exon similarity matches. (3) To obtain contiguity of coding regions, sequencing at three- to fourfold redundancy would be appropriate. From these results, we suggest that a useful intermediate product for genome sequencing might be obtained by three- to fourfold redundancy. Such a product would allow a large amount of biologically useful data to be extracted while postponing the majority of work involved in producing a high quality consensus sequence.

Animals↗

Structure and evolution of the Smith-Magenis syndrome repeat gene clusters, SMS-REPs.

An approximately 4-Mb genomic segment on chromosome 17p11.2, commonly deleted in patients with the Smith-Magenis syndrome (SMS) and duplicated in patients with dup(17)(p11.2p11.2) syndrome, is flanked by large, complex low-copy repeats (LCRs), termed proximal and distal SMS-REP. A third copy, the middle SMS-REP, is located between them. SMS-REPs are believed to mediate nonallelic homologous recombination, resulting in both SMS deletions and reciprocal duplications. To delineate the genomic structure and evolutionary origin of SMS-REPs, we constructed a bacterial artificial chromosome/P1 artificial chromosome contig spanning the entire SMS region, including the SMS-REPs, determined its genomic sequence, and used fluorescence in situ hybridization to study the evolution of SMS-REP in several primate species. Our analysis shows that both the proximal SMS-REP (approximately 256 kb) and the distal copy (approximately 176 kb) are located in the same orientation and derived from a progenitor copy, whereas the middle SMS-REP (approximately 241 kb) is inverted and appears to have been derived from the proximal copy. The SMS-REP LCRs are highly homologous (>98%) and contain at least 14 genes/pseudogenes each. SMS-REPs are not present in mice and were duplicated after the divergence of New World monkeys from pre-monkeys approximately 40-65 million years ago. Our findings potentially explain why the vast majority of SMS deletions and dup(17)(p11.2p11.2) occur at proximal and distal SMS-REPs and further support previous observations that higher-order genomic architecture involving LCRs arose recently during primate speciation and may predispose the human genome to both meiotic and mitotic rearrangements.

Abnormalities, Multiple↗

Whole-genome sequence assembly for mammalian genomes: Arachne 2.

We previously described the whole-genome assembly program Arachne, presenting assemblies of simulated data for small to mid-sized genomes. Here we describe algorithmic adaptations to the program, allowing for assembly of mammalian-size genomes, and also improving the assembly of smaller genomes. Three principal changes were simultaneously made and applied to the assembly of the mouse genome, during a six-month period of development: (1) Supercontigs (scaffolds) were iteratively broken and rejoined using several criteria, yielding a 64-fold increase in length (N50), and apparent elimination of all global misjoins; (2) gaps between contigs in supercontigs were filled (partially or completely) by insertion of reads, as suggested by pairing within the supercontig, increasing the N50 contig length by 50%; (3) memory usage was reduced fourfold. The outcome of this mouse assembly and its analysis are described in (Mouse Genome Sequencing Consortium 2002).

Animals↗

Identification of candidate coding region single nucleotide polymorphisms in 165 human genes using assembled expressed sequence tags.

Using assembled expressed sequence tags (ESTs) from 50 different cDNA libraries, we have identified contigs that represent the complete coding sequences of 850 known human genes, and have scanned these for high quality sequence substitutions. We report the identification and characteristics of 201 candidate single nucleotide polymorphisms found in the coding sequences (cSNPs) of 165 of these genes. Using a conservative calculation, coding region nucleotide diversity (the average number of differences between any pair of chromosomes) was found to be 3 per 10,000 bp based on this data. This analysis reveals that assembled ESTs from multiple libraries may provide a rich source of comparative sequences to search for cSNPs in the human genome.

Amino Acid Substitution↗

Frequent alternative splicing of human genes.

Alternative splicing can produce variant proteins and expression patterns as different as the products of different genes, yet the prevalence of alternative splicing has not been quantified. Here the spliced alignment algorithm was used to make a first inventory of exon-intron structures of known human genes using EST contigs from the TIGR Human Gene Index. The results on any one gene may be incomplete and will require verification, yet the overall trends are significant. Evidence of alternative splicing was shown in 35% of genes and the majority of splicing events occurred in 5' untranslated regions, suggesting wide occurrence of alternative regulation. Most of the alternative splices of coding regions generated additional protein domains rather than alternating domains.

5' Untranslated Regions↗

Duplications on human chromosome 22 reveal a novel Ret Finger Protein-like gene family with sense and endogenous antisense transcripts.

Analysis of 600 kb of sequence encompassing the beta-prime adaptin (BAM22) gene on human chromosome 22 revealed intrachromosomal duplications within 22q12-13 resulting in three active RFPL genes, two RFPL pseudogenes, and two pseudogenes of BAM22. The genomic sequence of BAM22vartheta1 shows a remarkable similarity to that of BAM22. The cDNA sequence comparison of RFPL1, RFPL2, and RFPL3 showed 95%-96% identity between the genes, which were most similar to the Ret Finger Protein gene from human chromosome 6. The sense RFPL transcripts encode proteins with the tripartite structure, composed of RING finger, coiled-coil, and B30-2 domains, which are characteristic of the RING-B30 family. Each of these domains are thought to mediate protein-protein interactions by promoting homo- or heterodimerization. The MID1 gene on Xp22 is also a member of the RING-B30 family and is mutated in Opitz syndrome (OS). The autosomal dominant form of OS shows linkage to 22q11-q12. We detected a polymorphic protein-truncating allele of RFPL1 in 8% of the population, which was not associated with the OS phenotype. We identified 6-kb and 1.2-kb noncoding antisense mRNAs of RFPL1S and RFPL3S antisense genes, respectively. The RFPL1S and RFPL3S genes cover substantial portions of their sense counterparts, which suggests that the function of RFPL1S and RFPL3S is a post-transcriptional regulation of the sense RFPL genes. We illustrate the role of intrachromosomal duplications in the generation of RFPL genes, which were created by a series of duplications and share an ancestor with the RING-B30 domain containing genes from the major histocompatibility complex region on human chromosome 6.

Adaptor Protein Complex 1↗

CAP3: A DNA sequence assembly program.

We describe the third generation of the CAP sequence assembly program. The CAP3 program includes a number of improvements and new features. The program has a capability to clip 5' and 3' low-quality regions of reads. It uses base quality values in computation of overlaps between reads, construction of multiple sequence alignments of reads, and generation of consensus sequences. The program also uses forward-reverse constraints to correct assembly errors and link contigs. Results of CAP3 on four BAC data sets are presented. The performance of CAP3 was compared with that of PHRAP on a number of BAC data sets. PHRAP often produces longer contigs than CAP3 whereas CAP3 often produces fewer errors in consensus sequences than PHRAP. It is easier to construct scaffolds with CAP3 than with PHRAP on low-pass data with forward-reverse constraints.

Algorithms↗

Development and evaluation of an automated annotation pipeline and cDNA annotation system.

Manual curation has long been held to be the "gold standard" for functional annotation of DNA sequence. Our experience with the annotation of more than 20,000 full-length cDNA sequences revealed problems with this approach, including inaccurate and inconsistent assignment of gene names, as well as many good assignments that were difficult to reproduce using only computational methods. For the FANTOM2 annotation of more than 60,000 cDNA clones, we developed a number of methods and tools to circumvent some of these problems, including an automated annotation pipeline that provides high-quality preliminary annotation for each sequence by introducing an "uninformative filter" that eliminates uninformative annotations, controlled vocabularies to accurately reflect both the functional assignments and the evidence supporting them, and a highly refined, Web-based manual annotation tool that allows users to view a wide array of sequence analyses and to assign gene names and putative functions using a consistent nomenclature. The ultimate utility of our approach is reflected in the low rate of reassignment of automated assignments by manual curation. Based on these results, we propose a new standard for large-scale annotation, in which the initial automated annotations are manually investigated and then computational methods are iteratively modified and improved based on the results of manual curation.

Animals↗

Segmental duplications: organization and impact within the current human genome project assembly.

Segmental duplications play fundamental roles in both genomic disease and gene evolution. To understand their organization within the human genome, we have developed the computational tools and methods necessary to detect identity between long stretches of genomic sequence despite the presence of high copy repeats and large insertion-deletions. Here we present our analysis of the most recent genome assembly (January 2001) in which we focus on the global organization of these segments and the role they play in the whole-genome assembly process. Initially, we considered only large recent duplication events that fell well-below levels of draft sequencing error (alignments 90%-98% similar and > or =1 kb in length). Duplications (90%-98%; > or =1 kb) comprise 3.6% of all human sequence. These duplications show clustering and up to 10-fold enrichment within pericentromeric and subtelomeric regions. In terms of assembly, duplicated sequences were found to be over-represented in unordered and unassigned contigs indicating that duplicated sequences are difficult to assign to their proper position. To assess coverage of these regions within the genome, we selected BACs containing interchromosomal duplications and characterized their duplication pattern by FISH. Only 47% (106/224) of chromosomes positive by FISH had a corresponding chromosomal position by comparison. We present data that indicate that this is attributable to misassembly, misassignment, and/or decreased sequencing coverage within duplicated regions. Surprisingly, if we consider putative duplications >98% identity, we identify 10.6% (286 Mb) of the current assembly as paralogous. The majority of these alignments, we believe, represent unmerged overlaps within unique regions. Taken together the above data indicate that segmental duplications represent a significant impediment to accurate human genome assembly, requiring the development of specialized techniques to finish these exceptional regions of the genome. The identification and characterization of these highly duplicated regions represents an important step in the complete sequencing of a human reference genome.

Base Sequence↗

Using genomic resources to guide research directions. The arabinogalactan protein gene family as a test case.

Arabinogalactan proteins (AGPs) are extracellular hydroxyproline-rich proteoglycans implicated in plant growth and development. The protein backbones of AGPs are rich in proline/hydroxyproline, serine, alanine, and threonine. Most family members have less than 40% similarity; therefore, finding family members using Basic Local Alignment Search Tool searches is difficult. As part of our systematic analysis of AGP function in Arabidopsis, we wanted to make sure that we had identified most of the members of the gene family. We used the biased amino acid composition of AGPs to identify AGPs and arabinogalactan (AG) peptides in the Arabidopsis genome. Different criteria were used to identify the fasciclin-like AGPs. In total, we have identified 13 classical AGPs, 10 AG-peptides, three basic AGPs that include a short lysine-rich region, and 21 fasciclin-like AGPs. To streamline the analysis of genomic resources to assist in the planning of targeted experimental approaches, we have adopted a flow chart to maximize the information that can be obtained about each gene. One of the key steps is the reformatting of the Arabidopsis Functional Genomics Consortium microarray data. This customized software program makes it possible to view the ratio data for all Arabidopsis Functional Genomics Consortium experiments and as many genes as desired in a single spreadsheet. The results for reciprocal experiments are grouped to simplify analysis and candidate AGPs involved in development or biotic and abiotic stress responses are readily identified. The microarray data support the suggestion that different AGPs have different functions.

Acids↗

Comparison of RNA expression profiles based on maize expressed sequence tag frequency analysis and micro-array hybridization.

Assembly of 73,000 expressed sequence tags (ESTs) representing multiple organs and developmental stages of maize (Zea mays) identified approximately 22,000 tentative unique genes (TUGs) at the criterion of 95% identity. Based on sequence similarity, overlap between any two of nine libraries with more than 3,000 ESTs ranged from 4% to 20% of the constituent TUGs. The most abundant ESTs were recovered from only one or a minority of the libraries, and only 26 EST contigs had members from all nine EST sets (presumably representing ubiquitously expressed genes). For several examples, ESTs for different members of gene families were detected in distinct organs. To study this further, two types of micro-array slides were fabricated, one containing 5,534 ESTs from 10- to 14-d-old endosperm, and the other 4,844 ESTs from immature ear, estimated to represent about 2,800 and 2,500 unique genes, respectively. Each array type was hybridized with fluorescent cDNA targets prepared from endosperm and immature ear poly(A(+)) RNA. Although the 10- to 14-d-old postpollination endosperm TUGs showed only 12% overlap with immature ear TUGs, endosperm target hybridized with 94% of the ear TUGs, and ear target hybridized with 57% of the endosperm TUGs. Incomplete EST sampling of low-abundance transcripts contributes to an underestimate of shared gene expression profiles. Reassembly of ESTs at the criterion of 90% identity suggests how cross hybridization among gene family members can overestimate the overlap in genes expressed in micro-array hybridization experiments.

Contig Mapping↗

Chlamydomonas reinhardtii genome project. A guide to the generation and use of the cDNA information.

The National Science Foundation-funded Chlamydomonas reinhardtii genome project involves (a) construction and sequencing of cDNAs isolated from cells exposed to various environmental conditions, (b) construction of a high-density cDNA microarray, (c) generation of genomic contigs that are nucleated around specific physical and genetic markers, (d) generation of a complete chloroplast genome sequence and analyses of chloroplast gene expression, and (e) the creation of a Web-based resource that allows for easy access of the information in a format that can be readily queried. Phases of the project performed by the groups at the Carnegie Institution and Duke University involve the generation of normalized cDNA libraries, sequencing of cDNAs, analysis and assembly of these sequences to generate contigs and a set of predicted unique genes, and the use of this information to construct a high-density DNA microarray. In this paper, we discuss techniques involved in obtaining cDNA end-sequence information and the ways in which this information is assembled and analyzed. Descriptions of protocols for preparing cDNA libraries, assembling cDNA sequences and annotating the sequence information are provided (the reader is directed to Web sites for more detailed descriptions of these methods). We also discuss preliminary results in which the different cDNA libraries are used to identify genes that are potentially differentially expressed.

Animals↗

Nylon filter arrays reveal differential gene expression in proteoid roots of white lupin in response to phosphorus deficiency.

White lupin (Lupinus albus) adapts to phosphorus deficiency (-P) by the development of short, densely clustered lateral roots called proteoid (or cluster) roots. In an effort to better understand the molecular events mediating these adaptive responses, we have isolated and sequenced 2,102 expressed sequence tags (ESTs) from cDNA libraries prepared with RNA isolated at different stages of proteoid root development. Determination of overlapping regions revealed 322 contigs (redundant copy transcripts) and 1,126 singletons (single-copy transcripts) that compile to a total of 1,448 unique genes (unigenes). Nylon filter arrays with these 2,102 ESTs from proteoid roots were performed to evaluate global aspects of gene expression in response to -P stress. ESTs differentially expressed in P-deficient proteoid roots compared with +P and -P normal roots include genes involved in carbon metabolism, secondary metabolism, P scavenging and remobilization, plant hormone metabolism, and signal transduction.

Algorithms↗

Rapid genome evolution revealed by comparative sequence analysis of orthologous regions from four triticeae genomes.

Bread wheat (Triticum aestivum) is an allohexaploid species, consisting of three subgenomes (A, B, and D). To study the molecular evolution of these closely related genomes, we compared the sequence of a 307-kb physical contig covering the high molecular weight (HMW)-glutenin locus from the A genome of durum wheat (Triticum turgidum, AABB) with the orthologous regions from the B genome of the same wheat and the D genome of the diploid wheat Aegilops tauschii (Anderson et al., 2003; Kong et al., 2004). Although gene colinearity appears to be retained, four out of six genes including the two paralogous HMW-glutenin genes are disrupted in the orthologous region of the A genome. Mechanisms involved in gene disruption in the A genome include retroelement insertions, sequence deletions, and mutations causing in-frame stop codons in the coding sequences. Comparative sequence analysis also revealed that sequences in the colinear intergenic regions of these different genomes were generally not conserved. The rapid genome evolution in these regions is attributable mainly to the large number of retrotransposon insertions that occurred after the divergence of the three wheat genomes. Our comparative studies indicate that the B genome diverged prior to the separation of the A and D genomes. Furthermore, sequence comparison of two distinct types of allelic variations at the HMW-glutenin loci in the A genomes of different hexaploid wheat cultivars with the A genome locus of durum wheat indicates that hexaploid wheat may have more than one tetraploid ancestor.

Amino Acid Sequence↗

Rapid genome divergence at orthologous low molecular weight glutenin loci of the A and Am genomes of wheat.

To study genome evolution in wheat, we have sequenced and compared two large physical contigs of 285 and 142 kb covering orthologous low molecular weight (LMW) glutenin loci on chromosome 1AS of a diploid wheat species (Triticum monococcum subsp monococcum) and a tetraploid wheat species (Triticum turgidum subsp durum). Sequence conservation between the two species was restricted to small regions containing the orthologous LMW glutenin genes, whereas >90% of the compared sequences were not conserved. Dramatic sequence rearrangements occurred in the regions rich in repetitive elements. Dating of long terminal repeat retrotransposon insertions revealed different insertion events occurring during the last 5.5 million years in both species. These insertions are partially responsible for the lack of homology between the intergenic regions. In addition, the gene space was conserved only partially, because different predicted genes were identified on both contigs. Duplications and deletions of large fragments that might be attributable to illegitimate recombination also have contributed to the differentiation of this region in both species. The striking differences in the intergenic landscape between the A and A(m) genomes that diverged 1 to 3 million years ago provide evidence for a dynamic and rapid genome evolution in wheat species.

Base Sequence↗

Large intraspecific haplotype variability at the Rph7 locus results from rapid and recent divergence in the barley genome.

To study genome evolution and diversity in barley (Hordeum vulgare), we have sequenced and compared more than 300 kb of sequence spanning the Rph7 leaf rust disease resistance gene in two barley cultivars. Colinearity was restricted to five genic and two intergenic regions representing <35% of the two sequences. In each interval separating the seven conserved regions, the number and type of repetitive elements were completely different between the two homologous sequences, and a single gene was absent in one cultivar. In both cultivars, the nonconserved regions consisted of approximately 53% repetitive sequences mainly represented by long-terminal repeat retrotransposons that have inserted <1 million years ago. PCR-based analysis of intergenic regions at the Rph7 locus and at three other independent loci in 41 H. vulgare lines indicated large haplotype variability in the cultivated barley gene pool. Together, our data indicate rapid and recent divergence at homologous loci in the genome of H. vulgare, possibly providing the molecular mechanism for the generation of high diversity in the barley gene pool. Finally, comparative analysis of the gene composition in barley, wheat (Triticum aestivum), rice (Oryza sativa), and sorghum (Sorghum bicolor) suggested massive gene movements at the Rph7 locus in the Triticeae lineage.

Conserved Sequence↗

Comparative genomics of Brassica oleracea and Arabidopsis thaliana reveal gene loss, fragmentation, and dispersal after polyploidy.

We sequenced 2.2 Mb representing triplicated genome segments of Brassica oleracea, which are each paralogous with one another and homologous with a segmentally duplicated region of the Arabidopsis thaliana genome. Sequence annotation identified 177 conserved collinear genes in the B. oleracea genome segments. Analysis of synonymous base substitution rates indicated that the triplicated Brassica genome segments diverged from a common ancestor soon after divergence of the Arabidopsis and Brassica lineages. This conclusion was corroborated by phylogenetic analysis of protein families. Using A. thaliana as an outgroup, 35% of the genes inferred to be present when genome triplication occurred in the Brassica lineage have been lost, most likely via a deletion mechanism, in an interspersed pattern. Genes encoding proteins involved in signal transduction or transcription were not found to be significantly more extensively retained than those encoding proteins classified with other functions, but putative proteins predicted in the A. thaliana genome were underrepresented in B. oleracea. We identified one example of gene loss from the Arabidopsis lineage. We found evidence for the frequent insertion of gene fragments of nuclear genomic origin and identified four apparently intact genes in noncollinear positions in the B. oleracea and A. thaliana genomes.

Arabidopsis↗