Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Contig Mapping”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40Linked to original sources

PCAP: a whole-genome assembly program.

We describe a whole-genome assembly program named PCAP for processing tens of millions of reads. The PCAP program has several features to address efficiency and accuracy issues in assembly. Multiple processors are used to perform most time-consuming computations in assembly. A more sensitive method is used to avoid missing overlaps caused by sequencing errors. Repetitive regions of reads are detected on the basis of many overlaps with other reads, instead of many shorter word matches with other reads. Contaminated end regions of reads are identified and removed. Generation of a consensus sequence for a contig is based on an alignment of reads in the contig, in which both base quality values and coverage information are used to determine every consensus base. The PCAP program was tested on a mouse whole-genome data set of 30 million reads and a human Chromosome 20 data set of 1.7 million reads. The program is freely available for academic use.

Algorithms↗

Hierarchical scaffolding with Bambus.

The output of a genome assembler generally comprises a collection of contiguous DNA sequences (contigs) whose relative placement along the genome is not defined. A procedure called scaffolding is commonly used to order and orient these contigs using paired read information. This ordering of contigs is an essential step when finishing and analyzing the data from a whole-genome shotgun project. Most recent assemblers include a scaffolding module; however, users have little control over the scaffolding algorithm or the information produced. We thus developed a general-purpose scaffolder, called Bambus, which affords users significant flexibility in controlling the scaffolding parameters. Bambus was used recently to scaffold the low-coverage draft dog genome data. Most significantly, Bambus enables the use of linking data other than that inferred from mate-pair information. For example, the sequence of a completed genome can be used to guide the scaffolding of a related organism. We present several applications of Bambus: support for finishing, comparative genomics, analysis of the haplotype structure of genomes, and scaffolding of a mammalian genome at low coverage. Bambus is available as an open-source package from our Web site.

Algorithms↗

Sequence and comparative analysis of the mouse 1-megabase region orthologous to the human 11p15 imprinted domain.

A major barrier to conceptual advances in understanding the mechanisms and regulation of imprinting of a genomic region is our relatively poor understanding of the overall organization of genes and of the potentially important cis-acting regulatory sequences that lie in the nonexonic segments that make up 97% of the genome. Interspecies sequence comparison offers an effective approach to identify sequence from conserved functional elements. In this article we describe the successful use of this approach in comparing a approximately 1-Mb imprinted genomic domain on mouse chromosome 7 to its orthologous region on human 11p15.5. Within the region, we identified 112 exons of known genes as well as a novel gene identified uniquely in the mouse region, termed Msuit, that was found to be imprinted. In addition to these coding elements, we identified 33 CpG islands and 49 orthologous nonexonic, nonisland sequences that met our criteria as being conserved, and making up 4.1% of the total sequence. These conserved noncoding sequence elements were generally clustered near imprinted genes and the majority were between Igf2 and H19 or within Kvlqt1. Finally, the location of CpG islands provided evidence that suggested a two-island rule for imprinted genes. This study provides the first global view of the architecture of an entire imprinted domain and provides candidate sequence elements for subsequent functional analyses.

Amino Acid Sequence↗

Genomic analysis in the sting-2 quantitative trait locus for defensive behavior in the honey bee, Apis mellifera.

We have sequenced an 81-kb genomic region from the honey bee, Apis mellifera, associated with a quantitative trait locus (QTL) sting-2 for aggressive behavior. This sequence represents the first extensive study of the honey-bee genome structure encompassing putative genes in a QTL for a behavioral trait. Expression of 13 putative genes, as well as two transcripts that were present in a honey-bee EST database, was confirmed through reverse transcription analysis of mRNA from the honey-bee head. Whereas most transcripts exhibited little or no variation between European and Africanized honey-bee alleles, one transcript demonstrated significant nonsynonymous substitutions, deletions, and insertions. All 13 putative genes lacked similarity to known invertebrate or vertebrate proteins or transcripts. This observation may be reflective of the processes that determine the genomic evolution of an insect with social behavior and/or haplo-diploidy and are an indication of the unique nature of the honey-bee genome. These results make this sequence an invaluable research tool for the ongoing honey-bee whole-genome sequencing effort.

Animals↗

RePS: a sequence assembler that masks exact repeats identified from the shotgun data.

We describe a sequence assembler, RePS (repeat-masked Phrap with scaffolding), that explicitly identifies exact 20mer repeats from the shotgun data and removes them prior to the assembly. The established software is used to compute meaningful error probabilities for each base. Clone-end-pairing information is used to construct scaffolds that order and orient the contigs. We show with real data for human and rice that reasonable assemblies are possible even at coverages of only 4x to 6x, despite having up to 42.2% in exact repeats.

Cloning, Molecular↗

Automated finishing with autofinish.

Currently, the genome sequencing community is producing shotgun sequence data at a very high rate, but finishing (collecting additional directed sequence data to close gaps and improve the quality of the data) is not matching that rate. One reason for the difference is that shotgun sequencing is highly automated but finishing is not: Most finishing decisions, such as which directed reads to obtain and which specialized sequencing techniques to use, are made by people. If finishing rates are to increase to match shotgun sequencing rates, most finishing decisions also must be automated. The Autofinish computer program (which is part of the computer software package) does this by automatically choosing finishing reads. Autofinish is able to suggest most finishing reads required for completion of each sequencing project, greatly reducing the amount of human attention needed. sometimes completely finishes the project, with no human decisions required. It cannot solve the most complex problems, so we recommend that Autofinish be allowed to suggest reads for the first three rounds of finishing, and if the project still is not finished completely, a human finisher complete the work. We compared this Autofinish-Hybrid method of finishing against a human finisher in five different projects with a variety of shotgun depths by finishing each project twice--once with each method. This comparison shows that the Autofinish-Hybrid method saves many hours over a human finisher alone, while using roughly the same number and type of reads and closing gaps at roughly the same rate. Autofinish currently is in production use at several large sequencing centers. It is designed to be adaptable to the finishing strategy of the lab--it can finish using some or all of the following: resequencing reads, reverses, custom primer walks on either subclone templates or whole clone templates, PCR, or minilibraries. Autofinish has been used for finishing cDNA, genomic clones, and whole bacterial genomes (see http://www.phrap.org).

Contig Mapping↗

Mouse BAC ends quality assessment and sequence analyses.

A large-scale BAC end-sequencing project at The Institute for Genomic Research (TIGR) has generated one of the most extensive sets of sequence markers for the mouse genome to date. With a sequencing success rate of >80%, an average read length of 485 bp, and ABI3700 capillary sequencers, we have generated 449,234 nonredundant mouse BAC end sequences (mBESs) with 218 Mb total from 257,318 clones from libraries RPCI-23 and RPCI-24, representing 15x clone coverage, 7% sequence coverage, and a marker every 7 kb across the genome. A total of 191,916 BACs have sequences from both ends providing 12x genome coverage. The average Q20 length is 406 bp and 84% of the bases have phred quality scores > or = 20. RPCI-24 mBESs have more Q20 bases and longer reads on average than RPCI-23 sequences. ABI3700 sequencers and the sample tracking system ensure that > 95% of mBESs are associated with the right clone identifiers. We have found that a significant fraction of mBESs contains L1 repeats and approximately 48% of the clones have both ends with > or = 100 bp contiguous unique Q20 bases. About 3% mBESs match ESTs and > 70% of matches were conserved between the mouse and the human or the rat. Approximately 0.1% mBESs contain STSs. About 0.2% mBESs match human finished sequences and > 70% of these sequences have EST hits. The analyses indicate that our high-quality mouse BAC end sequences will be a valuable resource to the community.

Animals↗

Assembly of the working draft of the human genome with GigAssembler.

The data for the public working draft of the human genome contains roughly 400,000 initial sequence contigs in approximately 30,000 large insert clones. Many of these initial sequence contigs overlap. A program, GigAssembler, was built to merge them and to order and orient the resulting larger sequence contigs based on mRNA, paired plasmid ends, EST, BAC end pairs, and other information. This program produced the first publicly available assembly of the human genome, a working draft containing roughly 2.7 billion base pairs and covering an estimated 88% of the genome that has been used for several recent studies of the genome. Here we describe the algorithm used by GigAssembler.

Algorithms↗

Recent segmental duplications in the working draft assembly of the brown Norway rat.

We assessed the content, structure, and distribution of segmental duplications (> or =90% sequence identity, > or =5 kb length) within the published version of the Rattus norvegicus genome assembly (v.3.1). The overall fraction of duplicated sequence within the rat assembly (2.92%) is greater than that of the mouse (1%-1.2%) but significantly less than that of human ( approximately 5%). Duplications were nonuniformly distributed, occurring predominantly as tandem and tightly clustered intrachromosomal duplications. Regions containing extensive interchromosomal duplications were observed, particularly within subtelomeric and pericentromeric regions. We identified 41 discrete genomic regions greater than 1 Mb in size, termed "duplication blocks." These appear to have been the target of extensive duplication over millions of years of evolution. Gene content within duplicated regions ( approximately 1%) was lower than expected based on the genome representation. Interestingly, sequence contigs lacking chromosome assignment ("the unplaced chromosome") showed a marked enrichment for segmental duplication (45% of 75.2 Mb), indicating that segmental duplications have been problematic for sequence and assembly of the rat genome. Further targeted efforts are required to resolve the organization and complexity of these regions.

Animals↗

De novo repeat classification and fragment assembly.

Repetitive sequences make up a significant fraction of almost any genome, and an important and still open question in bioinformatics is how to represent all repeats in DNA sequences. We propose a new approach to repeat classification that represents all repeats in a genome as a mosaic of sub-repeats. Our key algorithmic idea also leads to new approaches to multiple alignment and fragment assembly. In particular, we show that our FragmentGluer assembler improves on Phrap and ARACHNE in assembly of BACs and bacterial genomes.

Algorithms↗

An intermediate grade of finished genomic sequence suitable for comparative analyses.

Although the cost of generating draft-quality genomic sequence continues to decline, refining that sequence by the process of "sequence finishing" remains expensive. Near-perfect finished sequence is an appropriate goal for the human genome and a small set of reference genomes; however, such a high-quality product cannot be cost-justified for large numbers of additional genomes, at least for the foreseeable future. Here we describe the generation and quality of an intermediate grade of finished genomic sequence (termed comparative-grade finished sequence), which is tailored for use in multispecies sequence comparisons. Our analyses indicate that this sequence is very high quality (with the residual gaps and errors mostly falling within repetitive elements) and reflects 99% of the total sequence. Importantly, comparative-grade sequence finishing requires approximately 40-fold less reagents and approximately 10-fold less personnel effort compared to the generation of near-perfect finished sequence, such as that produced for the human genome. Although applied here to finishing sequence derived from individual bacterial artificial chromosome (BAC) clones, one could envision establishing routines for refining sequences emanating from whole-genome shotgun sequencing projects to a similar quality level. Our experience to date demonstrates that comparative-grade sequence finishing represents a practical and affordable option for sequence refinement en route to comparative analyses.

Animals↗

Ancient haplotypes resulting from extensive molecular rearrangements in the wheat A genome have been maintained in species of three different ploidy levels.

Plant genomes, in particular grass genomes, evolve very rapidly. The closely related A genomes of diploid, tetraploid, and hexaploid wheat are derived from a common ancestor that lived <3 million years ago and represent a good model to study molecular mechanisms involved in such rapid evolution. We have sequenced and compared physical contigs at the Lr10 locus on chromosome 1AS from diploid (211 kb), tetraploid (187 kb), and hexaploid wheat (154 kb). A maximum of 33% of the sequences were conserved between two species. The sequences from diploid and tetraploid wheat shared all of the genes, including Lr10 and RGA2 and define a first haplotype (H1). The 130-kb intergenic region between Lr10 and RGA2 was conserved in size despite its activity as a hot spot for transposon insertion, which resulted in >70% of sequence divergence. The hexaploid wheat sequence lacks both Lr10 and RGA2 genes and defines a second haplotype, H2, which originated from ancient and extensive rearrangements. These rearrangements included insertions of retroelements and transposons deletions, as well as unequal recombination within elements. Gene disruption in haplotype H2 was caused by a deletion and subsequent large inversion. Gene conservation between H1 haplotypes, as well as conservation of rearrangements at the origin of the H2 haplotype at three different ploidy levels indicate that the two haplotypes are ancient and had a stable gene content during evolution, whereas the intergenic regions evolved rapidly. Polyploidization during wheat evolution had no detectable consequences on the structure and evolution of the two haplotypes.

Chromosome Deletion↗

Amplification generates modular diversity at an avirulence locus in the pathogen Phytophthora.

The destructive late blight pathogen Phytophthora infestans is notorious for its rapid adaptation to circumvent detection mediated by plant resistance (R) genes. We performed comparative genomic hybridization on microarrays (array-CGH) in a near genome-wide survey to identify genome rearrangements related to changes in virulence. Six loci with copy number variation were found, one of which involves an amplification colocalizing with a previously identified locus that confers avirulence in combination with either R gene R3b, R10, or R11. Besides array-CGH, we used three independent approaches to find candidate genes at the Avr3b-Avr10-Avr11 locus: positional cloning, cDNA-AFLP analysis, and Affymetrix array expression profiling. This resulted in one candidate, pi3.4, that encodes a protein of 1956 amino acids with regulatory domains characteristic for transcription factors. Amplification is restricted to the 3' end of the full-length gene but the amplified copies still contain the hallmarks of a regulatory protein. Sequence comparison showed that the amplification may generate modular diversity and assist in the assembly of novel full-length genes via unequal crossing-over. Analyses of P. infestans field isolates revealed that the pi3.4 amplification correlates with avirulence; isolates virulent on R3b, R10, and R11 plants lack the amplified gene cluster. The ancestral state of 3.4 in the Phytophthora lineage is a full-length, single-copy gene. In P. infestans, however, pi3.4 is a dynamic gene that is amplified and has moved to other locations. Modular diversity could be a novel mechanism for pathogens to quickly adapt to changes in the environment.

Chromosomes, Artificial, Bacterial↗

Genomic sequence and transcriptional profile of the boundary between pericentromeric satellites and genes on human chromosome arm 10p.

Contiguous finished sequence from highly duplicated pericentromeric regions of human chromosomes is needed if we are to understand the role of pericentromeric instability in disease, and in gene and karyotype evolution. Here, we have constructed a BAC contig spanning the transition from pericentromeric satellites to genes on the short arm of human chromosome 10, and used this to generate 1.4 Mb of finished genomic sequence. Combining RT-PCR, in silico gene prediction, and paralogy analysis, we can identify two domains within the sequence. The proximal 600 kb consists of satellite-rich pericentromerically duplicated DNA which is transcript poor, containing only three unspliced transcripts. In contrast, the distal 850 kb contains four known genes (ZNF248, ZNF25, ZNF33A, and ZNF37A) and up to 32 additional transcripts of unknown function. This distal region also contains seven out of the eight intrachromosomal duplications within the sequence, including the p arm copy of the approximately 250-kb duplication which gave rise to ZNF33A and ZNF33B. By sequencing orthologs of the duplicated ZNF33 genes we have established that ZNF33A has diverged significantly at residues critical for DNA binding but ZNF33B has not, indicating that ZNF33B has remained constrained by selection for ancestral gene function. These results provide further evidence of gene formation within intrachromosomal duplications, but indicate that recent interchromosomal duplications at this centromere have involved transcriptionally inert, satellite rich DNA, which is likely to be heterochromatic. This suggests that any novel gene structures formed by these interchromosomal events would require relocation to a more open chromatin environment to be expressed.

Amino Acid Sequence↗

Sequence analysis of a functional Drosophila centromere.

Centromeres are the site for kinetochore formation and spindle attachment and are embedded in heterochromatin in most eukaryotes. The repeat-rich nature of heterochromatin has hindered obtaining a detailed understanding of the composition and organization of heterochromatic and centromeric DNA sequences. Here, we report the results of extensive sequence analysis of a fully functional centromere present in the Drosophila Dp1187 minichromosome. Approximately 8.4% (31 kb) of the highly repeated satellite DNA (AATAT and TTCTC) was sequenced, representing the largest data set of Drosophila satellite DNA sequence to date. Sequence analysis revealed that the orientation of the arrays is uniform and that individual repeats within the arrays mostly differ by rare, single-base polymorphisms. The entire complex DNA component of this centromere (69.7 kb) was sequenced and assembled. The 39-kb "complex island" Maupiti contains long stretches of a complex A+T rich repeat interspersed with transposon fragments, and most of these elements are organized as direct repeats. Surprisingly, five single, intact transposons are directly inserted at different locations in the AATAT satellite arrays. We find no evidence for centromere-specific sequences within this centromere, providing further evidence for sequence-independent, epigenetic determination of centromere identity and function in higher eukaryotes. Our results also demonstrate that the sequence composition and organization of large regions of centric heterochromatin can be determined, despite the presence of repeated DNA.

Animals↗

Gene discovery in the apicomplexa as revealed by EST sequencing and assembly of a comparative gene database.

Large-scale EST sequencing projects for several important parasites within the phylum Apicomplexa were undertaken for the purpose of gene discovery. Included were several parasites of medical importance (Plasmodium falciparum, Toxoplasma gondii) and others of veterinary importance (Eimeria tenella, Sarcocystis neurona, and Neospora caninum). A total of 55192 ESTs, deposited into dbEST/GenBank, were included in the analyses. The resulting sequences have been clustered into nonredundant gene assemblies and deposited into a relational database that supports a variety of sequence and text searches. This database has been used to compare the gene assemblies using BLAST similarity comparisons to the public protein databases to identify putative genes. Of these new entries, approximately 15%-20% represent putative homologs with a conservative cutoff of p < 10(-9), thus identifying many conserved genes that are likely to share common functions with other well-studied organisms. Gene assemblies were also used to identify strain polymorphisms, examine stage-specific expression, and identify gene families. An interesting class of genes that are confined to members of this phylum and not shared by plants, animals, or fungi, was identified. These genes likely mediate the novel biological features of members of the Apicomplexa and hence offer great potential for biological investigation and as possible therapeutic targets.

Animals↗