Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Contig Mapping”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

Identification of mixups among DNA sequencing plates.

MOTIVATION: During the process of high-throughput genome sequencing there are opportunities for mixups of reagents and data associated with particular projects. The sequencing templates or sequence data generated for an assembly may become contaminated with reagents or sequences from another project, resulting in poorer quality and inaccurate assemblies. RESULTS: We have developed a system to assess sequence assemblies and monitor for laboratory mixups. We describe several methods for testing the consistency of assemblies and resolving mixed ones. We use statistical tests to evaluate the distribution of sequencing reads from different plates into contigs, and a graph-based approach to resolve situations where data has been inappropriately combined. While these methods have been designed for use in a high-throughput DNA sequencing environment processing thousands of clones, they can be applied in any situation where distinct sequencing projects are performed at redundant coverage.

Algorithms↗

Assembly of fingerprint contigs: parallelized FPC.

SUMMARY: One of the more common uses of the program FingerPrint Contigs (FPC) is to assemble random restriction digest 'fingerprints' of overlapping genomic clones into contigs. To improve the rate of assembling contigs from large fingerprint databases we have adapted FPC so that it can be run in parallel on multiple processors and servers. The current version of 'parallelized FPC' has been used in our laboratory to assemble mammalian BAC fingerprint databases, each containing more than 300000 BAC fingerprints. AVAILABILITY: This parallelized version of FPC is available under the GNU GPL licence, and can be downloaded from ftp://ftp.bcgsc.bc.ca/pub/fpcd.

Algorithms↗

PGAAS: a prokaryotic genome assembly assistant system.

MOTIVATION: In order to accelerate the finishing phase of genome assembly, especially for the whole genome shotgun approach of prokaryotic species, we have developed a software package designated prokaryotic genome assembly assistant system (PGAAS). The approach upon which PGAAS is based is to confirm the order of contigs and fill gaps between contigs through peptide links obtained by searching each contig end with BLASTX against protein databases. RESULTS: We used the contig dataset of the cyanobacterium Synechococcus sp. strain PCC7002 (PCC7002), which was sequenced with six-fold coverage and assembled using the Phrap package. The subject database is the protein database of the cyanobacterium, Synechocystis sp. strain PCC6803 (PCC6803). We found more than 100 non-redundant peptide segments which can link at least 2 contigs. We tested one pair of linked contigs by sequencing and obtained satisfactory result. PGAAS provides a graphic user interface to show the bridge peptides and pier contigs. We integrated Primer3 into our package to design PCR primers at the adjacent ends of the pier contigs. AVAILABILITY: We tested PGAAS on a Linux (Redhat 6.2) PC machine. It is developed with free software (MySQL, PHP and Apache). The whole package is distributed freely and can be downloaded as UNIX compress file: ftp://ftp.cbi.pku.edu.cn/pub/software/unix/pgaas1.0.tar.gz. The package is being continually updated.

Algorithms↗

ViralQC: a tool for assessing completeness and contamination of predicted viral contigs.

MOTIVATION: Viruses represent the most abundant biological entities on Earth, playing vital roles in diverse ecosystems. Cataloging viruses across various environments is essential for understanding their properties and functions. Metagenomic sequencing has emerged as the most comprehensive method for virus discovery. However, distinguishing viral sequences from the vast background of microbial organisms in metagenomic data remains a significant challenge. Existing tools experience varying degrees of false positive rates due to noise in sequencing and assembly, and the integration of proviruses into microbial genomes. This highlights the urgent need for an accurate and efficient method to evaluate the quality of viral contigs. RESULTS: To address these challenges, we introduce ViralQC, a tool designed to assess the quality of viral contigs or bins. ViralQC identifies microbial contamination within putative viral sequences using an ensemble framework powered by DNA and protein foundation models and estimates completeness by analyzing protein organization. We evaluated ViralQC on multiple datasets and compared its performance against the state-of-the-art tool, CheckV. Leveraging both DNA and protein foundation models, ViralQC achieves higher sensitivity on contamination detection for contigs longer than 10 kbp while maintaining comparable accuracy. Additionally, ViralQC delivers more accurate estimation on contigs with completeness > 50%. AVAILABILITY: The source code of ViralQC is available via: https://github.com/ChengPENG-wolf/ViralQC.

Software↗

Selection of oligonucleotide probes for protein coding sequences.

MOTIVATION: Large arrays of oligonucleotide probes have become popular tools for analyzing RNA expression. However to date most oligo collections contain poorly validated sequences or are biased toward untranslated regions (UTRs). Here we present a strategy for picking oligos for microarrays that focus on a design universe consisting exclusively of protein coding regions. We describe the constraints in oligo design that are imposed by this strategy, as well as a software tool that allows the strategy to be applied broadly. RESULT: In this work we sequentially apply a variety of simple filters to candidate sequences for oligo probes. The primary filter is a rejection of probes that contain contiguous identity with any other sequence in the sample universe that exceeds a pre-established threshold length. We find that rejection of oligos that contain 15 bases of perfect match with other sequences in the design universe is a feasible strategy for oligo selection for probe arrays designed to interrogate mammalian RNA populations. Filters to remove sequences with low complexity and predicted poor probe accessibility narrow the candidate probe space only slightly. Rejection based on global sequence alignment is performed as a secondary, rather than primary, test, leading to an algorithm that is computationally efficient. Splice isoforms pose unique challenges and we find that isoform prevalence will for the most part have to be determined by analysis of the patterns of hybridization of partially redundant oligonucleotides. AVAILABILITY: The oligo design program OligoPicker and its source code are freely available at our website.

Algorithms↗

Fragment assembly with short reads.

MOTIVATION: Current DNA sequencing technology produces reads of about 500-750 bp, with typical coverage under 10x. New sequencing technologies are emerging that produce shorter reads (length 80-200 bp) but allow one to generate significantly higher coverage (30x and higher) at low cost. Modern assembly programs and error correction routines have been tuned to work well with current read technology but were not designed for assembly of short reads. RESULTS: We analyze the limitations of assembling reads generated by these new technologies and present a routine for base-calling in reads prior to their assembly. We demonstrate that while it is feasible to assemble such short reads, the resulting contigs will require significant (if not prohibitive) finishing efforts. AVAILABILITY: Available from the web at http://www.cse.ucsd.edu/groups/bioinformatics/software.html

Algorithms↗

ESTminer: a Web interface for mining EST contig and cluster databases.

UNLABELLED: ESTminer is a Web application and database schema for interactive mining of expressed sequence tag (EST) contig and cluster datasets. The Web interface contains a query frame that allows the selection of contigs/clusters with specific cDNA library makeup or a threshold number of members. The results are displayed as color-coded tree nodes, where the color indicates the fractional size of each cDNA library component. The nodes are expandable, revealing library statistics as well as EST or contig members, with links to sequence data, GenBank records or user configurable links. Also, the interface allows 'queries within queries' where the result set of a query is further filtered by the subsequent query. AVAILABILITY: ESTminer is implemented in Java/JSP and the package, including MySQL and Oracle schema creation scripts, is available from http://cggc.agtec.uga.edu/Data/download.asp CONTACT: agingle@uga.edu.

Algorithms↗

A graph based algorithm for generating EST consensus sequences.

MOTIVATION: EST sequences constitute an abundant, yet error prone resource for computational biology. Expressed sequences are important in gene discovery and identification, and they are also crucial for the discovery and classification of alternative splicing. An important challenge when processing EST sequences is the reconstruction of mRNA by assembling EST clusters into consensus sequences. RESULTS: In contrast to the more established assembly tools, we propose an algorithm that constructs a graph over sequence fragments of fixed size, and produces consensus sequences as traversals of this graph. We provide a tool implementing this algorithm, and perform an experiment where the consensus sequences produced by our implementation, as well as by currently available tools, are compared to mRNA. The results show that our proposed algorithm in a majority of the cases produces consensus of higher quality than the established sequence assemblers and at a competitive speed. AVAILABILITY: The source code for the implementation is available under a GPL license from http://www.ii.uib.no/~ketil/bioinformatics/ CONTACT: ketil@ii.uib.no.

Algorithms↗

Analysis of a human fungiform papillae cDNA library and identification of taste-related genes.

Various genes related to early events in human gustation have recently been discovered, yet a thorough understanding of taste transduction is hampered by gaps in our knowledge of the signaling chain. As a first step toward gaining additional insight, the expression specificity of genes in human taste tissue needs to be determined. To this end, a fungiform papillae cDNA library has been generated and analyzed. For validation of the library, taste-related gene probes were used to detect known molecules. Subsequently, DNA sequence analysis was performed to identify further candidates. Of 987 clones sequenced, clustering results in 288 contigs. Comparison of these contigs with genomic databases reveals that 207 contigs (71.9%) match known genes, 16 (5.6%) match hypothetical genes, eight (2.8%) match repetitive sequences and 57 (19.8%) have no or low similarity to annotated genes. The results indicate that despite a high level of redundancy, this human fungiform cDNA library contains specific taste markers and is valuable for investigation of both known and novel taste-related genes.

Computational Biology↗

Sequence analysis of the genome of the unicellular cyanobacterium Synechocystis sp. strain PCC6803. II. Sequence determination of the entire genome and assignment of potential protein-coding regions.

The sequence determination of the entire genome of the Synechocystis sp. strain PCC6803 was completed. The total length of the genome finally confirmed was 3,573,470 bp, including the previously reported sequence of 1,003,450 bp from map position 64% to 92% of the genome. The entire sequence was assembled from the sequences of the physical map-based contigs of cosmid clones and of lambda clones and long PCR products which were used for gap-filling. The accuracy of the sequence was guaranteed by analysis of both strands of DNA through the entire genome. The authenticity of the assembled sequence was supported by restriction analysis of long PCR products, which were directly amplified from the genomic DNA using the assembled sequence data. To predict the potential protein-coding regions, analysis of open reading frames (ORFs), analysis by the GeneMark program and similarity search to databases were performed. As a result, a total of 3,168 potential protein genes were assigned on the genome, in which 145 (4.6%) were identical to reported genes and 1,257 (39.6%) and 340 (10.8%) showed similarity to reported and hypothetical genes, respectively. The remaining 1,426 (45.0%) had no apparent similarity to any genes in databases. Among the potential protein genes assigned, 128 were related to the genes participating in photosynthetic reactions. The sum of the sequences coding for potential protein genes occupies 87% of the genome length. By adding rRNA and tRNA genes, therefore, the genome has a very compact arrangement of protein- and RNA-coding regions. A notable feature on the gene organization of the genome was that 99 ORFs, which showed similarity to transposase genes and could be classified into 6 groups, were found spread all over the genome, and at least 26 of them appeared to remain intact. The result implies that rearrangement of the genome occurred frequently during and after establishment of this species.

Bacterial Proteins↗

Expressed sequence tags from immature female sexual organ of a liverwort, Marchantia polymorpha.

A total of 970 expressed sequence tag (EST) clones were generated from immature female sexual organ of a liverwort, Marchantia polymorpha. The 376 ESTs resulted in 123 redundant groups, thus the total number of unique sequences in the EST set was 717. Database search by BLAST algorithm showed that 302 of the unique sequences shared significant similarities to known nucleotide or amino acid sequences. Six unique sequences showed significant similarities to genes that are involved in flower development and sexual reproduction, such as cynarase, fimbriata-associated protein and S-receptor kinase genes. The remaining unique 415 sequences have no significant similarity with any database-registered genes or proteins. The redundant 123 ESTs implied the presence of gene families and abundant transcripts of unknown identity. Analyses of the coding sequences of 61 unique sequences, which contained no ambiguous bases in the predicted coding regions, highly homologous to known sequences at the amino acid level with a similarity score greater than 400, and with stop codons at similar positions as their possible orthologues, indicated the presence of biased codon usage and higher GC content within the coding sequences (50.4%) than that within 3' flanking sequences (41.9%).

Amino Acid Sequence↗

Genomic structure and localization of the human protein phosphatase 2A BRgamma regulatory subunit.

In the course of our analysis of genomic sequence from the human chromosome 4p16.1 region harboring both the Wolfram and Ellis van Creveld syndrome genes we have identified a sequence with high homology (98% at the amino acid level) to the rat cDNA coding for the protein phosphatase 2A BRgamma (PP2ABRgamma) regulatory subunit. Although the human cDNAs for both the BRalpha and BRbeta isoforms have been described previously, the BRgamma subunit has not yet been identified in humans. Here we describe the precise genomic organization and genetic localization of the human PP2ABRgamma gene.

Amino Acid Sequence↗

Sequence analysis of a total of three megabases of DNA in two regions of chromosome 8p.

Large-scale sequencing of genomic regions and in silico gene trapping together represent a highly efficient and powerful approach for identifying novel genes. We performed megabase-level sequence analyses of two genomic regions on human chromosome 8p (8p11.2 and 8p22-->p21.3), after covering those segments with sequence-ready contigs composed of 74 cosmids, 14 BACs, and three PAC clones. We determined continuous nucleotide sequences of 1,856,753 bases on 8p11.2 and 1,210,381 bases on 8p22-->p21.3 by combining the shotgun and primer-walking methods. In silico gene trapping identified four novel genes in the 8p11.2 region and, in the 8p22-->p21.3 region, six known genes (PRLTS, PCM1, MTAMR7, HCAT2, HFREP-1 and PHP) and three novel genes. The distribution of Alu and LINE1 repetitive elements and the densities of predicted exons were different in each region, and Alu-rich portions contained more exonic sequences than LINE1-rich areas.

Amino Acid Sequence↗

Sequence-ready 1-Mb YAC, BAC and cosmid contigs covering the distal imprinted region of mouse chromosome 7.

We have constructed approximately 1-Mb contigs of yeast artificial chromosome (YAC), bacterial artificial chromosome (BAC) and cosmid clones covering the imprinted region in mouse chromosome band 7F4/F5. This region is syntenic to human chromosome 11p15.5, which is associated with Beckwith-Wiedemann syndrome (BWS) and certain childhood and adult tumors. These contigs provide the basis for genomic sequencing, identification of genes and their regulatory elements, and functional studies in transgenic and knockout mice, which should be of help to understand not only the mechanisms of imprinting but also the molecular events involved in the genesis of BWS and tumors.

Animals↗

Comparison of expressed sequence tags from male and female sexual organs of Marchantia polymorpha.

A total of 935 expressed sequence tags (ESTs) from male immature sexual organ were determined, of which 600 ESTs were assembled into 110 non-redundant groups, resulting in 445 unique EST sequences. Of these, 244 sequences shared significant similarities to known nucleotide or amino acid sequences in other organisms. The remaining 201 unique sequences showed no significant matches and thus are likely to be novel transcripts. ESTs from male and female immature sexual organs of a liverwort, Marchantia polymorpha, were compared to characterize gene expression patterns during sex differentiation. Ninety-nine male ESTs turned out to be common genes found also in the female library. Interestingly, one of the ESTs found only in male shows a significant similarity to the transformer-2 gene involved in sex determination in Drosophila. In female, several unique lectin ESTs were found that are not present in the male library.

Amino Acid Sequence↗

Analysis of expressed sequence tags from two starvation, time-of-day-specific libraries of Neurospora crassa reveals novel clock-controlled genes.

In an effort to determine genes that are expressed in mycelial cultures of Neurospora crassa over the course of the circadian day, we have sequenced 13,000 cDNA clones from two time-of-day-specific libraries (morning and evening library) generating approximately 20,000 sequences. Contig analysis allowed the identification of 445 unique expressed sequence tags (ESTs) and 986 ESTs present in multiple cDNA clones. For approximately 50% of the sequences (710 of 1431), significant matches to sequences in the National Center for Biotechnology Information database (of known or unknown function) were detected. About 50% of the ESTs (721 of 1431) showed no similarity to previously identified genes. We hybridized Northern blots with probes derived from 26 clones chosen from contigs identified by multiple cDNA clones and EST sequences. Using these sequences, the representation of genes among the morning and evening sequences, respectively, in most cases does not reflect their expression patterns over the course of the day. Nevertheless, we were able to identify four new clock-controlled genes. On the basis of these data we predict that a significant proportion of the expressed Neurospora genes may be regulated by the circadian clock. The mRNA levels of all four genes peak in the subjective morning as is the case with previously identified ccgs.

Blotting, Northern↗

Characterization of the S-locus region of almond (Prunus dulcis): analysis of a somaclonal mutant and a cosmid contig for an S haplotype.

Almond has a self-incompatibility system that is controlled by an S locus consisting of the S-RNase gene and an unidentified "pollen S gene." An almond cultivar "Jeffries," a somaclonal mutant of "Nonpareil" (S(c)S(d)), has a dysfunctional S(c) haplotype both in pistil and pollen. Immunoblot and genomic Southern blot analyses detected no S(c) haplotype-specific signal in Jeffries. Southern blot showed that Jeffries has an extra copy of the S(d) haplotype. These results indicate that at least two mutations had occurred to generate Jeffries: (1) deletion of the S(c) haplotype and (2) duplication of the S(d) haplotype. To analyze the extent of the deletion in Jeffries and gain insight into the physical limit of the S locus region, approximately 200 kbp of a cosmid contig for the S(c) haplotype was constructed. Genomic Southern blot analyses showed that the deletion in Jeffries extends beyond the region covered by the contig. Most cosmid end probes, except those near the S(c)-RNase gene, cross-hybridized with DNA fragments from different S haplotypes. This suggests that regions away from the S(c)-RNase gene can recombine between different S haplotypes, implying that the cosmid contig extends to the borders of the S locus.

Base Sequence↗

High resolution deletion analysis of constitutional DNA from neurofibromatosis type 2 (NF2) patients using microarray-CGH.

Neurofibromatosis type 2 (NF2) is an autosomal dominant disorder whose hallmark is bilateral vestibular schwannoma. It displays a pronounced clinical heterogeneity with mild to severe forms. The NF2 tumor suppressor (merlin/schwannomin) has been cloned and extensively analyzed for mutations in patients with different clinical variants of the disease. Correlation between the type of the NF2 gene mutation and the patient phenotype has been suggested to exist. However, several independent studies have shown that a fraction of NF2 patients with various phenotypes have constitutional deletions that partly or entirely remove one copy of the NF2 gene. The purpose of this study was to examine a 7 Mb interval in the vicinity of the NF2 gene in a large series of NF2 patients in order to determine the frequency and extent of deletions. A total of 116 NF2 patients were analyzed using high-resolution array-comparative genomic hybridization (CGH) on an array covering at least 90% of this region of 22q around the NF2 locus. Deletions, which remove one copy of the entire gene or are predicted to truncate the schwannomin protein, were detected in 8 severe, 10 moderate and 6 mild patients. This result does not support the correlation between the type of mutation affecting the NF2 gene and the disease phenotype. This work also demonstrates the general usefulness of the array-CGH methodology for rapid and comprehensive detection of small (down to 40 kb) heterozygous and/or homozygous deletions occurring in constitutional or tumor-derived DNA.

Adolescent↗