Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Understanding the adaptation of Halobacterium species NRC-1 to its extreme environment through computational analysis of its genome sequence.

The genome of the halophilic archaeon Halobacterium sp. NRC-1 and predicted proteome have been analyzed by computational methods and reveal characteristics relevant to life in an extreme environment distinguished by hypersalinity and high solar radiation: (1) The proteome is highly acidic, with a median pI of 4.9 and mostly lacking basic proteins. This characteristic correlates with high surface negative charge, determined through homology modeling, as the major adaptive mechanism of halophilic proteins to function in nearly saturating salinity. (2) Codon usage displays the expected GC bias in the wobble position and is consistent with a highly acidic proteome. (3) Distinct genomic domains of NRC-1 with bacterial character are apparent by whole proteome BLAST analysis, including two gene clusters coding for a bacterial-type aerobic respiratory chain. This result indicates that the capacity of halophiles for aerobic respiration may have been acquired through lateral gene transfer. (4) Two regions of the large chromosome were found with relatively lower GC composition and overrepresentation of IS elements, similar to the minichromosomes. These IS-element-rich regions of the genome may serve to exchange DNA between the three replicons and promote genome evolution. (5) GC-skew analysis showed evidence for the existence of two replication origins in the large chromosome. This finding and the occurrence of multiple chromosomes indicate a dynamic genome organization with eukaryotic character.

Adaptation, Biological↗

Completion of the Norwalk virus genome sequence.

Norwalk virus (NV) is the prototype human calicivirus, and causes epidemic outbreaks of acute gastroenteritis. The sequence and predicted genome organization of NV and a NV-like virus [Southampton virus (SHV)] suggested they are similar viruses at the nucleotide and amino acid level, although SHV was reported to be antigenically distinct from NV. A recent review described the discovery of an additional 12 nucleotides at the 5' end of SHV and prompted us to investigate the possibility of additional nucleotides at the 5' end of the NV genome. The results obtained by homopolymeric tailing of NV cDNA with dCTP and dATP showed 12 additional nucleotides also are present on the NV genomic RNA. These data are important with respect to the biology of the virus, and make the genome sequence of NV complete.

Base Sequence↗

Genomic sequencing by ligation-mediated PCR.

Genomic sequencing permits studies of in vivo DNA methylation and protein-DNA interactions, but its use has been limited due to the complexity of the mammalian genome. Ligation-mediated PCR (LMPCR) is a sensitive genomic sequencing procedure that generates high quality, reproducible sequence ladders starting with only 1 microgram of uncloned mammalian DNA per reaction. This genomic sequencing procedure can be adapted for various methylation, in vivo footprinting and DNA adduct mapping procedures. We provide a detailed protocol for genomic sequencing by LMPCR and discuss the principles and applications of the method.

DNA Adducts↗

WindowMasker: window-based masker for sequenced genomes.

MOTIVATION: Matches to repetitive sequences are usually undesirable in the output of DNA database searches. Repetitive sequences need not be matched to a query, if they can be masked in the database. RepeatMasker/Maskeraid (RM), currently the most widely used software for DNA sequence masking, is slow and requires a library of repetitive template sequences, such as a manually curated RepBase library, that may not exist for newly sequenced genomes. RESULTS: We have developed a software tool called WindowMasker (WM) that identifies and masks highly repetitive DNA sequences in a genome, using only the sequence of the genome itself. WM is orders of magnitude faster than RM because WM uses a few linear-time scans of the genome sequence, rather than local alignment methods that compare each library sequence with each piece of the genome. We validate WM by comparing BLAST outputs from large sets of queries applied to two versions of the same genome, one masked by WM, and the other masked by RM. Even for genomes such as the human genome, where a good RepBase library is available, searching the database as masked with WM yields more matches that are apparently non-repetitive and fewer matches to repetitive sequences. We show that these results hold for transcribed regions as well. WM also performs well on genomes for which much of the sequence was in draft form at the time of the analysis. AVAILABILITY: WM is included in the NCBI C++ toolkit. The source code for the entire toolkit is available at ftp://ftp.ncbi.nih.gov/toolbox/ncbi_tools++/CURRENT/. Once the toolkit source is unpacked, the instructions for building WindowMasker application in the UNIX environment can be found in file src/app/winmasker/README.build. SUPPLEMENTARY INFORMATION: Supplementary data are available at ftp://ftp.ncbi.nlm.nih.gov/pub/agarwala/windowmasker/windowmasker_suppl.pdf

Algorithms↗

Archaeal adaptation to higher temperatures revealed by genomic sequence of Thermoplasma volcanium.

The complete genomic sequence of the archaeon Thermoplasma volcanium, possessing optimum growth temperature (OGT) of 60 degrees C, is reported. By systematically comparing this genomic sequence with the other known genomic sequences of archaea, all possessing higher OGT, a number of strong correlations have been identified between characteristics of genomic organization and the OGT. With increasing OGT, in the genomic DNA, frequency of clustering purines and pyrimidines into separate dinucleotides rises (e.g., by often forming AA and TT, whereas avoiding TA and AT). Proteins coded in a genome are divided into two distinct subpopulations possessing isoelectric points in different ranges (i.e., acidic and basic), and with increasing OGT the size of the basic subpopulation becomes larger. At the metabolic level, genes coding for enzymes mediating pathways for synthesizing some coenzymes, such as heme, start missing. These findings provide insights into the design of individual genomic components, as well as principles for coordinating changes in these designs for the adaptation to new environments.

Adaptation, Physiological↗

The Genome Sequence DataBase: towards an integrated functional genomics resource.

During 1998 the primary focus of the Genome Sequence DataBase (GSDB; http://www.ncgr.org/gsdb ) located at the National Center for Genome Resources (NCGR) has been to improve data quality, improve data collections, and provide new methods and tools to access and analyze data. Data quality has been improved by extensive curation of certain data fields necessary for maintaining data collections and for using certain tools. Data quality has also been increased by improvements to the suite of programs that import data from the International Nucleotide Sequence Database Collaboration (IC). The Sequence Tag Alignment and Consensus Knowledgebase (STACK), a database of human expressed gene sequences developed by the South African National Bioinformatics Institute (SANBI), became available within the last year, allowing public access to this valuable resource of expressed sequences. Data access was improved by the addition of the Sequence Viewer, a platform-independent graphical viewer for GSDB sequence data. This tool has also been integrated with other searching and data retrieval tools. A BLAST homology search service was also made available, allowing researchers to search all of the data, including the unique data, that are available from GSDB. These improvements are designed to make GSDB more accessible to users, extend the rich searching capability already present in GSDB, and to facilitate the transition to an integrated system containing many different types of biological data.

Animals↗

Complete genome sequence of the broad-host-range vibriophage KVP40: comparative genomics of a T4-related bacteriophage.

The complete genome sequence of the T4-like, broad-host-range vibriophage KVP40 has been determined. The genome sequence is 244,835 bp, with an overall G+C content of 42.6%. It encodes 386 putative protein-encoding open reading frames (CDSs), 30 tRNAs, 33 T4-like late promoters, and 57 potential rho-independent terminators. Overall, 92.1% of the KVP40 genome is coding, with an average CDS size of 587 bp. While 65% of the CDSs were unique to KVP40 and had no known function, the genome sequence and organization show specific regions of extensive conservation with phage T4. At least 99 KVP40 CDSs have homologs in the T4 genome (Blast alignments of 45 to 68% amino acid similarity). The shared CDSs represent 36% of all T4 CDSs but only 26% of those from KVP40. There is extensive representation of the DNA replication, recombination, and repair enzymes as well as the viral capsid and tail structural genes. KVP40 lacks several T4 enzymes involved in host DNA degradation, appears not to synthesize the modified cytosine (hydroxymethyl glucose) present in T-even phages, and lacks group I introns. KVP40 likely utilizes the T4-type sigma-55 late transcription apparatus, but features of early- or middle-mode transcription were not identified. There are 26 CDSs that have no viral homolog, and many did not necessarily originate from Vibrio spp., suggesting an even broader host range for KVP40. From these latter CDSs, an NAD salvage pathway was inferred that appears to be unique among bacteriophages. Features of the KVP40 genome that distinguish it from T4 are presented, as well as those, such as the replication and virion gene clusters, that are substantially conserved.

Bacteriophage T4↗

META-DIFF: a k-mer-based pipeline that detects differentially abundant sequences in metagenomics whole genome sequencing.

Traditional case-control metagenomic studies are constrained by their dependence on taxonomic and functional databases. Because annotation occurs before differential analysis, they are limited to known elements and keep function and taxonomy separate. Although binning strategies have emerged to reconstruct genomes and mitigate this issue, they still require an assembly step, preventing the use of all available sequencing data. Here, we introduce META-DIFF, a pipeline based on differentially abundant k-mers independently of any prior annotation. From those k-mers, it reconstructs longer sequences and provides biological context, as well as the best set of unitigs to discriminate between conditions. Across both taxonomy-centric and functionally-centric benchmarks, it showed robust performance and displayed great reproducibility. It also behaved more conservatively than did other univariate methodologies, i.e. it maintained a high precision at the expense of recall, particularly in conditions of low fold-change and limited sequencing depth. The efficacy of META-DIFF was further validated through its application to a real-world colorectal cancer dataset, which produced both confirmatory and novel results compared with those of previous publications. The pipeline is able to exploit all reads and identify differentially abundant elements, including unknown DNA, prior to annotation. With the guidelines provided, META-DIFF provides users with great exploratory power to unravel microbiome changes.

Metagenomics↗

Construction of a BAC library and generation of BAC end sequence-tagged connectors for genome sequencing of the African malaria mosquito Anopheles gambiae.

A Bacterial Artificial Chromosome (BAC) genomic DNA library of Anopheles gambiae, the major human malaria vector in sub-Saharan Africa, was constructed and characterized. This library (ND-TAM) is composed of 30,720 BAC clones in eighty 384-well plates. The estimated average insert size of the library is 133 kb, with an overall genome coverage of approximately 14-fold. The ends of approximately two-thirds of the clones in the library were sequenced, yielding 32,340 pair-mate ends. A statistical analysis (G-test) of the results of PCR screening of the library indicated a random distribution of BACs in the genome, although one gap encompassing the white locus on the X-chromosome was identified. Furthermore, combined with another previously constructed BAC library (ND-1), ~2,000 BACs have been physically mapped by polytene chromosomal in situ hybridization. These BAC end pair mates and physically mapped BACs have been useful for both the assembly of a fully sequenced A. gambiae genome and for linking the assembled sequence to the three polytene chromosomes. This ND-TAM library is now publicly available at both http://www.malaria.mr4.org/mr4pages/index.html/ and http://hbz.tamu.edu/, providing a valuable resource to the mosquito research community.

Animals↗

Genome sequence comparisons: hurdles in the fast lane to functional genomics.

An important computational technique for extracting the wealth of information hidden in human genomic sequence data is to compare the sequence with that from the corresponding region of the mouse genome, looking for segments that are conserved over evolutionary time. Moreover, the approach generalises to comparison of sequences from any two related species. The underlying rationale (which is abundantly confirmed by observation) is that a random mutation in a functional region is usually deleterious to the organism, and hence unlikely to become fixed in the population, whereas mutations in a non-functional region are free to accumulate over time. The potential value of this approach is so attractive that the public and private projects to sequence the human genome are now turning to sequencing the mouse, and you will soon be able to compare the human and mouse sequences of your favourite genomic region. We are currently witnessing an explosion of computer tools for comparative analysis of two genomic sequences. Here the capabilities of two new network servers for comparing genomic sequences from any pair of closely related species are sketched. The Syntenic Gene Prediction Program SGP-I utilises sequence comparisons to enhance the ability to locate protein coding segments in genomic data. PipMaker attempts to determine all conserved genomic regions, regardless of their function.

Animals↗

A complexity reduction algorithm for analysis and annotation of large genomic sequences.

DNA is a universal language encrypted with biological instruction for life. In higher organisms, the genetic information is preserved predominantly in an organized exon/intron structure. When a gene is expressed, the exons are spliced together to form the transcript for protein synthesis. We have developed a complexity reduction algorithm for sequence analysis (CRASA) that enables direct alignment of cDNA sequences to the genome. This method features a progressive data structure in hierarchical orders to facilitate a fast and efficient search mechanism. CRASA implementation was tested with already annotated genomic sequences in two benchmark data sets and compared with 15 annotation programs (10 ab initio and 5 homology-based approaches) against the EST database. By the use of layered noise filters, the complexity of CRASA-matched data was reduced exponentially. The results from the benchmark tests showed that CRASA annotation excelled in both the sensitivity and specificity categories. When CRASA was applied to the analysis of human Chromosomes 21 and 22, an additional 83 potential genes were identified. With its large-scale processing capability, CRASA can be used as a robust tool for genome annotation with high accuracy by matching the EST sequences precisely to the genomic sequences.

Algorithms↗

Computational detection of prokaryotic core promoters in genomic sequences.

The high-throughput sequencing of microbial genomes has resulted in the relatively rapid accumulation of an enormous amount of genomic sequence data. In this context, the problem posed by the detection of promoters in genomic DNA sequences via computational methods has attracted considerable research attention in recent years. This paper addresses the development of a predictive model, known as the dependence decomposition weight matrix model (DDWMM), which was designed to detect the core promoter region, including the -10 region and the transcription start sites (TSSs), in prokaryotic genomic DNA sequences. This is an issue of some importance with regard to genome annotation efforts. Our predictive model captures the most significant dependencies between positions (allowing for non-adjacent as well as adjacent dependencies) via the maximal dependence decomposition (MDD) procedure, which iteratively decomposes data sets into subsets, based on the significant dependence between positions in the promoter region to be modeled. Such dependencies may be intimately related to biological and structural concerns, since promoter elements are present in a variety of combinations, which are separated by various distances. In this respect, the DDWMM may prove to be appropriate with regard to the detection of core promoter regions and TSSs in long microbial genomic contigs. In order to demonstrate the effectiveness of our predictive model, we applied 10-fold cross-validation experiments on the 607 experimentally-verified promoter sequences, which evidenced good performance in terms of sensitivity.

Base Sequence↗

Enrichment of regulatory signals in conserved non-coding genomic sequence.

MOTIVATION: Whole genome shotgun sequencing strategies generate sequence data prior to the application of assembly methodologies that result in contiguous sequence. Sequence reads can be employed to indicate regions of conservation between closely related species for which only one genome has been assembled. Consequently, by using pairwise sequence alignments methods it is possible to identify novel, non-repetitive, conserved segments in non-coding sequence that exist between the assembled human genome and mouse whole genome shotgun sequencing fragments. Conserved non-coding regions identify potentially functional DNA that could be involved in transcriptional regulation. RESULTS: Local sequence alignment methods were applied employing mouse fragments and the assembled human genome. In addition, transcription factor binding sites were detected by aligning their corresponding positional weight matrices to the sequence regions. These methods were applied to a set of transcripts corresponding to 502 genes associated with a variety of different human diseases taken from the Online Mendelian Inheritance in Man database. Using statistical arguments we have shown that conserved non-coding segments contain an enrichment of transcription factor binding sites when compared to the sequence background in which the conserved segments are located. This enrichment of binding sites was not observed in coding sequence. Conserved non-coding segments are not extensively repeated in the genome and therefore their identification provides a rapid means of finding genes with related conserved regions, and consequently potentially related regulatory mechanism. Conserved segments in upstream regions are found to contain binding sites that are co-localized in a manner consistent with experimentally known transcription factor pairwise co-occurrences and afford the identification of novel co-occurring Transcription Factor (TF) pairs. This study provides a methodology and more evidence to suggest that conserved non-coding regions are biologically significant since they contain a statistical enrichment of regulatory signals and pairs of signals that enable the construction of regulatory models for human genes. CONTACT: samuel.levy@celera.com.

Algorithms↗

The complete genome sequence of a dog: a perspective.

A complete, high-quality reference sequence of a dog genome was recently produced by a team of researchers led by the Broad Institute, achieving another major milestone in deciphering the genomic landscape of mammalian organisms. The genome sequence provides an indispensable resource for comparative analysis and novel insights into dog and human evolution and history. Together with the survey sequence of a poodle previously published in 2003, the two dog genome sequences allowed identification of more than 2.5 million single nucleotide polymorphisms within and between dog breeds, which can be used in evolutionary analysis, behavioral studies and disease gene mapping.(1)

Animals↗

Arabidopsis genome sequence as a tool for functional genomics in tomato.

Tomato is a well-established model organism for studying many biological processes including resistance and susceptibility to pathogens and the development and ripening of fleshy fruits. The availability of the complete Arabidopsis genome sequence will expedite map-based cloning in tomato on the basis of chromosomal synteny between the two species, and will facilitate the functional analysis of tomato genes.

Arabidopsis↗

Comparing the performance of exome and genome sequencing for rare disease diagnostics: A randomized implementation effectiveness trial.

PURPOSE: Exome sequencing (ES) and genome sequencing (GS) can improve rare disease diagnosis but are not routinely available in many jurisdictions. To inform implementation, we report on a randomized implementation effectiveness trial comparing ES and GS. METHODS: Eligible trios were randomized to receive ES or GS in the same clinically accredited laboratory. Patient-level data on diagnostic utility and turnaround times were collected. Outcomes were compared statistically between clinically important subgroups. RESULTS: Of 1048 patients, 68.5% had syndromic intellectual disability/developmental delay (ID/DD) and 20.5% had multisystem disorders without ID/DD. Most had prior genetic test(s) that were nondiagnostic (95.5%), and of these, 91.6% included chromosome microarray. Diagnostic yields were 33.8% and 33.6%, for ES (n = 526) and GS (n = 522), respectively. Within sequencing groups, diagnostic results were more frequent among those with ID/DD than those without (P < .005). For routine (ie, nonexpedited) patients (n = 1020), 87.0% were reported in <12 weeks, and the mean turnaround time was 55.5 days (SD: 24.0). Turnaround time for ES and GS did not differ; however, result type (P < .001) and age of onset (P < .005) significantly affected turnaround time. CONCLUSION: Findings provide robust evidence of diagnostic utility and timeliness of ES and GS and will inform policy related to the organization, delivery, and reimbursement of clinical-grade genome diagnostics for rare diseases.

Adolescent↗

The Drosophila melanogaster genome sequencing and annotation projects: a status report.

The sequence and genome annotations of Drosophila melanogaster were initially published in late 1999 and early 2000. Since then, the Berkeley Drosophila Genome Project (BDGP) and FlyBase have improved the quality of the sequence and reviewed the annotations by hand, respectively, to produce an account of the fruit fly genome that is of the highest quality. This review discusses the main features of this process, both from the point of view of the biology revealed in the end result and in the development of software that has been central to this genome sequencing and annotation project.

Animals↗