Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Use of ordered deletions in genome sequencing.

Previous attempts to use the non-random approach for sequencing long DNA fragments have met with little success. As a result, nearly all genomic sequencing is done by the random (shotgun) approach, and the economy promised by the non-random approach has so far not materialized. Here we describe a simple system based on the use of ordered deletions that can be incorporated in the common strategies for genome sequencing. Long genomic fragments are cloned in the pAL-F cosmid and fragmented by digestion with specific restriction endonucleases. The digests are religated to subclone individual restriction fragments. The subclones are then subdivided by overlapping deletions and used for sequencing. We present the nucleotide sequences of two cosmid inserts from chromosome IV of Drosophila (containing the ci gene and the 5' end of the zfh-2 gene) that were determined by this method. This is the first report of successful sequencing of long genomic fragments by the use of overlapping deletions. Our calculations show that, with the present approach, sequence data can be acquired at a rate comparable to the shotgun approach but with significantly reduced numbers (approximately 30%) of sequencing runs. Hence, the use of ordered deletions should allow significant savings in both the amount and cost of sequencing work.

Animals↗

Genomic sequence sampling: a strategy for high resolution sequence-based physical mapping of complex genomes.

We present a simple and efficient method for constructing high resolution physical maps of large regions of genomic DNA based upon sampled sequencing. The physical map is constructed by ordering high density cosmid contigs and determining a sequence fragment from each end of every clone. The resulting map, which contains 30-50% of the complete DNA sequence, allows the identification of many genes and makes possible PCR amplification of virtually any part of the genome. We apply this strategy to the automated analysis of the genome of the primitive eukaryote Giardia lamblia and evaluate its applicability to the physical mapping and DNA sequencing of the human genome.

Amino Acid Sequence↗

Developmental expression pattern screen for genes predicted in the C. elegans genome sequencing project.

Maximum use should be made of information generated in the genome sequencing projects. Toward this end, we have initiated a genome sequence-based, expression pattern screen of genes predicted from the Caenorhabditis elegans genome sequence data. We examined beta-galactosidase expression patterns in C. elegans lines transformed with lacZ reporter gene fusions constructed using predicted C. elegans gene promoter regions. Of the predicted genes in the cosmids analysed so far, 67% are amenable to the approach and 54% of examined genes yielded a developmental expression pattern. Expression pattern information is being made generally available using computer databases.

Animals↗

Genome sequence of enterohaemorrhagic Escherichia coli O157:H7.

The bacterium Escherichia coli O157:H7 is a worldwide threat to public health and has been implicated in many outbreaks of haemorrhagic colitis, some of which included fatalities caused by haemolytic uraemic syndrome. Close to 75,000 cases of O157:H7 infection are now estimated to occur annually in the United States. The severity of disease, the lack of effective treatment and the potential for large-scale outbreaks from contaminated food supplies have propelled intensive research on the pathogenesis and detection of E. coli O157:H7 (ref. 4). Here we have sequenced the genome of E. coli O157:H7 to identify candidate genes responsible for pathogenesis, to develop better methods of strain detection and to advance our understanding of the evolution of E. coli, through comparison with the genome of the non-pathogenic laboratory strain E. coli K-12 (ref. 5). We find that lateral gene transfer is far more extensive than previously anticipated. In fact, 1,387 new genes encoded in strain-specific clusters of diverse sizes were found in O157:H7. These include candidate virulence factors, alternative metabolic capacities, several prophages and other new functions--all of which could be targets for surveillance.

Base Sequence↗

Thermoadaptation trait revealed by the genome sequence of thermophilic Geobacillus kaustophilus.

We present herein the first complete genome sequence of a thermophilic Bacillus-related species, Geobacillus kaustophilus HTA426, which is composed of a 3.54 Mb chromosome and a 47.9 kb plasmid, along with a comparative analysis with five other mesophilic bacillar genomes. Upon orthologous grouping of the six bacillar sequenced genomes, it was found that 1257 common orthologous groups composed of 1308 genes (37%) are shared by all the bacilli, whereas 839 genes (24%) in the G.kaustophilus genome were found to be unique to that species. We were able to find the first prokaryotic sperm protamine P1 homolog, polyamine synthase, polyamine ABC transporter and RNA methylase in the 839 unique genes; these may contribute to thermophily by stabilizing the nucleic acids. Contrasting results were obtained from the principal component analysis (PCA) of the amino acid composition and synonymous codon usage for highlighting the thermophilic signature of the G.kaustophilus genome. Only in the PCA of the amino acid composition were the Bacillus-related species located near, but were distinguishable from, the borderline distinguishing thermophiles from mesophiles on the second principal axis. Further analysis revealed some asymmetric amino acid substitutions between the thermophiles and the mesophiles, which are possibly associated with the thermoadaptation of the organism.

Adaptation, Physiological↗

GS-Aligner: a novel tool for aligning genomic sequences using bit-level operations.

A novel algorithm, GS-Aligner, that uses bit-level operations was developed for aligning genomic sequences. GS-Aligner is efficient in terms of both time and space for aligning two very long genomic sequences and for identifying genomic rearrangements such as translocations and inversions. It is suitable for aligning fairly divergent sequences such as human and mouse genomic sequences. It consists of several efficient components: bit-level coding, search for matching segments between the two sequences as alignment anchors, longest increasing subsequence (LIS), and optimal local alignment. Efforts have been made to reduce the execution time of the program to make it truly practical for aligning very long sequences. Empirical tests suggest that for relatively divergent sequences such as sequences from different mammalian orders or from a mammal and a nonmammalian vertebrate GS-Aligner performs better than existing methods. The program and data can be downloaded from http://pondside.uchicago.edu/~lilab/ and http://webcollab.iis.sinica.edu.tw/~biocom.

Algorithms↗

Further variability within the genus Crinivirus, as revealed by determination of the complete RNA genome sequence of Cucurbit yellow stunting disorder virus.

The complete nucleotide (nt) sequences of genomic RNAs 1 and 2 of Cucurbit yellow stunting disorder virus (CYSDV) were determined for the Spanish isolate CYSDV-AlLM. RNA1 is 9123 nt long and contains at least five open reading frames (ORFs). Computer-assisted analyses identified papain-like protease, methyltransferase, RNA helicase and RNA-dependent RNA polymerase domains in the first two ORFs of RNA1. This is the first study on the sequences of RNA1 from CYSDV. RNA2 is 7976 nt long and contains the hallmark gene array of the family Closteroviridae, characterized by ORFs encoding a heat shock protein 70 homologue, a 59 kDa protein, the major coat protein and a divergent copy of the coat protein. This genome organization resembles that of Sweet potato chlorotic stunt virus (SPCSV), Cucumber yellows virus (CuYV) and Lettuce infectious yellows virus (LIYV), the other three criniviruses sequenced completely to date. However, several differences were observed. The most striking novel features of CYSDV compared to SPCSV, CuYV and LIYV are a unique gene arrangement in the 3'-terminal region of RNA1, the identification in this region of an ORF potentially encoding a protein which has no homologues in any databases, and the prediction of an unusually long 5' non-coding region in RNA2. Additionally, the CYSDV genome resembles that of SPCSV in having very similar 3' regions in RNAs 1 and 2, although for CYSDV similarity in primary structures did not result in predictions of equivalent secondary structures. Overall, these data reinforce the view that the genus Crinivirus contains considerable genetic variation. Additionally, several subgenomic RNAs (sgRNAs) were detected in CYSDV-infected plants, suggesting that generation of sgRNAs is a strategy used by CYSDV for the expression of internal ORFs.

3' Untranslated Regions↗

Biology of Treponema pallidum: correlation of functional activities with genome sequence data.

Aspects of the biology of T. pallidum subsp. pallidum, the agent of syphilis, are examined in the context of a century of experimental studies and the recently determined genome sequence. T. pallidum and a group of closely related pathogenic spirochetes have evolved to become highly invasive, persistent pathogens with little toxigenic activity and an inability to survive outside the mammalian host. Analysis of the genome sequence confirms morphologic studies indicating the lack of lipopolysaccharide and lipid biosynthesis mechanisms, as well as a paucity of outer membrane protein candidates. The metabolic capabilities and adaptability of T. pallidum are minimal, and this relative deficiency is reflected by the absence of many pathways, including the tricarboxylic acid cycle, components of oxidative phosphorylation, and most biosynthetic pathways. Although multiplication of T. pallidum has been obtained in a tissue culture system, continuous in vitro culture has not been achieved. The balance of oxygen utilization and toxicity is key to the survival and growth of T. pallidum, and the genome sequence reveals a similarity to lactic acid bacteria that may be useful in understanding this relationship. The identification of relatively few genes potentially involved in pathogenesis reflects our lack of understanding of invasive pathogens relative to toxigenic organisms. The genome sequence will provide useful raw data for additional functional studies on the structure, metabolism, and pathogenesis of this enigmatic organism.

Animals↗

Gene recognition in eukaryotic DNA by comparison of genomic sequences.

MOTIVATION: Sequencing of complete eukaryotic genomes and large syntenic fragments of genomes makes it possible to apply genomic comparison for gene recognition. RESULTS: This paper describes a spliced alignment algorithm that aligns candidate exon chains of two homologous genomic sequence fragments from different species. The algorithm is implemented in Pro-Gen software. Unlike other algorithms, Pro-Gen does not assume conservation of the exon-intron structure. Amino acid sequences obtained by the formal translation of candidate exons are aligned instead of nucleotide sequences, which allows for distant comparisons. The algorithm was tested on a sample of human-mammal (mouse), human-vertebrate (Xenopus ) and human-invertebrate (Drosophila ) gene pairs. Surprisingly, the best results, 97-98% correlation between the actual and predicted genes, were obtained for more distant comparisons, whereas the correlation on the human-mouse sample was only 93%. The latter value increases to 95% if conservation of the exon-intron structure is assumed. This is caused by a large amount of sequence conservation in non-coding regions of the human and mouse genes probably due to regulatory elements. AVAILABILITY: Pro-Gen v. 3.0 is available to academic researchers free of charge at http://www.anchorgen.com/pro_gen/pro_gen.html.

Algorithms↗

Genome sequence of a serotype M28 strain of group a streptococcus: potential new insights into puerperal sepsis and bacterial disease specificity.

Puerperal sepsis, a major cause of death of young women in Europe in the 1800s, was due predominantly to the gram-positive pathogen group A Streptococcus. Studies conducted during past decades have shown that serotype M28 strains are the major group A Streptococcus organisms responsible for many of these infections. To begin to increase our understanding of their enrichment in puerperal sepsis, we sequenced the genome of a genetically representative strain. This strain has genes encoding a novel array of prophage virulence factors, cell-surface proteins, and other molecules likely to contribute to host-pathogen interactions. Importantly, genes for 7 inferred extracellular proteins are encoded by a 37.4-kb foreign DNA element that is shared with group B Streptococcus and is present in all serotype M28 strains. Proteins encoded by the 37.4-kb element were expressed extracellularly and in human infections. Acquisition of foreign genes has helped create a disease-specialist clone of this pathogen.

Antigens, Bacterial↗

Genome science: a video tour of the Washington University Genome Sequencing Center for high school and undergraduate students.

Sequencing of the human genome has ushered in a new era of biology. The technologies developed to facilitate the sequencing of the human genome are now being applied to the sequencing of other genomes. In 2004, a partnership was formed between Washington University School of Medicine Genome Sequencing Center's Outreach Program and Washington University Department of Biology Science Outreach to create a video tour depicting the processes involved in large-scale sequencing. "Sequencing a Genome: Inside the Washington University Genome Sequencing Center" is a tour of the laboratory that follows the steps in the sequencing pipeline, interspersed with animated explanations of the scientific procedures used at the facility. Accompanying interviews with the staff illustrate different entry levels for a career in genome science. This video project serves as an example of how research and academic institutions can provide teachers and students with access and exposure to innovative technologies at the forefront of biomedical research. Initial feedback on the video from undergraduate students, high school teachers, and high school students provides suggestions for use of this video in a classroom setting to supplement present curricula.

Feedback↗

GenePalette: a universal software tool for genome sequence visualization and analysis.

To make effective use of the growing host of complete genome sequences, biologists must have easy-to-use software tools that allow them to visualize, analyze, and modify genome data in an interactive and generalized manner. In an effort to bridge the gap between genome and researcher, we have created GenePalette (www.genepalette.org), a desktop application that can access any genome sequence and display the positions of various features [e.g., transcription factor binding sites (TFBSs)] relative to the introns and exons of annotated genes. Written in Java, GenePalette can run on all Java-supporting operating systems (Mac, PC, Unix, Linux). Annotated sequence encompassing the majority of public genome data is rapidly retrieved from GenBank or Ensembl. The software provides intuitive access to the selected genomic region through three interface components: a colorful graphical display showing a schematic of genes and features; an annotated sequence view in which features and genes are highlighted directly on the sequence; and the selectable raw sequence. The three interface components are fully integrated and presented on one page, permitting the user to move easily between representations at different levels of resolution, ranging from kilobases to individual nucleotides. GenePalette is a particularly powerful platform for analyzing the organization of cis-regulatory elements and designing wet-lab experiments to investigate them.

Computational Biology↗

Enzymatic amplification of beta-globin genomic sequences and restriction site analysis for diagnosis of sickle cell anemia.

Two new methods were used to establish a rapid and highly sensitive prenatal diagnostic test for sickle cell anemia. The first involves the primer-mediated enzymatic amplification of specific beta-globin target sequences in genomic DNA, resulting in the exponential increase (220,000 times) of target DNA copies. In the second technique, the presence of the beta A and beta S alleles is determined by restriction endonuclease digestion of an end-labeled oligonucleotide probe hybridized in solution to the amplified beta-globin sequences. The beta-globin genotype can be determined in less than 1 day on samples containing significantly less than 1 microgram of genomic DNA.

Alleles↗

Jak3 expression and genomic sequence in pediatric acute lymphoblastic leukemia.

Janus tyrosine kinase 3 (JAK3) is one of several key regulatory enzymes in B-cell precursors which is highly conserved between multiple species. The gene for Jak3 has been mapped to human chromosome 19p12-13.1 and encompasses 23 exons. Constitutively high levels of JAK3 activity may contribute to drug resistance and enhanced clonogenicity of leukemic B-cell precursors from children and infants with acute lymphoblastic leukemia (ALL). As part of a systematic effort to accurately determine the genomic sequence of Jak3 gene in normal and leukemic B-cell precursors, we sequenced a relatively short region of Jak3 spanning two introns, originally termed introns 10 and 11. This genomic sequence appeared in certain RT-PCR products from our analysis of Jak3 gene expression in pediatric, as well as infant, primary ALL cells. Unexpectedly, a gap in the original Jak3 genomic sequence was found in intron 10 across the sequence matching to an Alu element. Furthermore, the sequence obtained from intron 11 did not match at all to that previously reported, and the length of the intron was much larger than expected at 1.1 kb. Homology to Alu elements (three regions, 699 bp total) and a LINE2 element (one region, 189 bp total) were seen across the entire region covering exons 10-12 (2.1 kb total). Two potential single nucleotide polymorphisms (SNPs) were observed in intron 11. No apparent genomic mutation was found across this region in leukemic B-cell precursors from any of the ALL patients examined. This newly described sequence corrects the previous published genomic sequence from this region rather than identifying an insertion or translocation specific to these ALL cases. Our results significantly extend previous efforts to determine the genomic sequence of Jak3 and analyze its expression in childhood pro-B ALL and other forms of ALL.

B-Lymphocytes↗

The mitochondrial genome of the firefly, Pyrocoelia rufa: complete DNA sequence, genome organization, and phylogenetic analysis with other insects.

The complete nucleotide sequences of the mt genome from the firefly, Pyrococelia rufa (Coeleoptera: Lampyridae) was determined. The circular genome is 17,739-bp long, and contains a typical gene complement, order, and arrangement identical to Drosophila yacuba. The presence of 1,724-bp long intergenic spacer in the P. rufa mt genome is unique. The putative initiation codon for ND1 gene appears to be TTG, instead of frequently found ATN. All tRNAs showed stable canonical clover-leaf structure of other mt tRNAs, except for tRNA(Ser) (AGN), DHU arm of which could not form stable stem-loop structure. Phylogenetic analysis among insect orders confirmed a monophyletic Endopterygota, a monophyletic Mecopterida, a monophyletic Diptera, a monophyletic Lepidoptera, and a monophyletic Coleoptera, suggesting that the complete insect mt genome sequence has a resolving power in the diversification events within Endopterygota. However, internal relationships among three coleopteran species are not clear, and the inclusion of some insect orders (i.e., apterygotan T. gertschi) in the analysis provided inconsistent results compared to other molecular studies.

Animals↗