Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

The complete genome sequence of Mycobacterium avium subspecies paratuberculosis.

We describe here the complete genome sequence of a common clone of Mycobacterium avium subspecies paratuberculosis (Map) strain K-10, the causative agent of Johne's disease in cattle and other ruminants. The K-10 genome is a single circular chromosome of 4,829,781 base pairs and encodes 4,350 predicted ORFs, 45 tRNAs, and one rRNA operon. In silico analysis identified >3,000 genes with homologs to the human pathogen, M. tuberculosis (Mtb), and 161 unique genomic regions that encode 39 previously unknown Map genes. Analysis of nucleotide substitution rates with Mtb homologs suggest overall strong selection for a vast majority of these shared mycobacterial genes, with only 68 ORFs with a synonymous to nonsynonymous substitution ratio of >2. Comparative sequence analysis reveals several noteworthy features of the K-10 genome including: a relative paucity of the PE/PPE family of sequences that are implicated as virulence factors and known to be immunostimulatory during Mtb infection; truncation in the EntE domain of a salicyl-AMP ligase (MbtA), the first gene in the mycobactin biosynthesis gene cluster, providing a possible explanation for mycobactin dependence of Map; and Map-specific sequences that are likely to serve as potential targets for sensitive and specific molecular and immunologic diagnostic tests. Taken together, the availability of the complete genome sequence offers a foundation for the study of the genetic basis for virulence and physiology in Map and enables the development of new generations of diagnostic tests for bovine Johne's disease.

Animals↗

What's in the genome of a filamentous fungus? Analysis of the Neurospora genome sequence.

The German Neurospora Genome Project has assembled sequences from ordered cosmid and BAC clones of linkage groups II and V of the genome of Neurospora crassa in 13 and 12 contigs, respectively. Including additional sequences located on other linkage groups a total of 12 Mb were subjected to a manual gene extraction and annotation process. The genome comprises a small number of repetitive elements, a low degree of segmental duplications and very few paralogous genes. The analysis of the 3218 identified open reading frames provides a first overview of the protein equipment of a filamentous fungus. Significantly, N.crassa possesses a large variety of metabolic enzymes including a substantial number of enzymes involved in the degradation of complex substrates as well as secondary metabolism. While several of these enzymes are specific for filamentous fungi many are shared exclusively with prokaryotes.

Chromosome Mapping↗

Automated gene identification in large-scale genomic sequences.

Computational methods for gene identification in genomic sequences typically have two phases: coding region recognition and gene parsing. While there are a number of effective methods for recognizing coding regions (exons), parsing the recognized exons into proper gene structures, to a large extent, remains an unsolved problem. We have developed a computer program which can automatically parse the recognized exons into gene models that are most consistent with the available Expressed Sequence Tags (ESTs) and a set of biological heuristics, derived empirically. The gene modeling algorithm used in this program provides a general framework for applying EST information so the modeling accuracy improves as the amount of available EST information increases. Based on preliminary tests on a number of large DNA sequences, using the dbEST database, we have observed that the algorithm can (1) accurately model complicated multiple gene structures, including embedded genes, (2) identify falsely-recognized exons and locate missed exons by the initial exon recognition phase, and (3) make more accurate exon boundary predictions, if the necessary EST information is available. We have extended this EST-based gene modeling algorithm to model genes on unfinished DNA contigs at the end of the shotgun sequencing. This extended version can automatically determine the orientations and the relative order of the DNA contigs (with gaps between them) using the available ESTs as reference models, before the gene modeling phase.

Algorithms↗

Polypeptide hormone regulation of gene transcription: specific 5' genomic sequences are required for epidermal growth factor and phorbol ester regulation of prolactin gene expression.

A fusion gene containing 5' rat prolactin genomic sequences ligated to the structural portion of the rat growth hormone gene ( grl ) was introduced by DNA-mediated gene transfer into mammalian cells by using a chimeric plasmid vector. Clonal transfected cell lines produced a mRNA that used the authentic 5' initiation site and that was processed to the predicted size. The intracellular levels of this RNA product were increased 2.5- to 5-fold by exposure of the cells to epidermal growth factor (EGF) and 2- to 3-fold by exposure of the cells to a potent phorbol ester, phorbol 12-myristate 13-acetate, apparently due to regulation at the level of gene transcription. Substitution of the 5' prolactin DNA sequences by 5' growth hormone DNA sequences resulted in the loss of EGF inducibility. A genomic sequence in or near the 5' flanking portion of the prolactin gene therefore appears to confer polypeptide hormone transcriptional regulation upon the gene.

Base Sequence↗

Coelacanth genome sequence reveals the evolutionary history of vertebrate genes.

The coelacanth is one of the nearest living relatives of tetrapods. However, a teleost species such as zebrafish or Fugu is typically used as the outgroup in current tetrapod comparative sequence analyses. Such studies are complicated by the fact that teleost genomes have undergone a whole-genome duplication event, as well as individual gene-duplication events. Here, we demonstrate the value of coelacanth genome sequence by complete sequencing and analysis of the protocadherin gene cluster of the Indonesian coelacanth, Latimeria menadoensis. We found that coelacanth has 49 protocadherin cluster genes organized in the same three ordered subclusters, alpha, beta, and gamma, as the 54 protocadherin cluster genes in human. In contrast, whole-genome and tandem duplications have generated two zebrafish protocadherin clusters comprised of at least 97 genes. Additionally, zebrafish protocadherins are far more prone to homogenizing gene conversion events than coelacanth protocadherins, suggesting that recombination- and duplication-driven plasticity may be a feature of teleost genomes. Our results indicate that coelacanth provides the ideal outgroup sequence against which tetrapod genomes can be measured. We therefore present L. menadoensis as a candidate for whole-genome sequencing.

Animals↗

Comparative analysis of genome sequences of three isolates of Orf virus reveals unexpected sequence variation.

Orf virus (ORFV) is the type species of the Parapoxvirus genus. Here, we present the genomic sequence of the most well studied ORFV isolate, strain NZ2. The NZ2 genome is 138 kbp and contains 132 putative genes, 88 of which are present in all analyzed chordopoxviruses. Comparison of the NZ2 genome with the genomes of 2 other fully sequenced isolates of ORFV revealed that all 3 genomes carry each of the 132 genes, but there are substantial sequence variations between isolates in a significant number of genes, including 9 with inter-isolate amino acid sequence identity of only 38-79%. Each genome has an average of 64% G+C but each has a distinctive pattern of substantial deviation from the average within particular regions of the genome. The same pattern of variation was also seen in the genome of another parapoxvirus species and was clearly unlike the uniform patterns of G+C content seen in all other genera of chordopoxviruses. The availability of genomic sequences of three orf virus isolates allowed us to more accurately assess likely coding regions and thereby revise published data for 24 genes and to predict two previously unrecognized genes.

Base Composition↗

Use of a fluorescent-PCR reaction to detect genomic sequence copy number and transcriptional abundance.

We present a fluorescent-PCR-based technique to assay genomic sequence copy number and transcriptional abundance. This technique relies on the ability to follow fluorescent PCR progressively in real time during the exponential phase of the reaction so that quantitative PCR is accomplished. We demonstrated the ability of this technique to quantitate both known deletions and amplifications of loci that have been measured previously by other methods, and to measure transcriptional abundance. Using an efficient variant of the fluorescent-PCR technology, we can monitor transcription semiquantitatively. The ability to detect all amplifications and deletions at any single copy locus by PCR makes this the technique of choice to assay genomic sequence copy number anomalies in birth defects and cancers. The ability to detect variations in transcript abundance enables this technique to fashion a time and tissue analysis of transcription.

Chromosome Aberrations↗

Chromosomal localization and complete genomic sequence of the murine autoimmune regulator gene (Aire).

We have recently cloned the murine autoimmune regulator (Aire) gene, the homologue of human AIRE responsible for the autoimmune polyglandular syndrome type 1 (APS1) or autoimmune polyendocrinopathy candidiasis ectodermal dystrophy (APECED). Here, we report the genomic sequence (18,413 bp) for the entire Aire gene and its 5' flanking region, which contains putative regulatory sequences. Comparison of the genomic and cDNA sequences indicates that the Aire gene is composed of 14 exons and the coding sequence shares high similarities between mouse and human. The sizes of the homologous introns in the two species are conserved; however, the introns do not share significant sequence homologies except the sequences near the splice donor and acceptor sites. Sequence analyses of the 5' regulatory region and the complete coding region in three mouse strains (B6, NOD and SJL) did not reveal any sequence variation, suggesting sequence conservation between different inbred mouse strains. Using one of the six microsatellite markers identified by genomic sequencing and a B6 x Cast backcross mapping panel, we mapped the mouse Aire gene to chromosome 10, a syntenic region containing the Cdl18 and Pfkl genes on human chromosome 21q22.

Animals↗

Analysis of multiple genomic sequence alignments: a web resource, online tools, and lessons learned from analysis of mammalian SCL loci.

Comparative analysis of genomic sequences is becoming a standard technique for studying gene regulation. However, only a limited number of tools are currently available for the analysis of multiple genomic sequences. An extensive data set for the testing and training of such tools is provided by the SCL gene locus. Here we have expanded the data set to eight vertebrate species by sequencing the dog SCL locus and by annotating the dog and rat SCL loci. To provide a resource for the bioinformatics community, all SCL sequences and functional annotations, comprising a collation of the extensive experimental evidence pertaining to SCL regulation, have been made available via a Web server. A Web interface to new tools specifically designed for the display and analysis of multiple sequence alignments was also implemented. The unique SCL data set and new sequence comparison tools allowed us to perform a rigorous examination of the true benefits of multiple sequence comparisons. We demonstrate that multiple sequence alignments are, overall, superior to pairwise alignments for identification of mammalian regulatory regions. In the search for individual transcription factor binding sites, multiple alignments markedly increase the signal-to-noise ratio compared to pairwise alignments.

Animals↗

The genome sequence of the obligately chemolithoautotrophic, facultatively anaerobic bacterium Thiobacillus denitrificans.

The complete genome sequence of Thiobacillus denitrificans ATCC 25259 is the first to become available for an obligately chemolithoautotrophic, sulfur-compound-oxidizing, beta-proteobacterium. Analysis of the 2,909,809-bp genome will facilitate our molecular and biochemical understanding of the unusual metabolic repertoire of this bacterium, including its ability to couple denitrification to sulfur-compound oxidation, to catalyze anaerobic, nitrate-dependent oxidation of Fe(II) and U(IV), and to oxidize mineral electron donors. Notable genomic features include (i) genes encoding c-type cytochromes totaling 1 to 2 percent of the genome, which is a proportion greater than for almost all bacterial and archaeal species sequenced to date, (ii) genes encoding two [NiFe]hydrogenases, which is particularly significant because no information on hydrogenases has previously been reported for T. denitrificans and hydrogen oxidation appears to be critical for anaerobic U(IV) oxidation by this species, (iii) a diverse complement of more than 50 genes associated with sulfur-compound oxidation (including sox genes, dsr genes, and genes associated with the AMP-dependent oxidation of sulfite to sulfate), some of which occur in multiple (up to eight) copies, (iv) a relatively large number of genes associated with inorganic ion transport and heavy metal resistance, and (v) a paucity of genes encoding organic-compound transporters, commensurate with obligate chemolithoautotrophy. Ultimately, the genome sequence of T. denitrificans will enable elucidation of the mechanisms of aerobic and anaerobic sulfur-compound oxidation by beta-proteobacteria and will help reveal the molecular basis of this organism's role in major biogeochemical cycles (i.e., those involving sulfur, nitrogen, and carbon) and groundwater restoration.

Bacterial Proteins↗

Impact of the first Streptomyces genome sequence on the discovery and production of bioactive substances.

An important addition to the field of bacterial genomics is the recent publication of the complete genome sequence of Streptomyces coelicolor. This strain has been for some decades the model organism for streptomycetes and other filamentous actinomycetes, Gram-positive bacteria highly valuable for their ability to produce thousands of bioactive metabolites, many of which have found important applications in medicine and agriculture. We discuss here the impacts that the S. coelicolor genome sequence is likely to have on the production of bioactive metabolites by current industrial strains, on the possible development of future superhost(s) for the production of valuable drugs, and on the search for new bioactive substances from microbial sources.

Anti-Bacterial Agents↗

Isolation of rapidly evolving genomic sequences: construction of a differential library and identification of a human DNA fragment that does not hybridize to chimpanzee DNA.

A differential library enriched in rapidly evolving human genomic sequences was obtained by phenol-enhanced hybridization of human genomic DNA with an excess of chimpanzee DNA. A DNA fragment 110 bp in length that did not hybridize to either chimpanzee or other primate DNA was identified in this library. It was shown to be a substantially diverged member of the human beta satellite family of tandem repeats. The genomic sequences homologous to the fragment were located on the short arms of human acrocentric chromosomes by in situ hybridization. The human-specific fragment failed to hybridize with RNA from different human tissues. The human-specific fragment exhibits a remarkable level of DNA polymorphism in humans and may be used in the identification of human tissue samples, in the selection of human/rodent somatic cell hybrids containing human acrocentric chromosomes, and in the mapping of these chromosomes.

Animals↗

Complete genome sequence of bacteriophage T5.

The 121,752-bp genome sequence of bacteriophage T5 was determined; the linear, double-stranded DNA is nicked in one of the strands and has large direct terminal repeats of 10,139 bp (8.3%) at both ends. The genome structure is consistently arranged according to its lytic life cycle. Of the 168 potential open reading frames (ORFs), 61 were annotated; these annotated ORFs are mainly enzymes involved in phage DNA replication, repair, and nucleotide metabolism. At least five endonucleases that believed to help inducing nicks in T5 genomic DNA, and a DNA ligase gene was found to be split into two separate ORFs. Analysis of T5 early promoters suggests a probable motif AAA{3, 4 T}nTTGCTT{17, 18 n}TATAATA{12, 13 W}{10 R} for strong promoters that may strengthen the step modification of host RNA polymerase, and thus control transcription of phage DNA. The distinct protein domain profile and a mosaic genome structure suggest an origin from the common genetic pool.

Bacteriophages↗

Extensive mosaic structure revealed by the complete genome sequence of uropathogenic Escherichia coli.

We present the complete genome sequence of uropathogenic Escherichia coli, strain CFT073. A three-way genome comparison of the CFT073, enterohemorrhagic E. coli EDL933, and laboratory strain MG1655 reveals that, amazingly, only 39.2% of their combined (nonredundant) set of proteins actually are common to all three strains. The pathogen genomes are as different from each other as each pathogen is from the benign strain. The difference in disease potential between O157:H7 and CFT073 is reflected in the absence of genes for type III secretion system or phage- and plasmid-encoded toxins found in some classes of diarrheagenic E. coli. The CFT073 genome is particularly rich in genes that encode potential fimbrial adhesins, autotransporters, iron-sequestration systems, and phase-switch recombinases. Striking differences exist between the large pathogenicity islands of CFT073 and two other well-studied uropathogenic E. coli strains, J96 and 536. Comparisons indicate that extraintestinal pathogenic E. coli arose independently from multiple clonal lineages. The different E. coli pathotypes have maintained a remarkable synteny of common, vertically evolved genes, whereas many islands interrupting this common backbone have been acquired by different horizontal transfer events in each strain.

Acute Disease↗

Bacterial genome sequencing and drug discovery.

The availability of bacterial genome sequence information has opened up many new strategies for antibacterial drug hunting. There are obvious benefits for the identification and evaluation of new drug targets, but genomic-based technology is also beginning to provide new tools for the downstream, preclinical, optimisation of compounds. The greatest benefit from these new approaches lies in the ability to examine the entire genome (or several genomes) simultaneously and in total. In this way, one potential target can be evaluated against another, and either the total effects of functional impairment can be established or the effects of a compound can be compared across species.

Anti-Bacterial Agents↗

Use of bovine EST data and human genomic sequences to map 100 gene-specific bovine markers.

A system to use bovine EST data in conjunction with human genomic sequence to improve the bovine linkage map over the entire genome or on specific chromosomes was evaluated. Bovine EST sequence was used to provide primer sequences corresponding to bovine genes, while human genomic sequence directed primer design to flank introns and produce amplicons of appropriate size for efficient direct sequencing. The sequence tagged sites (STS) produced in this way from the four sires of the MARC reference families were examined for single nucleotide polymorphisms (SNPs) that could be used to map the corresponding genes. With this approach, along with a primer/extension mass spectrometry SNP genotyping assay, 100 ESTs were placed on the bovine genetic linkage map. The first 70 were chosen at random from bovine EST-human genomic comparisons. An additional 30 ESTs were successfully mapped to bovine Chromosome 19 (BTA19), and comparison of the resulting BTA19 map to the position of the corresponding human orthologs on the HSA17 draft sequences revealed differences in the spacing and order of genes. Over 80% of successful amplicons contained SNPs, indicating that this is an efficient approach to generating EST-associated genetic markers. We have demonstrated the feasibility of constructing a linkage map based on SNPs associated with ESTs and the plausibility of utilizing EST, comparative mapping information, and human sequence data to target regions of the bovine genome for SNP marker development.

Animals↗

A protein database constructed from low-coverage genomic sequence of Bacillus megaterium and its use for accelerated proteomic analysis.

Peptide mass fingerprint (PMF) matching is a high-throughput method used for protein spot identification in connection with two-dimensional gel electrophoresis (2DE). However, the success of PMF matching largely depends on whether the proteins to be identified exist in the database searched. Consequently, it is often necessary to apply other more sophisticated but also time-consuming technologies to generate sequence-tags for definitive protein identification. On the other hand, modern sequencing technologies are generating a large quantity of DNA sequences, first in unfinished form or with low genome coverage due to the time-consuming and thus limiting steps of finishing and annotation. We recently started to sequence the genome of Bacillus megaterium DSM 319, a bacterium of industrial interest. In this study, we demonstrate that a protein database generated from merely three-fold coverage, unfinished genomic sequences of this bacterium allows a fast and reliable protein spot identification solely based on PMF from high-throughput MALDI-TOF MS analysis. We further show that the strain-specific protein database from low coverage genomic sequence greatly outperforms the commonly used cross-species databases constructed from 13 completely sequenced Bacillus strains for protein spot identification via PMF.

Algorithms↗

Finishing a whole-genome shotgun: release 3 of the Drosophila melanogaster euchromatic genome sequence.

BACKGROUND: The Drosophila melanogaster genome was the first metazoan genome to have been sequenced by the whole-genome shotgun (WGS) method. Two issues relating to this achievement were widely debated in the genomics community: how correct is the sequence with respect to base-pair (bp) accuracy and frequency of assembly errors? And, how difficult is it to bring a WGS sequence to the accepted standard for finished sequence? We are now in a position to answer these questions. RESULTS: Our finishing process was designed to close gaps, improve sequence quality and validate the assembly. Sequence traces derived from the WGS and draft sequencing of individual bacterial artificial chromosomes (BACs) were assembled into BAC-sized segments. These segments were brought to high quality, and then joined to constitute the sequence of each chromosome arm. Overall assembly was verified by comparison to a physical map of fingerprinted BAC clones. In the current version of the 116.9 Mb euchromatic genome, called Release 3, the six euchromatic chromosome arms are represented by 13 scaffolds with a total of 37 sequence gaps. We compared Release 3 to Release 2; in autosomal regions of unique sequence, the error rate of Release 2 was one in 20,000 bp. CONCLUSIONS: The WGS strategy can efficiently produce a high-quality sequence of a metazoan genome while generating the reagents required for sequence finishing. However, the initial method of repeat assembly was flawed. The sequence we report here, Release 3, is a reliable resource for molecular genetic experimentation and computational analysis.

Animals↗