Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

A low rate of nucleotide changes in Escherichia coli K-12 estimated from a comparison of the genome sequences between two different substrains.

Two genome sequences of Escherichia coli K-12 substrains, one partial W3110 and one complete MG1655, have been determined by Japanese and American genome projects, respectively. In order to estimate the rate of nucleotide changes, we directly compared 2 Mb of the nucleotide sequences from these closely-related E. coli substrains. Given that the two substrains separated about 40 years ago, the rate of nucleotide changes was estimated to be less than 10(-7) per site per year. This rate was supported by a further comparison between partial genome sequences of E. coli and Shigella flexneri.

Escherichia coli↗

Whole-genome sequence of Listeria welshimeri reveals common steps in genome reduction with Listeria innocua as compared to Listeria monocytogenes.

We present the complete genome sequence of Listeria welshimeri, a nonpathogenic member of the genus Listeria. Listeria welshimeri harbors a circular chromosome of 2,814,130 bp with 2,780 open reading frames. Comparative genomic analysis of chromosomal regions between L. welshimeri, Listeria innocua, and Listeria monocytogenes shows strong overall conservation of synteny, with the exception of the translocation of an F(o)F(1) ATP synthase. The smaller size of the L. welshimeri genome is the result of deletions in all of the genes involved in virulence and of "fitness" genes required for intracellular survival, transcription factors, and LPXTG- and LRR-containing proteins as well as 55 genes involved in carbohydrate transport and metabolism. In total, 482 genes are absent from L. welshimeri relative to L. monocytogenes. Of these, 249 deletions are commonly absent in both L. welshimeri and L. innocua, suggesting similar genome evolutionary paths from an ancestor. We also identified 311 genes specific to L. welshimeri that are absent in the other two species, indicating gene expansion in L. welshimeri, including horizontal gene transfer. The species L. welshimeri appears to have been derived from early evolutionary events and an ancestor more compact than L. monocytogenes that led to the emergence of nonpathogenic Listeria spp.

Chromosomes, Bacterial↗

Genomic sequence and chromosomal location of human interleukin-11 gene (IL11).

The genomic sequence of human interleukin-11 (IL11) has been isolated based on its sequence homology with a cDNA clone encoding primate IL11. The human IL11 genomic sequence is 7 kb in length and consists of five exons and four introns. The IL11 gene has been localized to the long arm of human chromosome 19 at band 19q13.3-q13.4 by in situ hybridization. Several potential transcriptional control sequences that may play an important role in the control of IL11 gene expression have been identified within the 5'-flanking region of the human IL11 gene. The 5'-flanking region of the human IL11 gene contains sequences similar to those present in the 5'-regulatory regions of other cytokine genes. Four repeats of a 30-bp DNA segment were found in the fourth intron, and several copies of ATTTA and Alu repetitive sequences were identified in the 3'-noncoding region of the human IL11 gene. A DNA sequence (ACATGGCAAAACCC) that has 71% similarity with the IL1-responsive element of the IL6 gene was found in the 3'-flanking region of the human IL11 gene. The two polyadenylation sites located at nucleotide positions 6762 and 5591 correspond to the 2.5- and 1.5-kb IL11 transcripts expressed in IL1-induced PU-34 cells. The availability of the human IL11 genomic sequence should aid in studies of the regulatory mechanisms of IL11 gene expression in different cell types and the role of IL11 expression in the hematopoietic microenvironment.

Amino Acid Sequence↗

Prophage Finder: a prophage loci prediction tool for prokaryotic genome sequences.

Prophage loci often remain under-annotated or even unrecognized in prokaryotic genome sequencing projects. A PHP application, Prophage Finder, has been developed and implemented to predict prophage loci, based upon clusters of phage-related gene products encoded within DNA sequences. This application provides results detailing several facets of these clusters to facilitate rapid prediction and analysis of prophage sequences. Prophage Finder was tested using previously annotated prokaryotic genomic sequences with manually curated prophage loci as benchmarks. Additional analyses from Prophage Finder searches of several draft prokaryotic genome sequences are available through the Web site (http://bioinformatics.uwp.edu/~phage/DOEResults.php) to illustrate the potential of this application.

Bacteria↗

The Genome sequence of the SARS-associated coronavirus.

We sequenced the 29,751-base genome of the severe acute respiratory syndrome (SARS)-associated coronavirus known as the Tor2 isolate. The genome sequence reveals that this coronavirus is only moderately related to other known coronaviruses, including two human coronaviruses, HCoV-OC43 and HCoV-229E. Phylogenetic analysis of the predicted viral proteins indicates that the virus does not closely resemble any of the three previously known groups of coronaviruses. The genome sequence will aid in the diagnosis of SARS virus infection in humans and potential animal hosts (using polymerase chain reaction and immunological tests), in the development of antivirals (including neutralizing antibodies), and in the identification of putative epitopes for vaccine development.

3' Untranslated Regions↗

Frame: detection of genomic sequencing errors.

MOTIVATION: The underlying error rate for genomic sequencing sometimes results in the introduction of artificial frameshifts and in-frame stop codons into putative protein encoding genes. Severe errors are then introduced into the inferred transcripts through mis-translation or premature termination. RESULTS: We describe a system for screening segments of DNA for frameshift and in-frame stop errors in coding regions. The method is based on homology matching using blastx to compare all six reading frames of the query nucleotide sequence against selected protein sequence databases. Fragments of protein matching neighbouring regions of the query DNA are united and extended laterally to define candidate open reading frames, within which, frameshifts and stops are identified. Suitable targets include prokaryotic or other intron-free genomic sequence and complementary DNAs. As an example of its use, we report here two frameshifted ORFs that deviate from the original TIGR sequence annotations for the recently released Helicobacter pylori genome. AVAILABILITY: The tool is accessible via the URL http://www.sander.ebi.ac.uk/frame/. CONTACT: brown@ebi.ac.uk.

Amino Acid Sequence↗

Genome sequence of Yersinia pestis KIM.

We present the complete genome sequence of Yersinia pestis KIM, the etiologic agent of bubonic and pneumonic plague. The strain KIM, biovar Mediaevalis, is associated with the second pandemic, including the Black Death. The 4.6-Mb genome encodes 4,198 open reading frames (ORFs). The origin, terminus, and most genes encoding DNA replication proteins are similar to those of Escherichia coli K-12. The KIM genome sequence was compared with that of Y. pestis CO92, biovar Orientalis, revealing homologous sequences but a remarkable amount of genome rearrangement for strains so closely related. The differences appear to result from multiple inversions of genome segments at insertion sequences, in a manner consistent with present knowledge of replication and recombination. There are few differences attributable to horizontal transfer. The KIM and E. coli K-12 genome proteins were also compared, exposing surprising amounts of locally colinear "backbone," or synteny, that is not discernible at the nucleotide level. Nearly 54% of KIM ORFs are significantly similar to K-12 proteins, with conserved housekeeping functions. However, a number of E. coli pathways and transport systems and at least one global regulator were not found, reflecting differences in lifestyle between them. In KIM-specific islands, new genes encode candidate pathogenicity proteins, including iron transport systems, putative adhesins, toxins, and fimbriae.

Bacteriophages↗

Automated correction of genome sequence errors.

By using information from an assembly of a genome, a new program called AutoEditor significantly improves base calling accuracy over that achieved by previous algorithms. This in turn improves the overall accuracy of genome sequences and facilitates the use of these sequences for polymorphism discovery. We describe the algorithm and its application in a large set of recent genome sequencing projects. The number of erroneous base calls in these projects was reduced by 80%. In an analysis of over one million corrections, we found that AutoEditor made just one error per 8828 corrections. By substantially increasing the accuracy of base calling, AutoEditor can dramatically accelerate the process of finishing genomes, which involves closing all gaps and ensuring minimum quality standards for the final sequence. It also greatly improves our ability to discover single nucleotide polymorphisms (SNPs) between closely related strains and isolates of the same species.

Algorithms↗

Insights from the rat genome sequence.

The availability of the rat genome sequence, and detailed three-way comparison of the rat, mouse and human genomes, is revealing a great deal about mammalian genome evolution. Together with recent developments in cloning technologies, this heralds an important phase in rat research.

Animals↗

Molecular analysis of two complete rice tungro bacilliform virus genomic sequences from India.

The complete genomic sequences of two geographically distinct isolates of rice tungro bacilliform virus (RTBV) from India were determined. Both the sequences showed equal divergence from previously reported Southeast Asian isolates. Numerous insertions, deletions and substitutions, mostly in the intergenic regions, were found. The genome sizes were 7907 and 7934 bp respectively, 95 and 68 residues short of an infectious clone reported earlier. Between them, both the isolates showed high homology all along the genome, except for a 30-nucleotide insertion/deletion close to the 3' end of ORF III in one of them. Both the isolates indicated an unconventional start codon in ORF I, similar to the type isolate. In addition, as novel features, both the Indian isolates showed an unconventional start codon for ORF IV. Considering the low amounts of genome variability noticed in other RTBV isolates, the Indian isolates show that they have diverged sufficiently from the rest and should be considered belonging to a distinct strain.

Amino Acid Sequence↗

Whole-Genome Sequencing Reveals Population Structure, Genetic Diversity, and Selection Signatures in Kazakh Dromedary and Bactrian Camels.

Understanding the genomic basis of environmental adaptation is essential for the conservation and genetic improvement of domestic camels. In this study, we investigated the population structure, genetic diversity, and genomic variation potentially associated with environmental adaptation of Kazakh dromedary and Bactrian camels using whole-genome sequencing. Whole-genome sequencing data were generated for Kazakh camels (15 dromedaries and 16 Bactrian camels) and integrated with 131 publicly available genomes representing camel populations from the Arabian Peninsula, Iran, Xinjiang, Inner Mongolia, and Mongolian wild camels. Population structure, genetic diversity, and genome-wide selection were evaluated using principal component analysis, ADMIXTURE, nucleotide diversity, linkage disequilibrium, runs of homozygosity, genomic inbreeding (FROH), and selection scans based on FST, θπ ratio, and XP-EHH. Population genomic analyses revealed clear differentiation between dromedary and Bactrian camels, whereas Kazakh camel populations exhibited higher nucleotide diversity (θπ = 1.307-1.551 × 10-3), and lower genomic inbreeding (median FROH: 0.037-0.056) than Arabian populations. Genome-wide selection analyses identified MC4R as the prominent candidate gene in Kazakh dromedaries and RYR1 as a prominent candidate gene in Kazakh Bactrian camels. Functional enrichment analyses highlighted pathways related to energy metabolism, thermogenesis, calcium signaling, skeletal muscle function, mitochondrial activity, and oxidative stress response. These findings provide new insights into genomic variation potentially associated with environmental adaptation in Kazakh camels and offer valuable genomic resources for future conservation, breeding, and evolutionary studies.

MC4R↗

An intermediate grade of finished genomic sequence suitable for comparative analyses.

Although the cost of generating draft-quality genomic sequence continues to decline, refining that sequence by the process of "sequence finishing" remains expensive. Near-perfect finished sequence is an appropriate goal for the human genome and a small set of reference genomes; however, such a high-quality product cannot be cost-justified for large numbers of additional genomes, at least for the foreseeable future. Here we describe the generation and quality of an intermediate grade of finished genomic sequence (termed comparative-grade finished sequence), which is tailored for use in multispecies sequence comparisons. Our analyses indicate that this sequence is very high quality (with the residual gaps and errors mostly falling within repetitive elements) and reflects 99% of the total sequence. Importantly, comparative-grade sequence finishing requires approximately 40-fold less reagents and approximately 10-fold less personnel effort compared to the generation of near-perfect finished sequence, such as that produced for the human genome. Although applied here to finishing sequence derived from individual bacterial artificial chromosome (BAC) clones, one could envision establishing routines for refining sequences emanating from whole-genome shotgun sequencing projects to a similar quality level. Our experience to date demonstrates that comparative-grade sequence finishing represents a practical and affordable option for sequence refinement en route to comparative analyses.

Animals↗

The genomic sequence of defective interfering Semliki Forest virus (SFV) determines its ability to be replicated in mouse brain and to protect against a lethal SFV infection in vivo.

We have recently cloned and sequenced two genomes of defective interfering (DI) Semliki Forest virus (SFV), DI-6 (2146 nt), and DI-19 (1244 nt). These are similar in that both contain two large central deletions (encompassing the 5' part of the nsP1 gene and the 3' part of the nsP2 gene and all of the structural genes), and all the sequence of the latter is represented in the genome of SFV DI-6. RNA was transcribed from both and transfected into SFV-infected BHK-21 cells. RT-PCR analysis of tissue culture fluid harvested 18 h after transfection suggested that SFV DI virions had been rescued from the cloned genomes. Unlike the genomes of noncloned DI SFV, these genomes bred true for at least 7 serial passages. Cloned DI-6 and DI-19 viruses interfered to a similar extent with the multiplication of SFV in cultured cells, but only DI-19 protected mice from a lethal intranasal dose of SFV. Further investigation by RT-PCR analysis showed that DI-19 but not DI-6 genomes were replicated in mouse brain after direct intracerebral injection of DI virus together with an excess of infectious helper SFV. Thus the replication and hence antiviral activity of two closely related DI SFV genomes appears to be exquisitely sequence specific and cell specific. These findings mark a significant step on the way to using DI genomes as antivirals and also may explain why so few animal-protecting DI viruses have been identified.

Alphavirus Infections↗

The meaning and impact of the human genome sequence for microbiology.

The characterization of life is immeasurably enhanced by determination of complete genome sequences. For organisms that engage in intimate interactions with others, the genome sequence from one participant, and associated tools, provide unique insight into its partner. We discuss how the human genome sequence will further our understanding of microbial pathogens and commensals, and vice versa. We also propose criteria for implicating a host gene in microbial pathogenesis, and urge consideration of a'second human genome project'.

Bacteria↗

Random sheared fosmid library as a new genomic tool to accelerate complete finishing of rice (Oryza sativa spp. Nipponbare) genome sequence: sequencing of gap-specific fosmid clones uncovers new euchromatic portions of the genome.

The International Rice Genome Sequencing Project has recently announced the high-quality finished sequence that covers nearly 95% of the japonica rice genome representing 370 Mbp. Nevertheless, the current physical map of japonica rice contains 62 physical gaps corresponding to approximately 5% of the genome, that have not been identified/represented in the comprehensive array of publicly available BAC, PAC and other genomic library resources. Without finishing these gaps, it is impossible to identify the complete complement of genes encoded by rice genome and will also leave us ignorant of some 5% of the genome and its unknown functions. In this article, we report the construction and characterization of a tenfold redundant, 40 kbp insert fosmid library generated by random mechanical shearing. We demonstrated its utility in refining the physical map of rice by identifying and in silico mapping 22 gap-specific fosmid clones with particular emphasis on chromosomes 1, 2, 6, 7, 8, 9 and 10. Further sequencing of 12 of the gap-specific fosmid clones uncovered unique rice genome sequence that was not previously reported in the finished IRGSP sequence and emphasizes the need to complete finishing of the rice genome.

Base Sequence↗

Genome sequence and comparative microarray analysis of serotype M18 group A Streptococcus strains associated with acute rheumatic fever outbreaks.

Acute rheumatic fever (ARF), a sequelae of group A Streptococcus (GAS) infection, is the most common cause of preventable childhood heart disease worldwide. The molecular basis of ARF and the subsequent rheumatic heart disease are poorly understood. Serotype M18 GAS strains have been associated for decades with ARF outbreaks in the U.S. As a first step toward gaining new insight into ARF pathogenesis, we sequenced the genome of strain MGAS8232, a serotype M18 organism isolated from a patient with ARF. The genome is a circular chromosome of 1,895,017 bp, and it shares 1.7 Mb of closely related genetic material with strain SF370 (a sequenced serotype M1 strain). Strain MGAS8232 has 178 ORFs absent in SF370. Phages, phage-like elements, and insertion sequences are the major sources of variation between the genomes. The genomes of strain MGAS8232 and SF370 encode many of the same proven or putative virulence factors. Importantly, strain MGAS8232 has genes encoding many additional secreted proteins involved in human-GAS interactions, including streptococcal pyrogenic exotoxin A (scarlet fever toxin) and two uncharacterized pyrogenic exotoxin homologues, all phage-associated. DNA microarray analysis of 36 serotype M18 strains from diverse localities showed that most regions of variation were phages or phage-like elements. Two epidemics of ARF occurring 12 years apart in Salt Lake City, UT, were caused by serotype M18 strains that were genetically identical, or nearly so. Our analysis provides a critical foundation for accelerated research into ARF pathogenesis and a molecular framework to study the plasticity of GAS genomes.

Acute Disease↗

Evaluation of one-step amplicon-based targeted enrichment for SARS-CoV-2 whole-genome sequencing using the Midnight amplicon scheme.

Genomic surveillance proved invaluable during the COVID-19 pandemic for tracking SARS-CoV-2 variants and guiding outbreak responses, underscoring the ongoing need to reduce whole-genome sequencing (WGS) costs and improve workflow efficiency to ensure accessibility in resource limited settings. Here, we evaluated a one-step reverse transcription polymerase chain reaction (RT-PCR) approach using the Midnight V2 primer scheme for targeted amplification of the SARS-CoV-2 genome, assessed its compatibility with Illumina sequencing, and compared its performance to a well-established two-step method. Initially, we determined optimal RT-PCR reaction conditions using the Midnight V2 primer panel for the one-step RT-PCR kit and scaled reaction volumes for both RT-PCR and library preparation. Clinical specimens (n = 53) that had undergone routine WGS for surveillance purposes using the established two-step RT-PCR method were compared using the one-step RT-PCR assay. For samples with genome completeness greater than 70%, both methods gave comparable results with similar sequence coverage and 100% concordance for lineage assignment. Further investigation revealed a higher percentage of reads aligning to the SARS-CoV-2 genome with a greater depth of coverage using the one-step method compared to the two-step method. Finally, analysis of scaled one-step and library reaction volumes revealed significant cost savings for samples undergoing WGS. Overall, the results presented here verify the accuracy and reproducibility of one-step targeted amplification and offer an efficient and cost-effective workflow for routine SARS-CoV-2 genomic surveillance.

Humans↗