Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Direct cloning and sequence analysis of enzymatically amplified genomic sequences.

A method is described for directly cloning enzymatically amplified segments of genomic DNA into an M13 vector for sequence analysis. A 110-base pair fragment of the human beta-globin gene and a 242-base pair fragment of the human leukocyte antigen DQ alpha locus were amplified by the polymerase chain reaction method, a procedure based on repeated cycles of denaturation, primer annealing, and extension by DNA polymerase I. Oligonucleotide primers with restriction endonuclease sites added to their 5' ends were used to facilitate the cloning of the amplified DNA. The analysis of cloned products allowed the quantitative evaluation of the amplification method's specificity and fidelity. Given the low frequency of sequence errors observed, this approach promises to be a rapid method for obtaining reliable genomic sequences from nanogram amounts of DNA.

Base Sequence↗

Four basic symmetry types in the universal 7-cluster structure of microbial genomic sequences.

Coding information is the main source of heterogeneity (non-randomness) in the sequences of microbial genomes. The heterogeneity corresponds to a cluster structure in triplet distributions of relatively short genomic fragments (200-400 bp). We found a universal 7-cluster structure in microbial genomic sequences and explained its properties. We show that codon usage of bacterial genomes is a multi-linear function of their genomic G+C-content with high accuracy. Based on the analysis of 143 completely sequenced bacterial genomes available in Genbank in August 2004, we show that there are four "pure" types of the 7-cluster structure observed. All 143 cluster animated 3D-scatters are collected in a database which is made available on our web-site (http://www.ihes.fr/~zinovyev/7clusters). The findings can be readily introduced into software for gene prediction, sequence alignment or microbial genomes classification.

Codon↗

Insights in metabolism and toxin production from the complete genome sequence of Clostridium tetani.

The decryption of prokaryotic genome sequences progresses rapidly and provides the scientific community with an enormous amount of information. Clostridial genome sequencing projects have been finished only recently, starting with the genome of the solvent-producing Clostridium acetobutylicum in 2001. A lot of attention has been devoted to the genomes of pathogenic clostridia. In 2002, the genome sequence of C. perfringens, the causative agent of gas gangrene, has been released. Currently in the finishing stage and prior to publication are the genomes of the foodborne botulism-causing C. botulinum and of C. difficile, the causative agent of a wide spectrum of clinical manifestations such as antibiotic-associated diarrhea. Our team sequenced the genome of neuropathogenic C. tetani, a Gram-positive spore-forming bacterium predominantly found in the soil. In deep wound infections it occasionally causes spastic paralysis in humans and vertebrate animals, known as tetanus disease, by the secretion of potent neurotoxin, designated tetanus toxin. The toxin blocks the release of neurotransmitters from presynaptic membranes of interneurons of the spinal cord and the brainstem, thus preventing muscle relaxation. Fortunately, this disease is successfully controlled through immunization with tetanus toxoid, a formaldehyde-treated tetanus toxin, but nevertheless, an estimated 400,000 cases still occur each year, mainly of neonatal tetanus. The World Health Organization has stated that neonatal tetanus is the second leading cause of death from vaccine preventable diseases among children worldwide. This minireview focuses on an analysis of the genome sequence of C. tetani E88, a vaccine production strain, which is a toxigenic non-sporulating variant of strain Massachusetts. The genome consists of a 2,799,250 bp chromosome encoding 2618 open reading frames. The tetanus toxin is encoded on a 74,082 kb plasmid, containing 61 genes. Additional virulence-related factors as well as an insight into the metabolic strategy of C. tetani with regard to its pathogenic phenotype will be presented. The information from other clostridial genomes by means of comparative analysis will also be explored.

Journal Article↗

RiceGAAS: an automated annotation system and database for rice genome sequence.

An extensive effort of the International Rice Genome Sequencing Project (IRGSP) has resulted in rapid accumulation of genome sequence, and >137 Mb has already been made available to the public domain as of August 2001. This requires a high-throughput annotation scheme to extract biologically useful and timely information from the sequence data on a regular basis. A new automated annotation system and database called Rice Genome Automated Annotation System (RiceGAAS) has been developed to execute a reliable and up-to-date analysis of the genome sequence as well as to store and retrieve the results of annotation. The system has the following functional features: (i) collection of rice genome sequences from GenBank; (ii) execution of gene prediction and homology search programs; (iii) integration of results from various analyses and automatic interpretation of coding regions; (iv) re-execution of analysis, integration and automatic interpretation with the latest entries in reference databases; (v) integrated visualization of the stored data using web-based graphical view. RiceGAAS also has a data submission mechanism that allows public users to perform fully automated annotation of their own sequences. The system can be accessed at http://RiceGAAS.dna.affrc.go.jp/.

Automation↗

THGS: a web-based database of Transmembrane Helices in Genome Sequences.

Transmembrane Helices in Genome Sequences (THGS) is an interactive web-based database, developed to search the transmembrane helices in the user-interested gene sequences available in the Genome Database (GDB). The proposed database has provision to search sequence motifs in transmembrane and globular proteins. In addition, the motif can be searched in the other sequence databases (Swiss-Prot and PIR) or in the macromolecular structure database, Protein Data Bank (PDB). Further, the 3D structure of the corresponding queried motif, if it is available in the solved protein structures deposited in the Protein Data Bank, can also be visualized using the widely used graphics package RASMOL. All the sequence databases used in the present work are updated frequently and hence the results produced are up to date. The database THGS is freely available via the world wide web and can be accessed at http:// pranag.physics.iisc.ernet.in/thgs/ or http://144.16. 71.10/thgs/.

Animals↗

Densities, length proportions, and other distributional features of repetitive sequences in the human genome estimated from 430 megabases of genomic sequence.

The densities of repetitive elements in the human genome were calculated in each GC content class using non-overlapping windows of 50kb. The density of Alu is two to three times higher in GC-rich regions than in AT-rich regions, while the opposite is true for LINE1. In contrast, LINE2 and other elements, such as DNA transposons, are more uniformly distributed in the genome. The number of Alus in the human genome was estimated to be 1.4 million, higher than previous estimates. About 40% of the autosomes and approximately 51% of the X and Y chromosomes are occupied by repetitive elements. In total, the human genome is estimated to contain more than 4 million repetitive elements. The GC contents (%) of repetitive elements and their flanking regions were also calculated. The GC contents of almost all kinds of repeats are positively correlated with the window GC contents, suggesting that a repetitive sequence is subject to the same mutation pressure as its surrounding regions, so it tends to have the same GC content as its surrounding regions. This observation supports the regional mutation hypothesis. The only two exceptions are AluYa and AluYb8, the two youngest Alu subfamilies. The GC content of AluYb8 is negatively correlated with that of its surrounding regions, while AluYa shows no correlation, suggesting different insertion patterns for these two young Alu subfamilies. This suggestion was supported by the fact that the average genetic distance between members of AluYb8 in each GC window class is positively correlated with the GC content of the window, but no correlation was found for AluYa. AluYa is more frequent in Y chromosome than in other chromosomes; the same is true for LTR retroviruses. This pattern might be correlated with the evolutionary history of Y chromosome.

Alu Elements↗

The psychrophilic lifestyle as revealed by the genome sequence of Colwellia psychrerythraea 34H through genomic and proteomic analyses.

The completion of the 5,373,180-bp genome sequence of the marine psychrophilic bacterium Colwellia psychrerythraea 34H, a model for the study of life in permanently cold environments, reveals capabilities important to carbon and nutrient cycling, bioremediation, production of secondary metabolites, and cold-adapted enzymes. From a genomic perspective, cold adaptation is suggested in several broad categories involving changes to the cell membrane fluidity, uptake and synthesis of compounds conferring cryotolerance, and strategies to overcome temperature-dependent barriers to carbon uptake. Modeling of three-dimensional protein homology from bacteria representing a range of optimal growth temperatures suggests changes to proteome composition that may enhance enzyme effectiveness at low temperatures. Comparative genome analyses suggest that the psychrophilic lifestyle is most likely conferred not by a unique set of genes but by a collection of synergistic changes in overall genome content and amino acid composition.

Amino Acids↗

Complete genome sequence of Caulobacter crescentus.

The complete genome sequence of Caulobacter crescentus was determined to be 4,016,942 base pairs in a single circular chromosome encoding 3,767 genes. This organism, which grows in a dilute aquatic environment, coordinates the cell division cycle and multiple cell differentiation events. With the annotated genome sequence, a full description of the genetic network that controls bacterial differentiation, cell growth, and cell cycle progression is within reach. Two-component signal transduction proteins are known to play a significant role in cell cycle progression. Genome analysis revealed that the C. crescentus genome encodes a significantly higher number of these signaling proteins (105) than any bacterial genome sequenced thus far. Another regulatory mechanism involved in cell cycle progression is DNA methylation. The occurrence of the recognition sequence for an essential DNA methylating enzyme that is required for cell cycle regulation is severely limited and shows a bias to intergenic regions. The genome contains multiple clusters of genes encoding proteins essential for survival in a nutrient poor habitat. Included are those involved in chemotaxis, outer membrane channel function, degradation of aromatic ring compounds, and the breakdown of plant-derived carbon sources, in addition to many extracytoplasmic function sigma factors, providing the organism with the ability to respond to a wide range of environmental fluctuations. C. crescentus is, to our knowledge, the first free-living alpha-class proteobacterium to be sequenced and will serve as a foundation for exploring the biology of this group of bacteria, which includes the obligate endosymbiont and human pathogen Rickettsia prowazekii, the plant pathogen Agrobacterium tumefaciens, and the bovine and human pathogen Brucella abortus.

Adaptation, Biological↗

Costs and cost-effectiveness of returning secondary findings from genomic sequencing based on the return of additional findings in the 100,000 Genomes Project.

PURPOSE: To assess costs and cost-effectiveness of returning additional findings from genome sequencing using data from the 100,000 Genomes Project (100kGP). METHODS: A model-based cost-utility analysis combining yield, consent rates, and cost data from the 100kGP with published estimates of downstream costs and quality-adjusted life years expected to accrue over a lifetime, after the identification of a pathogenic variant. RESULTS: The cost of returning additional findings to participants in the 100kGP was £7.1m or £81 per participant, with a yield of 0.85% for consented participants. The estimated lifetime incremental cost per participant was £125 and quality-adjusted life years 0.004, giving an incremental cost-effectiveness ratio of £28,830. Implementing a policy of returning additional findings is unlikely to be cost-effective (ie, 13%) at a willingness-to-pay threshold of £20,000. A short-term cost of returning findings of £43 per participant or lower (compared with the base case of £81) would result in an incremental cost-effectiveness ratio of less than £20,000. Alternatively, cost-effectiveness may be improved by returning additional findings to younger patient populations. CONCLUSION: Return of additional findings following genome sequencing for this group of conditions may not be a cost-effective use of health care system resources. Our cost-effectiveness outcomes rely on published estimates and should be validated through long-term follow-up data.

Humans↗

Expressed sequence tags: alternative or complement to whole genome sequences?

Over three million sequences from approximately 200 plant species have been deposited in the publicly available plant expressed sequence tag (EST) sequence databases. Many of the ESTs have been sequenced as an alternative to complete genome sequencing or as a substrate for cDNA array-based expression analyses. This creates a formidable resource from both biodiversity and gene-discovery standpoints. Bioinformatics-based sequence analysis tools have extended the scope of EST analysis into the fields of proteomics, marker development and genome annotation. Although EST collections are certainly no substitute for a whole genome scaffold, this "poor man's genome" resource forms the core foundations for various genome-scale experiments within the as yet unsequenceable plant genomes.

Computational Biology↗

A tiled amplicon protocol for culture-free whole-genome sequencing of M. tuberculosis from clinical specimens.

Whole-genome sequencing of Mycobacterium tuberculosis can be a valuable tool for TB surveillance and treatment, providing insights into transmission patterns and comprehensive drug susceptibility testing. However, the slow growth of M. tuberculosis means traditional culture-based sequencing methods can take weeks to return results, which has limited the widespread adoption of these techniques and limited their use in clinical decision-making. Tiled amplicon sequencing is a fast, reliable, and cost-effective method of whole-genome sequencing that can be done directly on clinical specimens and has been implemented at scale in academic and public health laboratories across the world; it was the cornerstone of SARS-CoV-2 sequencing and has been adapted for a wide range of viral pathogens. However, similar methods are not yet available for far larger bacterial genomes. Extending this approach to M. tuberculosis would significantly reduce the cost, labor, and turnaround time for whole-genome sequencing. We designed a tiled amplicon panel consisting of 5,128 primers that covers the entire M. tuberculosis genome, the largest tiled amplicon sequencing panel we are aware of to date. Applying our amplicon panels to clinical samples of sputum, we show the ability to recover whole-genome bacterial sequences without the need for culture. The resulting sequence data can be used to determine M. tuberculosis lineage and reliably identify markers of drug resistance. Using this approach in clinical settings could reduce the time needed for comprehensive drug susceptibility testing from weeks to days and enable genomic epidemiology to be performed at scale, even in resource-limited settings.IMPORTANCEWe have developed and tested an amplicon panel, TB-seq, for the priority pathogen Mycobacterium tuberculosis, demonstrating recovery of near-full genomes directly from patient sputum, including mixed and low-concentration samples. This approach significantly reduces the turnaround time for this slow-growing bacterium while maintaining high accuracy in detecting clinically relevant mutations, including those associated with drug resistance. Given the global burden of tuberculosis and the critical need for faster diagnostic solutions, we believe our method has the potential to improve clinical decision-making and public health strategies.

Mycobacterium tuberculosis↗

MGAlignIt: A web service for the alignment of mRNA/EST and genomic sequences.

Splicing is a biological phenomenon that removes the non-coding sequence from the transcripts to produce a mature transcript suitable for translation. To study this phenomenon, information on the intron-exon arrangement of a gene is essential, usually obtained by aligning mRNA/EST sequences to their cognate genomic sequences. MGAlign is a novel, rapid, memory efficient and practical method for aligning mRNA/EST and genome sequences. We present here a freely available web service, MGAlignIt (http://origin.bic.nus.edu.sg/mgalign/mgalignit), based on MGAlign. Besides the alignment itself, this web service allows users to effectively visualize the alignment in a graphical manner and to perform limited analysis on the alignment output. The server also permits the alignment to be saved in several forms, both graphical and text, suitable for further processing and analysis by other programs.

Algorithms↗

Exploiting the complete yeast genome sequence.

The completion of the genome sequence of the budding yeast Saccharomyces cerevisiae marks the dawn of an exciting new era in eukaryotic biology that will bring with it a new understanding of yeast, other model organisms, and human beings. This body of sequence data benefits yeast researchers by obviating the need for piecemeal sequencing of genes, and allows researchers working with other organisms to tap into experimental advantages inherent in the yeast system and learn from functionally characterized yeast gene products which are their proteins of interest. In addition, the yeast post-genome sequence era is serving as a testing ground for powerful new technologies, and proven experimental approaches are being applied for the first time in a comprehensive fashion on a complete eukaryotic gene repertoire.

Animals↗

A unified benchmark of supervised and retrieval-based methods for viral genomic sequence classification.

The rapid growth of genomic sequencing demands fast, accurate, and scalable analysis methods. In viral genomic classification, expanding labeled reference collections can make supervised models costly to update and dependent on fixed label sets, motivating retrieval-based genomic classification as a simpler, more flexible alternative. We present a unified benchmark of supervised and retrieval-based methods for viral genomic sequence classification across three viral classification tasks: hepatitis C virus (HCV) genotyping, COVID-19 discrimination, and human papillomavirus (HPV) genotyping. We compare standard sequence encodings (one-hot, k-mers, FCGR) with dense embeddings (dna2vec, DNABERT). For each representation, we evaluate supervised classifiers (Random Forest, Decision Tree, XGBoost) and retrieval-based classification, where sequence vectors are indexed with FAISS and labels are assigned via similarity-weighted k-NN. Furthermore, we benchmark multiple FAISS index types (Flat, IVF, HNSW, IVFPQ, OPQ) to characterize accuracy-speed-memory trade-offs at scale. The results show that XGBoost and retrieval using Flat or IVF indexes achieve strong classification performance under different computational profiles. Compressed indexes such as IVFPQ and OPQ substantially reduce memory usage, although their accuracy loss depends on the dataset and representation. Overall, supervised XGBoost provides a favorable accuracy-size trade-off, while retrieval-based classification remains competitive and allows labeled reference sequences to be incorporated without retraining a global classifier. This benchmark provides practical guidance for selecting sequence representations, classifiers, and vector-search indexes under different accuracy, memory, and update requirements.

Genome, Viral↗

ZIPcnv: accurate and efficient inference of copy number variations from shallow whole-genome sequencing.

MOTIVATION: Shallow whole-genome sequencing (sWGS), a rapid and cost-effective sequencing technology, has gradually been widely adopted for CNV analyses. However, with genome‑wide coverage of only 0.1-5×, sWGS data display a pronounced zero‑inflation phenomenon-a large fraction of loci has zero sequencing reads. Zero inflation causes read counts to fluctuate by several‑fold between adjacent windows. As a result, random upward blips in coverage can be misinterpreted as copy‑number gains (false positives), and true deletions often become indistinguishable from pervasive zero‑coverage noise. In addition, existing CNV detection tools developed for sWGS data often struggle to adapt across different CNV sizes. These combined effects severely constrain the accuracy of CNV inference. RESULTS: To address above challenges, we propose ZIPcnv, a novel CNV detection tool specifically designed for sWGS data. First, we apply a segment sliding window to smooth the raw read depth signal, which transforms the original zero-inflated statistical characteristics into approximately normal distribution characteristics. We then design a statistical process model that robustly detects persistent shifts under high background noise using a cumulative sum strategy, classifying genomic regions into candidate and non-candidate CNV regions. Finally, dynamic sliding windows are used for one-pass detection of CNVs of varying lengths, with window size adapting to the CNV region size. We evaluated the performance of ZIPcnv on simulated data and 190 real whole-genome sequencing samples. Experimental results show that ZIPcnv consistently outperforms currently popular CNV detection tools. AVAILABILITY AND IMPLEMENTATION: The ZIPcnv source code is freely available at https://github.com/Nevermore233/ZIPcnv.

DNA Copy Number Variations↗

The complete genome sequence of the lactic acid bacterium Lactococcus lactis ssp. lactis IL1403.

Lactococcus lactis is a nonpathogenic AT-rich gram-positive bacterium closely related to the genus Streptococcus and is the most commonly used cheese starter. It is also the best-characterized lactic acid bacterium. We sequenced the genome of the laboratory strain IL1403, using a novel two-step strategy that comprises diagnostic sequencing of the entire genome and a shotgun polishing step. The genome contains 2,365,589 base pairs and encodes 2310 proteins, including 293 protein-coding genes belonging to six prophages and 43 insertion sequence (IS) elements. Nonrandom distribution of IS elements indicates that the chromosome of the sequenced strain may be a product of recent recombination between two closely related genomes. A complete set of late competence genes is present, indicating the ability of L. lactis to undergo DNA transformation. Genomic sequence revealed new possibilities for fermentation pathways and for aerobic respiration. It also indicated a horizontal transfer of genetic information from Lactococcus to gram-negative enteric bacteria of Salmonella-Escherichia group.

Amino Acids↗