Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

Genome sequence comparison and scenarios for gene rearrangements: a test case.

As large portions of related genomes are being sequenced, methods for comparing complete or nearly complete genomes, as opposed to comparing individual genes, are becoming progressively more important. A major, widespread phenomenon in genome evolution is the rearrangement of genes and gene blocks. There is, however, no consistent method for genome sequence comparison combined with the reconstruction of the evolutionary history of highly rearranged genomes. We developed a schema for genome sequence comparison that includes three successive steps: (i) comparison of all proteins encoded in different genomes and generation of genomic similarity plots; (ii) construction of an alphabet of conserved genes and gene blocks; and (iii) generation of most parsimonious genome rearrangement scenarios. The approach is illustrated by a comparison of the herpesvirus genomes that constitute the largest set of relatively long, complete genome sequences available to date. Herpesviruses have from 70 to about 200 genes; comparison of the amino acid sequences encoded in these genes results in an alphabet of about 30 conserved genes comprising 7 conserved blocks that are rearranged in the genomes of different herpesviruses. Algorithms to analyze rearrangements of multiple genomes were developed and applied to the derivation of most parsimonious scenarios of herpesvirus evolution under different evolutionary models. The developed approaches to genome comparison will be applicable to the comparative analysis of bacterial and eukaryotic genomes as soon as their sequences become available.

DNA, Viral↗

The first filamentous fungal genome sequences: Aspergillus leads the way for essential everyday resources or dusty museum specimens?

The published Aspergillus genome sequences (A. nidulans, A. fumigatus, A. oryzae) and further sequence data from A. clavatus, Neosartorya fischeri, A. flavus, A. niger, A. parasiticus and A. terreus are the first from a group of related filamentous fungi. They indicate the gains possible from genomic approaches, but also problems that arise after the sequences are finished. Benefits include a greater understanding of genome structure and evolution, insights into gene regulation, predictions of new factors that may be relevant to pathogenicity and the discovery of novel enzymes with biotechnological value. Areas where further developments are needed include gene and structure-function predictions, methods for comparative genome analysis and the interfaces for access to genome information. In addition, strategies for continued maintenance and updating need to be developed at the start of the post-genomic era to increase the value of genome sequences into the future.

Aspergillus↗

The first complete genome sequence of Ammi majus latent virus from the new natural host culantro.

A potyvirus (isolate AMLV-CQ) infecting culantro (Eryngium foetidum L.) imported from Vietnam was identified by RT-PCR. The complete genome sequence of AMLV-CQ was determined to be 9,549 nucleotides in length. It contains a large open reading frame encoding a 3,082-amino-acid putative polyprotein, flanked by 5´ and 3´ untranslated regions (UTRs) of 77 and 226 nt, respectively. AMLV-CQ is closely related to five other completely sequenced potyviruses, sharing 68-69% nucleotide and 69-70% amino acid sequence identity. However, the coat protein (CP) gene shares 89% nucleotide and 93% amino acid sequence identity with that of a partially sequenced potyvirus, Ammi majus latent virus (isolate AMLV-WF17). These results suggest that AMLV-CQ and AMLV-WF17 are isolates of the same species. To our knowledge, this is the first report of a complete genome sequence of an AMLV isolate, and culantro was identified as a new natural host for this virus. In addition, a one-step RT-PCR assay was developed that provides a rapid, robust, and highly sensitive approach for the detection of AMLV.

Eryngium↗

The complete genomic sequence of Nocardia farcinica IFM 10152.

We determined the genomic sequence of Nocardia farcinica IFM 10152, a clinical isolate, and revealed the molecular basis of its versatility. The genome consists of a single circular chromosome of 6,021,225 bp with an average G+C content of 70.8% and two plasmids of 184,027 (pNF1) and 87,093 (pNF2) bp with average G+C contents of 67.2% and 68.4%, respectively. The chromosome encoded 5,674 putative protein-coding sequences, including many candidate genes for virulence and multidrug resistance as well as secondary metabolism. Analyses of paralogous protein families suggest that gene duplications have resulted in a bacterium that can survive not only in soil environments but also in animal tissues, resulting in disease.

Amino Acid Sequence↗

What we can learn from the Mycobacterium tuberculosis genome sequencing projects.

A major milestone in tuberculosis research occurred in June 1998 with the report of the genomic sequence of Mycobacterium tuberculosis H37Rv. The complete determination of the 4411529 base pairs of the M. tuberculosis genome opens avenues for new scientific opportunities in basic science and clinical research on tuberculosis. In this paper we will review the findings presented by the complete genome and discuss the impact that this sequence will have on the areas of comparative genomics, bacterial pathogenesis, and diagnostics development for tuberculosis. Indirect benefits of the complete genome sequence are anticipated in the areas of drug development and vaccine development as future discoveries in bacterial pathogenesis and immunology accrue.

DNA, Bacterial↗

Comparison of outer membrane protein genes omp and pmp in the whole genome sequences of Chlamydia pneumoniae isolates from Japan and the United States.

Chlamydia pneumoniae is a widespread pathogen of the respiratory tract that is also associated with atherosclerosis. The whole genome sequence was determined for a Japanese isolate, C. pneumoniae strain J138. The sequence predicted a variety of genes encoding outer membrane proteins (OMPs) including ompA and porB, another 10 predicted omp genes, and 27 pmp genes. All were detected in the whole genome sequence of strain CWL029, a strain isolated and sequenced in the United States. A comparative study of the OMPs of the two strains revealed a nucleotide sequence identity of 89.6%-100% (deduced amino acid sequence identity, 71.1%-100%). The overall genomic organization and location of genes are identical in both strains. Thus, a few unique sequences of the OMPs may be essential for specific attributes that define the differential biology of two C. pneumoniae strains.

Bacterial Outer Membrane Proteins↗

The complete Mokola virus genome sequence: structure of the RNA-dependent RNA polymerase.

The genome sequence of the rabies-related virus Mokola virus (genus Lyssavirus) has been completed by sequencing the L gene, which consists of 6384 nucleotides encoding a 2127 amino acid polymerase. Alignment of the Mokola virus L protein with other polymerases from the virus order Mononegavirales defined three domains: a divergent NH2-terminal domain, a highly conserved central domain carrying most of the functional motifs and a COOH-terminal domain with alternating conserved and divergent regions. A statistical study outlined the stringency of conservation of glycine, acidic (D, E) and basic (K, R, H) amino acids in polymerases, particularly as key residues of the conserved motifs.

Amino Acid Sequence↗

Association of common and rare variants with Alzheimer's disease in more than 13,000 diverse individuals with whole-genome sequencing from the Alzheimer's Disease Sequencing Project.

INTRODUCTION: Alzheimer's disease (AD) is a common disorder of the elderly that is both highly heritable and genetically heterogeneous. METHODS: We investigated the association of AD with both common variants and aggregates of rare coding and non-coding variants in 13,371 individuals of diverse ancestry with whole genome sequencing (WGS) data. RESULTS: Pooled-population analyses of all individuals identified genetic variants at apolipoprotein E (APOE) and BIN1 associated with AD (p&#xa0;<&#xa0;5&#xa0;&#xd7;&#xa0;10-8). Subgroup-specific analyses identified a haplotype on chromosome 14 including PSEN1 associated with AD in Hispanics, further supported by aggregate testing of rare coding and non-coding variants in the region. Common variants in LINC00320 were observed associated with AD in Black individuals (p&#xa0;=&#xa0;1.9&#xa0;&#xd7;&#xa0;10-9). Finally, we observed rare non-coding variants in the promoter of TOMM40 distinct of APOE in pooled-population analyses (p&#xa0;=&#xa0;7.2&#xa0;&#xd7;&#xa0;10-8). DISCUSSION: We observed that complementary pooled-population and subgroup-specific analyses offered unique insights into the genetic architecture of AD. HIGHLIGHTS: We determine the association of genetic variants with Alzheimer's disease (AD) using 13,371 individuals of diverse ancestry with whole genome sequencing (WGS) data. We identified genetic variants at apolipoprotein E (APOE), BIN1, PSEN1, and LINC00320 associated with AD. We observed rare non-coding variants in the promoter of TOMM40 distinct of APOE.

Humans↗

Extracting phylogenetic information from whole-genome sequencing projects: the lactic acid bacteria as a test case.

The availability of an ever increasing number of complete genome sequences of diverse prokaryotic taxa has led to the introduction of novel approaches to infer phylogenetic relationships among bacteria. In the present study the sequences of the 16S rRNA gene and nine housekeeping genes were compared with the fraction of shared putative orthologous protein-encoding genes, conservation of gene order, dinucleotide relative abundance and codon usage among 11 genomes of species belonging to the lactic acid bacteria. In general there is a good correlation between the results obtained with various approaches, although it is clear that there is a stronger phylogenetic signal in some data sets than in others, and that different parameters have different taxonomic resolutions. It appears that trees based on different kinds of information derived from whole-genome sequencing projects do not provide much additional information about the phylogenetic relationships among bacterial taxa compared to more traditional alignment-based methods. Nevertheless, it is expected that the study of these novel forms of information will have its value in taxonomy, to determine which genes are shared, when genes or sets of genes were lost in evolutionary history, to detect the presence of horizontally transferred genes and/or confirm or enhance the phylogenetic signal derived from traditional methods. Although these conclusions are based on a relatively small data set, they are largely in agreement with other studies and it is anticipated that similar trends will be observed when comparing other genomes.

Bacterial Proteins↗

Determination of the recognition sequence of Mycobacterium smegmatis topoisomerase I on mycobacterial genomic sequences.

Mycobacterium smegmatis topoisomerase I has several distinctive features. The absence of the zinc finger motif found in other prokaryotic type I topoisomerases and the ability of the enzyme to recognise single-stranded and duplex DNA are unique characteristics of the enzyme. We have mapped the strong topoisomerase sites of the enzyme on genomic DNA sequences from Mycobacterium tuberculosis and M.smegmatis. The enzyme does not nick DNA in random fashion and DNA cleavage occurred at a few specific sites. Mapping of these sites revealed conservation of a pentanucleotide motif CG/TCT/T at the cleavage site (/ represents the cleavage site). The enzyme binds and cleaves consensus oligo-nucleotides having this sequence motif. The protein exhibits a very high preference for C or a G residue at the +2 position with respect to the cleavage site. Based on earlier and the present studies we propose that the enzyme functions in vivo mainly at these specific sites to carry out topological reactions.

Base Sequence↗

Complete genome sequence of the methanogenic archaeon, Methanococcus jannaschii.

The complete 1.66-megabase pair genome sequence of an autotrophic archaeon, Methanococcus jannaschii, and its 58- and 16-kilobase pair extrachromosomal elements have been determined by whole-genome random sequencing. A total of 1738 predicted protein-coding genes were identified; however, only a minority of these (38 percent) could be assigned a putative cellular role with high confidence. Although the majority of genes related to energy production, cell division, and metabolism in M. jannaschii are most similar to those found in Bacteria, most of the genes involved in transcription, translation, and replication in M. jannaschii are more similar to those found in Eukaryotes.

Amino Acid Sequence↗

Impact of human genome sequencing for in silico target discovery.

The year 2000 stands as a landmark in modern biology: the first draft of the human genome sequence has been completed. For the pharmaceutical industry, this achievement provides tremendous opportunities because the genomic sequence exposes all human drug targets for therapeutic intervention. The challenge for the pharmaceutical companies is to exploit this definitive resource for the identification of potential molecular targets, rapid characterization of their function and validation of their involvement in disease pathology. Bioinformatics approaches provide increasingly crucial tools to systematically support this exploratory target drug discovery activity.

Journal Article↗

Genome structure of the Lactobacillus temperate phage phi g1e: the whole genome sequence and the putative promoter/repressor system.

The complete genome sequence of a Lactobacillus temperate phage phi g1e was established. The double-stranded DNA is composed of 42,259 bp, and encodes for sixty-two possible open reading frames (ORF) as well as several potential regulatory sequences. Based on comparative analysis with other related proteins of the Lactobacillus and Lactococcus phages as well as the Escherichia coli phages (such as lambda), functions were putatively assigned to several phi g1e ORFs: cng and cpg (encoding for repressors), hel (helicase), ntp (NTPase), and several ORFs (e.g., minor capsid proteins). An about 1000-bp DNA region of phi g1e containing cpg and cng was inferred to function as a promoter/repressor system for the phi g1e lysogenic and lytic pathway.

Amino Acid Sequence↗

Coding-complete genome sequence of grapevine leafroll-associated virus 13 from grapevine in California.

In this study, we report the coding-complete genome sequence of Grapevine leafroll-associated virus 13 (GLRaV-13), isolate CA8881, detected in Vitis vinifera in California, USA. The genome sequence exhibited over 95% nucleotide identity with previously reported GLRaV-13 isolates and contributed to better understanding of the genetic diversity of ampeloviruses infecting grapevine.

California↗

Kelp fly virus: a novel group of insect picorna-like viruses as defined by genome sequence analysis and a distinctive virion structure.

The complete genomic sequence of kelp fly virus (KFV), originally isolated from the kelp fly, Chaetocoelopa sydneyensis, has been determined. Analyses of its genomic and structural organization and phylogeny show that it belongs to a hitherto undescribed group within the picorna-like virus superfamily. The single-stranded genomic RNA of KFV is 11,035 nucleotides in length and contains a single large open reading frame encoding a polypeptide of 3,436 amino acids with 5' and 3' untranslated regions of 384 and 343 nucleotides, respectively. The predicted amino acid sequence of the polypeptide shows that it has three regions. The N-terminal region contains sequences homologous to the baculoviral inhibitor of apoptosis repeat domain, an inhibitor of apoptosis commonly found in animals and in viruses with double-stranded DNA genomes. The second region contains at least two capsid proteins. The third region has three sequence motifs characteristic of replicase proteins of many plant and animal viruses, including a helicase, a 3C chymotrypsin-like protease, and an RNA-dependent RNA polymerase. Phylogenetic analysis of the replicase motifs shows that KFV forms a distinct and distant taxon within the picorna-like virus superfamily. Cryoelectron microscopy and image reconstruction of KFV to a resolution of 15 A reveals an icosahedral structure, with each of its 12 fivefold vertices forming a turret from the otherwise smooth surface of the 20-A-thick capsid. The architecture of the KFV capsid is unique among the members of the picornavirus superfamily for which structures have previously been determined.

Amino Acid Sequence↗

A large-scale comparison of genomic sequences: one promising approach.

We introduce a novel, linguistic-like method of genome analysis. We propose a natural approach to characterizing genomic sequences based on occurrences of fixed length words from a predefined, sufficiently large set of words (strings over the alphabet [A, C, G, T]). A measure based on this approach is called compositional spectrum and is actually a histogram of imperfect word occurrences. Our results assert that the compositional spectrum is an overall characteristic of a long sequence i.e., a complete genome or an uninterrupted part of a chromosome. This attribute is manifested in the similarity of spectra obtained on different stretches of the same genome, and simultaneously in a broad range of dissimilarities between spectral representations of different genomes. High flexibility characterizes this approach due to imperfect matching and as a result sets of relatively long words can be considered. The proposed approach may have various applications in intra- and intergenomic sequence comparisons.

Algorithms↗

Simultaneous amplification and detection of specific hepatitis B virus and hepatitis C virus genomic sequences in serum samples.

A sensitive and specific two-stage polymerase chain reaction (PCR) technique was developed for the simultaneous amplification and detection of specific genomic sequences of hepatitis B virus (HBV) and hepatitis C virus (HCV) in serum samples. Initially, HCV-RNA was reverse transcribed to cDNA. This cDNA and DNA from HBV were then co-amplified using primer pairs derived from conserved regions of HBV and HCV nucleotide sequences. The specificity of PCR products was confirmed by liquid hybridization analysis using 32P end-labeled oligomer probes specific for the target HBV and HCV nucleotide sequences. Independent human serum samples, positive and negative by PCR for both HBV-DNA and HCV-RNA, were used as controls. We tested sera from nine donors, of which seven were reactive for HBsAg, anti-HBc, and anti-HCV (multiantigen test), one of whom was reactive for anti-HCV and anti-HBc, and one of whom was reactive for HBsAg and anti-HBc. The assay detected HBV- and HCV-specific genomic sequences in eight of eight sera reactive for both HBV and HCV serological markers and also in the serum that was reactive for HBV markers only.

Base Sequence↗

Comparative genome sequencing for discovery of novel polymorphisms in Bacillus anthracis.

Comparison of the whole-genome sequence of Bacillus anthracis isolated from a victim of a recent bioterrorist anthrax attack with a reference reveals 60 new markers that include single nucleotide polymorphisms (SNPs), inserted or deleted sequences, and tandem repeats. Genome comparison detected four high-quality SNPs between the two sequenced B. anthracis chromosomes and seven differences among different preparations of the reference genome. These markers have been tested on a collection of anthrax isolates and were found to divide these samples into distinct families. These results demonstrate that genome-based analysis of microbial pathogens will provide a powerful new tool for investigation of infectious disease outbreaks.

Animals↗