Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “complete genome sequence”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Complete genome sequence of the shrimp white spot bacilliform virus.

We report the first complete genome sequence of a marine invertebrate virus. White spot bacilliform virus (WSBV; or white spot syndrome virus) is a major shrimp pathogen with a high mortality rate and a wide host range. Its double-stranded circular DNA genome of 305,107 bp contains 181 open reading frames (ORFs). Nine homologous regions containing 47 repeated minifragments that include direct repeats, atypical inverted repeat sequences, and imperfect palindromes were identified. This is the largest animal virus that has been completely sequenced. Although WSBV is morphologically similar to insect baculovirus, the two viruses are not detectably related at the amino acid level. Rather, some WSBV genes are more homologous to eukaryotic genes than viral genes. In fact, sequence analysis indicates that WSBV differs from all known viruses, although a few genes display a weak homology to herpesvirus genes. Most of the ORFs encode proteins that bear no homology to any known proteins, either suggesting that WSBV represents a novel class of viruses or perhaps implying a significant evolutionary distance between marine and terrestrial viruses. The most unique feature of WSBV is the presence of an intact collagen gene, a gene encoding an extracellular matrix protein of animal cells that has never been found in any viruses. Determination of the genome of WSBV will facilitate a better understanding of the molecular mechanism underlying the pathogenesis of the WSBV virus and will also provide useful information concerning the evolution and divergence of marine and terrestrial animal viruses at the molecular level.

Amino Acid Sequence↗

Transcript map and complete genomic sequence for the 310 kb region of minimal allele loss on chromosome segment 11p15.5 in non-small-cell lung cancer.

Molecular, functional, and clinical analyses strongly suggest that chromosome segment 11p15.5 contains a gene involved in lung cancer pathogenesis. The critical region of allele loss is 310 kb in size. We used our contig of P1-phage artificial chromosome (PAC) clones together with newly identified bacterial artificial chromosome (BAC) clones and the draft human genome sequence to complete a contiguous string of 380 407 bp. Three PAC clones that span the region were used to identify transcripts by exon trapping. Computational gene prediction algorithms were used to query the sequence for potential genes and exons. Screening for expression was performed with tissue-specific and cell line derived mRNA arrays. The region contains the complete SSA/Ro52 and RRM1 genes, exons 7-12 of the GOK gene, and the psirad pseudo-gene. A cluster of six nearly identical genes with an intact open reading frame (ORF) of 585 bp that share 75% identity with the HSPC182 gene was found. In addition, five putative novel genes were identified. Sequence tagged sites (STS) and polymorphic markers were used to screen 117 lung cancer cell lines for homozygous deletions and none were identified. These data provide the basis for the identification of a lung cancer suppressor gene on 11p15.5.

Base Sequence↗

The complete genome sequence of an El Amar isolate of plum pox virus (PPV) and its phylogenetic relationship to other PPV strains.

The genomic sequence of an El Amar isolate of plum pox virus (PPV) from Egypt was determined by sequencing overlapping cDNA fragments. This is the first complete sequence of a member of the El Amar (EA) strain of PPV. The genome consists of 9791 nt, excluding a poly(A) tail at the 3' terminus. The complete nt sequence of PPV EA is 79-80%, 80%, 77%, and 77% homologous with isolates of strains D/M, Rec (BOR3), C, and W, respectively. The polyprotein identity ranged from 87-91%. Phylogenetic analysis using the complete genome sequence of PPV EA confirmed its strain status. No significant recombination signals were identified using PhylPro and SimPlot scans of the PPV EA sequence, however an interesting recombination signal was identified in the P1/HC-Pro region of PPV W3174.

Base Sequence↗

The complete genome sequence of the Streptomyces temperate phage straight phiC31: evolutionary relationships to other viruses.

The completed genome sequence of the temperate Streptomyces phage straight phiC31 is reported. straight phiC31 contains genes that are related by sequence similarities to several other dsDNA phages infecting many diverse bacterial hosts, including Escherichia, Arthrobacter, Mycobacterium, Rhodobacter, Staphylococcus, Bacillus, Streptococcus, Lactobacillus and Lactococcus. These observations provide further evidence that dsDNA phages from diverse bacterial hosts are related and have had access to a common genetic pool. Analysis of the late genes was particularly informative. The sequences of the head assembly proteins (portal, head protease and major capsid) were conserved between straight phiC31, coliphage HK97, staphylococcal phage straight phiPVL, two Rhodobacter capsulatus prophages and two Mycobacterium tuberculosis prophages. These phages and prophages (where non-defective) from evolutionarily diverse hosts are, therefore, likely to share a common head assembly mechanism i.e. that of HK97. The organisation of the tail genes in straight phiC31 is highly reminiscent of tail regions from other phage genomes. The unusual organisation of the putative lysis genes in straight phiC31 is discussed, and speculations are made as to the roles of some inessential early gene products. Similarities between certain phage gene products and eukaryotic dsDNA virus proteins were noted, in particular, the primase/helicases and the terminases (large subunits). Furthermore, the complete sequence clarifies the overall transcription map of the phage during lytic growth and the positions of elements involved in the maintenance of lysogeny.

Amino Acid Sequence↗

Statistical properties of open reading frames in complete genome sequences.

Some statistical properties of open reading frames in all currently available complete genome sequences are analyzed (seventeen prokatyotic genomes, and 16 chromosome sequences from the yeast genome). The size distribution of open reading frames is characterized by various techniques, such as quantile tables, QQ-plots, rank-size plots (Zipf's plots), and spatial densities. The issue of the influence of CG% on the size distribution is addressed. When yeast chromosomes are compared with archaeal and eubacterial genomes, they tend to have more long open reading frames. There is little or no evidence to reject the null hypothesis that open reading frames on six different reading frames and two strands distribute similarly. A topic of current interest, the base composition asymmetry in open reading frames between the two strands, is studied using regression analysis. The base composition asymmetry at three codon positions is analyzed separately. It was shown in these genome sequences that the first codon position is G- and A-rich (i.e. purine-rich); there is a co-existence of A- and T-rich branches at the second codon position; and the third codon position is weakly T-rich.

Base Sequence↗

Comparison of the complete genome sequences of Pseudomonas syringae pv. syringae B728a and pv. tomato DC3000.

The complete genomic sequence of Pseudomonas syringae pv. syringae B728a (Pss B728a) has been determined and is compared with that of P. syringae pv. tomato DC3000 (Pst DC3000). The two pathovars of this economically important species of plant pathogenic bacteria differ in host range and other interactions with plants, with Pss having a more pronounced epiphytic stage of growth and higher abiotic stress tolerance and Pst DC3000 having a more pronounced apoplastic growth habitat. The Pss B728a genome (6.1 Mb) contains a circular chromosome and no plasmid, whereas the Pst DC3000 genome is 6.5 mbp in size, composed of a circular chromosome and two plasmids. Although a high degree of similarity exists between the two sequenced Pseudomonads, 976 protein-encoding genes are unique to Pss B728a when compared with Pst DC3000, including large genomic islands likely to contribute to virulence and host specificity. Over 375 repetitive extragenic palindromic sequences unique to Pss B728a when compared with Pst DC3000 are widely distributed throughout the chromosome except in 14 genomic islands, which generally had lower GC content than the genome as a whole. Content of the genomic islands varies, with one containing a prophage and another the plasmid pKLC102 of Pseudomonas aeruginosa PAO1. Among the 976 genes of Pss B728a with no counterpart in Pst DC3000 are those encoding for syringopeptin, syringomycin, indole acetic acid biosynthesis, arginine degradation, and production of ice nuclei. The genomic comparison suggests that several unique genes for Pss B728a such as ectoine synthase, DNA repair, and antibiotic production may contribute to the epiphytic fitness and stress tolerance of this organism.

Bacterial Proteins↗

Complete genome sequence of garlic latent virus, a member of the carlavirus family.

The complete genome sequence of the garlic latent virus (GLV) has been determined. The whole GLV genome consists of 8,353 nucleotides, excluding the 3'-end poly(A)+ tail, and contains six open-reading frames (ORFs). Putative proteins that were encoded by the reading frames contain the motifs that were conserved in carlavirus-specific RNA replicases, NTP-dependent DNA helicases, two viral membrane-bound proteins, a viral coat protein, and a zinc-finger. Overall, the GLV genome shows structural features that are common in carlaviruses. An in vitro translation analysis revealed that the zinc-finger protein is not produced as a transframe protein with the coat protein by ribosomal frameshifting. A Northern blot analysis showed that GLV-specific probes hybridized to garlic leaf RNA fragments of about 2.6 and 1.5 kb long, in addition to the 8.5 kb whole genome. The two subgenomic RNAs might be encapsidated into smaller viral particles. In garlic plants, 700 nm long flexuous rod-shaped virus particles were observed in the immunoelectron microscopy using polyclonal antibodies against the GLV coat proteins.

Amino Acid Sequence↗

The coding-complete genome sequence of Arabidopsis latent virus 1 from Iraq.

Here, we report the coding-complete genome sequence of Arabidopsis latent virus 1 (ArLV1; Comovirus arabidopsis) from Iraq. The virus was identified by metatranscriptomic sequencing of asymptomatic cucumber (Cucumis sativus) leaves harboring thrips collected from commercial greenhouses. The bipartite genome comprises RNA1 (5,553 nt) and RNA2 (3,582 nt).

Arabidopsis latent virus 1↗

Complete genomic sequence of the temperate bacteriophage PhiAT3 isolated from Lactobacillus casei ATCC 393.

The complete genomic sequence of a temperate bacteriophage PhiAT3 isolated from Lactobacillus (Lb.) casei ATCC 393 is reported. The phage consists of a linear DNA genome of 39,166 bp, an isometric head of 53 nm in diameter, and a flexible, noncontractile tail of approximately 200 nm in length. The number of potential open reading frames on the phage genome is 53. There are 15 unpaired nucleotides at both 5' ends of the PhiAT3 genome, indicating that the phage uses a cos-site for DNA packaging. The PhiAT3 genome was grouped into five distinct functional clusters: DNA packaging, morphogenesis, lysis, lysogenic/lytic switch, and replication. The amino acid sequences at the NH2-termini of some major proteins were determined. An in vivo integration assay for the PhiAT3 integrase (Int) protein in several lactobacilli was conducted by constructing an integration vector including PhiAT3 int and the attP (int-attP) region. It was found that PhiAT3 integrated at the tRNAArg gene locus of Lactobacillus rhamnosus HN 001, similar to that observed in its native host, Lb. casei ATCC 393.

Bacteriophages↗

Complete genome sequence of an aerobic hyper-thermophilic crenarchaeon, Aeropyrum pernix K1.

The complete sequence of the genome of an aerobic hyper-thermophilic crenarchaeon, Aeropyrum pernix K1, which optimally grows at 95 degrees C, has been determined by the whole genome shotgun method with some modifications. The entire length of the genome was 1,669,695 bp. The authenticity of the entire sequence was supported by restriction analysis of long PCR products, which were directly amplified from the genomic DNA. As the potential protein-coding regions, a total of 2,694 open reading frames (ORFs) were assigned. By similarity search against public databases, 633 (23.5%) of the ORFs were related to genes with putative function and 523 (19.4%) to the sequences registered but with unknown function. All the genes in the TCA cycle except for that of alpha-ketoglutarate dehydrogenase were included, and instead of the alpha-ketoglutarate dehydrogenase gene, the genes coding for the two subunits of 2-oxoacid:ferredoxin oxidoreductase were identified. The remaining 1,538 ORFs (57.1%) did not show any significant similarity to the sequences in the databases. Sequence comparison among the assigned ORFs suggested that a considerable member of ORFs were generated by sequence duplication. The RNA genes identified were a single 16S-23S rRNA operon, two 5S rRNA genes and 47 tRNA genes including 14 genes with intron structures. All the assigned ORFs and RNA coding regions occupied 89.12% of the whole genome. The data presented in this paper are available on the internet homepage (http://www.mild.nite.go.jp).

Archaea↗

Complete genome sequence of Acinetobacter bacterium strain BZX-2 isolated from hybrid sturgeon.

We report the complete genome sequence of Acinetobacter sp. strain BZX-2, isolated from a hybrid sturgeon (Acipenser baerii ♀ × Acipenser schrenckii ♂). The genome consists of a 3,819,278-bp chromosome with 3,566 predicted protein-coding genes. The genomic characteristics and antibiotic resistance genes identified provide a basis for the prevention and control of sturgeon diseases.

Acinetobacter↗

Use of the complete genome sequence information of Haemophilus influenzae strain Rd to investigate lipopolysaccharide biosynthesis.

The availability of the complete 1.83-megabase-pair sequence of the Haemophilus influenzae strain Rd genome has facilitated significant progress in investigating the biology of H.influenzae lipopolysaccharide (LPS), a major virulence determinant of this human pathogen. By searching the H. influenzae genomic database, with sequences of known LPS biosynthetic genes from other organisms, we identified and then cloned 25 candidate LPS genes. Construction of mutant strains and characterization of the LPS by reactivity with monoclonal antibodies, PAGE fractionation patterns and electrospray mass spectrometry comparative analysis have confirmed a potential role in LPS biosynthesis for the majority of these candidate genes. Virulence studies in the infant rat have allowed us to estimate the minimal LPS structure required for intravascular dissemination. This study is one of the first to demonstrate the rapidity, economy and completeness with which novel biological information can be accessed once the complete genome sequence of an organism is available.

Animals↗

The first complete genome sequence of Ammi majus latent virus from the new natural host culantro.

A potyvirus (isolate AMLV-CQ) infecting culantro (Eryngium foetidum L.) imported from Vietnam was identified by RT-PCR. The complete genome sequence of AMLV-CQ was determined to be 9,549 nucleotides in length. It contains a large open reading frame encoding a 3,082-amino-acid putative polyprotein, flanked by 5´ and 3´ untranslated regions (UTRs) of 77 and 226 nt, respectively. AMLV-CQ is closely related to five other completely sequenced potyviruses, sharing 68-69% nucleotide and 69-70% amino acid sequence identity. However, the coat protein (CP) gene shares 89% nucleotide and 93% amino acid sequence identity with that of a partially sequenced potyvirus, Ammi majus latent virus (isolate AMLV-WF17). These results suggest that AMLV-CQ and AMLV-WF17 are isolates of the same species. To our knowledge, this is the first report of a complete genome sequence of an AMLV isolate, and culantro was identified as a new natural host for this virus. In addition, a one-step RT-PCR assay was developed that provides a rapid, robust, and highly sensitive approach for the detection of AMLV.

Eryngium↗

Complete genome sequences of three effective nitrogen-fixing strains of Bradyrhizobium ottawaense from Canada.

We report complete genome sequences of three nitrogen-fixing Bradyrhizobium ottawaense strains isolated from soybeans in Canada. Each ~9.0 Mb genome (chromosome and plasmid) harbors predicted genes for nodulation, nitrogen fixation, N2O mitigation, phosphate solubilization, iron acquisition, phytohormone production, and stress tolerance, highlighting their potential for sustainable agriculture.

Bradyrhizobium ottawaense↗