Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “complete genome sequence”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Complete genome sequence of the industrial bacterium Bacillus licheniformis and comparisons with closely related Bacillus species.

BACKGROUND: Bacillus licheniformis is a Gram-positive, spore-forming soil bacterium that is used in the biotechnology industry to manufacture enzymes, antibiotics, biochemicals and consumer products. This species is closely related to the well studied model organism Bacillus subtilis, and produces an assortment of extracellular enzymes that may contribute to nutrient cycling in nature. RESULTS: We determined the complete nucleotide sequence of the B. licheniformis ATCC 14580 genome which comprises a circular chromosome of 4,222,336 base-pairs (bp) containing 4,208 predicted protein-coding genes with an average size of 873 bp, seven rRNA operons, and 72 tRNA genes. The B. licheniformis chromosome contains large regions that are colinear with the genomes of B. subtilis and Bacillus halodurans, and approximately 80% of the predicted B. licheniformis coding sequences have B. subtilis orthologs. CONCLUSIONS: Despite the unmistakable organizational similarities between the B. licheniformis and B. subtilis genomes, there are notable differences in the numbers and locations of prophages, transposable elements and a number of extracellular enzymes and secondary metabolic pathway operons that distinguish these species. Differences include a region of more than 80 kilobases (kb) that comprises a cluster of polyketide synthase genes and a second operon of 38 kb encoding plipastatin synthase enzymes that are absent in the B. licheniformis genome. The availability of a completed genome sequence for B. licheniformis should facilitate the design and construction of improved industrial strains and allow for comparative genomics and evolutionary studies within this group of Bacillaceae.

Anti-Bacterial Agents↗

Complete genome sequence and analyses of the subgenomic RNAs of sweet potato chlorotic stunt virus reveal several new features for the genus Crinivirus.

The complete nucleotide sequences of genomic RNA1 (9,407 nucleotides [nt]) and RNA2 (8,223 nt) of Sweet potato chlorotic stunt virus (SPCSV; genus Crinivirus, family Closteroviridae) were determined, revealing that SPCSV possesses the second largest identified positive-strand single-stranded RNA genome among plant viruses after Citrus tristeza virus. RNA1 contains two overlapping open reading frames (ORFs) that encode the replication module, consisting of the putative papain-like cysteine proteinase, methyltransferase, helicase, and polymerase domains. RNA2 contains the Closteroviridae hallmark gene array represented by a heat shock protein homologue (Hsp70h), a protein of 50 to 60 kDa depending on the virus, the major coat protein, and a divergent copy of the coat protein. This grouping resembles the genome organization of Lettuce infectious yellows virus (LIYV), the only other crinivirus for which the whole genomic sequence is available. However, in striking contrast to LIYV, the two genomic RNAs of SPCSV contained nearly identical 208-nt-long 3' terminal sequences, and the ORF for a putative small hydrophobic protein present in LIYV RNA2 was found at a novel position in SPCSV RNA1. Furthermore, unlike any other plant or animal virus, SPCSV carried an ORF for a putative RNase III-like protein (ORF2 on RNA1). Several subgenomic RNAs (sgRNAs) were detected in SPCSV-infected plants, indicating that the sgRNAs formed from RNA1 accumulated earlier in infection than those of RNA2. The 5' ends of seven sgRNAs were cloned and sequenced by an approach that provided compelling evidence that the sgRNAs are capped in infected plants, a novel finding for members of the Closteroviridae.

Amino Acid Sequence↗

Complete genome sequence of the Lactococcus lactis temperate phage phiLC3: comparative analysis of phiLC3 and its relatives in lactococci and streptococci.

Complete genome sequencing of the P335 temperate Lactococcus lactis bacteriophage phiLC3 (32, 172 bp) revealed fifty-one open reading frames (ORFs). Four ORFs did not show any homology to other proteins in the database and twenty-one ORFs were assigned a putative biological function. phiLC3 contained a unique replication module and orf201 was identified as the putative replication initiator protein-encoding gene. phiLC3 was closely related to the L. lactis r1t phage (73% DNA identity). Similarity was also shared with other lactococcal P335 phages and the Streptococcus pyogenes prophages 370.3, 8232.4 and 315.5 over the non-structural genes and the genes involved in DNA packaging/phage morphogenesis, respectively. phiLC3 contained small homologous regions distributed among lactococcal phages suggesting that these regions might be involved in mediating genetic exchange. Two regions of 30 and 32 bp were conserved among the streptococcal and lactococcal r1t-like phages. These two regions, as well as other homologous regions, were located at mosaic borders and close to putative transcriptional terminators indicating that such regions together might attract recombination. The conserved regions found among lactococcal and streptococcal phages might be used for identification of phages/prophages/prophage remnants in their hosts.

Base Sequence↗

Deciphering the biology of Mycobacterium tuberculosis from the complete genome sequence.

Countless millions of people have died from tuberculosis, a chronic infectious disease caused by the tubercle bacillus. The complete genome sequence of the best-characterized strain of Mycobacterium tuberculosis, H37Rv, has been determined and analysed in order to improve our understanding of the biology of this slow-growing pathogen and to help the conception of new prophylactic and therapeutic interventions. The genome comprises 4,411,529 base pairs, contains around 4,000 genes, and has a very high guanine + cytosine content that is reflected in the biased amino-acid content of the proteins. M. tuberculosis differs radically from other bacteria in that a very large portion of its coding capacity is devoted to the production of enzymes involved in lipogenesis and lipolysis, and to two new families of glycine-rich proteins with a repetitive structure that may represent a source of antigenic variation.

Chromosome Mapping↗

Complete genomic sequence of nitrogen-fixing symbiotic bacterium Bradyrhizobium japonicum USDA110.

The complete nucleotide sequence of the genome of a symbiotic bacterium Bradyrhizobium japonicum USDA110 was determined. The genome of B. japonicum was a single circular chromosome 9,105,828 bp in length with an average GC content of 64.1%. No plasmid was detected. The chromosome comprises 8317 potential protein-coding genes, one set of rRNA genes and 50 tRNA genes. Fifty-two percent of the potential protein genes showed sequence similarity to genes of known function and 30% to hypothetical genes. The remaining 18% had no apparent similarity to reported genes. Thirty-four percent of the B. japonicum genes showed significant sequence similarity to those of both Mesorhizobium loti and Sinorhizobium meliloti, while 23% were unique to this species. A presumptive symbiosis island 681 kb in length, which includes a 410-kb symbiotic region previously reported by Göttfert et al., was identified. Six hundred fifty-five putative protein-coding genes were assigned in this region, and the functions of 301 genes, including those related to symbiotic nitrogen fixation and DNA transmission, were deduced. A total of 167 genes for transposases/104 copies of insertion sequences were identified in the genome. It was remarkable that 100 out of 167 transposase genes are located in the presumptive symbiotic island. DNA segments of 4 to 97 kb inserted into tRNA genes were found at 14 locations in the genome, which generates partial duplication of the target tRNA genes. These observations suggest plasticity of the B. japonicum genome, which is probably due to complex genome rearrangements such as horizontal transfer and insertion of various DNA elements, and to homologous recombination.

Bradyrhizobium↗

Coding-complete genome sequence of grapevine leafroll-associated virus 13 from grapevine in California.

In this study, we report the coding-complete genome sequence of Grapevine leafroll-associated virus 13 (GLRaV-13), isolate CA8881, detected in Vitis vinifera in California, USA. The genome sequence exhibited over 95% nucleotide identity with previously reported GLRaV-13 isolates and contributed to better understanding of the genetic diversity of ampeloviruses infecting grapevine.

California↗

The complete genome sequence of severe acute respiratory syndrome coronavirus strain HKU-39849 (HK-39).

The complete genomic nucleotide sequence (29.7kb) of a Hong Kong severe acute respiratory syndrome (SARS) coronavirus (SARS-CoV) strain HK-39 is determined. Phylogenetic analysis of the genomic sequence reveals it to be a distinct member of the Coronaviridae family. 5' RACE assay confirms the presence of at least six subgenomic transcripts all containing the predicted intergenic sequences. Five open reading frames (ORFs), namely ORF1a, 1b, S, M, and N, are found to be homologues to other CoV members, and three more unknown ORFs (X1, X2, and X3) are unparalleled in all other known CoV species. Optimal alignment and computer analysis of the homologous ORFs has predicted the characteristic structural and functional domains on the putative genes. The overall nucleotides conservation of the homologous ORFs is low (<5%) compared with other known CoVs, implying that HK-39 is a newly emergent SARS-CoV phylogenetically distant from other known members. SimPlot analysis supports this finding, and also suggests that this novel virus is not a product of a recent recombinant from any of the known characterized CoVs. Together, these results confirm that HK-39 is a novel and distinct member of the Coronaviridae family, with unknown origin. The completion of the genomic sequence of the virus will assist in tracing its origin.

3' Untranslated Regions↗

Analysis of the complete genome sequence of acute bee paralysis virus shows that it belongs to the novel group of insect-infecting RNA viruses.

The complete genome sequence of acute bee paralysis virus (ABPV) was determined. The 9470 nucleotide, polyadenylated RNA genome encoded two open reading frames (ORF1 and ORF2), which were separated by 184 nucleotides. The deduced amino acid sequence of the 5' ORF1 (nucleotides 605 to 6325) showed significant similarity to the RNA-dependent RNA polymerase, helicase, and protease domains of viruses from the picornavirus, comovirus, calicivirus, and sequivirus families, as well as to a novel group of insect-infecting RNA viruses. The 3' ORF2 (nucleotides 6509-9253) was proposed as encoding a capsid polyprotein with three major structural proteins (35, 33, and 24 kDa) and a minor protein (9.4 kDa). This was confirmed by N-terminal sequence analysis of two of these proteins. The overall genome structure of ABPV showed similarities to those of Drosophila C virus, Plautia stali intestine virus, Rhopalosiphum padi virus, and Himetobi P virus, which have been classified into a novel group of picorna-like insect-infecting RNA viruses called cricket paralysis-like viruses. It is suggested that ABPV belongs to the cricket paralysis-like viruses.

Amino Acid Sequence↗

Hepatitis C virus complete genome sequences identified from China representing subtypes 6k and 6n and a novel, as yet unassigned subtype within genotype 6.

Here, the complete genome sequences for three hepatitis C virus (HCV) variants identified from China and belonging to genotype 6 are reported: km41, km42 and gz52557. Their entire genome lengths were 9430, 9441 and 9448 nt, respectively; the 5' untranslated regions (UTRs) contained 341, 342 and 339 nt, followed by single open reading frames of 9045, 9045 and 9057 nt, respectively; the 3' UTRs, up to the poly(U) tracts, were 41, 51 and 52 nt, respectively. Phylogenetic analyses showed that km41 is classified into subtype 6k and km42 into subtype 6n. Although gz52557 clustered distantly with subtype 6g, it appeared to belong to a distinct subtype. Analysis with 53 and 105 partial core and NS5B region sequences, respectively, representing 17 subtypes from 6a to 6q and three unassigned isolates of genotype 6 in co-analyses demonstrated that gz52557 was equidistant from all of these isolates, indicating that it belongs to a novel subtype. However, based on a recent consensus that three or more examples are required for a new HCV subtype designation, it is suggested that gz52557 remains unassigned to any subtype.

China↗

Vibrio cholerae phage K139: complete genome sequence and comparative genomics of related phages.

In this report, we characterize the complete genome sequence of the temperate phage K139, which morphologically belongs to the Myoviridae phage family (P2 and 186). The prophage genome consists of 33,106 bp, and the overall GC content is 48.9%. Forty-four open reading frames were identified. Homology analysis and motif search were used to assign possible functions for the genes, revealing a close relationship to P2-like phages. By Southern blot screening of a Vibrio cholerae strain collection, two highly K139-related phage sequences were detected in non-O1, non-O139 strains. Combinatorial PCR analysis revealed almost identical genome organizations. One region of variable gene content was identified and sequenced. Additionally, the tail fiber genes were analyzed, leading to the identification of putative host-specific sequence variations. Furthermore, a K139-encoded Dam methyltransferase was characterized.

Bacteriophages↗

Complete genome sequence of the methanogenic archaeon, Methanococcus jannaschii.

The complete 1.66-megabase pair genome sequence of an autotrophic archaeon, Methanococcus jannaschii, and its 58- and 16-kilobase pair extrachromosomal elements have been determined by whole-genome random sequencing. A total of 1738 predicted protein-coding genes were identified; however, only a minority of these (38 percent) could be assigned a putative cellular role with high confidence. Although the majority of genes related to energy production, cell division, and metabolism in M. jannaschii are most similar to those found in Bacteria, most of the genes involved in transcription, translation, and replication in M. jannaschii are more similar to those found in Eukaryotes.

Amino Acid Sequence↗

Complete genome sequence and comparative genomics of Shigella flexneri serotype 2a strain 2457T.

We determined the complete genome sequence of Shigella flexneri serotype 2a strain 2457T (4,599,354 bp). Shigella species cause >1 million deaths per year from dysentery and diarrhea and have a lifestyle that is markedly different from those of closely related bacteria, including Escherichia coli. The genome exhibits the backbone and island mosaic structure of E. coli pathogens, albeit with much less horizontally transferred DNA and lacking 357 genes present in E. coli. The strain is distinctive in its large complement of insertion sequences, with several genomic rearrangements mediated by insertion sequences, 12 cryptic prophages, 372 pseudogenes, and 195 S. flexneri-specific genes. The 2457T genome was also compared with that of a recently sequenced S. flexneri 2a strain, 301. Our data are consistent with Shigella being phylogenetically indistinguishable from E. coli. The S. flexneri-specific regions contain many genes that could encode proteins with roles in virulence. Analysis of these will reveal the genetic basis for aspects of this pathogenic organism's distinctive lifestyle that have yet to be explained.

Base Sequence↗

Complete genome sequence of lymphocystis disease virus isolated from China.

Lymphocystis diseases in fish throughout the world have been extensively described. Here we report the complete genome sequence of lymphocystis disease virus isolated in China (LCDV-C), an LCDV isolated from cultured flounder (Paralichthys olivaceus) with lymphocystis disease in China. The LCDV-C genome is 186,250 bp, with a base composition of 27.25% G+C. Computer-assisted analysis revealed 240 potential open reading frames (ORFs) and 176 nonoverlapping putative viral genes, which encode polypeptides ranging from 40 to 1,193 amino acids. The percent coding density is 67%, and the average length of each ORF is 702 bp. A search of the GenBank database using the 176 individual putative genes revealed 103 homologues to the corresponding ORFs of LCDV-1 and 73 potential genes that were not found in LCDV-1 and other iridoviruses. Among the 73 genes, there are 8 genes that contain conserved domains of cellular genes and 65 novel genes that do not show any significant homology with the sequences in public databases. Although a certain extent of similarity between putative gene products of LCDV-C and corresponding proteins of LCDV-1 was revealed, no colinearity was detected when their ORF arrangements and coding strategies were compared to each other, suggesting that a high degree of genetic rearrangements between them has occurred. And a large number of tandem and overlapping repeated sequences were observed in the LCDV-C genome. The deduced amino acid sequence of the major capsid protein (MCP) presents the highest identity to those of LCDV-1 and other iridoviruses among the LCDV-C gene products. Furthermore, a phylogenetic tree was constructed based on the multiple alignments of nine MCP amino acid sequences. Interestingly, LCDV-C and LCDV-1 were clustered together, but their amino acid identity is much less than that in other clusters. The unexpected levels of divergence between their genomes in size, gene organization, and gene product identity suggest that LCDV-C and LCDV-1 shouldn't belong to a same species and that LCDV-C should be considered a species different from LCDV-1.

Animals↗

Complete genome sequence of Rhodococcus qingshengii strain A3-8.

A chemostat culture was constructed with phenol and forest soil as an inoculum. We report the complete genome sequence of Rhodococcus qingshengii strain A3-8, which was isolated from the culture. The genome consists of a chromosome (6,436,695 bp) and a linear plasmid pA38 (257,365 bp).

Rhodococcus↗