Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “complete genome sequence”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Complete genomic sequence of the Amsacta moorei entomopoxvirus: analysis and comparison with other poxviruses.

The genome of the genus B entomopoxvirus from Amsacta moorei (AmEPV) was sequenced and found to contain 232,392 bases with 279 unique open reading frames (ORFs) of greater than 60 amino acids. The central core of the viral chromosome is flanked by 9.4-kb inverted terminal repeats (ITRs), each of which contains 13 ORFs, raising the total number of ORFs within the viral chromosome to 292. ORFs with no known homology to other poxvirus genes were shown to constitute 33.6% of the viral genome. Approximately 28.6% of the AmEPV genome encodes homologs of the mammalian poxvirus colinear core genes, which are found dispersed throughout the AmEPV chromosome. There is also no significant gene order conservation between AmEPV and the orthopteran genus B poxvirus of Melanoplus sanguinipes (MsEPV). Novel AmEPV genes include those encoding a putative ABC transporter and a Kunitz-motif protease inhibitor. The most unusual feature of the AmEPV genome relates to the viral encoded poly(A) polymerase. In all other poxviruses this heterodimeric enzyme consists of a single large and a single small subunit. However, AmEPV appears to encode one large and two distinct small poly(A) polymerase subunits. AmEPV is one of the few entomopoxviruses which can be grown and manipulated in cell culture. The complete genomic sequence of AmEPV paves the way for an understanding and comparison of the molecular properties and pathogenesis between the entomopoxviruses of insects and the more intensively studied vertebrate poxviruses.

ATP-Binding Cassette Transporters↗

Complete genome sequence of the entomopathogenic and metabolically versatile soil bacterium Pseudomonas entomophila.

Pseudomonas entomophila is an entomopathogenic bacterium that, upon ingestion, kills Drosophila melanogaster as well as insects from different orders. The complete sequence of the 5.9-Mb genome was determined and compared to the sequenced genomes of four Pseudomonas species. P. entomophila possesses most of the catabolic genes of the closely related strain P. putida KT2440, revealing its metabolically versatile properties and its soil lifestyle. Several features that probably contribute to its entomopathogenic properties were disclosed. Unexpectedly for an animal pathogen, P. entomophila is devoid of a type III secretion system and associated toxins but rather relies on a number of potential virulence factors such as insecticidal toxins, proteases, putative hemolysins, hydrogen cyanide and novel secondary metabolites to infect and kill insects. Genome-wide random mutagenesis revealed the major role of the two-component system GacS/GacA that regulates most of the potential virulence factors identified.

Animals↗

Completion of the sequence of a cetacean morbillivirus and comparative analysis of the complete genome sequences of four morbilliviruses.

The gene encoding the large (L) protein and the genome termini of the dolphin strain of cetacean morbillivirus (CeMV) were sequenced. The CeMV genome is 15702 nucleotides long and has been compared with other available morbillivirus genome sequences in regards to the "rule of six" and the "phase" of any particular nucleotide, defined as its position within a given hexamer, which here is defined as a group of six nucleotides starting from the 3' end of the genomic RNA. With exception of the position of the start of the F gene, the phase of the transcription start sites of each gene is strictly conserved between the morbilliviruses, but each gene is in a different phase. The lengths of gene transcripts differ between viruses by multiples of six nucleotides with exception of the M and F transcripts. The differences between the various morbilliviruses result from deletions or insertions of multiples of six nucleotides in the 3' and 5' UTRs of the different viral genes. The four bases were distributed non-randomly over the six positions in the hexamer boxes. However, the distribution patterns of each of the four bases indicated that multiples of three were more prevalent than those of six nucleotides. This reflected the positions of nucleotides in codons and codon usage in the reading frames. The L protein of CeMV was found to be 2183 amino acids in length and similar to that of MV and RPV. The CeMV L protein sequence was found to be equidistant between those of the CDV/PDV and MV/RPV subgroups of the morbilliviruses. This concurs with the analyses carried out on the other structural proteins.

3' Untranslated Regions↗

Complete genomic sequences of the GP5 protein gene of bluetongue virus serotype 11 and 17.

The complete nucleotide sequences of full-length copies of genomic segment 6 or M3 of US bluetongue virus serotype 11 and 17 consisted of 1638 nucleotides. The plus-strand contained an open reading frame for a protein of 526 amino acids which was equivalent to about 59,000 Da, similar to the molecular weight of GP5 as determined by SDS-PAGE analysis. This long open reading frame was flanked by a 5' non-coding region of 29 nucleotides and a 3' non-coding region of 28 bases. When the predicted amino acid sequences of GP5 of BTV-11 and -17 were aligned and compared with those of BTV-2, -10, -13, -1AU and -1SA, four major highly conserved domains interrupted by several variable regions were detected. The potential significance of these discrete domains is discussed. Evolutionary and phylogenetic characteristics of these US BTV serotypes were consistent with our finding concerning BTV-1AU and -1SA.

Amino Acid Sequence↗

Complete genome sequences and phylogenetic analysis of West Nile virus strains isolated from the United States, Europe, and the Middle East.

The complete nucleotide sequences of eight West Nile (WN) virus strains (Egypt 1951, Romania 1996-MQ, Italy 1998-equine, New York 1999-equine, MD 2000-crow265, NJ 2000MQ5488, NY 2000-grouse3282, and NY 2000-crow3356) were determined. Phylogenetic trees were constructed from the aligned nucleotide sequences of these eight viruses along with all other previously published complete WN virus genome sequences. The phylogenetic trees revealed the presence of two genetic lineages of WN viruses. Lineage 1 WN viruses have been isolated from the northeastern United States, Europe, Israel, Africa, India, Russia, and Australia. Lineage 2 WN viruses have been isolated only in sub-Saharan Africa and Madagascar. Lineage 1 viruses can be further subdivided into three monophyletic clades.

Animals↗

[Initial analysis of complete genome sequences of SARS coronavirus].

Multiple sequence alignment among 12 complete SARS coronavirus (SARS-CoV) sequences reveals that the major parts of 29708 b of the genomes have 99.82% identical bases. Forty two nucleotide mismatches were found in addition to the five and six gaps in two genomes. Among them, 28 mismatches result in changes of amino acid in the encoded proteins. Analysis of the changes implies possible effect on the Spike and Membrane protein of the virus, while most of the other changes seem not very significant to alter the structure and function of the proteins. These results have been released on the anti-sars web site maintained by the Centre of Bioinformatics, Peking University (antisars.cbi.pku.edu.cn) and may be of help for further experimental study.

Amino Acid Sequence↗

Bovine enterovirus 2: complete genomic sequence and molecular modelling of a reference strain and a wild-type isolate from endemically infected US cattle.

Bovine enteroviruses are members of the family Picornaviridae, genus Enterovirus. Whilst little is known about their pathogenic potential, they are apparently endemic in some cattle and cattle environments. Only one of the two current serotypes has been sequenced completely. In this report, the entire genome sequences of bovine enterovirus 2 (BEV-2) strain PS87 and a recent isolate from an endemically infected herd in Maryland, USA (Wye3A) are presented. The recent isolate clearly segregated phylogenetically with sequences representing the BEV-2 serotype, as did other isolates from the endemic herd. The Wye3A isolate shared 82 % nucleotide sequence identity with the PS87 strain and 68 % identity with a BEV-1 strain (VG5-27). Comparison of BEV-2 and BEV-1 deduced protein sequences revealed 72-73 % identity and showed that most differences were single amino acid changes or single deletions, with the exception of the VP1 protein, where both BEV-2 sequences were 7 aa shorter than that of BEV-1. Homology modelling of the capsid proteins of BEV-2 against protein database entries for picornaviruses indicated six significant differences among bovine enteroviruses and other members of the family Picornaviridae. Five of these were on the 'rim' of the proposed enterovirus receptor-binding site or 'canyon' (VP1) and one was near the base of the canyon (VP3). Two of these regions varied enough to distinguish BEV-2 from BEV-1 strains. This is the first report and analysis of full-length sequences for BEV-2. Continued analysis of these wild-type strains should yield useful information for genotyping enteroviruses and modelling enterovirus capsid structure.

Animals↗

Complete genomic sequence of viral hemorrhagic septicemia virus, a fish rhabdovirus.

The complete nucleotide sequence of the fish rhabdovirus viral hemorrhagic septicemia virus (VHSV) has been determined. The genome comprises 11158 bases and contains six long open reading frames encoding the nucleoprotein N, phosphoprotein P, matrix protein M, glycoprotein G, nonstructural viral protein NV, and polymerase L. Genes are arranged in the order 3'-N-P-M-G-NV-L-5'. The exact 3' and 5' ends were determined after RNA-oligonucleotide ligation or RACE. They show inverse complementarity as in other rhabdovirus genomes. Nucleotide and deduced amino acid sequences exhibit significant homology to corresponding sequences in the related fish rhabdovirus infectious hematopoietic necrosis virus.

Amino Acid Sequence↗

The complete genome sequence for a Turkish isolate of Wheat dwarf virus (WDV) from barley confirms the presence of two distinct WDV strains.

The complete genome for a barley isolate of Wheat dwarf virus (WDV) from Tekirdağ, Turkey, WDV-Bar[TR], was isolated and sequenced. The genome was found to be 2739 nucleotides long, which is shorter than wheat-infecting WDV isolates, and with a genome organization typical for mastreviruses. The complete genome of WDV-Bar[TR] showed 83-84% nucleotide identity to wheat isolates of WDV, with the non-coding regions SIR and LIR least conserved (72-74% identity). The deduced amino acid sequences for Rep and RepA were most conserved (92-93%), while CP and MP were less conserved (87% and 79-80%, respectively). The identity to other mastrevirus species was significantly lower. In phylogenetic analyses, the WDV isolates formed a distinct clade, well separated from the other mastreviruses with the wheat isolates grouping closely together. Phylogenetic analyses of WDV-Bar[TR], the partial sequence for another Turkish barley isolate (WDV-Bar[TR2]) and published WDV sequences further supported the division of WDV into two distinct strains. The barley strain could also be divided into three subtypes based on relationships and geographic origin. This study shows the first complete published sequence for a barley isolate of WDV.

Amino Acid Sequence↗

Insect mitochondrial genomics 2: The complete mitochondrial genome sequence of a giant stonefly, Pteronarcys princeps, asymmetric directional mutation bias, and conserved plecopteran A+T-region elements.

Mitochondrial (mt) genome sequences of insects are receiving renewed attention in molecular phylogentic studies, studies of mt-genome rearrangement, and other unusual molecular phenomena, such as translational frameshifting. At present, the basal neopteran lineages are poorly represented by mt-genome sequences. Complete mt-genome sequences are available in the databases for only the Orthoptera and Blatteria; 9 orders are unrepresented. Here, we present the complete mt-genome sequence of a giant stonefly, Pteronarcys princeps (Plecoptera; Pteronarcyidae). The 16,004 bp genome is typical in its genome content, gene organisation, and nucleotide composition. The genome shows evidence of strand-specific mutational biases, correlated with the time between the initiation of leading and the initiation of lagging strand replication. Comparisons with other insects reveal that this trend is seen in other insect groups, but is not universally consistent among sampled mt-genomes. The A+T region is compared with that of 2 stoneflies in the family Peltoperlidae. Conserved stem-loop structures and sequence blocks are noted between these distantly related families.

Animals↗

Tandem clusters of membrane proteins in complete genome sequences.

The distribution of genes coding for membrane proteins was investigated in 16 complete genomes: 4 archaea, 11 bacteria, and 1 eukaryote. Membrane proteins were identified by our new method of predicting transmembrane segments () after the removal of amino-terminal signal peptides. Interestingly, about half of the membrane protein genes in each genome were found to be located next to another, forming tandem clusters. Roughly 10%-30% of the tandem clusters were conserved among organisms, and most of the conserved tandem clusters belonged to one of the three functional groups, namely, transporters, the electron transport system, and cell motility. A tandem cluster sometimes contained paralogous membrane proteins, in which case the cluster size and the number of transmembrane segments could be related to a functional category, especially to transporters. In addition to the clustering of membrane proteins, the clustering of membrane proteins and ATP-binding proteins in the complete genomes was also analyzed. Although this clustering was not statistically significant, it was useful to identify candidate membrane protein partners of isolated ATP-binding protein components in the ABC transporters. Possible implications of tandem cluster organization of membrane protein genes are discussed including the complex formation and other functional coupling of protein products and the mechanism of protein translocation to the cell membrane.

ATP-Binding Cassette Transporters↗

The complete genome sequence of pepper severe mosaic virus and comparison with other potyviruses.

The complete nucleotide sequence of pepper severe mosaic virus (PepSMV) was determined. The viral genome consisted of 9890 nucleotides, excluding a poly (A) tract at the 3' end of the genome. The PepSMV RNA genome encoded a single polyprotein of 3085 amino acid residues, resulting in ten functionally distinct potyviral proteins. The lengths of the 5' nontranslated region (NTR) and the 3' NTR were 164 and 468 nucleotides, respectively. The genome organization of the virus was typical for members of the genus Potyvirus in the family Potyviridae. The coat protein amino acid sequence identity between PepSMV and the other 45 potyviruses ranged from 53.4 to 79.7%. Sequence alignments and phylogenetic analyses of the potyviral polyprotein sequences revealed that PepSMV was the closest to potato virus Y (PVY) and closely related to members of the PVY subgroup. Our genome sequence data clearly confirmed that PepSMV belongs to a separate species in the genus Potyvirus.

3' Untranslated Regions↗

Complete genomic sequence of Powassan virus: evaluation of genetic elements in tick-borne versus mosquito-borne flaviviruses.

The complete nucleotide sequence of the positive-stranded RNA genome of the tick-borne flavivirus Powassan (10,839 nucleotides) was elucidated and the amino acid sequence of all viral proteins was derived. Based on this sequence as well as serological data, Powassan virus represents the most divergent member of the tick-borne serocomplex within the genus flaviviruses, family Flaviviridae. The primary nucleotide sequence and potential RNA secondary structures of the Powassan virus genome as well as the protein sequences and the reactivities of the virion with a panel of monoclonal antibodies were compared to other tick-borne and mosquito-borne flaviviruses. These analyses corroborated significant differences between tick-borne and mosquito-borne flaviviruses, but also emphasized structural elements that are conserved among both vector groups. The comparisons among tick-borne flaviviruses revealed conserved sequence elements that might represent important determinants of the tick-borne flavivirus phenotype.

Amino Acid Sequence↗

Genome sequence completed of Alcanivorax borkumensis, a hydrocarbon-degrading bacterium that plays a global role in oil removal from marine systems.

In this paper, we provide background to the genome sequencing project of Alcanivorax borkumensis, which is a marine bacterium that uses exclusively petroleum oil hydrocarbons as sources of carbon and energy (therefore designated "hydrocarbonoclastic"). It is found in low numbers in all oceans of the world and in high numbers in oil-contaminated waters. Its ubiquity and unusual physiology suggest it is globally important in the removal of hydrocarbons from polluted marine systems. A functional genomics analysis of Alcanivorax borkumensis strain SK2 was recently initiated, and its genome sequence has just been completed. Annotation of the genome, metabolome modelling, and functional genomics, will soon reveal important insights into the genomic basis of the properties and physiology of this fascinating and globally important bacterium.

Biodegradation, Environmental↗

Complete genomic sequence of bacteriophage B3, a Mu-like phage of Pseudomonas aeruginosa.

Bacteriophage B3 is a transposable phage of Pseudomonas aeruginosa. In this report, we present the complete DNA sequence and annotation of the B3 genome. DNA sequence analysis revealed that the B3 genome is 38,439 bp long with a G+C content of 63.3%. The genome contains 59 proposed open reading frames (ORFs) organized into at least three operons. Of these ORFs, the predicted proteins from 41 ORFs (68%) display significant similarity to other phage or bacterial proteins. Many of the predicted B3 proteins are homologous to those encoded by the early genes and head genes of Mu and Mu-like prophages found in sequenced bacterial genomes. Only two of the predicted B3 tail proteins are homologous to other well-characterized phage tail proteins; however, several Mu-like prophages and transposable phage D3112 encode approximately 10 highly similar proteins in their predicted tail gene regions. Comparison of the B3 genomic organization with that of Mu revealed evidence of multiple genetic rearrangements, the most notable being the inversion of the proposed B3 immunity/early gene region, the loss of Mu-like tail genes, and an extreme leftward shift of the B3 DNA modification gene cluster. These differences illustrate and support the widely held view that tailed phages are genetic mosaics arising by the exchange of functional modules within a diverse genetic pool.

Base Sequence↗

Hepatitis E virus: complete genome sequence and phylogenetic analysis of a Nepali isolate.

The complete nucleotide sequence of the Nepali strain TK15/92 of hepatitis E (HEV) was determined. It showed the highest sequence homology with the Burmese B1 strain, but closer evolutionary relatedness to the Indian strains. Difficulties in reverse-transcribing and amplifying the hypervariable region in ORF1 suggested that strong secondary structures might be intrinsically responsible for the high mutational rate observed in this region of the HEV genome.

Base Sequence↗

Complete genome sequence of the plant commensal Pseudomonas fluorescens Pf-5.

Pseudomonas fluorescens Pf-5 is a plant commensal bacterium that inhabits the rhizosphere and produces secondary metabolites that suppress soilborne plant pathogens. The complete sequence of the 7.1-Mb Pf-5 genome was determined. We analyzed repeat sequences to identify genomic islands that, together with other approaches, suggested P. fluorescens Pf-5's recent lateral acquisitions include six secondary metabolite gene clusters, seven phage regions and a mobile genomic island. We identified various features that contribute to its commensal lifestyle on plants, including broad catabolic and transport capabilities for utilizing plant-derived compounds, the apparent ability to use a diversity of iron siderophores, detoxification systems to protect from oxidative stress, and the lack of a type III secretion system and toxins found in related pathogens. In addition to six known secondary metabolites produced by P. fluorescens Pf-5, three novel secondary metabolite biosynthesis gene clusters were also identified that may contribute to the biocontrol properties of P. fluorescens Pf-5.

Base Sequence↗

Nucleotide sequence of the 5'-terminus of Newcastle disease virus and assembly of the complete genomic sequence: agreement with the "rule of six".

We have determined the sequences of the 5' ends of three strains of Newcastle disease virus, permitting the assembly of the entire genomic sequence, which amounts to 15,186 nucleotides. This length is in agreement with the rule of six, which has been shown to determine replication efficiency in similar viruses. Comparison of the extreme 5' end of the trailer sequence with that of the 3'-terminal leader sequence of the virus reveals a high degree of complementarity. Variation between the 5'-terminal sequences of the different strains reveals the presence of alternative L gene polyadenylation signals, leading to correspondingly different trailer lengths.

Animals↗