Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “complete genome sequence”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

The complete genome sequence of the avian pathogen Mycoplasma gallisepticum strain R(low).

The complete genome of Mycoplasma gallisepticum strain R(low) has been sequenced. The genome is composed of 996,422 bp with an overall G+C content of 31 mol%. It contains 742 putative coding DNA sequences (CDSs), representing a 91 % coding density. Function has been assigned to 469 of the CDSs, while 150 encode conserved hypothetical proteins and 123 remain as unique hypothetical proteins. The genome contains two copies of the rRNA genes and 33 tRNA genes. The origin of replication has been localized based on sequence analysis in the region of the dnaA gene. The vlhA family (previously termed pMGA) contains 43 genes distributed among five loci containing 8, 2, 9, 12 and 12 genes. This family of genes constitutes 10.4% (103 kb) of the total genome. Two CDSs were identified immediately downstream of gapA and crmA encoding proteins that share homology to cytadhesins GapA and CrmA. Based on motif analysis it is predicted that 80 genes encode lipoproteins and 149 proteins contain multiple transmembrane domains. The authors have identified 75 proteins putatively involved in transport of biomolecules, 12 transposases, and a number of potential virulence factors. The completion of this sequence has spawned multiple projects directed at defining the biological basis of M. gallisepticum.

Animals↗

Complete genome sequence of Streptococcus vaginalis strain UMB8616 isolated from the bladder of a female with urge urinary incontinence.

Streptococcus vaginalis is a recently identified bacterial species closely related to Streptococcus anginosus. It has been isolated from the human urogenital tract. We report the complete genome sequence of S. vaginalis UMB8616 (=ATCC TSD-371 = CCUG 77169 = DSM 115471) isolated from the bladder of a human female with urge urinary incontinence.

Streptococcus↗

The complete genome sequence of the carcinogenic bacterium Helicobacter hepaticus.

Helicobacter hepaticus causes chronic hepatitis and liver cancer in mice. It is the prototype enterohepatic Helicobacter species and a close relative of Helicobacter pylori, also a recognized carcinogen. Here we report the complete genome sequence of H. hepaticus ATCC51449. H. hepaticus has a circular chromosome of 1,799,146 base pairs, predicted to encode 1,875 proteins. A total of 938, 953, and 821 proteins have orthologs in H. pylori, Campylobacter jejuni, and both pathogens, respectively. H. hepaticus lacks orthologs of most known H. pylori virulence factors, including adhesins, the VacA cytotoxin, and almost all cag pathogenicity island proteins, but has orthologs of the C. jejuni adhesin PEB1 and the cytolethal distending toxin (CDT). The genome contains a 71-kb genomic island (HHGI1) and several genomic islets whose G+C content differs from the rest of the genome. HHGI1 encodes three basic components of a type IV secretion system and other virulence protein homologs, suggesting a role of HHGI1 in pathogenicity. The genomic variability of H. hepaticus was assessed by comparing the genomes of 12 H. hepaticus strains with the sequenced genome by microarray hybridization. Although five strains, including all those known to have caused liver disease, were indistinguishable from ATCC51449, other strains lacked between 85 and 229 genes, including large parts of HHGI1, demonstrating extensive variation of genome content within the species.

Cell Movement↗

The R protein of SARS-CoV: analyses of structure and function based on four complete genome sequences of isolates BJ01-BJ04.

The R (replicase) protein is the uniquely defined non-structural protein (NSP) responsible for RNA replication, mutation rate or fidelity, regulation of transcription in coronaviruses and many other ssRNA viruses. Based on our complete genome sequences of four isolates (BJ01-BJ04) of SARS-CoV from Beijing, China, we analyzed the structure and predicted functions of the R protein in comparison with 13 other isolates of SARS-CoV and 6 other coronaviruses. The entire ORF (open-reading frame) encodes for two major enzyme activities, RNA-dependent RNA polymerase (RdRp) and proteinase activities. The R polyprotein undergoes a complex proteolytic process to produce 15 function-related peptides. A hydrophobic domain (HOD) and a hydrophilic domain (HID) are newly identified within NSP1. The substitution rate of the R protein is close to the average of the SARS-CoV genome. The functional domains in all NSPs of the R protein give different phylogenetic results that suggest their different mutation rate under selective pressure. Eleven highly conserved regions in RdRp and twelve cleavage sites by 3CLP (chymotrypsin-like protein) have been identified as potential drug targets. Findings suggest that it is possible to obtain information about the phylogeny of SARS-CoV, as well as potential tools for drug design, genotyping and diagnostics of SARS.

Amino Acid Sequence↗

The complete genome sequence of Chromobacterium violaceum reveals remarkable and exploitable bacterial adaptability.

Chromobacterium violaceum is one of millions of species of free-living microorganisms that populate the soil and water in the extant areas of tropical biodiversity around the world. Its complete genome sequence reveals (i) extensive alternative pathways for energy generation, (ii) approximately 500 ORFs for transport-related proteins, (iii) complex and extensive systems for stress adaptation and motility, and (iv) widespread utilization of quorum sensing for control of inducible systems, all of which underpin the versatility and adaptability of the organism. The genome also contains extensive but incomplete arrays of ORFs coding for proteins associated with mammalian pathogenicity, possibly involved in the occasional but often fatal cases of human C. violaceum infection. There is, in addition, a series of previously unknown but important enzymes and secondary metabolites including paraquat-inducible proteins, drug and heavy-metal-resistance proteins, multiple chitinases, and proteins for the detoxification of xenobiotics that may have biotechnological applications.

Adaptation, Physiological↗

Complete genome sequences of Chandipura and Isfahan vesiculoviruses.

Chandipura virus (CHPV) and Isfahan virus (ISFV) are two members of the genus Vesiculovirus from Asia. Both are arthropod-transmitted and are able to infect humans, but neither causes vesicular stomatitis in livestock. The complete genome sequence for each virus has been determined. The negative-sense RNA genome comprises 11,119 nt (CHPV) or 11,088 nt (ISFV). The most variable of the non-transcribed regions is the intergenic spacer at the G-L gene junction (4 bases in ISFV, 20 in CHPV). Phylogenetic analysis of deduced protein sequences shows that although CHPV and ISFV are distinct viruses, they are more related to each other than either is to the New World vesicular stomatitis viruses (VSV). The South American virus, Piry virus, is more closely related to the Asian viruses ISFV and CHPV, than it is to VSV.

Amino Acid Sequence↗

Complete genome sequence of a multidrug-resistant Proteus mirabilis clinical isolate harboring 22 antimicrobial resistance genes including blaCTX-M-15.

Proteus mirabilis causes urinary tract infections and wound infections and frequently exhibits multidrug resistance, complicating patient treatment. Here, we describe the complete genome sequence of a multidrug-resistant P. mirabilis wound isolate from 2025, providing insight into the repertoire of antimicrobial resistance genes in a recent P. mirabilis clinical isolate.

Proteus mirabilis↗

Complete genome sequence and phylogenetic analysis of hepatitis B virus (HBV) isolated from Mongolian patients with chronic HBV infection.

Although there is a report of a high rate of hepatitis B virus (HBV) infection in Mongolia, the entire nucleotide sequence of HBV circulating among Mongolian patients has not been reported. To obtain the complete nucleotide sequence of the Mongolian HBV, viral DNA was extracted from sera of patients with HBV infection. Six Mongolian HBV strains were amplified by PCR. Complete genomic sequences were determined for two Mongolian HBV isolates, MBT181 and MMU36. The entire genome of Mongolian HBV isolates was 3,182 bp long and genetic distance between Mongolian HBV isolates was 3.2%. Precore stop codon resulting from a guanine to adenine mutation at nucleotide 1,896 was detected in MBT181 strain. Based on phylogenetic analysis, the six Mongolian isolates were classified as genotype D.

Amino Acid Substitution↗

SDR and MDR: completed genome sequences show these protein families to be large, of old origin, and of complex nature.

Short-chain dehydrogenases/reductases (SDR) and medium-chain dehydrogenases/reductases (MDR) are protein families originally distinguished from characterisations of alcohol dehydrogenase of these two types. Screening of completed genome sequences now reveals that both these families are large, wide-spread and complex. In Escherichia coli alone, there are no fewer than 17 MDR forms, identified as open reading frames, considerably extending previously known MDR relationships in prokaryotes and including ethanol-active alcohol dehydrogenase. In entire databanks, 1056 SDR and 537 MDR forms are currently known, extending the multiplicity further. Complexity is also large, with several enzyme activity types, subgroups and evolutionary patterns. Repeated duplications can be traced for the alcohol dehydrogenases, with independent enzymogenesis of ethanol activity, showing a general importance of this enzyme activity.

Evolution, Molecular↗

The complete genome sequence of Francisella tularensis, the causative agent of tularemia.

Francisella tularensis is one of the most infectious human pathogens known. In the past, both the former Soviet Union and the US had programs to develop weapons containing the bacterium. We report the complete genome sequence of a highly virulent isolate of F. tularensis (1,892,819 bp). The sequence uncovers previously uncharacterized genes encoding type IV pili, a surface polysaccharide and iron-acquisition systems. Several virulence-associated genes were located in a putative pathogenicity island, which was duplicated in the genome. More than 10% of the putative coding sequences contained insertion-deletion or substitution mutations and seemed to be deteriorating. The genome is rich in IS elements, including IS630 Tc-1 mariner family transposons, which are not expected in a prokaryote. We used a computational method for predicting metabolic pathways and found an unexpectedly high proportion of disrupted pathways, explaining the fastidious nutritional requirements of the bacterium. The loss of biosynthetic pathways indicates that F. tularensis is an obligate host-dependent bacterium in its natural life cycle. Our results have implications for our understanding of how highly virulent human pathogens evolve and will expedite strategies to combat them.

Base Sequence↗

Complete genome sequence analysis of an iridovirus isolated from the orange-spotted grouper, Epinephelus coioides.

Orange-spotted grouper iridovirus (OSGIV) was the causative agent of serious systemic diseases with high mortality in the cultured orange-spotted grouper, Epinephelus coioides. Here we report the complete genome sequence of OSGIV. The OSGIV genome consists of 112,636 bp with a G+C content of 54%. 121 putative open reading frames (ORF) were identified with coding capacities for polypeptides varying from 40 to 1168 amino acids. The majority of OSGIV shared homologies to other iridovirus genes. Phylogenetic analysis of the major capsid protein, ATPase, cytosine DNA methyl transferase and DNA polymerase indicated that OSGIV was closely related to infectious spleen and kidney necrosis virus (ISKNV) and rock bream iridovirus (RBIV), but differed from lymphocytisvirus and ranavirus. The determination of the genome of OSGIV will facilitate a better understanding of the molecular mechanism underlying the pathogenesis of the OSGIV and may provide useful information to develop diagnosis method and strategies to control outbreak of OSGIV.

Animals↗

Complete genome sequence of Fer-de-Lance virus reveals a novel gene in reptilian paramyxoviruses.

The complete RNA genome sequence of the archetype reptilian paramyxovirus, Fer-de-Lance virus (FDLV), has been determined. The genome is 15,378 nucleotides in length and consists of seven nonoverlapping genes in the order 3' N-U-P-M-F-HN-L 5', coding for the nucleocapsid, unknown, phospho-, matrix, fusion, hemagglutinin-neuraminidase, and large polymerase proteins, respectively. The gene junctions contain highly conserved transcription start and stop signal sequences and tri-nucleotide intergenic regions similar to those of other Paramyxoviridae. The FDLV P gene expression strategy is like that of rubulaviruses, which express the accessory V protein from the primary transcript and edit a portion of the mRNA to encode P and I proteins. There is also an overlapping open reading frame potentially encoding a small basic protein in the P gene. The gene designated U (unknown), encodes a deduced protein of 19.4 kDa that has no counterpart in other paramyxoviruses and has no similarity with sequences in the National Center for Biotechnology Information database. Active transcription of the U gene in infected cells was demonstrated by Northern blot analysis, and bicistronic N-U mRNA was also evident. The genomes of two other snake paramyxovirus genotypes were also found to have U genes, with 11 to 16% nucleotide divergence from the FDLV U gene. Pairwise comparisons of amino acid identities and phylogenetic analyses of all deduced FDLV protein sequences with homologous sequences from other Paramyxoviridae indicate that FDLV represents a new genus within the subfamily Paramyxoviridae. We suggest the name Ferlavirus for the new genus, with FDLV as the type species.

Amino Acid Sequence↗

NPC1: Complete genomic sequence, mutation analysis, and characterization of haplotypes.

Niemann-Pick type C disease (NP-C) is a rare, autosomal recessive lipid storage disorder. At least 96% of all NP-C patients link to NPC1 which encodes for a lysosomally-targeted protein. We describe the complete genomic sequence of 57,052 kb corresponding to the transcribed region of human NPC1 including several exonic and intronic single nucleotide polymorphisms (SNPs). Sequencing of all exons, splice sites, and the promoter region of NPC1 in 12 unrelated Caucasian NP-C patients revealed nine novel and four known most likely disease-causing mutations. Ten unique mutations found only once in 24 disease alleles were observed in patients being compound heterozygous for two different mutations. Two of the three missense mutations identified more than once were observed in a total of four patients homozygous for the respective mutation along with homozygosity for the underlying haplotype. The patients were offspring of most likely nonconsanguineous couples. Based upon genotyping exonic SNPs c.2572A>G (I858V; g.45020A>G) and c.2793C>T (N931N; g.45686C>T) and segregation analysis we characterized the haplotype of all 24 NPC1 alleles and of 138 alleles of healthy Caucasian control subjects. All four permutations between the two SNPs were identified in the control alleles: 2572A-2793C (50%), 2572G-2793T (41%), 2572G-2793C (5%), and 2572A-2793T (4%). These data are suggestive for an ancestral intragenic recombination within a genomic fragment of <666 bp. While 17 of 24 NP-C alleles (71%) shared haplotype 2572G-2793T, this haplotype accounted for only 41% in the controls (p=0.007; 2-sided Fisher exact test) suggesting the possibility of an influence of the haplotypic background on expression of missense mutations in NPC1.

Carrier Proteins↗

Complete genomic sequencing shows that polioviruses and members of human enterovirus species C are closely related in the noncapsid coding region.

The 65 human enterovirus serotypes are currently classified into five species: Poliovirus (3 serotypes), Human enterovirus A (HEV-A) (12 serotypes), HEV-B (37 serotypes), HEV-C (11 serotypes), and HEV-D (2 serotypes). Coxsackie A virus (CAV) serotypes 1, 11, 13, 15, 17, 18, 19, 20, 21, 22, and 24 constitute HEV-C. We have determined the complete genome sequences for the remaining nine HEV-C serotypes and compared them with the complete sequences of CAV21, CAV24, and the polioviruses. The viruses were most diverse in the capsid region (4 to 36% amino acid difference). A high degree of capsid sequence conservation (96% amino acid identity) suggests that CAV15 and CAV18 should be classified as strains of CAV11 and CAV13, respectively. In the 3CD region, CAV1, CAV19, and CAV22 differed from one another by only 1.2 to 1.4% and CAV11, CAV13, CAV17, CAV20, CAV21, CAV24, and the polioviruses differed from one another by only 1.2 to 3.6%. The two groups, however, differed from one another by 14.6 to 16.2%. The polioviruses as a group were monophyletic only in the capsid region. Only one group of serotypes (CAV1, CAV19, and CAV22) was consistently monophyletic in multiple genome regions. Incongruities among phylogenetic trees based on different genome regions strongly suggest that recombination has occurred between the polioviruses, CAV11, CAV13, CAV17, and CAV20. The close relationship among the polioviruses and CAV11, CAV13, CAV17, CAV20, CAV21, and CAV24 and the uniqueness of CAV1, CAV19, and CAV22 suggest that revisions should be made to the classification of these viruses.

Amino Acid Sequence↗

The complete genomic sequence of a novel member of the genus Caulimovirus isolated from Dregea volubilis.

A novel caulimovirus was identified from diseased leaves of Dregea volubilis exhibiting yellowing and vein-associated chlorosis in Yuanjiang County, Yunnan Province, China. The virus was tentatively named Dregea volubilis caulimovirus 1 (DVCaV1). The complete genome sequence of DVCaV1, determined by de novo assembly of high-throughput sequencing data, comprises 8,160&#xa0;bp of circular double-stranded DNA containing two intergenic regions and seven open reading frames (ORFs). These ORFs encode (in order) a movement protein (MP), an aphid transmission factor (ATF), a virion-associated protein (VAP), a coat protein (CP), a polymerase polyprotein (Pol, containing protease, reverse transcriptase, and RNase H domains), a transactivator/viroplasmin (TAV) protein, and a hypothetical protein of unknown function. Sequence comparisons revealed the highest nucleotide similarity with strawberry vein banding virus (SVBV; NC_001725). Phylogenetic analysis confirmed DVCaV1 as a member of the genus Caulimovirus, with SVBV as its closest known relative. According to current ICTV species demarcation criteria for the genus Caulimovirus (host range and >&#x2009;20% nucleotide sequence divergence in the polymerase region), DVCaV1 represents a novel species. This is, to our knowledge, the first report of a caulimovirus detected in naturally symptomatic Dregea volubilis.

Genome, Viral↗

Complete genomic sequence of a dengue type 2 virus from the French West Indies.

Severe forms of dengue fever, dengue haemorrhagic fever, and dengue shock syndrome, were not prominent in the Americas until the epidemic of Cuba in 1981. Since that time, they have spread to other countries in Central and South America, correlating with the spread of dengue type 2 viruses related to Southeast Asian strains. We report here the complete genomic sequence of a dengue type 2 virus isolated during the epidemic in La Martinique in 1998. This constitutes the first complete genetic characterization of a dengue virus strain from French West Indies, and also the first molecular identification in this region of a dengue 2 strain phylogenetically related to the emerging American type 2 dengue viruses.

Amino Acid Substitution↗

Complete genome sequences of the SARS-CoV: the BJ Group (Isolates BJ01-BJ04).

Beijing has been one of the epicenters attacked most severely by the SARS-CoV (severe acute respiratory syndrome-associated coronavirus) since the first patient was diagnosed in one of the city's hospitals. We now report complete genome sequences of the BJ Group, including four isolates (Isolates BJ01, BJ02, BJ03, and BJ04) of the SARS-CoV. It is remarkable that all members of the BJ Group share a common haplotype, consisting of seven loci that differentiate the group from other isolates published to date. Among 42 substitutions uniquely identified from the BJ group, 32 are non-synonymous changes at the amino acid level. Rooted phylogenetic trees, proposed on the basis of haplotypes and other sequence variations of SARS-CoV isolates from Canada, USA, Singapore, and China, gave rise to different paradigms but positioned the BJ Group, together with the newly discovered GD01 (GD-Ins29) in the same clade, followed by the H-U Group (from Hong Kong to USA) and the H-T Group (from Hong Kong to Toronto), leaving the SP Group (Singapore) more distant. This result appears to suggest a possible transmission path from Guangdong to Beijing/Hong Kong, then to other countries and regions.

Genome, Viral↗

Repeat-associated phase variable genes in the complete genome sequence of Neisseria meningitidis strain MC58.

Phase variation, mediated through variation in the length of simple sequence repeats, is recognized as an important mechanism for controlling the expression of factors involved in bacterial virulence. Phase variation is associated with most of the currently recognized virulence determinants of Neisseria meningitidis. Based upon the complete genome sequence of the N. meningitidis serogroup B strain MC58, we have identified tracts of potentially unstable simple sequence repeats and their potential functional significance determined on the basis of sequence context. Of the 65 potentially phase variable genes identified, only 13 were previously recognized. Comparison with the sequences from the other two pathogenic Neisseria sequencing projects shows differences in the length of the repeats in 36 of the 65 genes identified, including 25 of those not previously known to be phase variable. Six genes that did not have differences in the length of the repeat instead had polymorphisms such that the gene would not be expected to be phase variable in at least one of the other strains. A further 12 candidates did not have homologues in either of the other two genome sequences. The large proportion of these genes that are associated with frameshifts and with differences in repeat length between the neisserial genome sequences is further corroborative evidence that they are phase variable. The number of potentially phase variable genes is substantially greater than for any other species studied to date, and would allow N. meningitidis to generate a very large repertoire of phenotypes through expression of these genes in different combinations. Novel phase variable candidates identified in the strain MC58 genome sequence include a spectrum of genes encoding glycosyltransferases, toxin related products, and metabolic activities as well as several restriction/modification and bacteriocin-related genes and a number of open reading frames (ORFs) for which the function is currently unknown. This suggests that the potential role of phase variation in mediating bacterium-host interactions is much greater than has been appreciated to date. Analysis of the distribution of homopolymeric tract lengths indicates that this species has sequence-specific mutational biases that favour the instability of sequences associated with phase variation.

Bacterial Proteins↗