Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “complete genome sequence”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Comparison of the complete genomic sequence of the border disease virus, BD31, to other pestiviruses.

The genus Pestivirus is composed of hog cholera virus (HCV) [also known as classical swine fever virus (CSFV)], bovine viral diarrhea virus (BVDV), and border disease virus (BDV). Complete sequences have been published for HCV (or CSFV) and the two genotypes of BVDV (BVDV1 and BVDV2). In this study the complete sequence of the border disease virus (BDV), BD31, was determined. BD31 was isolated from a lamb with hairy shaker syndrome and is the BDV type virus offered by ATCC (ATCC VR-996). The genome was 12268 nucleotides long and had a single large open reading frame (ORF) beginning at nucleotide 357 and ending at nucleotide 12045. The sequence identity of the predicted amino acid sequence of BD31 and other published pestivirus sequences varied from 71% to 78%. Phylogenetic analysis of available complete genomic sequences segregated pestiviruses into two branches. One branch contained BD31 and HCV (or CSFV) isolates while the other branch contained BVDV1 and BVDV2 isolates. Pestiviruses from the same branch were similar in the length of the 5' and 3' untranslated regions (UTR). When complete genomic sequences were compared among BD31, HCV (or CSFV), BVDV1 and BVDV2, the highest sequence identity was observed in the 5' UTR. Within the ORF, the highest sequence identity was observed in the genomic region coding for the nonstructural viral polypeptide p80.

Amino Acid Sequence↗

Complete genomic sequence and phylogenetic relatedness of hepatitis B virus isolates from Iran.

Hepatitis B virus (HBV) is one of the main etiological agents of acute and chronic liver disease that is still a major public health problem in the world. Numerous HBV isolates have grouped into eight genotypes, A to H, based on the complete genome sequence. To date, no study has been carried out on the complete HBV genome sequence in Iran. The objective of this study was to investigate the complete genome sequence organization and phylogenetic analysis of the five HBV strains, which obtained from Iranian chronic infected patients. Results showed that Iranian strains were closely related to each other, with 97-100% nucleotide similarity. Phylogenetic analysis based on the complete genome sequences and the precore/core gene sequences revealed that all strains were of genotype D, sub-genotype D1 with bootstrap value 100 and 99%, respectively. The S gene encoded Arg122, Pro127, and Lys160 corresponding to subtype ayw2. Iranian HBV isolates had closely related with Turkish HBV strains. All strains had a nucleotide length of 3,182 base pair (bp) except IR-P4 strain, with a 3,185 bp in length and with a unique Phe89 insertion in the X gene. The intragenotypic divergence of the complete genome sequence of Iranian strains was 1.8% and the intergenotypic in genotype D was 3.8% and with the other genotypes was 7.9-15.4%. In conclusion, this study revealed that the HBV genotype D, sub-genotype D1, subtype ayw2 dominates in the Iranian infected patients. A single Phe89 insertion in the X gene of the one Iranian strain with an unforeseen length of 3185 bp was identified.

Adult↗

[Complete genome sequence analysis of the Hantavirus Z10 strain].

OBJECTIVE: Study on the complete genome sequence of Hantavirus Z10 strain which has been applied for inactivated vaccine production in China, to assess its molecular characteristics and the diversity with other hantaviruses. METHODS: The total RNA were prepared from Z10 virus infected cells and the RT-PCR products was cloned into T vector, sequenced and analyzed by using DNASTAR software. RESULTS: The Z10 complete genome, L segment is 6,553, M segment is 3,615, S segment is 1,701 nucleotides in length, with a single open reading frame encoding 2,151, 1,135, 429 amino acids respectively. Sequence homology comparison showed that the 3 segment nucleotide of Z10 strain were close to HTN type virus, but only 83.6-87.4% homology with other HTN viruses at the nucleotide level. The phylogenetic analysis was made on their nucleotide and amino acid sequences. CONCLUSION: The results firstly demonstrates that Z10 strain is a new subtype of the Hantaan(HTN) type.

Amino Acid Sequence↗

The complete genomic sequence of strain ROS/HUVLV-100, a representative Russian Crimean Congo hemorrhagic fever virus strain.

The complete genomic sequence (minus primer-generated ends) of the laboratory-adapted Crimean Congo hemorrhagic fever virus (CCHFV) strain ROS/HUVLV-100, isolated in 2003 from the blood of a deceased female from the Rostov region of southern European Russia, was determined by direct sequencing of overlapping reverse transcription/polymerase chain reaction amplified products. The size of the ROS/HUVLV-100 genome is 19.2 kilobases--individual genome segments are similar in size and sequence features to previously reported "Europe-1" group CCHFV strains. The low-passage ROS/HUVLV-100 strain is the first Russian Crimean Congo hemorrhagic fever virus isolate for which complete sequence information is available, and this work reports the first complete genomic CCHFV sequence determined from a single viral RNA preparation in the same laboratory.

Female↗

Determination and analysis of the complete genomic sequence of avian hepatitis E virus (avian HEV) and attempts to infect rhesus monkeys with avian HEV.

Avian hepatitis E virus (avian HEV), recently identified from a chicken with hepatitis-splenomegaly syndrome in the United States, is genetically and antigenically related to human and swine HEVs. In this study, sequencing of the genome was completed and an attempt was made to infect rhesus monkeys with avian HEV. The full-length genome of avian HEV, excluding the poly(A) tail, is 6654 bp in length, which is about 600 bp shorter than that of human and swine HEVs. Similar to human and swine HEV genomes, the avian HEV genome consists of a short 5' non-coding region (NCR) followed by three partially overlapping open reading frames (ORFs) and a 3'NCR. Avian HEV shares about 50 % nucleotide sequence identity over the complete genome, 48-51 % identity in ORF1, 46-48 % identity in ORF2 and only 29-34 % identity in ORF3 with human and swine HEV strains. Significant genetic variations such as deletions and insertions, particularly in ORF1 of avian HEV, were observed. However, motifs in the putative functional domains of ORF1, such as the helicase and methyltransferase, were relatively conserved between avian HEV and mammalian HEVs, supporting the conclusion that avian HEV is a member of the genus Hepevirus. Phylogenetic analysis revealed that avian HEV represents a branch distinct from human and swine HEVs. Swine HEV infects non-human primates and possibly humans and thus may be zoonotic. An attempt was made to determine whether avian HEV also infects across species by experimentally inoculating two rhesus monkeys with avian HEV. Evidence of virus infection was not observed in the inoculated monkeys as there was no seroconversion, viraemia, faecal virus shedding or serum liver enzyme elevation. The results from this study confirmed that avian HEV is related to, but distinct from, human and swine HEVs; however, unlike swine HEV, avian HEV is probably not transmissible to non-human primates.

Amino Acid Sequence↗

The utility of complete genome sequences in the study of pathogenic bacteria.

The availability of complete genome sequences is a revolution in the study of microorganisms. A fully annotated genome sequence provides an interactive tool for scientists and influences the approach and focus of research. In this article I discuss the impact of genome sequencing projects of bacteria. Much useful data have been obtained but the experimental methods needed to fully exploit the information continue to develop. Some of the approaches and particular applications relevant to bacteria of clinical importance are discussed.

Bacteria↗

Complete genome sequence of an aerobic thermoacidophilic crenarchaeon, Sulfolobus tokodaii strain7.

The complete genomic sequence of an aerobic thermoacidophilic crenarchaeon, Sulfolobus tokodaii strain7 which optimally grows at 80 degrees C, at low pH, and under aerobic conditions, has been determined by the whole genome shotgun method with slight modifications. The genomic size was 2,694,756 bp long and the G + C content was 32.8%. The following RNA-coding genes were identified: a single 16S-23S rRNA cluster, one 5S rRNA gene and 46 tRNA genes (including 24 intron-containing tRNA genes). The repetitive sequences identified were SR-type repetitive sequences, long dispersed-type repetitive sequences and Tn-like repetitive elements. The genome contained 2826 potential protein-coding regions (open reading frames, ORFs). By similarity search against public databases, 911 (32.2%) ORFs were related to functional assigned genes, 921 (32.6%) were related to conserved ORFs of unknown function, 145 (5.1%) contained some motifs, and remaining 849 (30.0%) did not show any significant similarity to the registered sequences. The ORFs with functional assignments included the candidate genes involved in sulfide metabolism, the TCA cycle and the respiratory chain. Sequence comparison provided evidence suggesting the integration of plasmid, rearrangement of genomic structure, and duplication of genomic regions that may be responsible for the larger genomic size of the S. tokodaii strain7 genome. The genome contained eukaryote-type genes which were not identified in other archaea and lacked the CCA sequence in the tRNA genes. The result suggests that this strain is closer to eukaryotes among the archaea strains so far sequenced. The data presented in this paper are also available on the internet homepage (http://www.bio.nite.go.jp/E-home/genome_list-e.html/).

Archaeal Proteins↗

Complete genomic sequence of the virulent Salmonella bacteriophage SP6.

We report the complete genome sequence of enterobacteriophage SP6, which infects Salmonella enterica serovar Typhimurium. The genome contains 43,769 bp, including a 174-bp direct terminal repeat. The gene content and organization clearly place SP6 in the coliphage T7 group of phages, but there is approximately 5 kb at the right end of the genome that is not present in other members of the group, and the homologues of T7 genes 1.3 through 3 appear to have undergone an unusual reorganization. Sequence analysis identified 10 putative promoters for the SP6-encoded RNA polymerase and seven putative rho-independent terminators. The terminator following the gene encoding the major capsid subunit has a termination efficiency of about 50% with the SP6-encoded RNA polymerase. Phylogenetic analysis of phages related to SP6 provided clear evidence for horizontal exchange of sequences in the ancestry of these phages and clearly demarcated exchange boundaries; one of the recombination joints lies within the coding region for a phage exonuclease. Bioinformatic analysis of the SP6 sequence strongly suggested that DNA replication occurs in large part through a bidirectional mechanism, possibly with circular intermediates.

Amino Acid Sequence↗

Complete genome sequence of the acetic acid bacterium Gluconobacter oxydans.

Gluconobacter oxydans is unsurpassed by other organisms in its ability to incompletely oxidize a great variety of carbohydrates, alcohols and related compounds. Furthermore, the organism is used for several biotechnological processes, such as vitamin C production. To further our understanding of its overall metabolism, we sequenced the complete genome of G. oxydans 621H. The chromosome consists of 2,702,173 base pairs and contains 2,432 open reading frames. In addition, five plasmids were identified that comprised 232 open reading frames. The sequence data can be used for metabolic reconstruction of the pathways leading to industrially important products derived from sugars and alcohols. Although the respiratory chain of G. oxydans was found to be rather simple, the organism contains many membrane-bound dehydrogenases that are critical for the incomplete oxidation of biotechnologically important substrates. Moreover, the genome project revealed the unique biochemistry of G. oxydans with respect to the process of incomplete oxidation.

Acetic Acid↗

Prediction of transcriptional regulatory sites in the complete genome sequence of Escherichia coli K-12.

MOTIVATION: As one of the best-characterized free-living organisms, Escherichia coli and its recently completed genomic sequence offer a special opportunity to exploit systematically the variety of regulatory data available in the literature in order to make a comprehensive set of regulatory predictions in the whole genome. RESULTS: The complete genome sequence of E.coli was analyzed for the binding of transcriptional regulators upstream of coding sequences. The biological information contained in RegulonDB (Huerta, A.M. et al., Nucleic Acids Res.,26,55-60, 1998) for 56 different transcriptional proteins was the support to implement a stringent strategy combining string search and weight matrices. We estimate that our search included representatives of 15-25% of the total number of regulatory binding proteins in E.coli. This search was performed on the set of 4288 putative regulatory regions, each 450 bp long. Within the regions with predicted sites, 89% are regulated by one protein and 81% involve only one site. These numbers are reasonably consistent with the distribution of experimental regulatory sites. Regulatory sites are found in 603 regions corresponding to 16% of operon regions and 10% of intra-operonic regions. Additional evidence gives stronger support to some of these predictions, including the position of the site, biological consistency with the function of the downstream gene, as well as genetic evidence for the regulatory interaction. The predictions described here were incorporated into the map presented in the paper describing the complete E.coli genome (Blattner,F.R. et al., Science, 277, 1453-1461, 1997). AVAILABILITY: The complete set of predictions in GenBank format is available at the url: http://www. cifn.unam.mx/Computational_Biology/E.coli-predictions CONTACT: ecoli-reg@cifn.unam.mx, collado@cifn.unam.mx

Bacterial Proteins↗

Complete genome sequences of cellular life forms: glimpses of theoretical evolutionary genomics.

The availability of complete genome sequences of cellular life forms creates the opportunity to explore the functional content of the genomes and evolutionary relationships between them at a new qualitative level. With the advent of these sequences, the construction of a minimal gene set sufficient for sustaining cellular life and reconstruction of the genome of the last common ancestor of bacteria, eukaryotes, and archaea become realistic, albeit challenging, research projects. A version of the minimal gene set for modern-type cellular life derived by comparative analysis of two bacterial genomes, those of Haemophilus influenzae and Mycoplasma genitalium, consists of approximately 250 genes. A comparison of the protein sequences encoded in these genes with those of the proteins encoded in the complete yeast genome suggests that the last common ancestor of all extant life might have had an RNA genome.

Bacterial Proteins↗

Complete genome sequence of Bacillus subtilis strain S-LA1, a potential plant probiotic endophyte from the medicinal plant Leucas aspera.

Bacillus subtilis strain S-LA1 is an endophytic bacterium isolated from Leucas aspera roots that harbors a 4.2 Mbp genome predicted to encode several traits for nutrient acquisition, plant growth promotion, and plant probiotic efficacy. Genomic characterization underscores its potential as a microbial resource supporting sustainable agriculture and crop disease management strategies.

Bacillus↗

Complete genome sequence of a novel alternavirus infecting Fusarium falciforme.

We present the complete genome sequence of a novel alternavirus, tentatively named "Fusarium falciforme alternavirus 1 (FfAV1)", isolated from Fusarium falciforme. The host, F. falciforme strain Fod375, was isolated from a soil sample in Spain in 2012 and was found to be infected with a virus containing a tetra-segmented double-stranded (ds) RNA genome. The genome segments, designated as dsRNA1 (3529 bp), dsRNA2 (2641 bp), dsRNA3 (2459 bp), and dsRNA4 (1471 bp), each possess a single open reading frame (ORF). The protein predicted from dsRNA1 contains the typical domains of an RNA-dependent RNA polymerase (RdRP) homologous to those of previously reported alternaviruses, while the protein predicted from dsRNA3 shows homology to alternavirus capsid proteins. The proteins encoded by dsRNA2 and dsRNA4 are of unknown function. All predicted proteins exhibited the highest sequence identity with their counterparts in Hebei alternavirus and Marquandomyces marquandii alternavirus 1. Phylogenetic analysis supported the placement of this FfAV1 isolate within the genus Alternavirus. Considering these results, we propose that FfAV1, along with the two closely related unassigned alternaviruses, represents a new species within the genus.

Genome, Viral↗

The complete genomic sequence of Mycoplasma penetrans, an intracellular bacterial pathogen in humans.

The complete genomic sequence of an intracellular bacterial pathogen, Mycoplasma penetrans HF-2 strain, was determined. The HF-2 genome consists of a 1 358 633 bp single circular chromosome containing 1038 predicted coding sequences (CDSs), one set of rRNA genes and 30 tRNA genes. Among the 1038 CDSs, 264 predicted proteins are common to the Mycoplasmataceae sequenced thus far and 463 are M.penetrans specific. The genome contains the two-component system but lacks the essential cellular gene, uridine kinase. The relatively large genome of M.penetrans HF-2 among mycoplasma species may be accounted for by both its rich core proteome and the presence of a number of paralog families corresponding to 25.4% of all CDSs. The largest paralog family is the p35 family, which encodes surface lipoproteins including the major antigen, P35. A total of 44 genes for p35 and p35 homologs were identified and 30 of them form one large cluster in the chromosome. The genetic tree of p35 paralogs suggests the occurrence of dynamic chromosomal rearrangement in paralog formation during evolution. Thus, M.penetrans HF-2 may have acquired diverse repertoires of antigenic variation-related genes to allow its persistent infection in humans.

Antigenic Variation↗