Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Oat blue dwarf marafivirus resembles the tymoviruses in sequence, genome organization, and expression strategy.

The complete nucleotide sequence and genome organization of oat blue dwarf marafivirus (OBDV) were determined. The 6509 nucleotide RNA genome encodes a putative 227-kDa polyprotein (p227) with sequence motifs similar to the methyltransferase, papain-like protease, helicase, and polymerase motifs present in the nonstructural proteins of other positive strand RNA viruses. The 3' end of the open reading frame (ORF) that encodes p227 (ORF 227) also encodes the two capsid proteins: a 24-kDa capsid protein is presumably cleaved from the p227 polyprotein, whereas the 21-kDa capsid protein appears to be translated from a subgenomic RNA (sgRNA). Encoded amino acid and nucleotide sequence comparisons, as well as the OBDV genome expression strategy, show that OBDV closely resembles the tymoviruses. OBDV differs from the tymoviruses in its general biology, in its lack of a putative movement gene that overlaps the replication-associated genes, and in its fusion of the capsid gene sequences to the major ORF. OBDV also possesses a 3' poly(A) tail, as compared to the tRNA-like structures found in most tymoviral genomes. Due to the strong similarities in genome sequence and expression strategy, OBDV, and presumably the other marafiviruses, should be considered a member of the tymovirus lineage of the alpha-like plant viruses.

Amino Acid Sequence↗

The genome sequence of the probiotic intestinal bacterium Lactobacillus johnsonii NCC 533.

Lactobacillus johnsonii NCC 533 is a member of the acidophilus group of intestinal lactobacilli that has been extensively studied for their "probiotic" activities that include, pathogen inhibition, epithelial cell attachment, and immunomodulation. To gain insight into its physiology and identify genes potentially involved in interactions with the host, we sequenced and analyzed the 1.99-Mb genome of L. johnsonii NCC 533. Strikingly, the organism completely lacked genes encoding biosynthetic pathways for amino acids, purine nucleotides, and most cofactors. In apparent compensation, a remarkable number of uncommon and often duplicated amino acid permeases, peptidases, and phosphotransferase-type transporters were discovered, suggesting a strong dependency of NCC 533 on the host or other intestinal microbes to provide simple monomeric nutrients. Genome analysis also predicted an abundance (>12) of large and unusual cell-surface proteins, including fimbrial subunits, which may be involved in adhesion to glycoproteins or other components of mucin, a characteristic expected to affect persistence in the gastrointestinal tract (GIT). Three bile salt hydrolases and two bile acid transporters, proteins apparently critical for GIT survival, were also detected. In silico genome comparisons with the >95% complete genome sequence of the closely related Lactobacillus gasseri revealed extensive synteny punctuated by clear-cut insertions or deletions of single genes or operons. Many of these regions of difference appear to encode metabolic or structural components that could affect the organisms competitiveness or interactions with the GIT ecosystem.

Biological Transport↗

Assignment of homology to genome sequences using a library of hidden Markov models that represent all proteins of known structure.

Of the sequence comparison methods, profile-based methods perform with greater selectively than those that use pairwise comparisons. Of the profile methods, hidden Markov models (HMMs) are apparently the best. The first part of this paper describes calculations that (i) improve the performance of HMMs and (ii) determine a good procedure for creating HMMs for sequences of proteins of known structure. For a family of related proteins, more homologues are detected using multiple models built from diverse single seed sequences than from one model built from a good alignment of those sequences. A new procedure is described for detecting and correcting those errors that arise at the model-building stage of the procedure. These two improvements greatly increase selectivity and coverage. The second part of the paper describes the construction of a library of HMMs, called SUPERFAMILY, that represent essentially all proteins of known structure. The sequences of the domains in proteins of known structure, that have identities less than 95 %, are used as seeds to build the models. Using the current data, this gives a library with 4894 models. The third part of the paper describes the use of the SUPERFAMILY model library to annotate the sequences of over 50 genomes. The models match twice as many target sequences as are matched by pairwise sequence comparison methods. For each genome, close to half of the sequences are matched in all or in part and, overall, the matches cover 35 % of eukaryotic genomes and 45 % of bacterial genomes. On average roughly 15% of genome sequences are labelled as being hypothetical yet homologous to proteins of known structure. The annotations derived from these matches are available from a public web server at: http://stash.mrc-lmb.cam.ac.uk/SUPERFAMILY. This server also enables users to match their own sequences against the SUPERFAMILY model library.

Amino Acid Sequence↗

The genome sequence of Yaba-like disease virus, a yatapoxvirus.

The genome sequence of Yaba-like disease virus (YLDV), an unclassified member of the yatapoxvirus genus, has been determined. Excluding the terminal hairpin loops, the YLDV genome is 144,575 bp in length and contains inverted terminal repeats (ITRs) of 1883 bp. Within 20 nucleotides of the termini, there is a sequence that is conserved in other poxviruses and is required for the resolution of concatemeric replicative DNA intermediates. The nucleotide composition of the genome is 73% A+T, but the ITRs are only 63% A+T. The genome contains 151 tightly packed open reading frames (ORFs) that either are > or =180 nucleotides in length or are conserved in other poxviruses. ORFs within 23 kb of each end are transcribed toward the termini, whereas ORFs within the central region of the genome are encoded on either DNA strand. In the central region ORFs have a conserved position, orientation, and sequence compared with vaccinia virus ORFs and encode many enzymes, transcription factors, or structural proteins. In contrast, ORFs near the termini are more divergent and in seven cases are without counterparts in other poxviruses. The YLDV genome encodes several predicted immunomodulators; examples include two proteins with similarity to CC chemokine receptors and predicted secreted proteins with similarity to MHC class I antigen, OX-2, interleukin-10/mda-7, poxvirus growth factor, serpins, and a type I interferon-binding protein. Phylogenic analyses indicated that YLDV is very closely related to yaba monkey tumor virus, but outside the yatapoxvirus genus YLDV is more closely related to swinepox virus and leporipoxviruses than to other chordopoxvirus genera.

Adjuvants, Immunologic↗

The genome sequence of Yersinia pestis bacteriophage phiA1122 reveals an intimate history with the coliphage T3 and T7 genomes.

The genome sequence of bacteriophage phiA1122 has been determined. phiA1122 grows on almost all isolates of Yersinia pestis and is used by the Centers for Disease Control and Prevention as a diagnostic agent for the causative agent of plague. phiA1122 is very closely related to coliphage T7; the two genomes are colinear, and the genome-wide level of nucleotide identity is about 89%. However, a quarter of the phiA1122 genome, one that includes about half of the morphogenetic and maturation functions, is significantly more closely related to coliphage T3 than to T7. It is proposed that the yersiniophage phiA1122 recombined with a close relative of the Y. enterocolitica phage phiYeO3-12 to yield progeny phages, one of which became the classic T3 coliphage of Demerec and Fano (M. Demerec and U. Fano, Genetics 30:119-136, 1945).

Amino Acid Sequence↗

Complete genomic sequence of the virulent Salmonella bacteriophage SP6.

We report the complete genome sequence of enterobacteriophage SP6, which infects Salmonella enterica serovar Typhimurium. The genome contains 43,769 bp, including a 174-bp direct terminal repeat. The gene content and organization clearly place SP6 in the coliphage T7 group of phages, but there is approximately 5 kb at the right end of the genome that is not present in other members of the group, and the homologues of T7 genes 1.3 through 3 appear to have undergone an unusual reorganization. Sequence analysis identified 10 putative promoters for the SP6-encoded RNA polymerase and seven putative rho-independent terminators. The terminator following the gene encoding the major capsid subunit has a termination efficiency of about 50% with the SP6-encoded RNA polymerase. Phylogenetic analysis of phages related to SP6 provided clear evidence for horizontal exchange of sequences in the ancestry of these phages and clearly demarcated exchange boundaries; one of the recombination joints lies within the coding region for a phage exonuclease. Bioinformatic analysis of the SP6 sequence strongly suggested that DNA replication occurs in large part through a bidirectional mechanism, possibly with circular intermediates.

Amino Acid Sequence↗

Perspectives: sequence data base searching in the era of large-scale genomic sequencing.

Large-scale sequencing of human and model organism genomes will have a profound impact on our ability to use sequence data base searching to predict the biochemical functions of sequences of interest. Despite the great value of more sequences in the data bases, a huge increase in data base size will also have adverse effects on data base searches. Upcoming problems will include (1) greatly increased search times, (2) an increase in background noise of high-scoring but biologically irrelevant matches, (3) inaccurate coding region prediction, leading to problems in protein data base searching, and (4) limited first-pass sequence annotation, making it difficult to determine the biological relevance of data base hits. Improved data base annotation tools and construction of smaller data bases of representative and highly-annotated sequences for first-pass analyses will be essential to deal with the impending flood of new genomic sequence.

Animals↗

Bacteriophage Mu genome sequence: analysis and comparison with Mu-like prophages in Haemophilus, Neisseria and Deinococcus.

We report the complete 36,717 bp genome sequence of bacteriophage Mu and provide an analysis of the sequence, both with regard to the new genes and other genetic features revealed by the sequence itself and by a comparison to eight complete or nearly complete Mu-like prophage genomes found in the genomes of a diverse group of bacteria. The comparative studies confirm that members of the Mu-related family of phage genomes are genetically mosaic with respect to each other, as seen in other groups of phages such as the phage lambda-related group of phages of enteric hosts and the phage L5-related group of mycobacteriophages. Mu also possesses segments of similarity, typically gene-sized, to genomes of otherwise non-Mu-like phages. The comparisons show that some well-known features of the Mu genome, including the invertible segment encoding tail fiber sequences, are not present in most members of the Mu genome sequence family examined here, suggesting that their presence may be relatively volatile over evolutionary time. The head and tail-encoding structural genes of Mu have only very weak similarity to the corresponding genes of other well-studied phage types. However, these weak similarities, and in some cases biochemical data, can be used to establish tentative functional assignments for 12 of the head and tail genes. These assignments are strongly supported by the fact that the order of gene functions assigned in this way conforms to the strongly conserved order of head and tail genes established in a wide variety of other phages. We show that the Mu head assembly scaffolding protein is encoded by a gene nested in-frame within the C-terminal half of another gene that encodes the putative head maturation protease. This is reminiscent of the arrangement established for phage lambda.

Amino Acid Sequence↗

The complete genomic sequence of strain ROS/HUVLV-100, a representative Russian Crimean Congo hemorrhagic fever virus strain.

The complete genomic sequence (minus primer-generated ends) of the laboratory-adapted Crimean Congo hemorrhagic fever virus (CCHFV) strain ROS/HUVLV-100, isolated in 2003 from the blood of a deceased female from the Rostov region of southern European Russia, was determined by direct sequencing of overlapping reverse transcription/polymerase chain reaction amplified products. The size of the ROS/HUVLV-100 genome is 19.2 kilobases--individual genome segments are similar in size and sequence features to previously reported "Europe-1" group CCHFV strains. The low-passage ROS/HUVLV-100 strain is the first Russian Crimean Congo hemorrhagic fever virus isolate for which complete sequence information is available, and this work reports the first complete genomic CCHFV sequence determined from a single viral RNA preparation in the same laboratory.

Female↗

Whole-genome sequence variation among multiple isolates of Pseudomonas aeruginosa.

Whole-genome shotgun sequencing was used to study the sequence variation of three Pseudomonas aeruginosa isolates, two from clonal infections of cystic fibrosis patients and one from an aquatic environment, relative to the genomic sequence of reference strain PAO1. The majority of the PAO1 genome is represented in these strains; however, at least three prominent islands of PAO1-specific sequence are apparent. Conversely, approximately 10% of the sequencing reads derived from each isolate fail to align with the PAO1 backbone. While average sequence variation among all strains is roughly 0.5%, regions of pronounced differences were evident in whole-genome scans of nucleotide diversity. We analyzed two such divergent loci, the pyoverdine and O-antigen biosynthesis regions, by complete resequencing. A thorough analysis of isolates collected over time from one of the cystic fibrosis patients revealed independent mutations resulting in the loss of O-antigen synthesis alternating with a mucoid phenotype. Overall, we conclude that most of the PAO1 genome represents a core P. aeruginosa backbone sequence while the strains addressed in this study possess additional genetic material that accounts for at least 10% of their genomes. Approximately half of these additional sequences are novel.

Adolescent↗

The genome sequence of the common green lacewing, Chrysoperla carnea (Stephens, 1836).

We present a genome assembly from an individual female Chrysoperla carnea (a common green lacewing; Arthropoda; Insecta; Neuroptera; Chrysopidae). The genome sequence is 560 megabases in span. The majority of the assembly (95.70%) is scaffolded into six chromosomal pseudomolecules, with the X sex chromosome assembled. Gene annotation of this assembly by the NCBI Eukaryotic Genome Annotation Pipeline has identified 12,985 protein coding genes.

Chrysoperla carnea↗

Determination and analysis of the complete genomic sequence of avian hepatitis E virus (avian HEV) and attempts to infect rhesus monkeys with avian HEV.

Avian hepatitis E virus (avian HEV), recently identified from a chicken with hepatitis-splenomegaly syndrome in the United States, is genetically and antigenically related to human and swine HEVs. In this study, sequencing of the genome was completed and an attempt was made to infect rhesus monkeys with avian HEV. The full-length genome of avian HEV, excluding the poly(A) tail, is 6654 bp in length, which is about 600 bp shorter than that of human and swine HEVs. Similar to human and swine HEV genomes, the avian HEV genome consists of a short 5' non-coding region (NCR) followed by three partially overlapping open reading frames (ORFs) and a 3'NCR. Avian HEV shares about 50 % nucleotide sequence identity over the complete genome, 48-51 % identity in ORF1, 46-48 % identity in ORF2 and only 29-34 % identity in ORF3 with human and swine HEV strains. Significant genetic variations such as deletions and insertions, particularly in ORF1 of avian HEV, were observed. However, motifs in the putative functional domains of ORF1, such as the helicase and methyltransferase, were relatively conserved between avian HEV and mammalian HEVs, supporting the conclusion that avian HEV is a member of the genus Hepevirus. Phylogenetic analysis revealed that avian HEV represents a branch distinct from human and swine HEVs. Swine HEV infects non-human primates and possibly humans and thus may be zoonotic. An attempt was made to determine whether avian HEV also infects across species by experimentally inoculating two rhesus monkeys with avian HEV. Evidence of virus infection was not observed in the inoculated monkeys as there was no seroconversion, viraemia, faecal virus shedding or serum liver enzyme elevation. The results from this study confirmed that avian HEV is related to, but distinct from, human and swine HEVs; however, unlike swine HEV, avian HEV is probably not transmissible to non-human primates.

Amino Acid Sequence↗

Optimized genomic sequencing as a tool for the study of cytosine methylation in the regulatory region of the chicken vitellogenin II gene.

Some critical parameters of genomic sequencing are described. As an example we show a 200-nucleotide (nt) sequence of the estradiol-regulated avian vitellogenin gene II upstream region containing four CpG nt pairs. The two CpG's at positions -612 and -618 are in the sequence binding estradiol-receptor complex. While all four CpG's are methylated in erythrocytes, they are hypomethylated in the DNA of estradiol-responsive organs of egg-laying hens. A simple electroblot DNA transfer system which gives no distortion of DNA bands and quantitative transfer of denatured DNA from the gel to the filter membrane is described. Using Gene Screen membranes, a maximal hybridization signal was obtained when 30-50% of the input DNA was stably bound to the filters. Hybridization background signal and exposure time could be largely reduced by using highly purified fractionated DNA. Using a 90-120 nt long homogeneous single-stranded DNA probe of high specific activity it was possible to read a genomic sequence of up to 200 nt. The resolution was further improved by reducing the extent of chemical modifications of the DNA during the Maxam-Gilbert sequencing reactions.

Animals↗

Drosophila genomic sequence annotation using the BLOCKS+ database.

A simple and general homology-based method for gene finding was applied to the 2.9-Mb Drosophila melanogaster Adh region, the target sequence of the Genome Annotation Assessment Project (GASP). Each strand of the entire sequence was used as query of the BLOCKS+ database of conserved regions of proteins. This led to functional assignments for more than one-third of the genes and two-thirds of the transposons. Considering the enormous size of the query, the fact that only two false-positive matches were reported emphasizes the high selectivity of protein family-based methods for gene finding. We used the search results to improve BLOCKS+ by identifying compositionally biased blocks. Our results confirm that protein family databases can be used effectively in automated sequence annotation efforts.

Alcohol Dehydrogenase↗

Genome sequence and attenuating mutations in West Nile virus isolate from Mexico.

The complete genome sequence of a Mexican West Nile virus isolate, TM171-03, included 46 nucleotide (0.42%) and 4 amino acid (0.11%) differences from the NY99 prototype. Mouse virulence differences between plaque-purified variants of TM171-03 with mutations at the E protein glycosylation motif suggest the emergence of an attenuating mutation.

Animals↗

Complete genome sequence of the marine planctomycete Pirellula sp. strain 1.

Pirellula sp. strain 1 ("Rhodopirellula baltica") is a marine representative of the globally distributed and environmentally important bacterial order Planctomycetales. Here we report the complete genome sequence of a member of this independent phylum. With 7.145 megabases, Pirellula sp. strain 1 has the largest circular bacterial genome sequenced so far. The presence of all genes required for heterolactic acid fermentation, key genes for the interconversion of C1 compounds, and 110 sulfatases were unexpected for this aerobic heterotrophic isolate. Although Pirellula sp. strain 1 has a proteinaceous cell wall, remnants of genes for peptidoglycan synthesis were found. Genes for lipid A biosynthesis and homologues to the flagellar L- and P-ring protein indicate a former Gram-negative type of cell wall. Phylogenetic analysis of all relevant markers clearly affiliates the Planctomycetales to the domain Bacteria as a distinct phylum, but a deepest branching is not supported by our analyses.

Adaptation, Physiological↗

Draft genome sequence of Breoghania corrubedonensis DSM 23382T.

We report the genome sequence of Breoghania corrubedonensis DSM 23382T isolated from oil-spill contaminated beach sand. The 5,331,589-bp genome with 63.62% G + C encodes 4,746 genes. This reference genome will facilitate experimental studies investigating B. corrubedonensis's role in oil-contaminated marine environments, particularly oil degradation or resistivity.

computational biology↗

Genome sequence of Oceanobacillus iheyensis isolated from the Iheya Ridge and its unexpected adaptive capabilities to extreme environments.

Oceanobacillus iheyensis HTE831 is an alkaliphilic and extremely halotolerant Bacillus-related species isolated from deep-sea sediment. We present here the complete genome sequence of HTE831 along with analyses of genes required for adaptation to highly alkaline and saline environments. The genome consists of 3.6 Mb, encoding many proteins potentially associated with roles in regulation of intracellular osmotic pressure and pH homeostasis. The candidate genes involved in alkaliphily were determined based on comparative analysis with three Bacillus species and two other Gram-positive species. Comparison with the genomes of other major Gram-positive bacterial species suggests that the backbone of the genus Bacillus is composed of approximately 350 genes. This second genome sequence of an alkaliphilic Bacillus-related species will be useful in understanding life in highly alkaline environments and microbial diversity within the ubiquitous bacilli.

Bacillus↗