Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Genome sequence of Chlamydophila caviae (Chlamydia psittaci GPIC): examining the role of niche-specific genes in the evolution of the Chlamydiaceae.

The genome of Chlamydophila caviae (formerly Chlamydia psittaci, GPIC isolate) (1 173 390 nt with a plasmid of 7966 nt) was determined, representing the fourth species with a complete genome sequence from the Chlamydiaceae family of obligate intracellular bacterial pathogens. Of 1009 annotated genes, 798 were conserved in all three other completed Chlamydiaceae genomes. The C.caviae genome contains 68 genes that lack orthologs in any other completed chlamydial genomes, including tryptophan and thiamine biosynthesis determinants and a ribose-phosphate pyrophosphokinase, the product of the prsA gene. Notable amongst these was a novel member of the virulence-associated invasin/intimin family (IIF) of Gram-negative bacteria. Intriguingly, two authentic frameshift mutations in the ORF indicate that this gene is not functional. Many of the unique genes are found in the replication termination region (RTR or plasticity zone), an area of frequent symmetrical inversion events around the replication terminus shown to be a hotspot for genome variation in previous genome sequencing studies. In C.caviae, the RTR includes several loci of particular interest including a large toxin gene and evidence of ancestral insertion(s) of a bacteriophage. This toxin gene, not present in Chlamydia pneumoniae, is a member of the YopT effector family of type III-secreted cysteine proteases. One gene cluster (guaBA-add) in the RTR is much more similar to orthologs in Chlamydia muridarum than those in the phylogenetically closest species C.pneumoniae, suggesting the possibility of horizontal transfer of genes between the rodent-associated Chlamydiae. With most genes observed in the other chlamydial genomes represented, C.caviae provides a good model for the Chlamydiaceae and a point of comparison against the human atherosclerosis-associated C.pneumoniae. This crucial addition to the set of completed Chlamydiaceae genome sequences is enabling dissection of the roles played by niche-specific genes in these important bacterial pathogens.

Adhesins, Bacterial↗

The Escherichia coli genome sequence: the end of an era or the start of the FUN?

Our dream of determining the entire Escherichia coli K12 genome sequence has been realized. This calls for new approaches for the analysis of gene expression and function in biology's best-understood organism. Comparison of the E. coli genome sequence with others will provide important taxonomic insights and have implications for the study of bacterial virulence. Approximately 20% of E. coli genes have been designated FUN genes, because they have no known function or homologies to sequence databases. FUN genes promise to have an exciting impact on bacterial research. The post-genome era requires novel strategies that address gene regulation at the level of the entire cell. These strategies need to supersede the reductionist approach to genetic analysis. Only then will the genome sequence lead us to an understanding of how a bacterial cell really works.

Escherichia coli↗

Genome sequence of the cyanobacterium Prochlorococcus marinus SS120, a nearly minimal oxyphototrophic genome.

Prochlorococcus marinus, the dominant photosynthetic organism in the ocean, is found in two main ecological forms: high-light-adapted genotypes in the upper part of the water column and low-light-adapted genotypes at the bottom of the illuminated layer. P. marinus SS120, the complete genome sequence reported here, is an extremely low-light-adapted form. The genome of P. marinus SS120 is composed of a single circular chromosome of 1,751,080 bp with an average G+C content of 36.4%. It contains 1,884 predicted protein-coding genes with an average size of 825 bp, a single rRNA operon, and 40 tRNA genes. Together with the 1.66-Mbp genome of P. marinus MED4, the genome of P. marinus SS120 is one of the two smallest genomes of a photosynthetic organism known to date. It lacks many genes that are involved in photosynthesis, DNA repair, solute uptake, intermediary metabolism, motility, phototaxis, and other functions that are conserved among other cyanobacteria. Systems of signal transduction and environmental stress response show a particularly drastic reduction in the number of components, even taking into account the small size of the SS120 genome. In contrast, housekeeping genes, which encode enzymes of amino acid, nucleotide, cofactor, and cell wall biosynthesis, are all present. Because of its remarkable compactness, the genome of P. marinus SS120 might approximate the minimal gene complement of a photosynthetic organism.

Adaptation, Physiological↗

Whole genome sequence data of an infectious molecular clone of the SIVagm TYO-1 strain.

We first sequenced a full genome of simian immunodeficiency virus isolated from African green monkey (SIVagm) but the clone sequenced was found not to be biologically active. We subsequently succeeded in reconstructing a full genome infectious molecular clone, named pSA212. The infectious pSA212 clone (known as the TYO-1 strain of SIVagm) has been distributed widely for research analysis of SIVagm but its genome has never been fully sequenced. Here, we report the whole genome sequence of the infectious pSA212.

Amino Acid Sequence↗

Functional genomics in Arabidopsis: large-scale insertional mutagenesis complements the genome sequencing project.

The ultimate goal of genome research on the model flowering plant Arabidopsis thaliana is the identification of all of the genes and understanding their functions. A major step towards this goal, the genome sequencing project, is nearing completion; however, functional studies of newly discovered genes have not yet kept up to this pace. Recent progress in large-scale insertional mutagenesis opens new possibilities for functional genomics in Arabidopsis. The number of T-DNA and transposon insertion lines from different laboratories will soon represent insertions into most Arabidopsis genes. Vast resources of gene knockouts are becoming available that can be subjected to different types of reverse genetics screens to deduce the functions of the sequenced genes.

Arabidopsis↗

Complete nucleotide sequences of ALV-related endogenous retroviruses available from the draft chicken genome sequence.

Complete nucleotide sequences of chicken endogenous retroviruses belonging to E33/E51 and EAV-0 groups have been analysed on the basis of the recently available draft genome sequence of red jungle fowl (Gallus gallus), the progenitor of domestic chicken (G.g. domesticus). It was shown that all these proviruses have deletions in the SU-coding domain of the env gene, involved in receptor recognition, whereas gag and pol genes appear to be intact. Phylogenetic analysis demonstrated that E33/E51 and EAV-0 groups are related to the ALV genus. An analysis of expression using chicken EST databases showed that these proviruses are transcriptionally active.

Alpharetrovirus↗

The Plasmodium vivax genome sequencing project.

With the successful completion of the project to sequence the Plasmodium falciparum genome, researchers are now turning their attention to other malaria parasite species. Here, an update on the Plasmodium vivax genome sequencing project is presented, as part of the Trends in Parasitology series of reviews expanding on various aspects of P. vivax research.

Animals↗

Integration of telomere sequences with the draft human genome sequence.

Telomeres are the ends of linear eukaryotic chromosomes. To ensure that no large stretches of uncharacterized DNA remain between the ends of the human working draft sequence and the ends of each chromosome, we would need to connect the sequences of the telomeres to the working draft sequence. But telomeres have an unusual DNA sequence composition and organization that makes them particularly difficult to isolate and analyse. Here we use specialized linear yeast artificial chromosome clones, each carrying a large telomere-terminal fragment of human DNA, to integrate most human telomeres with the working draft sequence. Subtelomeric sequence structure appears to vary widely, mainly as a result of large differences in subtelomeric repeat sequence abundance and organization at individual telomeres. Many subtelomeric regions appear to be gene-rich, matching both known and unknown expressed genes. This indicates that human subtelomeric regions are not simply buffers of nonfunctional 'junk DNA' next to the molecular telomere, but are instead functional parts of the expressed genome.

Chromosomes, Artificial, Bacterial↗

Toward a complete human genome sequence.

We have begun a joint program as part of a coordinated international effort to determine a complete human genome sequence. Our strategy is to map large-insert bacterial clones and to sequence each clone by a random shotgun approach followed by directed finishing. As of September 1998, we have identified the map positions of bacterial clones covering approximately 860 Mb for sequencing and completed >98 Mb ( approximately 3.3%) of the human genome sequence. Our progress and sequencing data can be accessed via the World Wide Web (http://webace.sanger.ac.uk/HGP/ or http://genome.wustl.edu/gsc/).

Contig Mapping↗

Draft genome sequence of the sexually transmitted pathogen Trichomonas vaginalis.

We describe the genome sequence of the protist Trichomonas vaginalis, a sexually transmitted human pathogen. Repeats and transposable elements comprise about two-thirds of the approximately 160-megabase genome, reflecting a recent massive expansion of genetic material. This expansion, in conjunction with the shaping of metabolic pathways that likely transpired through lateral gene transfer from bacteria, and amplification of specific gene families implicated in pathogenesis and phagocytosis of host proteins may exemplify adaptations of the parasite during its transition to a urogenital environment. The genome sequence predicts previously unknown functions for the hydrogenosome, which support a common evolutionary origin of this unusual organelle with mitochondria.

Animals↗

Protein identification from two-dimensional gel electrophoresis analysis of Klebsiella pneumoniae by combined use of mass spectrometry data and raw genome sequences.

Separation of proteins by two-dimensional gel electrophoresis (2-DE) coupled with identification of proteins through peptide mass fingerprinting (PMF) by matrix-assisted laser desorption ionization time-of-flight mass spectrometry (MALDI-TOF MS) is the widely used technique for proteomic analysis. This approach relies, however, on the presence of the proteins studied in public-accessible protein databases or the availability of annotated genome sequences of an organism. In this work, we investigated the reliability of using raw genome sequences for identifying proteins by PMF without the need of additional information such as amino acid sequences. The method is demonstrated for proteomic analysis of Klebsiella pneumoniae grown anaerobically on glycerol. For 197 spots excised from 2-DE gels and submitted for mass spectrometric analysis 164 spots were clearly identified as 122 individual proteins. 95% of the 164 spots can be successfully identified merely by using peptide mass fingerprints and a strain-specific protein database (ProtKpn) constructed from the raw genome sequences of K. pneumoniae. Cross-species protein searching in the public databases mainly resulted in the identification of 57% of the 66 high expressed protein spots in comparison to 97% by using the ProtKpn database. 10 dha regulon related proteins that are essential for the initial enzymatic steps of anaerobic glycerol metabolism were successfully identified using the ProtKpn database, whereas none of them could be identified by cross-species searching. In conclusion, the use of strain-specific protein database constructed from raw genome sequences makes it possible to reliably identify most of the proteins from 2-DE analysis simply through peptide mass fingerprinting.

Journal Article↗

DNA methylation mapping by tag-modified bisulfite genomic sequencing.

A tag-modified bisulfite genomic sequencing (tBGS) method employing direct cycle sequencing of polymerase chain reaction (PCR) products at kilobase scale, without conventional DNA fragment cloning, was developed for simplified evaluation of DNA methylation sites. The method entails subjecting bisulfite-modified genomic DNA to a second-round PCR amplification employing GC-tagged primers. Qualitative results from tBGS closely correlated with those from conventional BGS (R=0.935, p=0.002). In application, the intertissue and interindividual CpG methylation differences in promoter sequence for two genes, CYP1B1 and GSTP1, were then explored across four human tissue types (peripheral blood cells, exfoliated buccal cells, paired nontumor-tumor lung tissues), and two lung cell types in culture (normal NHBE and malignant A549). Predominantly conserved methylation maps for the two gene promoters were apparent across donors and tissues. At any given CpG site, variation in the degree of methylation could be determined by the relative height of C and T peaks in the sequencing trace. Methylation maps for the GSTP1 promoter diverged between NHBE (unmethylated) and A549 (completely methylated) cells in a previously unexplored upstream region, correlating with a 2.7-fold difference in GSTP1 mRNA expression (p<0.01). The tBGS method simplifies detailed methylation scanning of kilobase-scale genomic DNA, facilitating more ambitious genomic methylation mapping studies.

Base Sequence↗

Repeat-associated phase variable genes in the complete genome sequence of Neisseria meningitidis strain MC58.

Phase variation, mediated through variation in the length of simple sequence repeats, is recognized as an important mechanism for controlling the expression of factors involved in bacterial virulence. Phase variation is associated with most of the currently recognized virulence determinants of Neisseria meningitidis. Based upon the complete genome sequence of the N. meningitidis serogroup B strain MC58, we have identified tracts of potentially unstable simple sequence repeats and their potential functional significance determined on the basis of sequence context. Of the 65 potentially phase variable genes identified, only 13 were previously recognized. Comparison with the sequences from the other two pathogenic Neisseria sequencing projects shows differences in the length of the repeats in 36 of the 65 genes identified, including 25 of those not previously known to be phase variable. Six genes that did not have differences in the length of the repeat instead had polymorphisms such that the gene would not be expected to be phase variable in at least one of the other strains. A further 12 candidates did not have homologues in either of the other two genome sequences. The large proportion of these genes that are associated with frameshifts and with differences in repeat length between the neisserial genome sequences is further corroborative evidence that they are phase variable. The number of potentially phase variable genes is substantially greater than for any other species studied to date, and would allow N. meningitidis to generate a very large repertoire of phenotypes through expression of these genes in different combinations. Novel phase variable candidates identified in the strain MC58 genome sequence include a spectrum of genes encoding glycosyltransferases, toxin related products, and metabolic activities as well as several restriction/modification and bacteriocin-related genes and a number of open reading frames (ORFs) for which the function is currently unknown. This suggests that the potential role of phase variation in mediating bacterium-host interactions is much greater than has been appreciated to date. Analysis of the distribution of homopolymeric tract lengths indicates that this species has sequence-specific mutational biases that favour the instability of sequences associated with phase variation.

Bacterial Proteins↗

The malaria genome sequencing project.

An international consortium of genome centres, advanced development teams and funding agencies has begun the task of sequencing the genome of the parasite Plasmodium falciparum, the most important cause of human malaria. Sequencing is proceeding chromosome by chromosome, and the annotated sequence of chromosome 2 is nearly finished. With the continual release of sequence data as they are generated, malaria researchers have access to a steady stream of genomic sequences and will soon have the complete annotation of all of the estimated 5000-7000 P. falciparum genes. The task will then be how to best apply these data to the development of new anti-malarial drugs, vaccines and diagnostic tests. This review provides a brief overview of the Malaria Genome Sequencing Project and suggests potential directions for future malaria research.

Journal Article↗

Genomic sequencing reveals the structure of the Kcnk6 and map3k11 genes and their close vicinity to the sipa1 gene on mouse chromosome 19.

In this report we present the analysis of two overlapping mouse cosmid clones that contain the entire Kcnk6, Map3k11 and Pcnxl3 genes, as well as part of the Sipa1 gene. The sequence and genomic organisation of the Kcnk6 and Map3k11 genes are described in detail. Sipa1 and Map3k11, which have independently been mapped with low resolution to the centromeric region of mouse chromosome 19, are shown here to lie close to each other and to the Kcnk6 gene, which has not previously been mapped. This gene cluster maps to the vicinity of the Dancer (Dc) mutation, which involves inner ear abnormalities and circling phenotypes. Since potassium channels have been implicated in deafness disorders, we have analysed the Kcnk6 gene, which encodes a two-P domain potassium channel, in the Dc mutant. No Dc-causing mutation in the Kcnk6 coding region could be identified. However, we detected a polymorphism in the Kcnk6 gene that leads to a C-terminal extension of the encoded protein by eight amino acids.

Actins↗