Search PubMed⌕ Search

Biomedical subjects

C M Fraser

Publications and source records attributed to C M Fraser.

At least 37 records · Page 2Linked to original sources

Microbial genome sequencing.

Complete genome sequences of 30 microbial species have been determined during the past five years, and work in progress indicates that the complete sequences of more than 100 further microbial species will be available in the next two to four years. These results have revealed a tremendous amount of information on the physiology and evolution of microbial species, and should provide novel approaches to the diagnosis and treatment of infectious disease.

Anti-Infective Agents↗

DNA sequence of both chromosomes of the cholera pathogen Vibrio cholerae.

Here we determine the complete genomic sequence of the gram negative, gamma-Proteobacterium Vibrio cholerae El Tor N16961 to be 4,033,460 base pairs (bp). The genome consists of two circular chromosomes of 2,961,146 bp and 1,072,314 bp that together encode 3,885 open reading frames. The vast majority of recognizable genes for essential cell functions (such as DNA replication, transcription, translation and cell-wall biosynthesis) and pathogenicity (for example, toxins, surface antigens and adhesins) are located on the large chromosome. In contrast, the small chromosome contains a larger fraction (59%) of hypothetical genes compared with the large chromosome (42%), and also contains many more genes that appear to have origins other than the gamma-Proteobacteria. The small chromosome also carries a gene capture system (the integron island) and host 'addiction' genes that are typically found on plasmids; thus, the small chromosome may have originally been a megaplasmid that was captured by an ancestral Vibrio species. The V. cholerae genomic sequence provides a starting point for understanding how a free-living, environmental organism emerged to become a significant human bacterial pathogen.

Base Sequence↗

Theileria parva genomics reveals an atypical apicomplexan genome.

The discipline of genomics is setting new paradigms in research approaches to resolving problems in human and animal health. We propose to determine the genome sequence of Theileria parva, a pathogen of cattle, using the random shotgun approach pioneered at The Institute for Genomic Research (TIGR). A number of features of the T. parva genome make it particularly suitable for this approach. The G+C content of genomic DNA is about 31%, non-coding repetitive DNA constitutes less than 1% of total DNA and a framework for the 10-12 Mbp genome is available in the form of a physical map for all four chromosomes. Minisatellite sequences are the only dispersed repetitive sequences identified so far, but they are limited in distribution to 13 of 33 SfiI fragments. Telomere and sub-telomeric non-coding sequences occupy less than 10 kbp at each chromosomal end and there are only two units encoding cytoplasmic rRNAs. Three sets of distinct multicopy sequences encoding ORFs have been identified but it is not known if these are associated with expression of parasite antigenic diversity. Protein coding genes exhibit a bias in codon usage and introns when present are unusually short. Like other apicomplexan organisms, T. parva contains two extrachromosomal DNAs, a mitochondrial DNA and a plastid DNA molecule. By annotating the genome sequence, in combination with the use of microarray technology and comparative genomics, we expect to gain significant insights into unique aspects of the biology of T. parva. We believe that the data will underpin future research to aid in the identification of targets of protective CD8+ cell mediated immune responses, and parasite molecules involved in inducing reversible host leukocyte transformation and tumour-like behaviour of transformed parasitised cells.

Animals↗

Genome sequences of Chlamydia trachomatis MoPn and Chlamydia pneumoniae AR39.

The genome sequences of Chlamydia trachomatis mouse pneumonitis (MoPn) strain Nigg (1 069 412 nt) and Chlamydia pneumoniae strain AR39 (1 229 853 nt) were determined using a random shotgun strategy. The MoPn genome exhibited a general conservation of gene order and content with the previously sequenced C.trachomatis serovar D. Differences between C.trachomatis strains were focused on an approximately 50 kb 'plasticity zone' near the termination origins. In this region MoPn contained three copies of a novel gene encoding a >3000 amino acid toxin homologous to a predicted toxin from Escherichia coli O157:H7 but had apparently lost the tryptophan biosyntheis genes found in serovar D in this region. The C. pneumoniae AR39 chromosome was >99.9% identical to the previously sequenced C.pneumoniae CWL029 genome, however, comparative analysis identified an invertible DNA segment upstream of the uridine kinase gene which was in different orientations in the two genomes. AR39 also contained a novel 4524 nt circular single-stranded (ss)DNA bacteriophage, the first time a virus has been reported infecting C. pneumoniae. Although the chlamydial genomes were highly conserved, there were intriguing differences in key nucleotide salvage pathways: C.pneumoniae has a uridine kinase gene for dUTP production, MoPn has a uracil phosphororibosyl transferase, while C.trachomatis serovar D contains neither gene. Chromosomal comparison revealed that there had been multiple large inversion events since the species divergence of C.trachomatis and C.pneumoniae, apparently oriented around the axis of the origin of replication and the termination region. The striking synteny of the Chlamydia genomes and prevalence of tandemly duplicated genes are evidence of minimal chromosome rearrangement and foreign gene uptake, presumably owing to the ecological isolation of the obligate intracellular parasites. In the absence of genetic analysis, comparative genomics will continue to provide insight into the virulence mechanisms of these important human pathogens.

Animals↗

Complete genome sequence of Neisseria meningitidis serogroup B strain MC58.

The 2,272,351-base pair genome of Neisseria meningitidis strain MC58 (serogroup B), a causative agent of meningitis and septicemia, contains 2158 predicted coding regions, 1158 (53.7%) of which were assigned a biological role. Three major islands of horizontal DNA transfer were identified; two of these contain genes encoding proteins involved in pathogenicity, and the third island contains coding sequences only for hypothetical proteins. Insights into the commensal and virulence behavior of N. meningitidis can be gleaned from the genome, in which sequences for structural proteins of the pilus are clustered and several coding regions unique to serogroup B capsular polysaccharide synthesis can be identified. Finally, N. meningitidis contains more genes that undergo phase variation than any pathogen studied to date, a mechanism that controls their expression and contributes to the evasion of the host immune system.

Antigenic Variation↗

A novel lipothrixvirus, SIFV, of the extremely thermophilic crenarchaeon Sulfolobus.

We describe a novel lipothrixvirus, SIFV, of the crenarchaeotal archaeon Sulfolobus islandicus. SIFV (S. islandicus filamentous virus) has a linear virion with a linear double-stranded DNA genome. These two features coincide in several crenarchaeotal but not in any other viruses. The SIFV core is formed by a zipper-like array of DNA-associated protein subunits and is covered by a lipid envelope containing host lipids. We sequenced approximately 96% of the virus genome excepting the DNA termini, which were modified in an unusual, yet uncharacterized, manner. Both, the 5' and the 3' DNA termini were insensitive to enzymatic degradation and labelling. Two open reading frames (ORFs) of the SIFV genome are likely to encode helicases and resemble uncharacterized ORFs from other archaea in sequence. Three ORFs showed sequence similarity with each other and each contained a glycosyl transferase motif. Another ORF of the SIFV genome showed significant sequence similarity to the ORF a291 from the well characterized, spindle-shaped Sulfolobus virus SSV1. Due to its structure, SIFV is classified as a lipothrixvirus.

Amino Acid Sequence↗

Genome data: what do we learn?

Genome sequence information has continued to accumulate at a spectacular pace during the past year. Details of the sequence and gene content of human chromosome 22 were published. The sequencing and annotation of the first two Arabidopsis thaliana chromosomes was completed. The sequence of chromosome 3 from Plasmodium falciparum, the second sequenced malaria chromosome, was reported, as was that of chromosome 1 from Leishmania major. The complete genomic sequences of five microbes were reported. Approaches to using data from completely sequenced microbial genomes in phylogenetic studies are being explored, as is the application of microarrays to whole genome expression analysis.

Animals↗

Status of genome projects for nonpathogenic bacteria and archaea.

Since the first microbial genome was sequenced in 1995, 30 others have been completed and an additional 99 are known to be in progress. Although the early emphasis of microbial genomics was on human pathogens for obvious reasons, a significant number of sequencing projects have focused on nonpathogenic organisms, beginning with the release of the complete genome sequence of the archaeon Methanococcus jannaschii in 1996. The past 18 months have seen the completion of the genomes of several unusual organisms, including Thermotoga maritima, whose genome reveals extensive potential lateral transfer with archaea; Deinococcus radiodurans, the most radiation-resistant microorganism known; and Aeropyrum pernix, the first Crenarchaeota to be completely sequenced. Although the functional characterization of genomic data is still in its initial stages, it is likely that microbial genomics will have a significant impact on environmental, food, and industrial biotechnology as well as on genomic medicine.

Archaea↗

A bacterial genome in flux: the twelve linear and nine circular extrachromosomal DNAs in an infectious isolate of the Lyme disease spirochete Borrelia burgdorferi.

We have determined that Borrelia burgdorferi strain B31 MI carries 21 extrachromosomal DNA elements, the largest number known for any bacterium. Among these are 12 linear and nine circular plasmids, whose sequences total 610 694 bp. We report here the nucleotide sequence of three linear and seven circular plasmids (comprising 290 546 bp) in this infectious isolate. This completes the genome sequencing project for this organism; its genome size is 1 521 419 bp (plus about 2000 bp of undetermined telomeric sequences). Analysis of the sequence implies that there has been extensive and sometimes rather recent DNA rearrangement among a number of the linear plasmids. Many of these events appear to have been mediated by recombinational processes that formed duplications. These many regions of similarity are reflected in the fact that most plasmid genes are members of one of the genome's 161 paralogous gene families; 107 of these gene families, which vary in size from two to 41 members, contain at least one plasmid gene. These rearrangements appear to have contributed to a surprisingly large number of apparently non-functional pseudogenes, a very unusual feature for a prokaryotic genome. The presence of these damaged genes suggests that some of the plasmids may be in a period of rapid evolution. The sequence predicts 535 plasmid genes >/=300 bp in length that may be intact and 167 apparently mutationally damaged and/or unexpressed genes (pseudogenes). The large majority, over 90%, of genes on these plasmids have no convincing similarity to genes outside Borrelia, suggesting that they perform specialized functions.

Base Sequence↗

Enterococcus faecalis conjugative plasmid pAM373: complete nucleotide sequence and genetic analyses of sex pheromone response.

pAM373 is a 36.7 kb conjugative plasmid in Enterococcus faecalis that encodes a response to a peptide sex pheromone, cAM373, secreted by plasmid-free (recipient) strains of enterococci. It was identified over 15 years ago as one of five plasmids in E. faecalis strain RC73 and was of interest because a related pheromone activity could be detected in culture supernatants of Staphylococcus aureus and Streptococcus gordonii. Because of increased clinical concern relating to the possibility of mobilizing vancomycin resistance determinants from enterococci, where they are becoming common, into pathogens such as S. aureus, efforts were initiated to characterize pAM373 further. The results of a complete nucleotide sequence determination of pAM373, as well as a genetic analysis of key genes related to regulation of the pheromone response, are reported here. With regard to determinants related to conjugation, the plasmid has a structural organization similar to other known pheromone-responsive plasmids such as pAD1, pCF10 and pPD1; however, there are several unique features. Although there are significant homologues relating to a pheromone-binding surface protein (TraC) and a negatively regulating protein (TraA), there is an absence of a determinant equivalent to traB of pAD1 (reduces endogenous pheromone) and a determinant for surface-exclusion protein. The precursor structure of the inhibitor peptide iAM373 was identified, and its determinant (iam373) was found to be about 500 nt upstream of an apparent transcription terminator t1. Tn917-lac insertion analyses provided interesting insights into aspects of control of the pheromone response and showed that, although the traA product is sensitive to pheromone, it appears to act differently from the traA homologue of pAD1.

Adhesins, Bacterial↗

Comparative analysis of Chlamydia bacteriophages reveals variation localized to a putative receptor binding domain.

Three recently discovered ssDNA Chlamydia-infecting microviruses, phiCPG1, phiAR39, and Chp2, were compared with the previously characterized phage from avian C. psittaci, Chp1. Although the four bacteriophages share an identical arrangement of their five main genes, Chpl has diverged significantly in its nucleotide and protein sequences from the other three, which form a closely related group. The VP1 major viral capsid proteins of phiCPG1 and phiAR39 (from guinea pig-infecting C. psittaci and C. pneumoniae, respectively) are almost identical. However, VP1 of ovine C. psittaci phage Chp2 shows a high rate of nucleotide sequence change localized to a region encoding the "IN5" loop of the protein, thought to be a potential receptor-binding site. Phylogenetic analysis suggests that the ORF4 replication initiation protein is evolving faster than the other phage proteins. phiCPG1, phiAR39, and Chp2 are closely related to an ORF4 homolog inserted in the C. pneumoniae chromosome. This sequence analysis opens the way toward understanding the host-range and evolutionary history of these phages.

Amino Acid Sequence↗

Characterization of Porphyromonas gingivalis insertion sequence-like element ISPg5.

Porphyromonas gingivalis, a black-pigmented, gram-negative anaerobe, is found in periodontitis lesions, and its presence in subgingival plaque significantly increases the risk for periodontitis. In contrast to many bacterial pathogens, P. gingivalis strains display considerable variability, which is likely due to genetic exchange and intragenomic changes. To explore the latter possibility, we have studied the occurrence of insertion sequence (IS)-like elements in P. gingivalis W83 by utilizing a convenient and rapid method of capturing IS-like sequences and through analysis of the genome sequence of P. gingivalis strain W83. We adapted the method of Matsutani et al. (S. Matsutani, H. Ohtsubo, Y. Maeda, and E. Ohtsubo, J. Mol. Biol. 196:445-455, 1987) to isolate and clone rapidly annealing DNA sequences characteristic of repetitive regions within a genome. We show that in P. gingivalis strain W83, such sequences include (i) nucleotide sequence with homology to tRNA genes, (ii) a previously described IS element, and (iii) a novel IS-like element. Analysis of the P. gingivalis genome sequence for the distribution of the least used tetranucleotide, CTAG, identified regions in many of the initial 218 contigs which contained CTAG clusters. Examination of these CTAG clusters led to the discovery of 11 copies of the same novel IS-like element identified by the repeated sequence capture method of Matsutani et al. This new 1,512-bp IS-like element, designated ISPg5, has features of the IS3 family of IS elements. When a recombinant plasmid containing much of ISPg5 was used in Southern analysis of several P. gingivalis strains, including clinical isolates, diversity among strains was apparent. This suggests that ISPg5 and other IS elements may contribute to strain diversity and can be used for strain fingerprinting.

Amino Acid Sequence↗

Microbial genome sequencing 2000: new insights into physiology, evolution and expression analysis.

The complete genome sequence has been reported for 24 microbial organisms. The genome organization and gene content of these organisms has revealed an incredible diversity. Nearly half of the open reading frames identified by these sequencing projects are for potential genes with no known biological function. Efforts to make evolutionary sense and biological sense of the gene content of these organisms have been initiated. The greatest future challenge of genomics will be to determine function for the unknown genes.

Bacteria↗

Sequence and analysis of chromosome 2 of the plant Arabidopsis thaliana.

Arabidopsis thaliana (Arabidopsis) is unique among plant model organisms in having a small genome (130-140 Mb), excellent physical and genetic maps, and little repetitive DNA. Here we report the sequence of chromosome 2 from the Columbia ecotype in two gap-free assemblies (contigs) of 3.6 and 16 megabases (Mb). The latter represents the longest published stretch of uninterrupted DNA sequence assembled from any organism to date. Chromosome 2 represents 15% of the genome and encodes 4,037 genes, 49% of which have no predicted function. Roughly 250 tandem gene duplications were found in addition to large-scale duplications of about 0.5 and 4.5 Mb between chromosomes 2 and 1 and between chromosomes 2 and 4, respectively. Sequencing of nearly 2 Mb within the genetically defined centromere revealed a low density of recognizable genes, and a high density and diverse range of vestigial and presumably inactive mobile elements. More unexpected is what appears to be a recent insertion of a continuous stretch of 75% of the mitochondrial genome into chromosome 2.

Arabidopsis↗

Global transposon mutagenesis and a minimal Mycoplasma genome.

Mycoplasma genitalium with 517 genes has the smallest gene complement of any independently replicating cell so far identified. Global transposon mutagenesis was used to identify nonessential genes in an effort to learn whether the naturally occurring gene complement is a true minimal genome under laboratory growth conditions. The positions of 2209 transposon insertions in the completely sequenced genomes of M. genitalium and its close relative M. pneumoniae were determined by sequencing across the junction of the transposon and the genomic DNA. These junctions defined 1354 distinct sites of insertion that were not lethal. The analysis suggests that 265 to 350 of the 480 protein-coding genes of M. genitalium are essential under laboratory growth conditions, including about 100 genes of unknown function.

ATP-Binding Cassette Transporters↗

Genome sequence of the radioresistant bacterium Deinococcus radiodurans R1.

The complete genome sequence of the radiation-resistant bacterium Deinococcus radiodurans R1 is composed of two chromosomes (2,648,638 and 412,348 base pairs), a megaplasmid (177,466 base pairs), and a small plasmid (45,704 base pairs), yielding a total genome of 3,284, 156 base pairs. Multiple components distributed on the chromosomes and megaplasmid that contribute to the ability of D. radiodurans to survive under conditions of starvation, oxidative stress, and high amounts of DNA damage were identified. Deinococcus radiodurans represents an organism in which all systems for DNA repair, DNA damage export, desiccation and starvation recovery, and genetic redundancy are present in one cell.

Bacterial Proteins↗

Evidence for lateral gene transfer between Archaea and bacteria from genome sequence of Thermotoga maritima.

The 1,860,725-base-pair genome of Thermotoga maritima MSB8 contains 1,877 predicted coding regions, 1,014 (54%) of which have functional assignments and 863 (46%) of which are of unknown function. Genome analysis reveals numerous pathways involved in degradation of sugars and plant polysaccharides, and 108 genes that have orthologues only in the genomes of other thermophilic Eubacteria and Archaea. Of the Eubacteria sequenced to date, T. maritima has the highest percentage (24%) of genes that are most similar to archaeal genes. Eighty-one archaeal-like genes are clustered in 15 regions of the T. maritima genome that range in size from 4 to 20 kilobases. Conservation of gene order between T. maritima and Archaea in many of the clustered regions suggests that lateral gene transfer may have occurred between thermophilic Eubacteria and Archaea.

Archaea↗