Should non-peer-reviewed raw DNA sequence data release be forced on the scientific community?
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to J C Venter.
Explore the source record for details and available documents.
The human genome is thought to harbor 50,000 to 100,000 genes, of which about half have been sampled to date in the form of expressed sequence tags. An international consortium was organized to develop and map gene-based sequence tagged site markers on a set of two radiation hybrid panels and a yeast artificial chromosome library. More than 16,000 human genes have been mapped relative to a framework map that contains about 1000 polymorphic genetic markers. The gene map unifies the existing genetic and physical maps with the nucleotide and protein sequence databases in a fashion that should speed the discovery of genes underlying inherited human disease. The integrated resource is available through a site on the World Wide Web at http://www.ncbi.nlm.nih.gov/SCIENCE96/.
Explore the source record for details and available documents.
The whole genome sequence (1.83 Mbp) of Haemophilus influenzae strain Rd was searched to identify tandem oligonucleotide repeat sequences. Loss or gain of one or more nucleotide repeats through a recombination-independent slippage mechanism is known to mediate phase variation of surface molecules of pathogenic bacteria, including H. influenzae. This facilitates evasion of host defenses and adaptation to the varying microenvironments of the host. We reasoned that iterative nucleotides could identify novel genes relevant to microbe-host interactions. Our search of the Rd genome sequence identified 9 novel loci with multiple (range 6-36, mean 22) tandem tetranucleotide repeats. All were found to be located within putative open reading frames and included homologues of hemoglobin-binding proteins of Neisseria, a glycosyltransferase (lgtC gene product) of Neisseria, and an adhesin of Yersinia. These tetranucleotide repeat sequences were also shown to be present in two other epidemiologically different H. influenzae type b strains, although the number and distribution of repeats was different. Further characterization of the lgtC gene showed that it was involved in phenotypic switching of a lipopolysaccharide epitope and that this variable expression was associated with changes in the number of tetranucleotide repeats. Mutation of lgtC resulted in attenuated virulence of H. influenzae in an infant rat model of invasive infection. These data indicate the rapidity, economy, and completeness with which whole genome sequences can be used to investigate the biology of pathogenic bacteria.
The complete 1.66-megabase pair genome sequence of an autotrophic archaeon, Methanococcus jannaschii, and its 58- and 16-kilobase pair extrachromosomal elements have been determined by whole-genome random sequencing. A total of 1738 predicted protein-coding genes were identified; however, only a minority of these (38 percent) could be assigned a putative cellular role with high confidence. Although the majority of genes related to energy production, cell division, and metabolism in M. jannaschii are most similar to those found in Bacteria, most of the genes involved in transcription, translation, and replication in M. jannaschii are more similar to those found in Eukaryotes.
Existing approaches to sequencing the human genome are based on the assumption that each region to be sequenced must first be mapped. But there is a simpler strategy in which any number of laboratories can cooperate.
The presenilin 1 gene has recently been identified as the locus on chromosome 14 which is responsible for a large proportion of early onset, autosomal dominantly inherited Alzheimer's disease (AD). We have elucidated the intron/exon structure of the gene and designed intronic primers to enable direct sequencing of the entire coding region (10 exons) of the presenilin gene in a large number of families. This strategy has enabled us to find a further two novel mutations in the gene. We discuss the distribution of mutations and the proportions of autosomal dominant AD with a mean age of onset below 60 years caused by mutations in this gene.
The availability of the complete 1.83-megabase-pair sequence of the Haemophilus influenzae strain Rd genome has facilitated significant progress in investigating the biology of H.influenzae lipopolysaccharide (LPS), a major virulence determinant of this human pathogen. By searching the H. influenzae genomic database, with sequences of known LPS biosynthetic genes from other organisms, we identified and then cloned 25 candidate LPS genes. Construction of mutant strains and characterization of the LPS by reactivity with monoclonal antibodies, PAGE fractionation patterns and electrospray mass spectrometry comparative analysis have confirmed a potential role in LPS biosynthesis for the majority of these candidate genes. Virulence studies in the infant rat have allowed us to estimate the minimal LPS structure required for intravascular dissemination. This study is one of the first to demonstrate the rapidity, economy and completeness with which novel biological information can be accessed once the complete genome sequence of an organism is available.
The generation of large numbers of partial cDNA sequences, or expressed sequence tags (ESTs), has provided a method with which to sample a large number of genes from an organism. More than 25,000 Arabidopsis thaliana ESTs have been deposited in public databases, producing the largest collection of ESTs for any plant species. We describe here the application of a method of reducing redundancy and increasing information content in this collection by grouping overlapping ESTs representing the same gene into a "contig" or assembly. The increased information content of these assemblies allows more putative identifications to be assigned based on the results of similarity searches with nucleotide and protein databases. The results of this analysis indicate that sequence information is available for approximately 12,600 nonoverlapping ESTs from Arabidopsis. Comparison of the assemblies with 953 Arabidopsis coding sequences indicates that up to 57% of all Arabidopsis genes are represented by an EST. Clustering analysis of these sequences suggests that between 300 and 700 gene families are represented by between 700 and 2000 sequences in the EST database. A database of the assembled sequences, their putative identifications, and cellular roles is available through the World Wide Web.
The complete nucleotide sequence (580,070 base pairs) of the Mycoplasma genitalium genome, the smallest known genome of any free-living organism, has been determined by whole-genome random sequencing and assembly. A total of only 470 predicted coding regions were identified that include genes required for DNA replication, transcription and translation, DNA repair, cellular transport, and energy metabolism. Comparison of this genome to that of Haemophilus influenzae suggests that differences in genome content are reflected as profound differences in physiology and metabolic capacity between these two organisms.
Advances in the Human Genome Project are shaping the strategies for identifying the 50,000-100,000 human genes. High-resolution genetic maps of the human genome combined with sequencing herald an era of rapid regional definition of disease genes. However, only once their chromosome band location is known will the systematic partial sequencing of thousands of random cDNA clones provide the reagents for teh rapid assessment of the genes responsible for the inherited disorders. We now present an approach to the rapid determination of map position and therefore to the creation of a transcribed map of the human genome. Sensitive fluorescence in situ hybridization has been combined with high-resolution chromosome banding and random cDNA sequencing to map 41 cDNAs with an average insert size of <2 kb to single human chromosome bands. The result provide 15 new genes, with database and functional information, as candidates for human disease. These include the large extracellular signal-related kinase (HUMERK), the ERK activator kinase (PRKMK1), a new member of the RAS oncogene family, protein phosphatase 2 regulatory subunit B alpha isoform (PPP2R2A), and a novel human gene with very high homology to a plant membrane transport family. Further, an analysis of expressed genes associated with pseudogenes showed that by using these techniques, it is possible to detect accurately the transcribed locus within a multigene or processed pseudogene family in most cases. These findings suggest that direct cDNA mapping using fluorescence in situ hybridization provides an accurate and rapid approach to the definition of a transcribed map of the human genome. This low-cost, high-resolution (2-5 Mb) mapping greatly enhances the speed with which these genes can be subsequently assigned to contigs. This assignment provides a necessary first step in understanding the relationship of the genes to both acquired and inherited human diseases.
The naturally transformable, Gram-negative bacterium Haemophilus influenzae Rd preferentially takes up DNA of its own species by recognizing a 9-base pair sequence, 5'-AAGTGCGGT, carried in multiple copies in its chromosome. With the availability of the complete genome sequence, 1465 copies of the 9-base pair uptake site have been identified. Alignment of these sites unexpectedly reveals an extended consensus region of 29 base pairs containing the core 9-base pair region and two downstream 6-base pair A/T-rich regions, each spaced about one helix turn apart. Seventeen percent of the sites are in inverted repeat pairs, many of which are located downstream to gene termini and are capable of forming stem-loop structures in messenger RNA that might function as signals for transcription termination.
A directional size-selected cDNA library constructed from Schistosoma mansoni (Sm) adult worm RNA was used for the generation of expressed sequence tags (EST). From one or both ends of 429 distinct cDNA clones 607 EST were obtained. Of these, only 16% were previously known Sm genes. More than 22% of the clones had matches with entries for other organisms in the databases. These new Sm genes constituted a broad range of transcripts distributed among cytoplasmic structural and regulatory proteins, enzymes, membrane, nuclear and secretory proteins, and proteins with other functions. Almost 33% of the clones had no significant database matches and thus potentially represent Sm-specific genes. Among the latter, several clones, as judged by their redundancy in the library, appear to represent abundant transcripts. The data, taken as a whole, more than double the number of Sm genes identified by nucleotide sequencing and indicate the potential value of the adoption of genome sequencing strategies for the rapid increase in knowledge of complex disease-causing organisms.
Explore the source record for details and available documents.
We analyzed the 186,102 base pairs (bp) that constitute the entire DNA genome of a highly virulent variola virus isolated from Bangladesh in 1975. The linear, double-stranded molecule has relatively small (725 bp) inverted terminal repeat (ITR) sequences containing three 69-bp direct repeat elements, a 54-bp partial repeat element, and a 105-base telomeric end-loop that can be maximally base-paired to contain 17 mismatches. Proximal to the right-end ITR sequences are another seven 69-bp elements and a 53- and a 27-bp partial element. Sequence analysis showed 187 closely spaced open reading frames specifying putative major proteins containing > or = 65 amino acids. Most of the virus proteins correspond to proteins in current databases, including 150 proteins that have > 90% identity to major gene products encoded by vaccinia virus, the smallpox vaccine. Variola virus has a group of proteins that are truncated compared with vaccinia virus counterparts and a smaller group of proteins that are elongated. The terminal regions encode several novel proteins and variants of other poxvirus proteins that potentially augment variola virus transmissibility and virulence for its only natural host, humans.
High-throughput automated sequencing has enabled researchers to examine large numbers of clones from a cDNA library as a measure of the steady-state levels of mRNA species. The past year has witnessed many new applications of this technique to allow the qualitative and quantitative comparison of the changes in transcript levels from multiple genes.
Explore the source record for details and available documents.