Celera genome licensing terms spark concerns over 'monopoly'.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
MOTIVATION: Since the simultaneous publication of the human genome assembly by the International Human Genome Sequencing Consortium (HGSC) and Celera Genomics, several comparisons have been made of various aspects of these two assemblies. In this work, we set out to provide a more comprehensive comparative analysis of the two assemblies and their associated gene sets. RESULTS: The local sequence content for both draft genome assemblies has been similar since the early releases, however it took a year for the quality of the Celera assembly to approach that of HGSC, suggesting an advantage of HGSC's hierarchical shotgun (HS) sequencing strategy over Celera's whole genome shotgun (WGS) approach. While similar numbers of ab initio predicted genes can be derived from both assemblies, Celera's Otto approach consistently generated larger, more varied gene sets than the Ensembl gene build system. The presence of a non-overlapping gene set has persisted with successive data releases from both groups. Since most of the unique genes from either genome assembly could be mapped back to the other assembly, we conclude that the gene set discrepancies do not reflect differences in local sequence content but rather in the assemblies and especially the different gene-prediction methodologies.
A dispute has been raging behind the scenes for weeks over the conditions under which Celera Genomics is prepared to make its human genome sequence data publicly available. The argument went public on 6 December, when geneticist Michael Ashburner e-mailed an open letter to Science's board of reviewing editors and members of the press slamming an agreement on data release that Science had reached with Celera as a condition for accepting its paper for review. This spat is the latest round in an intense rivalry between Celera president J. Craig Venter and leaders of the Human Genome Project, a publicly funded consortium that has produced its own draft human genome sequence.
The basis of human growth and development has long been considered to be one of the great mysteries of science and mankind. The portal to understanding this mystery was achieved by the Human Genome Project and Celera Genomics in 2001, with their joint announcement of the sequencing of 99% of the human genome map. Current reproductive options, however, remain restricted to the prevention of transmitting an at-risk gene or genes, but do not include treatment or cure. It is anticipated that this state of "halfway technology" will continue for years to come. As such, the scientific and ethical issues associated with each of these reproductive options will continue to affect the decision making of at-risk individuals. As the omnipresent health care provider, nurses have a duty to know and disseminate accurate and current information about reproductive options for individuals at risk for transmission of a genetic disorder. Nurses also have a duty to advocate for and ensure the privacy and confidentiality of genetic information.
The recent publications in Nature and Science by the Human Genome Consortium and Celera Genomics, respectively, while being landmark achievements in themselves, have also given pause for thought. A definitive catalogue of human genes is still not available but the broad picture of how humans compare with lower organisms at the genomic level is becoming clearer. The full impact of these findings on the practice of medicine is hard to predict, but research being conducted now, in the early years of the 21st century, will form the basis of future advances in the diagnosis and treatment of disease. Exactly what this will entail is the subject of intense debate, but there are some common starting points that were discussed at this meeting in Munich. The main theme to emerge was the need to move beyond the human genome sequence towards an understanding of proteins and their interactions in complex biological pathways, thereby increasing opportunities for drug discovery through the identification of new targets. The majority of the talks were therefore devoted to the description of technological advances in the analysis of gene and protein expression (and interaction) and in the use of various methods of gene deletion in order to validate individual proteins as drug targets. Perhaps it will still be a few years before it will be possible to report on the application of genomic analyses to routine medical practice at the first point of care for patients but when that happens, the research efforts described here will have been worthwhile.
BACKGROUND: The availability of both mouse and human draft genomes has marked the beginning of a new era of comparative mammalian genomics. The two available mouse genome assemblies, from the public mouse genome sequencing consortium and Celera Genomics, were obtained using different clone libraries and different assembly methods. RESULTS: We present here a critical comparison of the two latest mouse genome assemblies. The utility of the combined genomes is further demonstrated by comparing them with the human 'golden path' and through a subsequent analysis of a resulting conserved sequence element (CSE) database, which allows us to identify over 6,000 potential novel genes and to derive independent estimates of the number of human protein-coding genes. CONCLUSION: The Celera and public mouse assemblies differ in about 10% of the mouse genome. Each assembly has advantages over the other: Celera has higher accuracy in base-pairs and overall higher coverage of the genome; the public assembly, however, has higher sequence quality in some newly finished bacterial artificial chromosome clone (BAC) regions and the data are freely accessible. Perhaps most important, by combining both assemblies, we can get a better annotation of the human genome; in particular, we can obtain the most complete set of CSEs, one third of which are related to known genes and some others are related to other functional genomic regions. More than half the CSEs are of unknown function. From the CSEs, we estimate the total number of human protein-coding genes to be about 40,000. This searchable publicly available online CSEdb will expedite new discoveries through comparative genomics.
Bipolar affective disorder is one of the most common mental illnesses with a population prevalence of approximately 1%. The disorder is genetically complex, with an increasing number of loci being implicated through genetic linkage studies. However, the specific genetic variations and molecules involved in bipolar susceptibility and pathogenesis are yet to be identified. Genetic linkage analysis has identified a bipolar disorder susceptibility locus on chromosome 4q35, and the interval harbouring this susceptibility gene has been narrowed to a size that is amenable to positional cloning. We have used the resources of the Human Genome Project (HGP) and Celera Genomics to identify overlapping sequenced BAC clones and sequence contigs that represent the region implicated by linkage analysis. A combination of bioinformatic tools and laboratory techniques have been applied to annotate this DNA sequence data and establish a comprehensive transcript map that spans approximately 5.5 Mb. This map encompasses the chromosome 4q35 bipolar susceptibility locus, which localises to a "most probable" candidate interval of approximately 2.3 Mb, within a more conservative candidate interval of approximately 5 Mb. Localised within this map are 11 characterised genes and eight novel genes of unknown function, which together provide a collection of candidate transcripts that may be investigated for association with bipolar disorder. Overall, this region was shown to be very gene-poor, with a high incidence of pseudogenes, and redundant and novel repetitive elements. Our analysis of the interval has demonstrated a significant difference in the extent to which the current HGP and Celera sequence data sets represent this region.
The cooperation of biochemistry with clinical medicine consists of two overlapping temporal phases. Phase 1 of the cooperation, which still is not finished, is characterized by joint work on the pathogenesis and diagnostics of systemic metabolic diseases, whereas in phase 2 the cooperation on tissue and cell specific as well as on molecular diseases is prevailing. In view of the conceptual revolution and shift in paradigm, which biochemistry and medicine are presently experiencing, the content of cooperation between the two disciplines will profoundly change. It will become deeply influenced by the results of the research into the human genome and human proteome. Biochemistry will strongly be occupied to relate the thousands of protein coding genes to the structure and function of the encoded proteins, and medicine will be concerned in finding new protein markers for diagnostics, to identify novel drug targets, and to investigate, for example, the proteomes of the variety of tumors to aid tumor classification, to mention only a few areas of interest which medicine will have in the progress of human genome research. The review summarizes the recent achievements in sequencing the human DNA as published in February 2001 by the International Human Genome Sequencing Consortium and Celera Genomics and discusses their significance in respect to the further development of molecular, in particular genetic, medicine as an interdisciplinary field of the modern clinical sciences. Only biochemistry can provide the conceptual and experimental basis for the causal understanding of biological mechanisms as encoded in the genome of an organism.
Explore the source record for details and available documents.
Two recent papers using different approaches reported draft sequences of the human genome. The international Human Genome Project (HGP) used the hierarchical shotgun approach, whereas Celera Genomics adopted the whole-genome shotgun (WGS) approach. Here, we analyze whether the latter paper provides a meaningful test of the WGS approach on a mammalian genome. In the Celera paper, the authors did not analyze their own WGS data. Instead, they decomposed the HGP's assembled sequence into a "perfect tiling path", combined it with their WGS data, and assembled the merged data set. To study the implications of this approach, we perform computational analysis and find that a perfect tiling path with 2-fold coverage is sufficient to recover virtually the entirety of a genome assembly. We also examine the manner in which the assembly was anchored to the human genome and conclude that the process primarily depended on the HGP's sequence-tagged site maps, BAC maps, and clone-based sequences. Our analysis indicates that the Celera paper provides neither a meaningful test of the WGS approach nor an independent sequence of the human genome. Our analysis does not imply that a WGS approach could not be successfully applied to assemble a draft sequence of a large mammalian genome, but merely that the Celera paper does not provide such evidence.
Previous comparative analysis has revealed a significant disparity between the predicted gene sets produced by the International Human Genome Sequencing Consortium (HGSC) and Celera Genomics. To determine whether the source of this discrepancy was due to underlying differences in the genomic sequences or different gene prediction methodologies, we analyzed both genome assemblies in parallel. Using the GENSCAN gene prediction algorithm, we generated predicted transcriptomes that could be directly compared. BLAST-based comparisons revealed a 20-30% difference between the transcriptomes. Further differences between the two genomes were revealed with protein domain PFAM analyses. These results suggest that fundamental differences between the two genome assemblies are likely responsible for a significant portion of the discrepancy between the transcript sets predicted by the two groups.
The Drosophila melanogaster genome consists of four chromosomes that contain 165 Mb of DNA, 120 Mb of which are euchromatic. The two Drosophila Genome Projects, in collaboration with Celera Genomics Systems, have sequenced the genome, complementing the previously established physical and genetic maps. In addition, the Berkeley Drosophila Genome Project has undertaken large-scale functional analysis based on mutagenesis by transposable P element insertions into autosomes. Here, we present a large-scale P element insertion screen for vital gene functions and a BAC tiling map for the X chromosome. A collection of 501 X-chromosomal P element insertion lines was used to map essential genes cytogenetically and to establish short sequence tags (STSs) linking the insertion sites to the genome. The distribution of the P element integration sites, the identified genes and transcription units as well as the expression patterns of the P-element-tagged enhancers is described and discussed.
A different kind of shake-up will hit the science establishment when the New Year dawns in earthquake-prone Japan, reports Nature in its lead story this week. Science kicks off self-referentially with a lead story about its decision to publish a paper on sequencing the human genome from Craig Venter of Celera Genomics.
Much of the available human genomic sequence data exist in a fragmentary draft state following the completion of the initial high-volume sequencing performed by the International Human Genome Sequencing Consortium (IHGSC) and Celera Genomics (CG). We compared six draft genome assemblies over a region of chromosome 4p (D4S394-D4S403), two consecutive releases by the IHGSC at University of California, Santa Cruz (UCSC), two consecutive releases from the National Centre for Biotechnology Information (NCBI), the public release from CG, and a hybrid assembly we have produced using IHGSC and CG sequence data. This region presents particular problems for genomic sequence assembly algorithms as it contains a large tandem repeat and is sparsely covered by draft sequences. The six assemblies differed both in terms of their relative coverage of sequence data from the region and in their estimated rates of misassembly. The CG assembly method attained the lowest level of misassembly, whereas NCBI and UCSC assemblies had the highest levels of coverage. All assemblies examined included <60% of the publicly available sequence from the region. At least 6% of the sequence data within the CG assembly for the D4S394-D4S403 region was not present in publicly available sequence data. We also show that even in a problematic region, existing software tools can be used with high-quality mapping data to produce genomic sequence contigs with a low rate of rearrangements.
CE fractions may also be collected and then subjected to additional analysis. Nanoliter fractions containing size or shape fractionated DNA fragments can be collected on moving affinity membranes (125) or into sample chambers (126). The exact timing of the collection steps is achieved by determining the velocity of each individual zone measured between two detection points near the end of the capillary. The DNA samples may subsequently be identified by probe hybridization, or by PCR-linked sequencing. Capillary fractions containing metabolites and derivatives of DNA and small DNA adducts can also be sampled, and then characterized directly by highly sensitive MALDI-TOF atomic analysis (112-118) and ESI-MS (118,119). The automation and integration of PCR and CE analysis (PCR-CE) on a microchip (3-12,96) will also contribute greatly to its adoption as the analysis tool of choice. Significantly, these tools will be applied for DNA sequencing (75,108), for genome mapping (65) and genotyping (42-46), for improved certainty in disease detection (3-6,107,120) and for DNA mutation analysis (2-12,27,58). Recent improvements in the design CAE arrays and associated equipment such as the radial CAE microplate and rotary confocal signal detection system (127) overcome some of the detection limitations of linear CAE and microchip devices and allow the parallel genotyping of 96 samples in about 120 s. The integration of microreactive capillary surface assays (128) and "in-capillary" analysis will also lead to further increases in the speed and sensitivity of CE-based analysis. The recent announcement of the completion of the first draft sequence of the 90% of the entire human genome within 6 mo by Celera Genomics by sequencing random DNA fragments using several hundred ABI 3700 machines (129) illustrates the enormous efficiency realized through the automation of DNA sequencing by CAE. Sequencing was performed at an average rate of approximately 6 x 10(9) bases/yr. The CAE machines will now be employed for a concerted resequencing of genome elements to create an extremely high-density polymorphism map of the entire genome (130). This map will be based principally on single nucleotide polymorphisms, and will catapult human medicine into a new era of closely detailed genetic trait mapping to identify the genetic basis of multi-gene diseases.
Two different strategies for determining the human genome are currently being pursued: one is the "clone-by-clone" approach, employed by the publicly funded project, and the other is the "whole genome shotgun assembler" approach, favored by researchers at Celera Genomics. An interim strategy employed at Celera, called compartmentalized shotgun assembly, makes use of preliminary data produced by both approaches. In this paper we describe the design, implementation and operation of the "compartmentalized shotgun assembler".