Search PubMed⌕ Search

Biomedical subjects

Pieter J de Jong

Publications and source records attributed to Pieter J de Jong.

At least 19 recordsLinked to original sources

Draft genome sequence of the sexually transmitted pathogen Trichomonas vaginalis.

We describe the genome sequence of the protist Trichomonas vaginalis, a sexually transmitted human pathogen. Repeats and transposable elements comprise about two-thirds of the approximately 160-megabase genome, reflecting a recent massive expansion of genetic material. This expansion, in conjunction with the shaping of metabolic pathways that likely transpired through lateral gene transfer from bacteria, and amplification of specific gene families implicated in pathogenesis and phagocytosis of host proteins may exemplify adaptations of the parasite during its transition to a urogenital environment. The genome sequence predicts previously unknown functions for the hydrogenosome, which support a common evolutionary origin of this unusual organelle with mitochondria.

Animals↗

BAC clones generated from sheared DNA.

BAC libraries generated from restriction-digested genomic DNA display representational bias and lack some sequences. To facilitate completion of genome projects, procedures have been developed to create BACs from DNA physically sheared to create fragments extending up to 200 kb. The DNA fragments were repaired to create blunt ends and ligated to a new BAC vector. This approach has been tested by generating BAC libraries from Drosophila DNA with insert lengths between 50 and 150 kb. The libraries lack chimeric clone problems as determined by mapping paired BAC-end sequences to the assembled fly genome sequence. The utility of "sheared" libraries was demonstrated by closure of a previous clone gap and by isolation of clones from telomeric regions, which were notably absent from previous Drosophila BAC libraries.

Animals↗

A high-resolution map of synteny disruptions in gibbon and human genomes.

Gibbons are part of the same superfamily (Hominoidea) as humans and great apes, but their karyotype has diverged faster from the common hominoid ancestor. At least 24 major chromosome rearrangements are required to convert the presumed ancestral karyotype of gibbons into that of the hominoid ancestor. Up to 28 additional rearrangements distinguish the various living species from the common gibbon ancestor. Using the northern white-cheeked gibbon (2n = 52) (Nomascus leucogenys leucogenys) as a model, we created a high-resolution map of the homologous regions between the gibbon and human. The positions of 100 synteny breakpoints relative to the assembled human genome were determined at a resolution of about 200 kb. Interestingly, 46% of the gibbon-human synteny breakpoints occur in regions that correspond to segmental duplications in the human lineage, indicating a common source of plasticity leading to a different outcome in the two species. Additionally, the full sequences of 11 gibbon BACs spanning evolutionary breakpoints reveal either segmental duplications or interspersed repeats at the exact breakpoint locations. No specific sequence element appears to be common among independent rearrangements. We speculate that the extraordinarily high level of rearrangements seen in gibbons may be due to factors that increase the incidence of chromosome breakage or fixation of the derivative chromosomes in a homozygous state.

Animals↗

Independent centromere formation in a capricious, gene-free domain of chromosome 13q21 in Old World monkeys and pigs.

BACKGROUND: Evolutionary centromere repositioning and human analphoid neocentromeres occurring in clinical cases are, very likely, two stages of the same phenomenon whose properties still remain substantially obscure. Chromosome 13 is the chromosome with the highest number of neocentromeres. We reconstructed the mammalian evolutionary history of this chromosome and characterized two human neocentromeres at 13q21, in search of information that could improve our understanding of the relationship between evolutionarily new centromeres, inactivated centromeres, and clinical neocentromeres. RESULTS: Chromosome 13 evolution was studied, using FISH experiments, across several diverse superordinal phylogenetic clades spanning >100 million years of evolution. The analysis revealed exceptional conservation among primates (hominoids, Old World monkeys, and New World monkeys), Carnivora (cat), Perissodactyla (horse), and Cetartiodactyla (pig). In contrast, the centromeres in both Old World monkeys and pig have apparently repositioned independently to a central location (13q21). We compared these results to the positions of two human 13q21 neocentromeres using chromatin immunoprecipitation and genomic microarrays. CONCLUSION: We show that a gene-desert region at 13q21 of approximately 3.9 Mb in size possesses an inherent potential to form evolutionarily new centromeres over, at least, approximately 95 million years of mammalian evolution. The striking absence of genes may represent an important property, making the region tolerant to the extensive pericentromeric reshuffling during subsequent evolution. Comparison of the pericentromeric organization of chromosome 13 in four Old World monkey species revealed many differences in sequence organization. The region contains clusters of duplicons showing peculiar features.

Animals↗

Construction of a California condor BAC library and first-generation chicken-condor comparative physical map as an endangered species conservation genomics resource.

To support genomic analysis of the endangered California condor (Gymnogyps californianus), a BAC library (CHORI-262) was generated using DNA from the blood of a female. The library consists of 89,665 recombinant BAC clones providing approximately 14-fold coverage of the presumed approximately 1.48-Gb genome. Taking advantage of recent progress in chicken genomics, we developed a first-generation comparative chicken-condor physical map using an overgo hybridization approach. The overgos were derived from chicken (164 probes) and New World vulture (8 probes) sequences. Screening a 2.8x subset of the total library resulted in 236 BAC-gene assignments with 2.5 positive BAC clones per successful probe. A preliminary comparative chicken-condor BAC-based map included 93 genes. Comparison of selected condor BAC sequences with orthologous chicken sequences suggested a high degree of conserved synteny between the two avian genomes. This work will aid in identification and characterization of candidate loci for the chondrodystrophy mutation to advance genetic management of this disease.

Animals↗

DNA sequence of human chromosome 17 and analysis of rearrangement in the human lineage.

Chromosome 17 is unusual among the human chromosomes in many respects. It is the largest human autosome with orthology to only a single mouse chromosome, mapping entirely to the distal half of mouse chromosome 11. Chromosome 17 is rich in protein-coding genes, having the second highest gene density in the genome. It is also enriched in segmental duplications, ranking third in density among the autosomes. Here we report a finished sequence for human chromosome 17, as well as a structural comparison with the finished sequence for mouse chromosome 11, the first finished mouse chromosome. Comparison of the orthologous regions reveals striking differences. In contrast to the typical pattern seen in mammalian evolution, the human sequence has undergone extensive intrachromosomal rearrangement, whereas the mouse sequence has been remarkably stable. Moreover, although the human sequence has a high density of segmental duplication, the mouse sequence has a very low density. Notably, these segmental duplications correspond closely to the sites of structural rearrangement, demonstrating a link between duplication and rearrangement. Examination of the main classes of duplicated segments provides insight into the dynamics underlying expansion of chromosome-specific, low-copy repeats in the human genome.

Animals↗

A human-horse comparative map based on equine BAC end sequences.

In an effort to increase the density of sequence-based markers for the horse genome we generated 9473 BAC end sequences (BESs) from the CHORI-241 BAC library with an average read length of 677 bp. BLASTN searches with the BESs revealed 4036 meaningful hits (E <or= 10(-5)) in the human genome that provide useful markers for the human-horse comparative map. The 4036 BLASTN hits allowed the anchoring of 3079 BAC clones to the human genome, on average one corresponding equine BAC clone per megabase of human DNA. We used the BLASTN anchored BESs for an in silico prediction of the gene content and chromosome assignment of comparatively mapped equine BAC clones. As a first verification of our in silico mapping strategy we placed 19 equine BESs with matches to HSA6 onto the RH map. All markers were assigned to the predicted localizations on ECA10, ECA20, and ECA31, respectively.

Animals↗

Decoding the fine-scale structure of a breast cancer genome and transcriptome.

A comprehensive understanding of cancer is predicated upon knowledge of the structure of malignant genomes underlying its many variant forms and the molecular mechanisms giving rise to them. It is well established that solid tumor genomes accumulate a large number of genome rearrangements during tumorigenesis. End Sequence Profiling (ESP) maps and clones genome breakpoints associated with all types of genome rearrangements elucidating the structural organization of tumor genomes. Here we extend the ESP methodology in several directions using the breast cancer cell line MCF-7. First, targeted ESP is applied to multiple amplified loci, revealing a complex process of rearrangement and co-amplification in these regions reminiscent of breakage/fusion/bridge cycles. Second, genome breakpoints identified by ESP are confirmed using a combination of DNA sequencing and PCR. Third, in vitro functional studies assign biological function to a rearranged tumor BAC clone, demonstrating that it encodes anti-apoptotic activity. Finally, ESP is extended to the transcriptome identifying four novel fusion transcripts and providing evidence that expression of fusion genes may be common in tumors. These results demonstrate the distinct advantages of ESP including: (1) the ability to detect all types of rearrangements and copy number changes; (2) straightforward integration of ESP data with the annotated genome sequence; (3) immortalization of the genome; (4) ability to generate tumor-specific reagents for in vitro and in vivo functional studies. Given these properties, ESP could play an important role in a tumor genome project.

Breast Neoplasms↗

Genetic analysis of completely sequenced disease-associated MHC haplotypes identifies shuffling of segments in recent human history.

The major histocompatibility complex (MHC) is recognised as one of the most important genetic regions in relation to common human disease. Advancement in identification of MHC genes that confer susceptibility to disease requires greater knowledge of sequence variation across the complex. Highly duplicated and polymorphic regions of the human genome such as the MHC are, however, somewhat refractory to some whole-genome analysis methods. To address this issue, we are employing a bacterial artificial chromosome (BAC) cloning strategy to sequence entire MHC haplotypes from consanguineous cell lines as part of the MHC Haplotype Project. Here we present 4.25 Mb of the human haplotype QBL (HLA-A26-B18-Cw5-DR3-DQ2) and compare it with the MHC reference haplotype and with a second haplotype, COX (HLA-A1-B8-Cw7-DR3-DQ2), that shares the same HLA-DRB1, -DQA1, and -DQB1 alleles. We have defined the complete gene, splice variant, and sequence variation contents of all three haplotypes, comprising over 259 annotated loci and over 20,000 single nucleotide polymorphisms (SNPs). Certain coding sequences vary significantly between different haplotypes, making them candidates for functional and disease-association studies. Analysis of the two DR3 haplotypes allowed delineation of the shared sequence between two HLA class II-related haplotypes differing in disease associations and the identification of at least one of the sites that mediated the original recombination event. The levels of variation across the MHC were similar to those seen for other HLA-disparate haplotypes, except for a 158-kb segment that contained the HLA-DRB1, -DQA1, and -DQB1 genes and showed very limited polymorphism compatible with identity-by-descent and relatively recent common ancestry (<3,400 generations). These results indicate that the differential disease associations of these two DR3 haplotypes are due to sequence variation outside this central 158-kb segment, and that shuffling of ancestral blocks via recombination is a potential mechanism whereby certain DR-DQ allelic combinations, which presumably have favoured immunological functions, can spread across haplotypes and populations.

Chromosome Mapping↗

A highly redundant BAC library of Atlantic salmon (Salmo salar): an important tool for salmon projects.

BACKGROUND: As farming of Atlantic salmon is growing as an aquaculture enterprise, the need to identify the genomic mechanisms for specific traits is becoming more important in breeding and management of the animal. Traits of importance might be related to growth, disease resistance, food conversion efficiency, color or taste. To identify genomic regions responsible for specific traits, genomic large insert libraries have previously proven to be of crucial importance. These large insert libraries can be screened using gene or genetic markers in order to identify and map regions of interest. Furthermore, large-scale mapping can utilize highly redundant libraries in genome projects, and hence provide valuable data on the genome structure. RESULTS: Here we report the construction and characterization of a highly redundant bacterial artificial chromosome (BAC) library constructed from a Norwegian aquaculture strain male of Atlantic salmon (Salmo salar). The library consists of a total number of 305,557 clones, in which approximately 299,000 are recombinants. The average insert size of the library is 188 kbp, representing 18-fold genome coverage. High-density filters each consisting of 18,432 clones spotted in duplicates have been produced for hybridization screening, and are publicly available 1. To characterize the library, 15 expressed sequence tags (ESTs) derived overgos and 12 oligo sequences derived from microsatellite markers were used in hybridization screening of the complete BAC library. Secondary hybridizations with individual probes were performed for the clones detected. The BACs positive for the EST probes were fingerprinted and mapped into contigs, yielding an average of 3 contigs for each probe. Clones identified using genomic probes were PCR verified using microsatellite specific primers. CONCLUSION: Identification of genes and genomic regions of interest is greatly aided by the availability of the CHORI-214 Atlantic salmon BAC library. We have demonstrated the library's ability to identify specific genes and genetic markers using hybridization, PCR and fingerprinting experiments. In addition, multiple fingerprinting contigs indicated a pseudo-tetraploidity of the Atlantic salmon genome. The highly redundant CHORI-214 BAC library is expected to be an important resource for mapping and sequencing of the Atlantic salmon genome.

Animals↗

A physical map of the genome of Atlantic salmon, Salmo salar.

A physical map of the Atlantic salmon (Salmo salar) genome was generated based on HindIII fingerprints of a publicly available BAC (bacterial artificial chromosome) library constructed from DNA isolated from a Norwegian male. Approximately 11.5 haploid genome equivalents (185,938 clones) were successfully fingerprinted. Contigs were first assembled via FPC using high-stringency (1e-16), and then end-to-end joins yielded 4354 contigs and 37,285 singletons. The accuracy of the contig assembly was verified by hybridization and PCR analysis using genetic markers. A subset of the BACs in the library contained few or no HindIII recognition sites in their insert DNA. BglI digestion fragment patterns of these BACs allowed us to identify three classes: (1) BACs containing histone genes, (2) BACs containing rDNA-repeating units, and (3) those that do not have BglI recognition sites. End-sequence analysis of selected BACs representing these three classes confirmed the identification of the first two classes and suggested that the third class contained highly repetitive DNA corresponding to tRNAs and related sequences.

Animals↗

Pooled genomic indexing of rhesus macaque.

Pooled genomic indexing (PGI) is a method for mapping collections of bacterial artificial chromosome (BAC) clones between species by using a combination of clone pooling and DNA sequencing. PGI has been used to map a total of 3858 BAC clones covering approximately 24% of the rhesus macaque (Macaca mulatta) genome onto 4178 homologous loci in the human genome. A number of intrachromosomal rearrangements were detected by mapping multiple segments within the individual rhesus BACs onto multiple disjoined loci in the human genome. Transversal pooling designs involving shuffled BAC arrays were employed for robust mapping even with modest DNA sequence read coverage. A further innovation, short-tag pooled genomic indexing (ST-PGI), was also introduced to further improve the economy of mapping by sequencing multiple, short, mapable tags within a single sequencing reaction.

Animals↗

Independent intrachromosomal recombination events underlie the pericentric inversions of chimpanzee and gorilla chromosomes homologous to human chromosome 16.

Analyses of chromosomal rearrangements that have occurred during the evolution of the hominoids can reveal much about the mutational mechanisms underlying primate chromosome evolution. We characterized the breakpoints of the pericentric inversion of chimpanzee chromosome 18 (PTR XVI), which is homologous to human chromosome 16 (HSA 16). A conserved 23-kb inverted repeat composed of satellites, LINE and Alu elements was identified near the breakpoints and could have mediated the inversion by bringing the chromosomal arms into close proximity with each other, thereby facilitating intrachromosomal recombination. The exact positions of the breakpoints may then have been determined by local DNA sequence homologies between the inversion breakpoints, including a 22-base pair direct repeat. The similarly located pericentric inversion of gorilla (GGO) chromosome XVI, was studied by FISH and PCR analysis. The p- and q-arm breakpoints of the inversions in PTR XVI and GGO XVI were found to occur at slightly different locations, consistent with their independent origin. Further, FISH studies of the homologous chromosomal regions in macaque and orangutan revealed that the region represented by HSA BAC RP11-696P19, which spans the inversion breakpoint on HSA 16q11-12, was derived from the ancestral primate chromosome homologous to HSA 1. After the divergence of orangutan from the other great apes approximately 12 million years ago (Mya), a duplication of the corresponding region occurred followed by its interchromosomal transposition to the ancestral chromosome 16q. Thus, the most parsimonious interpretation is that the gorilla and chimpanzee homologs exhibit similar but nonidentical derived pericentric inversions, whereas HSA 16 represents the ancestral form among hominoids.

Animals↗

A physical map of the chicken genome.

Strategies for assembling large, complex genomes have evolved to include a combination of whole-genome shotgun sequencing and hierarchal map-assisted sequencing. Whole-genome maps of all types can aid genome assemblies, generally starting with low-resolution cytogenetic maps and ending with the highest resolution of sequence. Fingerprint clone maps are based upon complete restriction enzyme digests of clones representative of the target genome, and ultimately comprise a near-contiguous path of clones across the genome. Such clone-based maps are used to validate sequence assembly order, supply long-range linking information for assembled sequences, anchor sequences to the genetic map and provide templates for closing gaps. Fingerprint maps are also a critical resource for subsequent functional genomic studies, because they provide a redundant and ordered sampling of the genome with clones. In an accompanying paper we describe the draft genome sequence of the chicken, Gallus gallus, the first species sequenced that is both a model organism and a global food source. Here we present a clone-based physical map of the chicken genome at 20-fold coverage, containing 260 contigs of overlapping clones. This map represents approximately 91% of the chicken genome and enables identification of chicken clones aligned to positions in other sequenced genomes.

Animals↗

Evolutionary history of chromosome 20.

The evolutionary history of human chromosome 20 in primates was investigated using a panel of human BAC/PAC probes spaced along the chromosome. Oligonucleotide primers derived from the sequence of each human clone were used to screen horse, cat, pig, and black lemur BAC libraries to assemble, for each species, a panel of probes mapping to chromosomal loci orthologous to the loci encompassed by the human BACs. This approach facilitated marker-order comparison aimed at defining marker arrangement in primate ancestor. To this goal, we also took advantage of the mouse and rat draft sequences. The almost perfect colinearity of chromosome 20 sequence in humans and mouse could be interpreted as evidence that their form was ancestral to primates. Contrary to this view, we found that horse, macaque, and two New World monkeys share the same marker-order arrangement from which the human and mouse forms can be derived, assuming similar but distinct inversions that fully account for the small difference in marker arrangement between humans and mouse. The evolutionary history of this chromosome unveiled also two centromere repositioning events in New World monkey species.

Animals↗

Mutations in a new member of the chromodomain gene family cause CHARGE syndrome.

CHARGE syndrome is a common cause of congenital anomalies affecting several tissues in a nonrandom fashion. We report a 2.3-Mb de novo overlapping microdeletion on chromosome 8q12 identified by array comparative genomic hybridization in two individuals with CHARGE syndrome. Sequence analysis of genes located in this region detected mutations in the gene CHD7 in 10 of 17 individuals with CHARGE syndrome without microdeletions, accounting for the disease in most affected individuals.

Abnormalities, Multiple↗

A set of BAC clones spanning the human genome.

Using the human bacterial artificial chromosome (BAC) fingerprint-based physical map, genome sequence assembly and BAC end sequences, we have generated a fingerprint-validated set of 32 855 BAC clones spanning the human genome. The clone set provides coverage for at least 98% of the human fingerprint map, 99% of the current assembled sequence and has an effective resolving power of 79 kb. We have made the clone set publicly available, anticipating that it will generally facilitate FISH or array-CGH-based identification and characterization of chromosomal alterations relevant to disease.

Base Sequence↗

Complete MHC haplotype sequencing for common disease gene mapping.

The future systematic mapping of variants that confer susceptibility to common diseases requires the construction of a fully informative polymorphism map. Ideally, every base pair of the genome would be sequenced in many individuals. Here, we report 4.75 Mb of contiguous sequence for each of two common haplotypes of the major histocompatibility complex (MHC), to which susceptibility to >100 diseases has been mapped. The autoimmune disease-associated-haplotypes HLA-A3-B7-Cw7-DR15 and HLA-A1-B8-Cw7-DR3 were sequenced in their entirety through a bacterial artificial chromosome (BAC) cloning strategy using the consanguineous cell lines PGF and COX, respectively. The two sequences were annotated to encompass all described splice variants of expressed genes. We defined the complete variation content of the two haplotypes, revealing >18,000 variations between them. Average SNP densities ranged from less than one SNP per kilobase to >60. Acquisition of complete and accurate sequence data over polymorphic regions such as the MHC from large-insert cloned DNA provides a definitive resource for the construction of informative genetic maps, and avoids the limitation of chromosome regions that are refractory to PCR amplification.

Autoimmune Diseases↗