Search PubMed⌕ Search

Biomedical subjects

Kazutoyo Osoegawa

Publications and source records attributed to Kazutoyo Osoegawa.

At least 19 recordsLinked to original sources

Draft genome sequence of the sexually transmitted pathogen Trichomonas vaginalis.

We describe the genome sequence of the protist Trichomonas vaginalis, a sexually transmitted human pathogen. Repeats and transposable elements comprise about two-thirds of the approximately 160-megabase genome, reflecting a recent massive expansion of genetic material. This expansion, in conjunction with the shaping of metabolic pathways that likely transpired through lateral gene transfer from bacteria, and amplification of specific gene families implicated in pathogenesis and phagocytosis of host proteins may exemplify adaptations of the parasite during its transition to a urogenital environment. The genome sequence predicts previously unknown functions for the hydrogenosome, which support a common evolutionary origin of this unusual organelle with mitochondria.

Animals↗

BAC clones generated from sheared DNA.

BAC libraries generated from restriction-digested genomic DNA display representational bias and lack some sequences. To facilitate completion of genome projects, procedures have been developed to create BACs from DNA physically sheared to create fragments extending up to 200 kb. The DNA fragments were repaired to create blunt ends and ligated to a new BAC vector. This approach has been tested by generating BAC libraries from Drosophila DNA with insert lengths between 50 and 150 kb. The libraries lack chimeric clone problems as determined by mapping paired BAC-end sequences to the assembled fly genome sequence. The utility of "sheared" libraries was demonstrated by closure of a previous clone gap and by isolation of clones from telomeric regions, which were notably absent from previous Drosophila BAC libraries.

Animals↗

A high-resolution map of synteny disruptions in gibbon and human genomes.

Gibbons are part of the same superfamily (Hominoidea) as humans and great apes, but their karyotype has diverged faster from the common hominoid ancestor. At least 24 major chromosome rearrangements are required to convert the presumed ancestral karyotype of gibbons into that of the hominoid ancestor. Up to 28 additional rearrangements distinguish the various living species from the common gibbon ancestor. Using the northern white-cheeked gibbon (2n = 52) (Nomascus leucogenys leucogenys) as a model, we created a high-resolution map of the homologous regions between the gibbon and human. The positions of 100 synteny breakpoints relative to the assembled human genome were determined at a resolution of about 200 kb. Interestingly, 46% of the gibbon-human synteny breakpoints occur in regions that correspond to segmental duplications in the human lineage, indicating a common source of plasticity leading to a different outcome in the two species. Additionally, the full sequences of 11 gibbon BACs spanning evolutionary breakpoints reveal either segmental duplications or interspersed repeats at the exact breakpoint locations. No specific sequence element appears to be common among independent rearrangements. We speculate that the extraordinarily high level of rearrangements seen in gibbons may be due to factors that increase the incidence of chromosome breakage or fixation of the derivative chromosomes in a homozygous state.

Animals↗

DNA sequence of human chromosome 17 and analysis of rearrangement in the human lineage.

Chromosome 17 is unusual among the human chromosomes in many respects. It is the largest human autosome with orthology to only a single mouse chromosome, mapping entirely to the distal half of mouse chromosome 11. Chromosome 17 is rich in protein-coding genes, having the second highest gene density in the genome. It is also enriched in segmental duplications, ranking third in density among the autosomes. Here we report a finished sequence for human chromosome 17, as well as a structural comparison with the finished sequence for mouse chromosome 11, the first finished mouse chromosome. Comparison of the orthologous regions reveals striking differences. In contrast to the typical pattern seen in mammalian evolution, the human sequence has undergone extensive intrachromosomal rearrangement, whereas the mouse sequence has been remarkably stable. Moreover, although the human sequence has a high density of segmental duplication, the mouse sequence has a very low density. Notably, these segmental duplications correspond closely to the sites of structural rearrangement, demonstrating a link between duplication and rearrangement. Examination of the main classes of duplicated segments provides insight into the dynamics underlying expansion of chromosome-specific, low-copy repeats in the human genome.

Animals↗

Vertebrate-type intron-rich genes in the marine annelid Platynereis dumerilii.

Previous genome comparisons have suggested that one important trend in vertebrate evolution has been a sharp rise in intron abundance. By using genomic data and expressed sequence tags from the marine annelid Platynereis dumerilii, we provide direct evidence that about two-thirds of human introns predate the bilaterian radiation but were lost from insect and nematode genomes to a large extent. A comparison of coding exon sequences confirms the ancestral nature of Platynereis and human genes. Thus, the urbilaterian ancestor had complex, intron-rich genes that have been retained in Platynereis and human.

Animals↗

A genome-wide comparison of recent chimpanzee and human segmental duplications.

We present a global comparison of differences in content of segmental duplication between human and chimpanzee, and determine that 33% of human duplications (> 94% sequence identity) are not duplicated in chimpanzee, including some human disease-causing duplications. Combining experimental and computational approaches, we estimate a genomic duplication rate of 4-5 megabases per million years since divergence. These changes have resulted in gene expression differences between the species. In terms of numbers of base pairs affected, we determine that de novo duplication has contributed most significantly to differences between the species, followed by deletion of ancestral duplications. Post-speciation gene conversion accounts for less than 10% of recent segmental duplication. Chimpanzee-specific hyperexpansion (> 100 copies) of particular segments of DNA have resulted in marked quantitative differences and alterations in the genome landscape between chimpanzee and human. Almost all of the most extreme differences relate to changes in chromosome structure, including the emergence of African great ape subterminal heterochromatin. Nevertheless, base per base, large segmental duplication events have had a greater impact (2.7%) in altering the genomic landscape of these two species than single-base-pair substitution (1.2%).

Animals↗

The genome sequence of Trypanosoma cruzi, etiologic agent of Chagas disease.

Whole-genome sequencing of the protozoan pathogen Trypanosoma cruzi revealed that the diploid genome contains a predicted 22,570 proteins encoded by genes, of which 12,570 represent allelic pairs. Over 50% of the genome consists of repeated sequences, such as retrotransposons and genes for large families of surface molecules, which include trans-sialidases, mucins, gp63s, and a large novel family (>1300 copies) of mucin-associated surface protein (MASP) genes. Analyses of the T. cruzi, T. brucei, and Leishmania major (Tritryp) genomes imply differences from other eukaryotes in DNA repair and initiation of replication and reflect their unusual mitochondrial DNA. Although the Tritryp lack several classes of signaling molecules, their kinomes contain a large and diverse set of protein kinases and phosphatases; their size and diversity imply previously unknown interactions and regulatory processes, which may be targets for intervention.

Animals↗

A highly redundant BAC library of Atlantic salmon (Salmo salar): an important tool for salmon projects.

BACKGROUND: As farming of Atlantic salmon is growing as an aquaculture enterprise, the need to identify the genomic mechanisms for specific traits is becoming more important in breeding and management of the animal. Traits of importance might be related to growth, disease resistance, food conversion efficiency, color or taste. To identify genomic regions responsible for specific traits, genomic large insert libraries have previously proven to be of crucial importance. These large insert libraries can be screened using gene or genetic markers in order to identify and map regions of interest. Furthermore, large-scale mapping can utilize highly redundant libraries in genome projects, and hence provide valuable data on the genome structure. RESULTS: Here we report the construction and characterization of a highly redundant bacterial artificial chromosome (BAC) library constructed from a Norwegian aquaculture strain male of Atlantic salmon (Salmo salar). The library consists of a total number of 305,557 clones, in which approximately 299,000 are recombinants. The average insert size of the library is 188 kbp, representing 18-fold genome coverage. High-density filters each consisting of 18,432 clones spotted in duplicates have been produced for hybridization screening, and are publicly available 1. To characterize the library, 15 expressed sequence tags (ESTs) derived overgos and 12 oligo sequences derived from microsatellite markers were used in hybridization screening of the complete BAC library. Secondary hybridizations with individual probes were performed for the clones detected. The BACs positive for the EST probes were fingerprinted and mapped into contigs, yielding an average of 3 contigs for each probe. Clones identified using genomic probes were PCR verified using microsatellite specific primers. CONCLUSION: Identification of genes and genomic regions of interest is greatly aided by the availability of the CHORI-214 Atlantic salmon BAC library. We have demonstrated the library's ability to identify specific genes and genetic markers using hybridization, PCR and fingerprinting experiments. In addition, multiple fingerprinting contigs indicated a pseudo-tetraploidity of the Atlantic salmon genome. The highly redundant CHORI-214 BAC library is expected to be an important resource for mapping and sequencing of the Atlantic salmon genome.

Animals↗

A physical map of the genome of Atlantic salmon, Salmo salar.

A physical map of the Atlantic salmon (Salmo salar) genome was generated based on HindIII fingerprints of a publicly available BAC (bacterial artificial chromosome) library constructed from DNA isolated from a Norwegian male. Approximately 11.5 haploid genome equivalents (185,938 clones) were successfully fingerprinted. Contigs were first assembled via FPC using high-stringency (1e-16), and then end-to-end joins yielded 4354 contigs and 37,285 singletons. The accuracy of the contig assembly was verified by hybridization and PCR analysis using genetic markers. A subset of the BACs in the library contained few or no HindIII recognition sites in their insert DNA. BglI digestion fragment patterns of these BACs allowed us to identify three classes: (1) BACs containing histone genes, (2) BACs containing rDNA-repeating units, and (3) those that do not have BglI recognition sites. End-sequence analysis of selected BACs representing these three classes confirmed the identification of the first two classes and suggested that the third class contained highly repetitive DNA corresponding to tRNAs and related sequences.

Animals↗

A physical map of the chicken genome.

Strategies for assembling large, complex genomes have evolved to include a combination of whole-genome shotgun sequencing and hierarchal map-assisted sequencing. Whole-genome maps of all types can aid genome assemblies, generally starting with low-resolution cytogenetic maps and ending with the highest resolution of sequence. Fingerprint clone maps are based upon complete restriction enzyme digests of clones representative of the target genome, and ultimately comprise a near-contiguous path of clones across the genome. Such clone-based maps are used to validate sequence assembly order, supply long-range linking information for assembled sequences, anchor sequences to the genetic map and provide templates for closing gaps. Fingerprint maps are also a critical resource for subsequent functional genomic studies, because they provide a redundant and ordered sampling of the genome with clones. In an accompanying paper we describe the draft genome sequence of the chicken, Gallus gallus, the first species sequenced that is both a model organism and a global food source. Here we present a clone-based physical map of the chicken genome at 20-fold coverage, containing 260 contigs of overlapping clones. This map represents approximately 91% of the chicken genome and enables identification of chicken clones aligned to positions in other sequenced genomes.

Animals↗

Human, mouse, and rat genome large-scale rearrangements: stability versus speciation.

Using paired-end sequences from bacterial artificial chromosomes, we have constructed high-resolution synteny and rearrangement breakpoint maps among human, mouse, and rat genomes. Among the >300 syntenic blocks identified are segments of over 40 Mb without any detected interspecies rearrangements, as well as regions with frequently broken synteny and extensive rearrangements. As closely related species, mouse and rat share the majority of the breakpoints and often have the same types of rearrangements when compared with the human genome. However, the breakpoints not shared between them indicate that mouse rearrangements are more often interchromosomal, whereas intrachromosomal rearrangements are more prominent in rat. Centromeres may have played a significant role in reorganizing a number of chromosomes in all three species. The comparison of the three species indicates that genome rearrangements follow a path that accommodates a delicate balance between maintaining a basic structure underlying all mammalian species and permitting variations that are necessary for speciation.

Animals↗

A set of BAC clones spanning the human genome.

Using the human bacterial artificial chromosome (BAC) fingerprint-based physical map, genome sequence assembly and BAC end sequences, we have generated a fingerprint-validated set of 32 855 BAC clones spanning the human genome. The clone set provides coverage for at least 98% of the human fingerprint map, 99% of the current assembled sequence and has an effective resolving power of 79 kb. We have made the clone set publicly available, anticipating that it will generally facilitate FISH or array-CGH-based identification and characterization of chromosomal alterations relevant to disease.

Base Sequence↗

TAHRE, a novel telomeric retrotransposon from Drosophila melanogaster, reveals the origin of Drosophila telomeres.

Drosophila telomeres do not have typical telomerase repeats. Instead, two families of non-LTR retrotransposons, HeT-A and TART, maintain telomere length by occasional transposition to the chromosome ends. Despite the work on Drosophila telomeres, its evolutionary origin remains controversial. Herein we describe a novel telomere-specific retroelement that we name TAHRE (Telomere-Associated and HeT-A-Related Element). The structure of the three telomere-specific elements indicates a common ancestor. These results suggest that preexisting transposable elements were recruited to perform the cellular function of telomere maintenance. A recruitment similar to that of a retrotransposal reverse transcriptase has been suggested as the common origin of telomerases.

Animals↗

Genomic analysis of Drosophila melanogaster telomeres: full-length copies of HeT-A and TART elements at telomeres.

The repetitive nature of heterochromatin hampers its analysis in general genome-sequencing projects. Specific studies are needed to extend the sequence into telomeric and centromeric heterochromatin. Drosophila telomeres lack the telomerase-generated repeats that are characteristic of other eukaryotic chromosomes. Instead, they consist of tandem arrays of HeT-A and TART elements. Herein, we present the genomic organization of the telomeres in the isogenic strain (y; cn bw sp) that was used for the Drosophila melanogaster sequencing project. The data indicate that the canonical features of telomere organization are widely conserved in evolution. In addition, we have identified full-length elements, likely competent elements, for HeT-A and TART.

Amino Acid Sequence↗

Complete MHC haplotype sequencing for common disease gene mapping.

The future systematic mapping of variants that confer susceptibility to common diseases requires the construction of a fully informative polymorphism map. Ideally, every base pair of the genome would be sequenced in many individuals. Here, we report 4.75 Mb of contiguous sequence for each of two common haplotypes of the major histocompatibility complex (MHC), to which susceptibility to >100 diseases has been mapped. The autoimmune disease-associated-haplotypes HLA-A3-B7-Cw7-DR15 and HLA-A1-B8-Cw7-DR3 were sequenced in their entirety through a bacterial artificial chromosome (BAC) cloning strategy using the consanguineous cell lines PGF and COX, respectively. The two sequences were annotated to encompass all described splice variants of expressed genes. We defined the complete variation content of the two haplotypes, revealing >18,000 variations between them. Average SNP densities ranged from less than one SNP per kilobase to >60. Acquisition of complete and accurate sequence data over polymorphic regions such as the MHC from large-insert cloned DNA provides a definitive resource for the construction of informative genetic maps, and avoids the limitation of chromosome regions that are refractory to PCR amplification.

Autoimmune Diseases↗

Genome sequence of the Brown Norway rat yields insights into mammalian evolution.

The laboratory rat (Rattus norvegicus) is an indispensable tool in experimental medicine and drug development, having made inestimable contributions to human health. We report here the genome sequence of the Brown Norway (BN) rat strain. The sequence represents a high-quality 'draft' covering over 90% of the genome. The BN rat sequence is the third complete mammalian genome to be deciphered, and three-way comparisons with the human and mouse genomes resolve details of mammalian evolution. This first comprehensive analysis includes genes and proteins and their relation to human disease, repeated sequences, comparative genome-wide studies of mammalian orthologous chromosomal regions and rearrangement breakpoints, reconstruction of ancestral karyotypes and the events leading to existing species, rates of variation, and lineage-specific and lineage-independent evolutionary events such as expansion of gene families, orthology relations and protein evolution.

Animals↗

Noncoding sequences conserved in a limited number of mammals in the SIM2 interval are frequently functional.

Cross-species DNA sequence comparison is a fundamental method for identifying biologically important elements, because functional sequences are evolutionarily conserved, wheres nonfunctional sequences drift. A recent genome-wide comparison of human and mouse DNA discovered over 200,000 conserved noncoding sequences with unknown function. Multispecies DNA comparison has been proposed as a method to prioritize these conserved noncoding sequences for functional analysis based on the hypothesis that elements present in many species are more likely to be functional than elements present in limited numbers of species. Here, we perform a comparative analysis of the single-minded 2 (SIM2) gene interval on human chromosome 21 with horse, cow, pig, dog, cat, and mouse DNA. We classify conserved sequences based on the number of mammals in which they are present, and experimentally test sequences in each class for function. As hypothesized, conserved sequences present in many mammals are frequently functional. Additionally, we demonstrate that sequences conserved in a limited number of mammals are also frequently functional. Examination of genomic deletions in chimpanzee and rhesus macaque DNA showed that several putatively functional conserved noncoding human sequences were absent in these primates. These findings suggest that functional conserved noncoding human sequences can be missing in other mammals, even closely related primate species.

Animals↗

The immunoglobulin lambda variable light-chain region in primates has been shaped by multiple, independent, small-scale and large-scale insertion/deletion events.

We analyzed genomes of nonhuman primates to determine the ancestral state of a 9.1-kb insertion/deletion polymorphism, located on human chromosome 22. The 9.1-kb+ allele was found in 16 chimpanzees, 3 bonobos, and 2 Bornean orangutans; however, 9 chimpanzees and 6 Sumatran orangutans showed neither the 9.1-kb+ nor the 9.1-kb- allele, but a novel allele, termed 9.1-kbnull. A clone from a chimpanzee BAC library carrying the 9.1-kbnull allele was sequenced: the BAC DNA aligns with the human chromosome 22 reference sequence except for a 75-kb region, suggesting that the 9.1-kbnull allele originated from a deletion. Furthermore, the 9.1-kb+ chromosomes of chimpanzees and bonobos contain a 1030-nucleotide sequence, absent in humans, that may result from a retro-transposition insertion in their common ancestor. Our results provide additional evidence that human chromosome 22 has undergone multiple small-scale and large-scale insertions and deletions since sharing a common ancestor with other primates.

Animals↗