Search PubMed⌕ Search

Biomedical subjects

Eric S Lander

Publications and source records attributed to Eric S Lander.

At least 73 records · Page 4Linked to original sources

PGC-1alpha-responsive genes involved in oxidative phosphorylation are coordinately downregulated in human diabetes.

DNA microarrays can be used to identify gene expression changes characteristic of human disease. This is challenging, however, when relevant differences are subtle at the level of individual genes. We introduce an analytical strategy, Gene Set Enrichment Analysis, designed to detect modest but coordinate changes in the expression of groups of functionally related genes. Using this approach, we identify a set of genes involved in oxidative phosphorylation whose expression is coordinately decreased in human diabetic muscle. Expression of these genes is high at sites of insulin-mediated glucose disposal, activated by PGC-1alpha and correlated with total-body aerobic capacity. Our results associate this gene set with clinically important variation in human metabolism and illustrate the value of pathway relationships in the analysis of genomic profiling experiments.

Animals↗

Whole-genome sequence assembly for mammalian genomes: Arachne 2.

We previously described the whole-genome assembly program Arachne, presenting assemblies of simulated data for small to mid-sized genomes. Here we describe algorithmic adaptations to the program, allowing for assembly of mammalian-size genomes, and also improving the assembly of smaller genomes. Three principal changes were simultaneously made and applied to the assembly of the mouse genome, during a six-month period of development: (1) Supercontigs (scaffolds) were iteratively broken and rejoined using several criteria, yielding a 64-fold increase in length (N50), and apparent elimination of all global misjoins; (2) gaps between contigs in supercontigs were filled (partially or completely) by insertion of reads, as suggested by pairing within the supercontig, increasing the N50 contig length by 50%; (3) memory usage was reduced fourfold. The outcome of this mouse assembly and its analysis are described in (Mouse Genome Sequencing Consortium 2002).

Animals↗

A molecular signature of metastasis in primary solid tumors.

Metastasis is the principal event leading to death in individuals with cancer, yet its molecular basis is poorly understood. To explore the molecular differences between human primary tumors and metastases, we compared the gene-expression profiles of adenocarcinoma metastases of multiple tumor types to unmatched primary adenocarcinomas. We found a gene-expression signature that distinguished primary from metastatic adenocarcinomas. More notably, we found that a subset of primary tumors resembled metastatic tumors with respect to this gene-expression signature. We confirmed this finding by applying the expression signature to data on 279 primary solid tumors of diverse types. We found that solid tumors carrying the gene-expression signature were most likely to be associated with metastasis and poor clinical outcome (P < 0.03). These results suggest that the metastatic potential of human tumors is encoded in the bulk of a primary tumor, thus challenging the notion that metastases arise from rare cells within a primary tumor that have the ability to metastasize.

Adenocarcinoma↗

The mosaic structure of variation in the laboratory mouse genome.

Most inbred laboratory mouse strains are known to have originated from a mixed but limited founder population in a few laboratories. However, the effect of this breeding history on patterns of genetic variation among these strains and the implications for their use are not well understood. Here we present an analysis of the fine structure of variation in the mouse genome, using single nucleotide polymorphisms (SNPs). When the recently assembled genome sequence from the C57BL/6J strain is aligned with sample sequence from other strains, we observe long segments of either extremely high (approximately 40 SNPs per 10 kb) or extremely low (approximately 0.5 SNPs per 10 kb) polymorphism rates. In all strain-to-strain comparisons examined, only one-third of the genome falls into long regions (averaging >1 Mb) of a high SNP rate, consistent with estimated divergence rates between Mus musculus domesticus and either M. m. musculus or M. m. castaneus. These data suggest that the genomes of these inbred strains are mosaics with the vast majority of segments derived from domesticus and musculus sources. These observations have important implications for the design and interpretation of positional cloning experiments.

Albinism↗

Identification of endoglin as a functional marker that defines long-term repopulating hematopoietic stem cells.

We describe a strategy to obtain highly enriched long-term repopulating (LTR) hematopoietic stem cells (HSCs) from bone marrow side-population (SP) cells by using a transgenic reporter gene driven by a stem cell enhancer. To analyze the gene-expression profile of the rare HSC population, we developed an amplification protocol termed "constant-ratio PCR," in which sample and control cDNAs are amplified in the same PCR. This protocol allowed us to identify genes differentially expressed in the enriched LTR-HSC population by oligonucleotide microarray analysis using as little as 1 ng of total RNA. Endoglin, an ancillary transforming growth factor beta receptor, was differentially expressed by the enriched HSCs. Importantly, endoglin-positive cells, which account for 20% of total SP cells, contain all the LTR-HSC activity within bone marrow SP. Our results demonstrate that endoglin, which plays important roles in angiogenesis and hematopoiesis, is a functional marker that defines LTR HSCs. Our overall strategy may be applicable for the identification of markers for other tissue-specific stem cells.

Animals↗

Detection of regulatory variation in mouse genes.

Functional polymorphism in genes can be classified as coding variation, altering the amino-acid sequence of the encoded protein, or regulatory variation, affecting the level or pattern of expression of the gene. Coding variation can be recognized directly from DNA sequence, and consequently its frequency and characteristics have been extensively described. By contrast, virtually nothing is known about the extent to which gene regulation varies in populations. Yet it is likely that regulatory variants are important in modulating gene function: alterations in gene regulation have been proposed to influence disease susceptibility and to have been the primary substrate for the evolution of species. Here, we report a systematic study to assess the extent of cis-acting regulatory variation in 69 genes across four inbred mouse strains. We find that at least four of these genes show allelic differences in expression level of 1.5-fold or greater, and that some of these differences are tissue specific. The results show that the impact of regulatory variants can be detected at a significant frequency in a genomic survey and suggest that such variation may have important consequences for organismal phenotype and evolution. The results indicate that larger-scale surveys in both mouse and human could identify a substantial number of genes with common regulatory variation.

Alleles↗

Detecting recent positive selection in the human genome from haplotype structure.

The ability to detect recent natural selection in the human population would have profound implications for the study of human history and for medicine. Here, we introduce a framework for detecting the genetic imprint of recent positive selection by analysing long-range haplotypes in human populations. We first identify haplotypes at a locus of interest (core haplotypes). We then assess the age of each core haplotype by the decay of its association to alleles at various distances from the locus, as measured by extended haplotype homozygosity (EHH). Core haplotypes that have unusually high EHH and a high population frequency indicate the presence of a mutation that rose to prominence in the human gene pool faster than expected under neutral evolution. We applied this approach to investigate selection at two genes carrying common variants implicated in resistance to malaria: G6PD and CD40 ligand. At both loci, the core haplotypes carrying the proposed protective mutation stand out and show significant evidence of selection. More generally, the method could be used to scan the entire genome for evidence of recent positive selection.

Africa↗

Abnormal gene expression in cloned mice derived from embryonic stem cell and cumulus cell nuclei.

To assess the extent of abnormal gene expression in clones, we assessed global gene expression by microarray analysis on RNA from the placentas and livers of neonatal cloned mice derived by nuclear transfer (NT) from both cultured embryonic stem cells and freshly isolated cumulus cells. Direct comparison of gene expression profiles of more than 10,000 genes showed that for both donor cell types approximately 4% of the expressed genes in the NT placentas differed dramatically in expression levels from those in controls and that the majority of abnormally expressed genes were common to both types of clones. Importantly, however, the expression of a smaller set of genes differed between the embryonic stem cell- and cumulus cell-derived clones. The livers of the cloned mice also showed abnormal gene expression, although to a lesser extent, and with a different set of affected genes, than seen in the placentas. Our results demonstrate frequent abnormal gene expression in clones, in which most expression abnormalities appear common to the NT procedure whereas others appear to reflect the particular donor nucleus.

Animals↗

Human genome sequence variation and the influence of gene history, mutation and recombination.

Variation in the human genome sequence is key to understanding susceptibility to disease in modern populations and the history of ancestral populations. Unlocking this information requires knowledge of the patterns and underlying causes of human sequence diversity. By applying a new population-genetic framework to two genome-wide polymorphism surveys, we find that the human genome contains sizeable regions (stretching over tens of thousands of base pairs) that have intrinsically high and low rates of sequence variation. We show that the primary determinant of these patterns is shared genealogical history. Only a fraction of the variation (at most 25%) is due to the local mutation rate. By measuring the average distance over which genealogical histories are typically preserved, these data provide the first genome-wide estimate of the average extent of correlation among variants (linkage disequilibrium). The results are best explained by extreme variability in the recombination rate at a fine scale, and provide the first empirical evidence that such recombination 'hot spots' are a general feature of the human genome and have a principal role in shaping genetic variation in the human population.

Animals↗

The structure of haplotype blocks in the human genome.

Haplotype-based methods offer a powerful approach to disease gene mapping, based on the association between causal mutations and the ancestral haplotypes on which they arose. As part of The SNP Consortium Allele Frequency Projects, we characterized haplotype patterns across 51 autosomal regions (spanning 13 megabases of the human genome) in samples from Africa, Europe, and Asia. We show that the human genome can be parsed objectively into haplotype blocks: sizable regions over which there is little evidence for historical recombination and within which only a few common haplotypes are observed. The boundaries of blocks and specific haplotypes they contain are highly correlated across populations. We demonstrate that such haplotype frameworks provide substantial statistical power in association studies of common genetic variation across each region. Our results provide a foundation for the construction of a haplotype map of the human genome, facilitating comprehensive genetic association studies of human disease.

Africa↗

On the sequencing of the human genome.

Two recent papers using different approaches reported draft sequences of the human genome. The international Human Genome Project (HGP) used the hierarchical shotgun approach, whereas Celera Genomics adopted the whole-genome shotgun (WGS) approach. Here, we analyze whether the latter paper provides a meaningful test of the WGS approach on a mammalian genome. In the Celera paper, the authors did not analyze their own WGS data. Instead, they decomposed the HGP's assembled sequence into a "perfect tiling path", combined it with their WGS data, and assembled the merged data set. To study the implications of this approach, we perform computational analysis and find that a perfect tiling path with 2-fold coverage is sufficient to recover virtually the entirety of a genome assembly. We also examine the manner in which the assembly was anchored to the human genome and conclude that the process primarily depended on the HGP's sequence-tagged site maps, BAC maps, and clone-based sequences. Our analysis indicates that the Celera paper provides neither a meaningful test of the WGS approach nor an independent sequence of the human genome. Our analysis does not imply that a WGS approach could not be successfully applied to assemble a draft sequence of a large mammalian genome, but merely that the Celera paper does not provide such evidence.

Chromosomes, Artificial, Bacterial↗

Prediction of central nervous system embryonal tumour outcome based on gene expression.

Embryonal tumours of the central nervous system (CNS) represent a heterogeneous group of tumours about which little is known biologically, and whose diagnosis, on the basis of morphologic appearance alone, is controversial. Medulloblastomas, for example, are the most common malignant brain tumour of childhood, but their pathogenesis is unknown, their relationship to other embryonal CNS tumours is debated, and patients' response to therapy is difficult to predict. We approached these problems by developing a classification system based on DNA microarray gene expression data derived from 99 patient samples. Here we demonstrate that medulloblastomas are molecularly distinct from other brain tumours including primitive neuroectodermal tumours (PNETs), atypical teratoid/rhabdoid tumours (AT/RTs) and malignant gliomas. Previously unrecognized evidence supporting the derivation of medulloblastomas from cerebellar granule cells through activation of the Sonic Hedgehog (SHH) pathway was also revealed. We show further that the clinical outcome of children with medulloblastomas is highly predictable on the basis of the gene expression profiles of their tumours at diagnosis.

Adolescent↗

Human macrophage activation programs induced by bacterial pathogens.

Understanding the response of innate immune cells to pathogens may provide insights to host defenses and the tactics used by pathogens to circumvent these defenses. We used DNA microarrays to explore the responses of human macrophages to a variety of bacteria. Macrophages responded to a broad range of bacteria with a robust, shared pattern of gene expression. The shared response includes genes encoding receptors, signal transduction molecules, and transcription factors. This shared activation program transforms the macrophage into a cell primed to interact with its environment and to mount an immune response. Further study revealed that the activation program is induced by bacterial components that are Toll-like receptor agonists, including lipopolysaccharide, lipoteichoic acid, muramyl dipeptide, and heat shock proteins. Pathogen-specific responses were also apparent in the macrophage expression profiles. Analysis of Mycobacterium tuberculosis-specific responses revealed inhibition of interleukin-12 production, suggesting one means by which this organism survives host defenses. These results improve our understanding of macrophage defenses, provide insights into mechanisms of pathogenesis, and suggest targets for therapeutic intervention.

Acetylmuramyl-Alanyl-Isoglutamine↗

Gene expression signatures define novel oncogenic pathways in T cell acute lymphoblastic leukemia.

Human T cell leukemias can arise from oncogenes activated by specific chromosomal translocations involving the T cell receptor genes. Here we show that five different T cell oncogenes (HOX11, TAL1, LYL1, LMO1, and LMO2) are often aberrantly expressed in the absence of chromosomal abnormalities. Using oligonucleotide microarrays, we identified several gene expression signatures that were indicative of leukemic arrest at specific stages of normal thymocyte development: LYL1+ signature (pro-T), HOX11+ (early cortical thymocyte), and TAL1+ (late cortical thymocyte). Hierarchical clustering analysis of gene expression signatures grouped samples according to their shared oncogenic pathways and identified HOX11L2 activation as a novel event in T cell leukemogenesis. These findings have clinical importance, since HOX11 activation is significantly associated with a favorable prognosis, while expression of TAL1, LYL1, or, surprisingly, HOX11L2 confers a much worse response to treatment. Our results illustrate the power of gene expression profiles to elucidate transformation pathways relevant to human leukemia.

Adaptor Proteins, Signal Transducing↗

Gene expression correlates of clinical prostate cancer behavior.

Prostate tumors are among the most heterogeneous of cancers, both histologically and clinically. Microarray expression analysis was used to determine whether global biological differences underlie common pathological features of prostate cancer and to identify genes that might anticipate the clinical behavior of this disease. While no expression correlates of age, serum prostate specific antigen (PSA), and measures of local invasion were found, a set of genes was identified that strongly correlated with the state of tumor differentiation as measured by Gleason score. Moreover, a model using gene expression data alone accurately predicted patient outcome following prostatectomy. These results support the notion that the clinical behavior of prostate cancer is linked to underlying gene expression differences that are detectable at the time of diagnosis.

Adult↗

Diffuse large B-cell lymphoma outcome prediction by gene-expression profiling and supervised machine learning.

Diffuse large B-cell lymphoma (DLBCL), the most common lymphoid malignancy in adults, is curable in less than 50% of patients. Prognostic models based on pre-treatment characteristics, such as the International Prognostic Index (IPI), are currently used to predict outcome in DLBCL. However, clinical outcome models identify neither the molecular basis of clinical heterogeneity, nor specific therapeutic targets. We analyzed the expression of 6,817 genes in diagnostic tumor specimens from DLBCL patients who received cyclophosphamide, adriamycin, vincristine and prednisone (CHOP)-based chemotherapy, and applied a supervised learning prediction method to identify cured versus fatal or refractory disease. The algorithm classified two categories of patients with very different five-year overall survival rates (70% versus 12%). The model also effectively delineated patients within specific IPI risk categories who were likely to be cured or to die of their disease. Genes implicated in DLBCL outcome included some that regulate responses to B-cell-receptor signaling, critical serine/threonine phosphorylation pathways and apoptosis. Our data indicate that supervised learning classification techniques can predict outcome in DLBCL and identify rational targets for intervention.

Antineoplastic Combined Chemotherapy Protocols↗

Using a genome-wide scan and meta-analysis to identify a novel IBD locus and confirm previously identified IBD loci.

Seven loci that potentially confer susceptibility to inflammatory bowel disease (IBD) or one of its subtypes have been identified to date; however, most are unconfirmed, and the complete set of loci contributing to disease susceptibility has not yet been determined. The authors aim to identify loci contributing to disease susceptibility in an IBD population from Canada and to compare their results in a systematic manner with those of previously published IBD data sets. The authors performed genome-wide linkage analysis on 63 IBD families from Nova Scotia, Canada. They then undertook a meta-analysis to combine the results of their study with those of the four previously published IBD genome-wide scans with complete data reported. Their genome-wide scan identified three regions of suggestive linkage to IBD: 11p, and The locus on chromosome 11p has not been previously reported. Meta-analysis of multiple scans revealed linked regions corresponding to the, and loci. Meta-analysis of linkage data is a powerful approach for identifying and confirming common susceptibility loci and specifically shows that, and are the major, common IBD susceptibility loci in the populations studied thus far.

Genetic Linkage↗

ARACHNE: a whole-genome shotgun assembler.

We describe a new computer system, called ARACHNE, for assembling genome sequence using paired-end whole-genome shotgun reads. ARACHNE has several key features, including an efficient and sensitive procedure for finding read overlaps, a procedure for scoring overlaps that achieves high accuracy by correcting errors before assembly, read merger based on forward-reverse links, and detection of repeat contigs by forward-reverse link inconsistency. To test ARACHNE, we created simulated reads providing approximately 10-fold coverage of the genomes of H. influenzae, S. cerevisiae, and D. melanogaster, as well as human chromosomes 21 and 22. The assemblies of these simulated reads yielded nearly complete coverage of the respective genomes, with a small number of contigs joined into a smaller number of supercontigs (or scaffolds). For example, analysis of the D. melanogaster genome yielded approximately 98% coverage with an N50 contig length of 324 kb and an N50 supercontig length of 5143 kb. The assembly accuracy was high, although not perfect: small errors occurred at a frequency of roughly 1 per 1 Mb (typically, deletion of approximately 1 kb in size), with a very small number of other misassemblies. The assembly was rapid: the Drosophila assembly required only 21 hours on a single 667 MHz processor and used 8.4 Gb of memory.

Algorithms↗