Search PubMed⌕ Search

Biomedical subjects

Eric S Lander

Publications and source records attributed to Eric S Lander.

At least 19 recordsLinked to original sources

Genetic evidence for complex speciation of humans and chimpanzees.

The genetic divergence time between two species varies substantially across the genome, conveying important information about the timing and process of speciation. Here we develop a framework for studying this variation and apply it to about 20 million base pairs of aligned sequence from humans, chimpanzees, gorillas and more distantly related primates. Human-chimpanzee genetic divergence varies from less than 84% to more than 147% of the average, a range of more than 4 million years. Our analysis also shows that human-chimpanzee speciation occurred less than 6.3 million years ago and probably more recently, conflicting with some interpretations of ancient fossils. Most strikingly, chromosome X shows an extremely young genetic divergence time, close to the genome minimum along nearly its entire length. These unexpected features would be explained if the human and chimpanzee lineages initially diverged, then later exchanged genes before separating permanently.

Animals↗

A bivalent chromatin structure marks key developmental genes in embryonic stem cells.

The most highly conserved noncoding elements (HCNEs) in mammalian genomes cluster within regions enriched for genes encoding developmentally important transcription factors (TFs). This suggests that HCNE-rich regions may contain key regulatory controls involved in development. We explored this by examining histone methylation in mouse embryonic stem (ES) cells across 56 large HCNE-rich loci. We identified a specific modification pattern, termed "bivalent domains," consisting of large regions of H3 lysine 27 methylation harboring smaller regions of H3 lysine 4 methylation. Bivalent domains tend to coincide with TF genes expressed at low levels. We propose that bivalent domains silence developmental genes in ES cells while keeping them poised for activation. We also found striking correspondences between genome sequence and histone methylation in ES cells, which become notably weaker in differentiated cells. These results highlight the importance of DNA sequence in defining the initial epigenetic landscape and suggest a novel chromatin-based mechanism for maintaining pluripotency.

Animals↗

Identification and classification of conserved RNA secondary structures in the human genome.

The discoveries of microRNAs and riboswitches, among others, have shown functional RNAs to be biologically more important and genomically more prevalent than previously anticipated. We have developed a general comparative genomics method based on phylogenetic stochastic context-free grammars for identifying functional RNAs encoded in the human genome and used it to survey an eight-way genome-wide alignment of the human, chimpanzee, mouse, rat, dog, chicken, zebra-fish, and puffer-fish genomes for deeply conserved functional RNAs. At a loose threshold for acceptance, this search resulted in a set of 48,479 candidate RNA structures. This screen finds a large number of known functional RNAs, including 195 miRNAs, 62 histone 3'UTR stem loops, and various types of known genetic recoding elements. Among the highest-scoring new predictions are 169 new miRNA candidates, as well as new candidate selenocysteine insertion sites, RNA editing hairpins, RNAs involved in transcript auto regulation, and many folds that form singletons or small functional RNA families of completely unknown function. While the rate of false positives in the overall set is difficult to estimate and is likely to be substantial, the results nevertheless provide evidence for many new human functional RNAs and present specific predictions to facilitate their further characterization.

3' Untranslated Regions↗

DNA sequence of human chromosome 17 and analysis of rearrangement in the human lineage.

Chromosome 17 is unusual among the human chromosomes in many respects. It is the largest human autosome with orthology to only a single mouse chromosome, mapping entirely to the distal half of mouse chromosome 11. Chromosome 17 is rich in protein-coding genes, having the second highest gene density in the genome. It is also enriched in segmental duplications, ranking third in density among the autosomes. Here we report a finished sequence for human chromosome 17, as well as a structural comparison with the finished sequence for mouse chromosome 11, the first finished mouse chromosome. Comparison of the orthologous regions reveals striking differences. In contrast to the typical pattern seen in mammalian evolution, the human sequence has undergone extensive intrachromosomal rearrangement, whereas the mouse sequence has been remarkably stable. Moreover, although the human sequence has a high density of segmental duplication, the mouse sequence has a very low density. Notably, these segmental duplications correspond closely to the sites of structural rearrangement, demonstrating a link between duplication and rearrangement. Examination of the main classes of duplicated segments provides insight into the dynamics underlying expansion of chromosome-specific, low-copy repeats in the human genome.

Animals↗

Reactive oxygen species have a causal role in multiple forms of insulin resistance.

Insulin resistance is a cardinal feature of type 2 diabetes and is characteristic of a wide range of other clinical and experimental settings. Little is known about why insulin resistance occurs in so many contexts. Do the various insults that trigger insulin resistance act through a common mechanism? Or, as has been suggested, do they use distinct cellular pathways? Here we report a genomic analysis of two cellular models of insulin resistance, one induced by treatment with the cytokine tumour-necrosis factor-alpha and the other with the glucocorticoid dexamethasone. Gene expression analysis suggests that reactive oxygen species (ROS) levels are increased in both models, and we confirmed this through measures of cellular redox state. ROS have previously been proposed to be involved in insulin resistance, although evidence for a causal role has been scant. We tested this hypothesis in cell culture using six treatments designed to alter ROS levels, including two small molecules and four transgenes; all ameliorated insulin resistance to varying degrees. One of these treatments was tested in obese, insulin-resistant mice and was shown to improve insulin sensitivity and glucose homeostasis. Together, our findings suggest that increased ROS levels are an important trigger for insulin resistance in numerous settings.

3T3-L1 Cells↗

Analysis of the DNA sequence and duplication history of human chromosome 15.

Here we present a finished sequence of human chromosome 15, together with a high-quality gene catalogue. As chromosome 15 is one of seven human chromosomes with a high rate of segmental duplication, we have carried out a detailed analysis of the duplication structure of the chromosome. Segmental duplications in chromosome 15 are largely clustered in two regions, on proximal and distal 15q; the proximal region is notable because recombination among the segmental duplications can result in deletions causing Prader-Willi and Angelman syndromes. Sequence analysis shows that the proximal and distal regions of 15q share extensive ancient similarity. Using a simple approach, we have been able to reconstruct many of the events by which the current duplication structure arose. We find that most of the intrachromosomal duplications seem to share a common ancestry. Finally, we demonstrate that some remaining gaps in the genome sequence are probably due to structural polymorphisms between haplotypes; this may explain a significant fraction of the gaps remaining in the human genome.

Animals↗

A lentiviral RNAi library for human and mouse genes applied to an arrayed viral high-content screen.

To enable arrayed or pooled loss-of-function screens in a wide range of mammalian cell types, including primary and nondividing cells, we are developing lentiviral short hairpin RNA (shRNA) libraries targeting the human and murine genomes. The libraries currently contain 104,000 vectors, targeting each of 22,000 human and mouse genes with multiple sequence-verified constructs. To test the utility of the library for arrayed screens, we developed a screen based on high-content imaging to identify genes required for mitotic progression in human cancer cells and applied it to an arrayed set of 5,000 unique shRNA-expressing lentiviruses that target 1,028 human genes. The screen identified several known and approximately 100 candidate regulators of mitotic progression and proliferation; the availability of multiple shRNAs targeting the same gene facilitated functional validation of putative hits. This work provides a widely applicable resource for loss-of-function screens, as well as a roadmap for its application to biological discovery.

Animals↗

Human chromosome 11 DNA sequence and analysis including novel gene identification.

Chromosome 11, although average in size, is one of the most gene- and disease-rich chromosomes in the human genome. Initial gene annotation indicates an average gene density of 11.6 genes per megabase, including 1,524 protein-coding genes, some of which were identified using novel methods, and 765 pseudogenes. One-quarter of the protein-coding genes shows overlap with other genes. Of the 856 olfactory receptor genes in the human genome, more than 40% are located in 28 single- and multi-gene clusters along this chromosome. Out of the 171 disorders currently attributed to the chromosome, 86 remain for which the underlying molecular basis is not yet known, including several mendelian traits, cancer and susceptibility loci. The high-quality data presented here--nearly 134.5 million base pairs representing 99.8% coverage of the euchromatic sequence--provide scientists with a solid foundation for understanding the genetic basis of these disorders and other biological phenomena.

Chromosomes, Human, Pair 11↗

A large family of ancient repeat elements in the human genome is under strong selection.

Although conserved noncoding elements (CNEs) constitute the majority of sequences under purifying selection in the human genome, they remain poorly understood. CNEs seem to be largely unique, with no large families of similar elements reported to date. Here, we search for CNEs among the ancestral repeat classes in the human genome and report the discovery of a large CNE family containing >900 members. This family belongs to the MER121 class of repeats. Although the MER121 family members show considerable sequence variation among one another, the individual copies show striking conservation in orthologous locations across the human, dog, mouse, and rat genomes. The element is also present and conserved in orthologous locations in the marsupial, but its genome-wide dispersal postdates the divergence from birds. The comparative genomic data indicate that MER121 does not encode a family of either protein-coding or RNA genes. Although the precise function of these elements remains unknown, the evidence suggests that this unusual family may play a cis-regulatory or structural role in mammalian genomes.

Animals↗

DNA sequence and analysis of human chromosome 8.

The International Human Genome Sequencing Consortium (IHGSC) recently completed a sequence of the human genome. As part of this project, we have focused on chromosome 8. Although some chromosomes exhibit extreme characteristics in terms of length, gene content, repeat content and fraction segmentally duplicated, chromosome 8 is distinctly typical in character, being very close to the genome median in each of these aspects. This work describes a finished sequence and gene catalogue for the chromosome, which represents just over 5% of the euchromatic human genome. A unique feature of the chromosome is a vast region of approximately 15 megabases on distal 8p that appears to have a strikingly high mutation rate, which has accelerated in the hominids relative to other sequenced mammals. This fast-evolving region contains a number of genes related to innate immunity and the nervous system, including loci that appear to be under positive selection--these include the major defensin (DEF) gene cluster and MCPH1, a gene that may have contributed to the evolution of expanded brain size in the great apes. The data from chromosome 8 should allow a better understanding of both normal and disease biology and genome evolution.

Animals↗

Searching for signals of evolutionary selection in 168 genes related to immune function.

Pathogens have played a substantial role in human evolution, with past infections shaping genetic variation at loci influencing immune function. We selected 168 genes known to be involved in the immune response, genotyped common single nucleotide polymorphisms across each gene in three population samples (CEPH Europeans from Utah, Han Chinese from Guangxi, and Yoruba Nigerians from Southwest Nigeria) and searched for evidence of selection based on four tests for non-neutral evolution: minor allele frequency (MAF), derived allele frequency (DAF), Fst versus heterozygosity and extended haplotype homozygosity (EHH). Six of the 168 genes show some evidence for non-neutral evolution in this initial screen, with two showing similar signals in independent data from the International HapMap Project. These analyses identify two loci involved in immune function that are candidates for having been subject to evolutionary selection, and highlight a number of analytical challenges in searching for selection in genome-wide polymorphism data.

Algorithms↗

The case for selection at CCR5-Delta32.

The C-C chemokine receptor 5, 32 base-pair deletion (CCR5-Delta32) allele confers strong resistance to infection by the AIDS virus HIV. Previous studies have suggested that CCR5-Delta32 arose within the past 1,000 y and rose to its present high frequency (5%-14%) in Europe as a result of strong positive selection, perhaps by such selective agents as the bubonic plague or smallpox during the Middle Ages. This hypothesis was based on several lines of evidence, including the absence of the allele outside of Europe and long-range linkage disequilibrium at the locus. We reevaluated this evidence with the benefit of much denser genetic maps and extensive control data. We find that the pattern of genetic variation at CCR5-Delta32 does not stand out as exceptional relative to other loci across the genome. Moreover using newer genetic maps, we estimated that the CCR5-Delta32 allele is likely to have arisen more than 5,000 y ago. While such results can not rule out the possibility that some selection may have occurred at C-C chemokine receptor 5 (CCR5), they imply that the pattern of genetic variation seen at CCR5-Delta32 is consistent with neutral evolution. More broadly, the results have general implications for the design of future studies to detect the signs of positive selection in the human genome.

Alleles↗

Reduced representation bisulfite sequencing for comparative high-resolution DNA methylation analysis.

We describe a large-scale random approach termed reduced representation bisulfite sequencing (RRBS) for analyzing and comparing genomic methylation patterns. BglII restriction fragments were size-selected to 500-600 bp, equipped with adapters, treated with bisulfite, PCR amplified, cloned and sequenced. We constructed RRBS libraries from murine ES cells and from ES cells lacking DNA methyltransferases Dnmt3a and 3b and with knocked-down (kd) levels of Dnmt1 (Dnmt[1(kd),3a-/-,3b-/-]). Sequencing of 960 RRBS clones from Dnmt[1(kd),3a-/-,3b-/-] cells generated 343 kb of non-redundant bisulfite sequence covering 66212 cytosines in the genome. All but 38 cytosines had been converted to uracil indicating a conversion rate of >99.9%. Of the remaining cytosines 35 were found in CpG and 3 in CpT dinucleotides. Non-CpG methylation was >250-fold reduced compared with wild-type ES cells, consistent with a role for Dnmt3a and/or Dnmt3b in CpA and CpT methylation. Closer inspection revealed neither a consensus sequence around the methylated sites nor evidence for clustering of residual methylation in the genome. Our findings indicate random loss rather than specific maintenance of methylation in Dnmt[1(kd),3a-/-,3b-/-] cells. Near-complete bisulfite conversion and largely unbiased representation of RRBS libraries suggest that random shotgun bisulfite sequencing can be scaled to a genome-wide approach.

Animals↗

Gene set enrichment analysis: a knowledge-based approach for interpreting genome-wide expression profiles.

Although genomewide RNA expression analysis has become a routine tool in biomedical research, extracting biological insight from such information remains a major challenge. Here, we describe a powerful analytical method called Gene Set Enrichment Analysis (GSEA) for interpreting gene expression data. The method derives its power by focusing on gene sets, that is, groups of genes that share common biological function, chromosomal location, or regulation. We demonstrate how GSEA yields insights into several cancer-related data sets, including leukemia and lung cancer. Notably, where single-gene analysis finds little similarity between two independent studies of patient survival in lung cancer, GSEA reveals many biological pathways in common. The GSEA method is embodied in a freely available software package, together with an initial database of 1,325 biologically defined gene sets.

Cell Line, Tumor↗

DNA sequence and analysis of human chromosome 18.

Chromosome 18 appears to have the lowest gene density of any human chromosome and is one of only three chromosomes for which trisomic individuals survive to term. There are also a number of genetic disorders stemming from chromosome 18 trisomy and aneuploidy. Here we report the finished sequence and gene annotation of human chromosome 18, which will allow a better understanding of the normal and disease biology of this chromosome. Despite the low density of protein-coding genes on chromosome 18, we find that the proportion of non-protein-coding sequences evolutionarily conserved among mammals is close to the genome-wide average. Extending this analysis to the entire human genome, we find that the density of conserved non-protein-coding sequences is largely uncorrelated with gene density. This has important implications for the nature and roles of non-protein-coding sequence elements.

Aneuploidy↗

A high-density screen for linkage in multiple sclerosis.

To provide a definitive linkage map for multiple sclerosis, we have genotyped the Illumina BeadArray linkage mapping panel (version 4) in a data set of 730 multiplex families of Northern European descent. After the application of stringent quality thresholds, data from 4,506 markers in 2,692 individuals were included in the analysis. Multipoint nonparametric linkage analysis revealed highly significant linkage in the major histocompatibility complex (MHC) on chromosome 6p21 (maximum LOD score [MLS] 11.66) and suggestive linkage on chromosomes 17q23 (MLS 2.45) and 5q33 (MLS 2.18). This set of markers achieved a mean information extraction of 79.3% across the genome, with a Mendelian inconsistency rate of only 0.002%. Stratification based on carriage of the multiple sclerosis-associated DRB1*1501 allele failed to identify any other region of linkage with genomewide significance. However, ordered-subset analysis suggested that there may be an additional locus on chromosome 19p13 that acts independent of the main MHC locus. These data illustrate the substantial increase in power that can be achieved with use of the latest tools emerging from the Human Genome Project and indicate that future attempts to systematically identify susceptibility genes for multiple sclerosis will have to involve large sample sizes and an association-based methodology.

Australia↗

An initial strategy for the systematic identification of functional elements in the human genome by low-redundancy comparative sequencing.

With the recent completion of a high-quality sequence of the human genome, the challenge is now to understand the functional elements that it encodes. Comparative genomic analysis offers a powerful approach for finding such elements by identifying sequences that have been highly conserved during evolution. Here, we propose an initial strategy for detecting such regions by generating low-redundancy sequence from a collection of 16 eutherian mammals, beyond the 7 for which genome sequence data are already available. We show that such sequence can be accurately aligned to the human genome and used to identify most of the highly conserved regions. Although not a long-term substitute for generating high-quality genomic sequences from many mammalian species, this strategy represents a practical initial approach for rapidly annotating the most evolutionarily conserved sequences in the human genome, providing a key resource for the systematic study of human genome function.

Animals↗

A high-resolution linkage-disequilibrium map of the human major histocompatibility complex and first generation of tag single-nucleotide polymorphisms.

Autoimmune, inflammatory, and infectious diseases present a major burden to human health and are frequently associated with loci in the human major histocompatibility complex (MHC). Here, we report a high-resolution (1.9 kb) linkage-disequilibrium (LD) map of a 4.46-Mb fragment containing the MHC in U.S. pedigrees with northern and western European ancestry collected by the Centre d'Etude du Polymorphisme Humain (CEPH) and the first generation of haplotype tag single-nucleotide polymorphisms (tagSNPs) that provide up to a fivefold increase in genotyping efficiency for all future MHC-linked disease-association studies. The data confirm previously identified recombination hotspots in the class II region and allow the prediction of numerous novel hotspots in the class I and class III regions. The region of longest LD maps outside the classic MHC to the extended class I region spanning the MHC-linked olfactory-receptor gene cluster. The extended haplotype homozygosity analysis for recent positive selection shows that all 14 outlying haplotype variants map to a single extended haplotype, which most commonly bears HLA-DRB1*1501. The SNP data, haplotype blocks, and tagSNPs analysis reported here have been entered into a multidimensional Web-based database (GLOVAR), where they can be accessed and viewed in the context of relevant genome annotation. This LD map allowed us to give coordinates for the extremely variable LD structure underlying the MHC.

Haplotypes↗