Search PubMed⌕ Search

Biomedical subjects

Eric S Lander

Publications and source records attributed to Eric S Lander.

At least 55 records · Page 3Linked to original sources

Assessing the impact of population stratification on genetic association studies.

Population stratification refers to differences in allele frequencies between cases and controls due to systematic differences in ancestry rather than association of genes with disease. It has been proposed that false positive associations due to stratification can be controlled by genotyping a few dozen unlinked genetic markers. To assess stratification empirically, we analyzed data from 11 case-control and case-cohort association studies. We did not detect statistically significant evidence for stratification but did observe that assessments based on a few dozen markers lack power to rule out moderate levels of stratification that could cause false positive associations in studies designed to detect modest genetic risk factors. After increasing the number of markers and samples in a case-cohort study (the design most immune to stratification), we found that stratification was in fact present. Our results suggest that modest amounts of stratification can exist even in well designed studies.

Case-Control Studies↗

Genetic dissection of complex traits with chromosome substitution strains of mice.

Chromosome substitution strains (CSSs) have been proposed as a simple and powerful way to identify quantitative trait loci (QTLs) affecting developmental, physiological, and behavioral processes. Here, we report the construction of a complete CSS panel for a vertebrate species. The CSS panel consists of 22 mouse strains, each of which carries a single chromosome substituted from a donor strain (A/J) onto a common host background (C57BL/6J). A survey of 53 traits revealed evidence for 150 QTLs affecting serum levels of sterols and amino acids, diet-induced obesity, and anxiety. These results demonstrate that CSSs greatly facilitate the detection and identification of genes that control the wide diversity of naturally occurring phenotypic variation in the A/J and C57BL/6J inbred strains.

Amino Acids↗

Proof and evolutionary analysis of ancient genome duplication in the yeast Saccharomyces cerevisiae.

Whole-genome duplication followed by massive gene loss and specialization has long been postulated as a powerful mechanism of evolutionary innovation. Recently, it has become possible to test this notion by searching complete genome sequence for signs of ancient duplication. Here, we show that the yeast Saccharomyces cerevisiae arose from ancient whole-genome duplication, by sequencing and analysing Kluyveromyces waltii, a related yeast species that diverged before the duplication. The two genomes are related by a 1:2 mapping, with each region of K. waltii corresponding to two regions of S. cerevisiae, as expected for whole-genome duplication. This resolves the long-standing controversy on the ancestry of the yeast genome, and makes it possible to study the fate of duplicated genes directly. Strikingly, 95% of cases of accelerated evolution involve only one member of a gene pair, providing strong support for a specific model of evolution, and allowing us to distinguish ancestral and derived functions.

Codon↗

Eric S. Lander.

Explore the source record for details and available documents.

Academies and Institutes↗

Methods in comparative genomics: genome correspondence, gene identification and regulatory motif discovery.

In Kellis et al. (2003), we reported the genome sequences of S. paradoxus, S. mikatae, and S. bayanus and compared these three yeast species to their close relative, S. cerevisiae. Genomewide comparative analysis allowed the identification of functionally important sequences, both coding and noncoding. In this companion paper we describe the mathematical and algorithmic results underpinning the analysis of these genomes. (1) We present methods for the automatic determination of genome correspondence. The algorithms enabled the automatic identification of orthologs for more than 90% of genes and intergenic regions across the four species despite the large number of duplicated genes in the yeast genome. The remaining ambiguities in the gene correspondence revealed recent gene family expansions in regions of rapid genomic change. (2) We present methods for the identification of protein-coding genes based on their patterns of nucleotide conservation across related species. We observed the pressure to conserve the reading frame of functional proteins and developed a test for gene identification with high sensitivity and specificity. We used this test to revisit the genome of S. cerevisiae, reducing the overall gene count by 500 genes (10% of previously annotated genes) and refining the gene structure of hundreds of genes. (3) We present novel methods for the systematic de novo identification of regulatory motifs. The methods do not rely on previous knowledge of gene function and in that way differ from the current literature on computational motif discovery. Based on genomewide conservation patterns of known motifs, we developed three conservation criteria that we used to discover novel motifs. We used an enumeration approach to select strongly conserved motif cores, which we extended and collapsed into a small number of candidate regulatory motifs. These include most previously known regulatory motifs as well as several noteworthy novel motifs. The majority of discovered motifs are enriched in functionally related genes, allowing us to infer a candidate function for novel motifs. Our results demonstrate the power of comparative genomics to further our understanding of any species. Our methods are validated by the extensive experimental knowledge in yeast and will be invaluable in the study of complex genomes like that of the human.

Algorithms↗

A complex interaction of imprinted and maternal-effect genes modifies sex determination in Odd Sex (Ods) mice.

The transgenic insertional mouse mutation Odd Sex (Ods) represents a model for the long-range regulation of Sox9. The mutation causes complete female-to-male sex reversal by inducing a male-specific expression pattern of Sox9 in XX Ods/+ embryonic gonads. We previously described an A/J strain-specific suppressor of Ods termed Odsm1(A). Here we show that phenotypic sex depends on a complex interaction between the suppressor and the transgene. Suppression can be achieved only if the transgene is transmitted paternally. In addition, the suppressor itself exhibits a maternal effect, suggesting that it may act on chromatin in the early embryo.

Animals↗

Integrated analysis of protein composition, tissue diversity, and gene regulation in mouse mitochondria.

Mitochondria are tailored to meet the metabolic and signaling needs of each cell. To explore its molecular composition, we performed a proteomic survey of mitochondria from mouse brain, heart, kidney, and liver and combined the results with existing gene annotations to produce a list of 591 mitochondrial proteins, including 163 proteins not previously associated with this organelle. The protein expression data were largely concordant with large-scale surveys of RNA abundance and both measures indicate tissue-specific differences in organelle composition. RNA expression profiles across tissues revealed networks of mitochondrial genes that share functional and regulatory mechanisms. We also determined a larger "neighborhood" of genes whose expression is closely correlated to the mitochondrial genes. The combined analysis identifies specific genes of biological interest, such as candidates for mtDNA repair enzymes, offers new insights into the biogenesis and ancestry of mammalian mitochondria, and provides a framework for understanding the organelle's contribution to human disease.

Animals↗

Position specific variation in the rate of evolution in transcription factor binding sites.

BACKGROUND: The binding sites of sequence specific transcription factors are an important and relatively well-understood class of functional non-coding DNAs. Although a wide variety of experimental and computational methods have been developed to characterize transcription factor binding sites, they remain difficult to identify. Comparison of non-coding DNA from related species has shown considerable promise in identifying these functional non-coding sequences, even though relatively little is known about their evolution. RESULTS: Here we analyse the genome sequences of the budding yeasts Saccharomyces cerevisiae, S. bayanus, S. paradoxus and S. mikatae to study the evolution of transcription factor binding sites. As expected, we find that both experimentally characterized and computationally predicted binding sites evolve slower than surrounding sequence, consistent with the hypothesis that they are under purifying selection. We also observe position-specific variation in the rate of evolution within binding sites. We find that the position-specific rate of evolution is positively correlated with degeneracy among binding sites within S. cerevisiae. We test theoretical predictions for the rate of evolution at positions where the base frequencies deviate from background due to purifying selection and find reasonable agreement with the observed rates of evolution. Finally, we show how the evolutionary characteristics of real binding motifs can be used to distinguish them from artefacts of computational motif finding algorithms. CONCLUSION: As has been observed for protein sequences, the rate of evolution in transcription factor binding sites varies with position, suggesting that some regions are under stronger functional constraint than others. This variation likely reflects the varying importance of different positions in the formation of the protein-DNA complex. The characterization of the pattern of evolution in known binding sites will likely contribute to the effective use of comparative sequence data in the identification of transcription factor binding sites and is an important step toward understanding the evolution of functional non-coding DNA.

Artifacts↗

An integrated haplotype map of the human major histocompatibility complex.

Numerous studies have clearly indicated a role for the major histocompatibility complex (MHC) in susceptibility to autoimmune diseases. Such studies have focused on the genetic variation of a small number of classical human-leukocyte-antigen (HLA) genes in the region. Although these genes represent good candidates, given their immunological roles, linkage disequilibrium (LD) surrounding these genes has made it difficult to rule out neighboring genes, many with immune function, as influencing disease susceptibility. It is likely that a comprehensive analysis of the patterns of LD and variation, by using a high-density map of single-nucleotide polymorphisms (SNPs), would enable a greater understanding of the nature of the observed associations, as well as lead to the identification of causal variation. We present herein an initial analysis of this region, using 201 SNPs, nine classical HLA loci, two TAP genes, and 18 microsatellites. This analysis suggests that LD and variation in the MHC, aside from the classical HLA loci, are essentially no different from those in the rest of the genome. Furthermore, these data show that multi-SNP haplotypes will likely be a valuable means for refining association signals in this region.

Chromosome Mapping↗

Phylogenetically and spatially conserved word pairs associated with gene-expression changes in yeasts.

BACKGROUND: Transcriptional regulation in eukaryotes often involves multiple transcription factors binding to the same transcription control region, and to understand the regulatory content of eukaryotic genomes it is necessary to consider the co-occurrence and spatial relationships of individual binding sites. The determination of conserved sequences (often known as phylogenetic footprinting) has identified individual transcription factor binding sites. We extend this concept of functional conservation to higher-order features of transcription control regions. RESULTS: We used the genome sequences of four yeast species of the genus Saccharomyces to identify sequences potentially involved in multifactorial control of gene expression. We found 989 potential regulatory 'templates': pairs of hexameric sequences that are jointly conserved in transcription regulatory regions and also exhibit non-random relative spacing. Many of the individual sequences in these templates correspond to known transcription factor binding sites, and the sets of genes containing a particular template in their transcription control regions tend to be differentially expressed in conditions where the corresponding transcription factors are known to be active. The incorporation of word pairs to define sequence features yields more specific predictions of average expression profiles and more informative regression models for genome-wide expression data than considering sequence conservation alone. CONCLUSIONS: The incorporation of both joint conservation and spacing constraints of sequence pairs predicts groups of target genes that are specific for common patterns of gene expression. Our work suggests that positional information, especially the relative spacing between transcription factor binding sites, may represent a common organizing principle of transcription control regions.

Base Sequence↗

IBD5 is a general risk factor for inflammatory bowel disease: replication of association with Crohn disease and identification of a novel association with ulcerative colitis.

Inflammatory bowel disease (IBD) refers to complex chronic relapsing autoimmune disorders of the gastrointestinal tract that have been traditionally classified into Crohn disease (CD) and ulcerative colitis (UC). We have previously reported that genetic variation within a 250-kb haplotype (IBD5) in the 5q31 cytokine gene cluster confers susceptibility to CD in a Canadian population. In the current study, we first replicated this association by examining 368 German trios with CD and demonstrating, by transmission/disequilibrium testing (TDT), that the same haplotype is associated with CD (chi2=5.97; P=.007). Our original association study focused on the role of IBD5 in CD; we next explored the potential contribution of this locus to UC susceptibility in 187 German trios. Given the TDT results in the present cohort with UC, IBD5 may also act as a susceptibility locus for UC (chi2=8.10; P=.002). We then examined locus-locus interactions between IBD5 and CARD15, a locus reported elsewhere to confer risk exclusively to CD. Our current results indicate that the two loci act independently to confer risk to CD but that these two loci may behave in an epistatic fashion to promote the development of UC. Moreover, IBD5 was not associated with particular clinical manifestations upon phenotypic stratification in the current cohort with CD. Taken together, our results suggest that IBD5 may act as a general risk factor for IBD, with loci such as CARD15 modifying the clinical characteristics of disease.

Colitis, Ulcerative↗

Sequencing and comparison of yeast species to identify genes and regulatory elements.

Identifying the functional elements encoded in a genome is one of the principal challenges in modern biology. Comparative genomics should offer a powerful, general approach. Here, we present a comparative analysis of the yeast Saccharomyces cerevisiae based on high-quality draft sequences of three related species (S. paradoxus, S. mikatae and S. bayanus). We first aligned the genomes and characterized their evolution, defining the regions and mechanisms of change. We then developed methods for direct identification of genes and regulatory motifs. The gene analysis yielded a major revision to the yeast gene catalogue, affecting approximately 15% of all genes and reducing the total count by about 500 genes. The motif analysis automatically identified 72 genome-wide elements, including most known regulatory motifs and numerous new motifs. We inferred a putative function for most of these motifs, and provided insights into their combinatorial interactions. The results have implications for genome analysis of diverse organisms, including the human.

Base Sequence↗

The genome sequence of the filamentous fungus Neurospora crassa.

Neurospora crassa is a central organism in the history of twentieth-century genetics, biochemistry and molecular biology. Here, we report a high-quality draft sequence of the N. crassa genome. The approximately 40-megabase genome encodes about 10,000 protein-coding genes--more than twice as many as in the fission yeast Schizosaccharomyces pombe and only about 25% fewer than in the fruitfly Drosophila melanogaster. Analysis of the gene set yields insights into unexpected aspects of Neurospora biology including the identification of genes potentially associated with red light photobiology, genes implicated in secondary metabolism, and important differences in Ca2+ signalling as compared with plants and animals. Neurospora possesses the widest array of genome defence mechanisms known for any eukaryotic organism, including a process unique to fungi called repeat-induced point mutation (RIP). Genome analysis suggests that RIP has had a profound impact on genome evolution, greatly slowing the creation of new genes through genomic duplication and resulting in a genome with an unusually low proportion of closely related genes.

Calcium Signaling↗

Identification of a gene causing human cytochrome c oxidase deficiency by integrative genomics.

Identifying the genes responsible for human diseases requires combining information about gene position with clues about biological function. The recent availability of whole-genome data sets of RNA and protein expression provides powerful new sources of functional insight. Here we illustrate how such data sets can expedite disease-gene discovery, by using them to identify the gene causing Leigh syndrome, French-Canadian type (LSFC, Online Mendelian Inheritance in Man no. 220111), a human cytochrome c oxidase deficiency that maps to chromosome 2p16-21. Using four public RNA expression data sets, we assigned to all human genes a "score" reflecting their similarity in RNA-expression profiles to known mitochondrial genes. Using a large survey of organellar proteomics, we similarly classified human genes according to the likelihood of their protein product being associated with the mitochondrion. By intersecting this information with the relevant genomic region, we identified a single clear candidate gene, LRPPRC. Resequencing identified two mutations on two independent haplotypes, providing definitive genetic proof that LRPPRC indeed causes LSFC. LRPPRC encodes an mRNA-binding protein likely involved with mtDNA transcript processing, suggesting an additional mechanism of mitochondrial pathophysiology. Similar strategies to integrate diverse genomic information can be applied likewise to other disease pathways and will become increasingly powerful with the growing wealth of diverse, functional genomics data.

Amino Acid Sequence↗

Meta-analysis of genetic association studies supports a contribution of common variants to susceptibility to common disease.

Association studies offer a potentially powerful approach to identify genetic variants that influence susceptibility to common disease, but are plagued by the impression that they are not consistently reproducible. In principle, the inconsistency may be due to false positive studies, false negative studies or true variability in association among different populations. The critical question is whether false positives overwhelmingly explain the inconsistency. We analyzed 301 published studies covering 25 different reported associations. There was a large excess of studies replicating the first positive reports, inconsistent with the hypothesis of no true positive associations (P < 10(-14)). This excess of replications could not be reasonably explained by publication bias and was concentrated among 11 of the 25 associations. For 8 of these 11 associations, pooled analysis of follow-up studies yielded statistically significant replication of the first report, with modest estimated genetic effects. Thus, a sizable fraction (but under half) of reported associations have strong evidence of replication; for these, false negative, underpowered studies probably contribute to inconsistent replication. We conclude that there are probably many common variants in the human genome with modest but real effects on common disease risk, and that studies using large samples will convincingly identify such variants.

Alleles↗

Pathogen discovery from human tissue by sequence-based computational subtraction.

We have recently reported a new pathogen discovery approach, "computational subtraction". With this approach, non-human transcripts are detected by sequencing cDNA libraries from infected tissue and eliminating those transcripts that match the human genome. We show now that this method is experimentally feasible. We generated a cDNA library from a tissue sample of post-transplant lymphoproliferative disorder (PTLD). 27,840 independent cDNA sequences were filtered by computational subtraction against the known human sequence to identify 32 nonmatching transcripts. Of these, 22 (0.1%) were found to be amplifiable from both infected and noninfected samples and were inferred to be human DNA not yet contained in the available human genome sequence. The remaining 10 sequences could be amplified only from Epstein-Barr virus (EBV)-infected tissues. All 10 corresponded to the known EBV sequence. This proof-of-principle experiment demonstrates that computational subtraction can detect pathogenic microbes in primary human-diseased tissue.

DNA, Complementary↗

Genetic dissection of lymphopenia from autoimmunity by introgression of mutated Ian5 gene onto the F344 rat.

Peripheral T cell lymphopenia (lyp) in the BioBreeding (BB) rat is linked to a frameshift mutation in Ian5, a member of the Immune Associated Nucleotide (Ian) gene family on rat chromosome 4. This lymphopenia leads to type 1 (insulin-dependent) diabetes mellitus (T1DM) at rates up to 100% when combined with the BB rat MHC RT1 u/u genotype. In order, to better study the lymphopenia phenotype without possible confounding effects of diabetes or other autoimmune disease, we generated congenic F344.lyp rats by introgression of lyp on diabetes-resistant MHC RT1 lv1/lv1 F344 rats. Analysis of thymic CD4 and CD8 T lymphocytes revealed no difference in the percentage of CD4(-)CD8(+)and CD4(+)CD8(-)subsets in lyp/lyp compared to +/+ F344 rats. The same subsets was however dramatically reduced in blood (P=0.005), spleen (P=0.019) and mesenteric lymph nodes (MLN) (P<0.0001). Compared to F344 +/+ rats double positive CD4(+)CD8(+)T cells were increased only in lyp/lyp spleen (P=0.034) while double negative CD4(-)CD8(-)were increased in thymus (P=0.033), spleen (P=0.012), MLN (P<0.0001), and peripheral blood (P<0.0001). There were no signs of inflammatory lesions in organs and tissues in F344.lyp/lyp rats examined at 120 days of age or older. We thus conclude that the lymphopenia phenotype was reconstituted by introgression of lyp on to F344 rats without subsequent development of organ-specific autoimmunity. The congenic F344.lyp rat should prove useful to dissect the mechanisms by which the Ian5 frameshift mutation affects T cell selection, differentiation and maturation without organ-specific autoimmunity.

Animals↗