Search PubMed⌕ Search

Biomedical subjects

Webb Miller

Publications and source records attributed to Webb Miller.

At least 19 recordsLinked to original sources

The structure and evolution of centromeric transition regions within the human genome.

An understanding of how centromeric transition regions are organized is a critical aspect of chromosome structure and function; however, the sequence context of these regions has been difficult to resolve on the basis of the draft genome sequence. We present a detailed analysis of the structure and assembly of all human pericentromeric regions (5 megabases). Most chromosome arms (35 out of 43) show a gradient of dwindling transcriptional diversity accompanied by an increasing number of interchromosomal duplications in proximity to the centromere. At least 30% of the centromeric transition region structure originates from euchromatic gene-containing segments of DNA that were duplicatively transposed towards pericentromeric regions at a rate of six-seven events per million years during primate evolution. This process has led to the formation of a minimum of 28 new transcripts by exon exaptation and exon shuffling, many of which are primarily expressed in the testis. The distribution of these duplicated segments is nonrandom among pericentromeric regions, suggesting that some regions have served as preferential acceptors of euchromatic DNA.

Animals↗

Improvements in the HbVar database of human hemoglobin variants and thalassemia mutations for population and sequence variation studies.

HbVar (http://globin.cse.psu.edu/globin/hbvar/) is a relational database developed by a multi-center academic effort to provide up-to-date and high quality information on the genomic sequence changes leading to hemoglobin variants and all types of thalassemia and hemoglobinopathies. Extensive information is recorded for each variant and mutation, including sequence alterations, biochemical and hematological effects, associated pathology, ethnic occurrence and references. In addition to the regular updates to entries, we report two significant advances: (i) The frequencies for a large number of mutations causing beta-thalassemia in at-risk populations have been extracted from the published literature and made available for the user to query upon. (ii) HbVar has been linked with the GALA (Genome Alignment and Annotation database, available at http://globin.cse.psu.edu/gala/) so that users can combine information on hemoglobin variants and thalassemia mutations with a wide spectrum of genomic data. It also expands the capacity to view and analyze the data, using tools within GALA and the University of California at Santa Cruz (UCSC) Genome Browser.

Databases, Genetic↗

Aligning multiple genomic sequences with the threaded blockset aligner.

We define a "threaded blockset," which is a novel generalization of the classic notion of a multiple alignment. A new computer program called TBA (for "threaded blockset aligner") builds a threaded blockset under the assumption that all matching segments occur in the same order and orientation in the given sequences; inversions and duplications are not addressed. TBA is designed to be appropriate for aligning many, but by no means all, megabase-sized regions of multiple mammalian genomes. The output of TBA can be projected onto any genome chosen as a reference, thus guaranteeing that different projections present consistent predictions of which genomic positions are orthologous. This capability is illustrated using a new visualization tool to view TBA-generated alignments of vertebrate Hox clusters from both the mammalian and fish perspectives. Experimental evaluation of alignment quality, using a program that simulates evolutionary change in genomic sequences, indicates that TBA is more accurate than earlier programs. To perform the dynamic-programming alignment step, TBA runs a stand-alone program called MULTIZ, which can be used to align highly rearranged or incompletely sequenced genomes. We describe our use of MULTIZ to produce the whole-genome multiple alignments at the Santa Cruz Genome Browser.

Animals↗

Regulatory potential scores from genome-wide three-way alignments of human, mouse, and rat.

We generalize the computation of the Regulatory Potential (RP) score from two-way alignments of human and mouse to three-way alignments of human, mouse, and rat. This requires overcoming technical challenges that arise because the complexity of the models underlying the score increases exponentially with the number of species. Despite the close evolutionary proximity of rat to mouse, we find that adding the rat sequence increases our ability to predict genomic sites that regulate gene transcription. A variant of the RP scoring scheme that accounts for local variation in neutral mutational patterns further improves our predictions.

Actins↗

Patterns of insertions and their covariation with substitutions in the rat, mouse, and human genomes.

The rates at which human genomic DNA changes by neutral substitution and insertion of certain families of transposable elements covary in large, megabase-sized segments. We used the rat, mouse, and human genomic DNA sequences to examine these processes in more detail in comparisons over both shorter (rat-mouse) and longer (rodent-primate) times, and demonstrated the generality of the covariation. Different families of transposable elements show distinctive insertion preferences and patterns of variation with substitution rates. SINEs are more abundant in GC-rich DNA, but the regional GC preference for insertion (monitored in young SINEs) differs between rodents and humans. In contrast, insertions in the rodent genomes are predominantly LINEs, which prefer to insert into AT-rich DNA in all three mammals. The insertion frequency of repeats other than SINEs correlates strongly positively with the frequency of substitutions in all species. However, correlations with SINEs show the opposite effects. The correlations are explained only in part by the GC content, indicating that other factors also contribute to the inherent tendency of DNA segments to change over evolutionary time.

Animals↗

zPicture: dynamic alignment and visualization tool for analyzing conservation profiles.

Comparative sequence analysis has evolved as an essential technique for identifying functional coding and noncoding elements conserved throughout evolution. Here, we introduce zPicture, an interactive Web-based sequence alignment and visualization tool for dynamically generating conservation profiles and identifying evolutionarily conserved regions (ECRs). zPicture is highly flexible, because critical parameters can be modified interactively, allowing users to differentially predict ECRs in comparisons of sequences of different phylogenetic distances and evolutionary rates. We demonstrate the application of this module to identify a known regulatory element in the HOXD locus, in which functional ECRs are difficult to discern against the highly conserved genomic background. zPicture also facilitates transcription factor binding-site analysis via the rVista tool portal. We present an example of the HBB complex when zPicture/rVista combination specifically pinpoints to two ECRs containing GATA-1, NF-E2, and TAL1/E47 binding sites that were identified previously as transcriptional enhancers. In addition, zPicture is linked to the UCSC Genome Browser, allowing users to automatically extract sequences and gene annotations for any recorded locus. Finally, we describe how this tool can be efficiently applied to the analysis of nonvertebrate genomes, including those of microbial organisms.

Animals↗

Comparative analysis of the alpha-like globin clusters in mouse, rat, and human chromosomes indicates a mechanism underlying breaks in conserved synteny.

We have sequenced and fully annotated a 65,871-bp region of mouse Chromosome 17 including the Hba-ps4 alpha-globin pseudogene. Comparative sequence analysis with the functional alpha-globin loci at human Chromosome 16p13.3 and mouse Chromosome 11 shows that this segment of mouse Chromosome 17 contains a group of three alpha-like pseudogenes (Hba-psm-Hba-ps4-Hba-q3), similar to the duplicated sets found at the functional mouse cluster on Chromosome 11. In addition, exons 7 to 12 of the mLuc7L gene are present just downstream from the pseudogene cluster, indicating that this clone contains the region in which human 16p13.3 switches in synteny between mouse Chromosomes 11 and 17. Comparison of the sequences around the alpha-like clusters on the two mouse chromosomes reveals the presence of conserved tandem repeats. We propose that these repetitive elements have played a role in the fragmentation of the mouse alpha cluster during evolution.

Animals↗

Gene length and proximity to neighbors affect genome-wide expression levels.

Steady-state levels of mRNA in cells theoretically depend on the rate and efficiency of transcription and posttranscriptional processing, on mRNA stability, on transcriptional interference from other genes, and on poorly defined long-range chromatin effects. Although each of these cellular processes has been studied in detail for a few genes, it is not possible to predict expression levels by simply examining gene sequences. In this report, we have used a bioinformatics approach to identify critical factors that influence expression levels. To simplify the problem, we have limited our analysis to the collection of genes expressed in all tissues, because such genes provide a unique opportunity to distinguish the role of general genomic features that constrain gene expression from the effect of tissue-specific factors. Using correlation and regression techniques, we have investigated the dependence between expression level and morphological parameters (distance to neighbors, gene, mRNA or 3'-UTR length, number of exons, etc.) that can be directly related to transcription, posttranscriptional processing, mRNA stability, or transcriptional interference. We found that, on a genome-wide scale, highly expressed genes are significantly farther from their closest neighboring genes, are smaller, contain a moderate number of exons, and produce shorter mRNAs with shorter 3'-UTRs. This confirms that transcriptional and posttranscriptional processes are highly interrelated and implies that transcriptional interference plays a role in determining steady-state levels of mRNA in cells.

Computational Biology↗

A combinatorial network of evolutionarily conserved myelin basic protein regulatory sequences confers distinct glial-specific phenotypes.

Myelin basic protein (MBP) is required for normal myelin compaction and is implicated in both experimental and human demyelinating diseases. In this study, as an initial step in defining the regulatory network controlling MBP transcription, we located and characterized the function of evolutionarily conserved regulatory sequences. Long-range human-mouse sequence comparison revealed over 1 kb of conserved noncoding MBP 5' flanking sequence distributed into four widely spaced modules ranging from 0.1 to 0.4 kb. We demonstrate first that a controlled strategy of transgenesis provides an effective means to assign and compare qualitative and quantitative in vivo regulatory programs. Using this strategy, single-copy reporter constructs, designed to evaluate the regulatory significance of modular and intermodular sequences, were introduced by homologous recombination into the mouse hprt (hypoxanthine-guanine phosphoribosyltransferase) locus. The proximal modules M1 and M2 confer comparatively low-level oligodendrocyte expression primarily limited to early postnatal development, whereas the upstream M3 confers high-level oligodendrocyte expression extending throughout maturity. Furthermore, constructs devoid of M3 fail to target expression to newly myelinating oligodendrocytes in the mature CNS. Mutation of putative Nkx6.2/Gtx sites within M3, although not eliminating oligodendrocyte targeting, significantly decreases transgene expression levels. High-level and continuous expression is conferred to myelinating or remyelinating Schwann cells by M4. In addition, when isolated from surrounding MBP sequences, M3 confers transient expression to Schwann cells elaborating myelin. These observations define the in vivo regulatory roles played by conserved noncoding MBP sequences and lead to a combinatorial model in which different regulatory modules are engaged during primary myelination, myelin maintenance, and remyelination.

Animals↗

Evolution's cauldron: duplication, deletion, and rearrangement in the mouse and human genomes.

This study examines genomic duplications, deletions, and rearrangements that have happened at scales ranging from a single base to complete chromosomes by comparing the mouse and human genomes. From whole-genome sequence alignments, 344 large (>100-kb) blocks of conserved synteny are evident, but these are further fragmented by smaller-scale evolutionary events. Excluding transposon insertions, on average in each megabase of genomic alignment we observe two inversions, 17 duplications (five tandem or nearly tandem), seven transpositions, and 200 deletions of 100 bases or more. This includes 160 inversions and 75 duplications or transpositions of length >100 kb. The frequencies of these smaller events are not substantially higher in finished portions in the assembly. Many of the smaller transpositions are processed pseudogenes; we define a "syntenic" subset of the alignments that excludes these and other small-scale transpositions. These alignments provide evidence that approximately 2% of the genes in the human/mouse common ancestor have been deleted or partially deleted in the mouse. There also appears to be slightly less nontransposon-induced genome duplication in the mouse than in the human lineage. Although some of the events we detect are possibly due to misassemblies or missing data in the current genome sequence or to the limitations of our methods, most are likely to represent genuine evolutionary events. To make these observations, we developed new alignment techniques that can handle large gaps in a robust fashion and discriminate between orthologous and paralogous alignments.

Animals↗

EnteriX 2003: Visualization tools for genome alignments of Enterobacteriaceae.

We describe EnteriX, a suite of three web-based visualization tools for graphically portraying alignment information from comparisons among several fixed and user-supplied sequences from related enterobacterial species, anchored on a reference genome (http://bio.cse.psu.edu/). The first visualization, Enteric, displays stacked pairwise alignments between a reference genome and each of the related bacteria, represented schematically as PIPs (Percent Identity Plots). Encoded in the views are large-scale genomic rearrangement events and functional landmarks. The second visualization, Menteric, computes and displays 1 Kb views of nucleotide-level multiple alignments of the sequences, together with annotations of genes, regulatory sites and conserved regions. The third, a Java-based tool named Maj, displays alignment information in two formats, corresponding roughly to the Enteric and Menteric views, and adds zoom-in capabilities. The uses of such tools are diverse, from examining the multiple sequence alignment to infer conserved sites with potential regulatory roles, to scrutinizing the commonalities and differences between the genomes for pathogenicity or phylogenetic studies. The EnteriX suite currently includes >15 enterobacterial genomes, generates views centered on four different anchor genomes and provides support for including user sequences in the alignments.

Computer Graphics↗

MultiPipMaker and supporting tools: Alignments and analysis of multiple genomic DNA sequences.

Analysis of multiple sequence alignments can generate important, testable hypotheses about the phylogenetic history and cellular function of genomic sequences. We describe the MultiPipMaker server, which aligns multiple, long genomic DNA sequences quickly and with good sensitivity (available at http://bio.cse.psu.edu/ since May 2001). Alignments are computed between a contiguous reference sequence and one or more secondary sequences, which can be finished or draft sequence. The outputs include a stacked set of percent identity plots, called a MultiPip, comparing the reference sequence with subsequent sequences, and a nucleotide-level multiple alignment. New tools are provided to search MultiPipMaker output for conserved matches to a user-specified pattern and for conserved matches to position weight matrices that describe transcription factor binding sites (singly and in clusters). We illustrate the use of MultiPipMaker to identify candidate regulatory regions in WNT2 and then demonstrate by transfection assays that they are functional. Analysis of the alignments also confirms the phylogenetic inference that horses are more closely related to cats than to cows.

Algorithms↗

Transcription-associated mutational asymmetry in mammalian evolution.

Although mutation is commonly thought of as a random process, evolutionary studies show that different types of nucleotide substitution occur with widely varying rates that presumably reflect biases intrinsic to mutation and repair mechanisms. A strand asymmetry, the occurrence of particular substitution types at higher rates than their complementary types, that is associated with DNA replication has been found in bacteria and mitochondria. A strand asymmetry that is associated with transcription and attributable to higher rates of cytosine deamination on the coding strand has been observed in enterobacteria. Here, we describe a qualitatively different transcription-associated strand asymmetry in mammals, which may be a byproduct of transcription-coupled repair in germline cells. This mutational asymmetry has acted over long periods of time to produce a compositional asymmetry, an excess of G+T over A+C on the coding strand, in most genes. The mutational and compositional asymmetries can be used to detect the orientations and approximate extents of transcribed regions.

Animals↗

Dancing with complement C4 and the RP-C4-CYP21-TNX (RCCX) modules of the major histocompatibility complex.

The number of the complement component C4 genes varies from 2 to 8 in a diploid genome among different human individuals. Three quarters of the C4 genes in Caucasian populations have the endogenous retrovirus, HERV-K(C4), in the ninth intron. The remainder does not. The C4 serum proteins are highly polymorphic and their concentrations vary from 100 to approximately 1000 microg/ml. There are two distinct classes of C4 protein, C4A and C4B, which have diversified to fulfill (a) the opsonization/immunoclearance purposes and (b) the well-known complement function in the killing of microbes by lysis and neutralization, respectively. Many infectious and autoimmune diseases are associated with complete or partial deficiency of C4A and/or C4B. The adverse effects of high C4 gene dosages, however, are just emerging, as the concepts of human C4 genetics are revised and accurate techniques are applied to distinguish partial deficiencies from differential expression caused by unequal C4A and C4B gene dosages and gene sizes. This review attempts to dissect the sophisticated genetics of complement C4A and C4B. The emphases are on the qualitative and quantitative diversities of C4 genotypes and phenotypes. The many allotypic variants and the processed products of human and mouse C4 proteins are described. The modular variation of C4 genes together with the serine/threonine nuclear kinase gene RP, the steroid 21-hydroxylase CYP21, and extracellular matrix protein TNX (RCCX modules) are investigated for the effects on homogenization of C4 protein polymorphisms, and on the unequal genetic crossovers that knocked out the functions of CYP21 and/or TNX. Furthermore, the influence of the endogenous retrovirus HERV-K(C4) on C4 gene expression and the dispersal of HERV-K(C4) family members in the human genome are discussed.

Animals↗

Multispecies comparative analysis of a mammalian-specific genomic domain encoding secretory proteins.

The mammalian-specific casein gene cluster comprises 3 or 4 evolutionarily related genes and 1 physically linked gene with a functional association. To gain a better understanding of the mechanisms regulating the entire casein cluster at the genomic level we initiated a multispecies comparative sequence analysis. Despite the high level of divergence at the coding level, these studies have identified uncharacterized family members within two species and the presence at orthologous positions of previously uncharacterized genes. Also the previous suggestion that the histatin/statherin gene family, located in this region, was primate specific was ruled out. All 11 genes identified in this region appear to encode secretory proteins. Conservation of a number of noncoding regions was observed; one coincides with an element previously suggested to be important for beta-casein gene expression in human and cow. The conserved regions might have biological importance for the regulation of genes in this genomic "neighborhood."

Amino Acid Sequence↗

Significance of interspecies matches when evolutionary rate varies.

We develop techniques to estimate the statistical significance of gap-free alignments between two genomic DNA sequences, using human-mouse alignments as an example. The sequences are assumed to be sufficiently similar that some but not all of the neutrally evolving regions (i.e., those under no evolutionary constraint) can be reliably aligned. Our goal is to model the situation in which the neutral rate of evolution, and hence the extent of the aligning intervals, varies across the genome. In some cases, this permits the weaker of two matches to be judged as less likely to have arisen by chance, provided it lies in a genomic interval with a high level of background divergence. We employ a hidden Markov model to capture variations in divergence rates and assign probability values to gap-free alignments using techniques of Dembo and Karlin, which are related to those used for the same purpose by BLAST. Our methods are illustrated in detail using a 1.49 Mb genomic region. Results obtained from the analysis of human chromosome 22 using these techniques are also provided.

Animals↗

GALA, a database for genomic sequence alignments and annotations.

We have developed a relational database to contain whole genome sequence alignments between human and mouse with extensive annotations of the human sequence. Complex queries are supported on recorded features, both directly and on proximity among them. Searches can reveal a wide variety of relationships, such as finding all genes expressed in a designated tissue that have a highly conserved noncoding sequence 5' to the start site. Other examples are finding single nucleotide polymorphisms that occur in conserved noncoding regions upstream of genes and identifying CpG islands that overlap the 5' ends of divergently transcribed genes. The database is available online at http://globin.cse.psu.edu/ and http://bio.cse.psu.edu/.

5' Untranslated Regions↗

Human-mouse alignments with BLASTZ.

The Mouse Genome Analysis Consortium aligned the human and mouse genome sequences for a variety of purposes, using alignment programs that suited the various needs. For investigating issues regarding genome evolution, a particularly sensitive method was needed to permit alignment of a large proportion of the neutrally evolving regions. We selected a program called BLASTZ, an independent implementation of the Gapped BLAST algorithm specifically designed for aligning two long genomic sequences. BLASTZ was subsequently modified, both to attain efficiency adequate for aligning entire mammalian genomes and to increase its sensitivity. This work describes BLASTZ, its modifications, the hardware environment on which we run it, and several empirical studies to validate its results.

Animals↗