Search PubMed⌕ Search

Biomedical subjects

W Miller

Publications and source records attributed to W Miller.

At least 55 records · Page 3Linked to original sources

Genome sequence comparisons: hurdles in the fast lane to functional genomics.

An important computational technique for extracting the wealth of information hidden in human genomic sequence data is to compare the sequence with that from the corresponding region of the mouse genome, looking for segments that are conserved over evolutionary time. Moreover, the approach generalises to comparison of sequences from any two related species. The underlying rationale (which is abundantly confirmed by observation) is that a random mutation in a functional region is usually deleterious to the organism, and hence unlikely to become fixed in the population, whereas mutations in a non-functional region are free to accumulate over time. The potential value of this approach is so attractive that the public and private projects to sequence the human genome are now turning to sequencing the mouse, and you will soon be able to compare the human and mouse sequences of your favourite genomic region. We are currently witnessing an explosion of computer tools for comparative analysis of two genomic sequences. Here the capabilities of two new network servers for comparing genomic sequences from any pair of closely related species are sketched. The Syntenic Gene Prediction Program SGP-I utilises sequence comparisons to enhance the ability to locate protein coding segments in genomic data. PipMaker attempts to determine all conserved genomic regions, regardless of their function.

Animals↗

PipMaker--a web server for aligning two genomic DNA sequences.

PipMaker (http://bio.cse.psu.edu) is a World-Wide Web site for comparing two long DNA sequences to identify conserved segments and for producing informative, high-resolution displays of the resulting alignments. One display is a percent identity plot (pip), which shows both the position in one sequence and the degree of similarity for each aligning segment between the two sequences in a compact and easily understandable form. Positions along the horizontal axis can be labeled with features such as exons of genes and repetitive elements, and colors can be used to clarify and enhance the display. The web site also provides a plot of the locations of those segments in both species (similar to a dot plot). PipMaker is appropriate for comparing genomic sequences from any two related species, although the types of information that can be inferred (e.g., protein-coding regions and cis-regulatory elements) depend on the level of conservation and the time and divergence rate since the separation of the species. Gene regulatory elements are often detectable as similar, noncoding sequences in species that diverged as much as 100-300 million years ago, such as humans and mice, Caenorhabditis elegans and C. briggsae, or Escherichia coli and Salmonella spp. PipMaker supports analysis of unfinished or "working draft" sequences by permitting one of the two sequences to be in unoriented and unordered contigs.

Animals↗

Comparative genome sequence analysis of the Bpa/Str region in mouse and Man.

The progress of human and mouse genome sequencing programs presages the possibility of systematic cross-species comparison of the two genomes as a powerful tool for gene and regulatory element identification. As the opportunities to perform comparative sequence analysis emerge, it is important to develop parameters for such analyses and to examine the outcomes of cross-species comparison. Our analysis used gene prediction and a database search of 430 kb of genomic sequence covering the Bpa/Str region of the mouse X chromosome, and 745 kb of genomic sequence from the homologous human X chromosome region. We identified 11 genes in mouse and 13 genes and two pseudogenes in human. In addition, we compared the mouse and human sequences using pairwise alignment and searches for evolutionary conserved regions (ECRs) exceeding a defined threshold of sequence identity. This approach aided the identification of at least four further putative conserved genes in the region. Comparative sequencing revealed that this region is a mosaic in evolutionary terms, with considerably more rearrangement between the two species than realized previously from comparative mapping studies. Surprisingly, this region showed an extremely high LINE and low SINE content, low G+C content, and yet a relatively high gene density, in contrast to the low gene density usually associated with such regions.

3-Hydroxysteroid Dehydrogenases↗

Genomic sequence analysis of the mouse Naip gene array.

A mouse locus called Lgn1 determines differences in macrophage permissiveness for the intracellular replication of Legionella pneumophila. The only regional candidate genes for this phenotype difference lie within a cluster of closely linked paralogs of the Neuronal Apoptosis Inhibitory Protein (Naip) gene. Previous genetic and physical mapping of the Lgn1 phenotype narrowed it to an interval containing only Naip2 and Naip5, suggesting that there is not complete functional overlap among the mouse Naip loci. In order to gather more information about polymorphisms among the Naip genes of the 129 mouse haplotype, we have determined the genomic sequence of a substantial portion of the 129 Naip gene array. We have constructed an evolutionary model for the expansion of the Naip gene array from a single progenitor Naip gene. This model predicts the presence of two distinct families of Naip paralogs: Naip1/2/3 and Naip4/5/6/7. Unlike the divergences among all the other Naip paralogs, the splits among Naip4, Naip5, Naip6, and Naip7 occurred relatively recently. The high degree of sequence conservation within the Naip4/5/6/7 family increases the likelihood of functional overlap among these genes.

Animals↗

Sequence and comparative analysis of the mouse 1-megabase region orthologous to the human 11p15 imprinted domain.

A major barrier to conceptual advances in understanding the mechanisms and regulation of imprinting of a genomic region is our relatively poor understanding of the overall organization of genes and of the potentially important cis-acting regulatory sequences that lie in the nonexonic segments that make up 97% of the genome. Interspecies sequence comparison offers an effective approach to identify sequence from conserved functional elements. In this article we describe the successful use of this approach in comparing a approximately 1-Mb imprinted genomic domain on mouse chromosome 7 to its orthologous region on human 11p15.5. Within the region, we identified 112 exons of known genes as well as a novel gene identified uniquely in the mouse region, termed Msuit, that was found to be imprinted. In addition to these coding elements, we identified 33 CpG islands and 49 orthologous nonexonic, nonisland sequences that met our criteria as being conserved, and making up 4.1% of the total sequence. These conserved noncoding sequence elements were generally clustered near imprinted genes and the majority were between Igf2 and H19 or within Kvlqt1. Finally, the location of CpG islands provided evidence that suggested a two-island rule for imprinted genes. This study provides the first global view of the architecture of an entire imprinted domain and provides candidate sequence elements for subsequent functional analyses.

Amino Acid Sequence↗

Divergent human and mouse orthologs of a novel gene (WBSCR15/Wbscr15) reside within the genomic interval commonly deleted in Williams syndrome.

Williams syndrome (WS) is a contiguous gene deletion disorder resulting in complex and intriguing clinical features. Detailed molecular characterization studies of the genomic segment on human chromosome 7q11.23 commonly deleted in WS have uncovered numerous genes, each of which is being actively studied for its possible role in the etiology of the syndrome. Our efforts have focused on the comparative mapping and sequencing of the WS region in human and mouse. In previous studies, we uncovered important differences in the long-range organization of these human and mouse genomic regions; in particular, the notable absence of large duplicated blocks of DNA in mouse that are present in human. Aided by available genomic sequence data, we have used a combination of gene-prediction programs and cDNA isolation to identify the human and mouse orthologs of a novel gene (WBSCR15 and Wbscr15, respectively) residing within the genomic segment commonly deleted in WS. Unlike the flanking genes, which are closely related in human and mouse, WBSCR15 and Wbscr15 are strikingly different with respect to their cDNA and corresponding protein sequences as well as tissue-expression pattern. Neither the WBSCR15- nor Wbscr15-encoded amino acid sequence shows a statistically significant similarity to any characterized protein. These findings reveal another interesting evolutionary difference between the human and mouse WS regions and provide an additional candidate gene to evaluate with respect to its possible role in the pathogenesis of WS.

Adaptor Proteins, Signal Transducing↗

Development of a fluorescent ligand-binding assay using the AcroWell filter plate.

One of the most powerful tools for receptor research and drug discovery is the use of receptor-ligand affinity screening of combinatorial libraries. Early work involved the use of radioactive ligands to identify a binding event; however, there are numerous limitations involved in the use of radioactivity for high throughput screening. These limitations have led to the creation of highly sensitive, nonradioactive alternatives to investigate receptor-ligand interactions. Pall Gelman Laboratory has introduced the AcroWell, a patented low-fluorescent-background membrane and sealing process together with a filter plate design that is compatible with robotic systems. Taken together, these allow the AcroWell 96-well filter plate to detect trace quantities of lanthanide-labeled ligands for cell-, bead-, or membrane-based assays using time-resolved fluorescence. Using europium-labeled galanin, we have demonstrated that saturation binding experiments can be performed with low-background fluorescence and signal-to-noise ratios that rival traditional radioisotopic techniques while maintaining biological integrity of the receptor-ligand interaction. In addition, the ability to discriminate between active and inactive compounds in a mock galanin screen is demonstrated with low well-to-well variability, allowing reliable determination of positive hits even for low-affinity interactions.

Cell Line↗

Characterization of the human and mouse unconventional myosin XV genes responsible for hereditary deafness DFNB3 and shaker 2.

Mutations in myosin XV are responsible for congenital profound deafness DFNB3 in humans and deafness and vestibular defects in shaker 2 mice. By combining direct cDNA analyses with a comparison of 95.2 kb of genomic DNA sequence from human chromosome 17p11.2 and 88.4 kb from the homologous region on mouse chromosome 11, we have determined the genomic and mRNA structures of the human (MYO15) and mouse (Myo15) myosin XV genes. Our results indicate that full-length myosin XV transcripts contain 66 exons, are >12 kb in length, and encode 365-kDa proteins that are unique among myosins in possessing very long approximately 1200-aa N-terminal extensions preceding their conserved motor domains. The tail regions of the myosin XV proteins contain two MyTH4 domains, two regions with similarity to the membrane attachment FERM domain, and a putative SH3 domain. Northern and dot blot analyses revealed that myosin XV is expressed in the pituitary gland in both humans and mice. Myosin XV transcripts were also observed by in situ hybridization within areas corresponding to the sensory epithelia of the cochlea and vestibular systems in the developing mouse inner ear. Immunostaining of adult mouse organ of Corti revealed that myosin XV protein is concentrated within the cuticular plate and stereocilia of cochlear sensory hair cells. These results indicate a likely role for myosin XV in the formation or maintenance of the unique actin-rich structures of inner ear sensory hair cells.

Alternative Splicing↗

Comparison of five methods for finding conserved sequences in multiple alignments of gene regulatory regions.

Conserved segments in DNA or protein sequences are strong candidates for functional elements and thus appropriate methods for computing them need to be developed and compared. We describe five methods and computer programs for finding highly conserved blocks within previously computed multiple alignments, primarily for DNA sequences. Two of the methods are already in common use; these are based on good column agreement and high information content. Three additional methods find blocks with minimal evolutionary change, blocks that differ in at most k positions per row from a known center sequence and blocks that differ in at most k positions per row from a center sequence that is unknown a priori. The center sequence in the latter two methods is a way to model potential binding sites for known or unknown proteins in DNA sequences. The efficacy of each method was evaluated by analysis of three extensively analyzed regulatory regions in mammalian beta-globin gene clusters and the control region of bacterial arabinose operons. Although all five methods have quite different theoretical underpinnings, they produce rather similar results on these data sets when their parameters are adjusted to best approximate the experimental data. The optimal parameters for the method based on information content varied little for different regulatory regions of the beta-globin gene cluster and hence may be extrapolated to many other regulatory regions. The programs based on maximum allowed mismatches per row have simple parameters whose values can be chosen a priori and thus they may be more useful than the other methods when calibration against known functional sites is not available.

Animals↗

Comparative sequence analysis of the mouse and human Lgn1/SMA interval.

Human chromosome 5q11.2-q13.3 and its ortholog on mouse chromosome 13 contain candidate genes for an inherited human neurodegenerative disorder called spinal muscular atrophy (SMA) and for an inherited mouse susceptibility to infection with Legionella pneumophila (Lgn1). These homologous genomic regions also have unusual repetitive organizations that create practical difficulties in mapping and raise interesting issues about the evolutionary origin of the repeats. In an attempt to analyze this region in detail, and as a way to identify additional candidate genes for these diseases, we have determined the sequence of 179 kb of the mouse Lgn1/SMA interval. We have analyzed this sequence using BLAST searches and various exon prediction programs to identify potential genes. Since these methods can generate false-positive exon declarations, our alignments of the mouse sequence with available human orthologous sequence allowed us to discriminate rapidly among this collection of potential coding regions by indicating which regions were well conserved and were more likely to represent actual coding sequence. As a result of our analysis, we accurately mapped two additional genes in the SMA interval that can be tested for involvement in the pathogenesis of SMA. While no new Lgn1 candidates emerged, we have identified new genetic markers that exclude Smn as an Lgn1 candidate. In addition to providing important resources for studying SMA and Lgn1, our data provide further evidence of the value of sequencing the mouse genome as a means to help with the annotation of the human genomic sequence and vice versa.

Animals↗

The mouse cornichon gene family.

As part of a large scale mouse Expressed Sequence Tag (EST) project to identify molecules involved in the initiation of mammalian development, a homolog of the Drosophila cornichon gene was detected as a mouse maternal transcript present in the two-cell embryo. Cornichon is a multigene family in the mouse: the new gene, Cnih, maps to mouse chromosome 10, another cornichon homolog, Cnil, maps to chromosome 14 and two additional cornichon-related loci, possibly pseudogenes, localize to chromosomes 3 and 10, respectively. Cnih encodes an open reading frame (ORF) of 144 amino acids that is 93% homologous (68% identical) to the Drosophila protein, whereas the ORF of Cnil contains two extra polypeptide regions not found in these other proteins. Transcripts of Cnih are highly abundant in the full grown oocyte and the ovulated unfertilized egg, while Cnil message is only detectable after activation of the embryonic genome at the eight-cell stage. In situ hybridization shows specific localization of Cnih transcripts to ovarian oocytes. The lack of cytoplasmic polyadenylation of the maternally inherited Cnih transcript suggests that Cnih mRNA is translated in the full grown oocyte before, but not after, ovulation. In Drosophila, cornichon is involved in the establishment of both anterior-posterior and dorso-ventral polarity via the epidermal growth factor (EGF)-receptor signaling pathway. Finding Cnih in the mammalian oocyte opens a new perspective on the investigation of EGF-signaling in the oocyte.

Amino Acid Sequence↗

The gene mutated in bare patches and striated mice encodes a novel 3beta-hydroxysteroid dehydrogenase.

X-linked dominant disorders that are exclusively lethal prenatally in hemizygous males have been described in human and mouse. None of the genes responsible has been isolated in either species. The bare patches (Bpa) and striated (Str) mouse mutations were originally identified in female offspring of X-irradiated males. Subsequently, additional independent alleles were described. We have previously mapped these X-linked dominant, male-lethal mutations to an overlapping region of 600 kb that is homologous to human Xq28 (ref. 4) and identified several candidate genes in this interval. Here we report mutations in one of these genes, Nsdhl, encoding an NAD(P)H steroid dehydrogenase-like protein, in two independent Bpa and three independent Str alleles. Quantitative analysis of sterols from tissues of affected Bpa mice support a role for Nsdhl in cholesterol biosynthesis. Our results demonstrate that Bpa and Str are allelic mutations and identify the first mammalian locus associated with an X-linked dominant, male-lethal phenotype. They also expand the spectrum of phenotypes associated with abnormalities of cholesterol metabolism.

3-Hydroxysteroid Dehydrogenases↗

Post-processing long pairwise alignments.

MOTIVATION: The local alignment problem for two sequences requires determining similar regions, one from each sequence, and aligning those regions. For alignments computed by dynamic programming, current approaches for selecting similar regions may have potential flaws. For instance, the criterion of Smith and Waterman can lead to inclusion of an arbitrarily poor internal segment. Other approaches can generate an alignment scoring less than some of its internal segments. RESULTS: We develop an algorithm that decomposes a long alignment into sub-alignments that avoid these potential imperfections. Our algorithm runs in time proportional to the original alignment's length. Practical applications to alignments of genomic DNA sequences are described.

Algorithms↗

Mapping of sugar and amino acid availability in soil around roots with bacterial sensors of sucrose and tryptophan

We developed a technique to map the availability of sugars and amino acids along live roots in an intact soil-root matrix with native microbial soil flora and fauna present. It will allow us to study interactions between root exudates and soil microorganisms at the fine spatial scale necessary to evaluate mechanisms of nitrogen cycling in the rhizosphere. Erwinia herbicola 299R harboring a promoterless ice nucleation reporter gene, driven by either of two nutrient-responsive promoters, was used as a biosensor. Strain 299RTice exhibits tryptophan-dependent ice nucleation activity, while strain 299R(p61RYice) expresses ice nucleation activity proportional to sucrose concentration in its environment. Both biosensors exhibited up to 100-fold differences in ice nucleation activity in response to varying substrate abundance in culture. The biosensors were introduced into the rhizosphere of the annual grass Avena barbata and, as a control, into bulk soil. Neither strain exhibited significant ice nucleation activity in the bulk soil. Both tryptophan and sucrose were detected in the rhizosphere, but they showed different spatial patterns. Tryptophan was apparently most abundant in soil around roots 12 to 16 cm from the tip, while sucrose was most abundant in soil near the root tip. The largest numbers of bacteria (determined by acridine orange staining and direct microscopy) occurred near root sections with the highest apparent sucrose or tryptophan exudation. High sucrose availability at the root tip is consistent with leakage of photosynthate from immature, rapidly growing root tissues, while tryptophan loss from older root sections may result from lateral root perforation of the root epidermis.

Journal Article↗

Comparative sequence of human and mouse BAC clones from the mnd2 region of chromosome 2p13.

The mnd2 mutation on mouse chromosome 6 produces a progressive neuromuscular disorder. To determine the gene content of the 400-kb mnd2 nonrecombinant region, we sequenced 108 kb of mouse genomic DNA and 92 kb of human genomic sequence from the corresponding region of chromosome 2p13.3. Three genes with the indicated sizes and intergenic distances were identified: D6Mm5e (>/=81 kb)-787 bp-DOK (2 kb)-845 bp-LOR2 (>/=6 kb). D6Mm5e is expressed in many tissues at very low abundance and the predicted 526-residue protein contains no known functional domains. DOK encodes the p62(dok) rasGAP binding protein involved in signal transduction. LOR2 encodes a novel lysyl oxidase-related protein of 757 amino acid residues. We describe a simple search protocol for identification of conserved internal exons in genomic sequence. Evolutionary conservation proved to be a useful criterion for distinguishing between authentic exons and artifactual products obtained by exon amplification, RT-PCR, and 5' RACE. Conserved noncoding sequence elements longer than 80 bp with >/=75% nucleotide sequence identity comprise approximately 1% of the genomic sequence in this region. Comparative analysis of this human and mouse genomic DNA sequence was an efficient method for gene identification and is independent of developmental stage or quantitative level of gene expression. [The sequence data described in this paper have been submitted to the GenBank data library under the following accession numbers: AC003061, mouse BAC clone 245c12; AC003065, human BAC clone h173(E10); AF053368, mouse Lor2 cDNA; AF084363, 108-kb contig from mouse BAC 245c12; AF084364, mouse D6Mm5e cDNA.]

Amino Acid Sequence↗