Search PubMed⌕ Search

Biomedical subjects

Yvonne J K Edwards

Publications and source records attributed to Yvonne J K Edwards.

9 recordsLinked to original sources

Highly conserved non-coding sequences are associated with vertebrate development.

In addition to protein coding sequence, the human genome contains a significant amount of regulatory DNA, the identification of which is proving somewhat recalcitrant to both in silico and functional methods. An approach that has been used with some success is comparative sequence analysis, whereby equivalent genomic regions from different organisms are compared in order to identify both similarities and differences. In general, similarities in sequence between highly divergent organisms imply functional constraint. We have used a whole-genome comparison between humans and the pufferfish, Fugu rubripes, to identify nearly 1,400 highly conserved non-coding sequences. Given the evolutionary divergence between these species, it is likely that these sequences are found in, and furthermore are essential to, all vertebrates. Most, and possibly all, of these sequences are located in and around genes that act as developmental regulators. Some of these sequences are over 90% identical across more than 500 bases, being more highly conserved than coding sequence between these two species. Despite this, we cannot find any similar sequences in invertebrate genomes. In order to begin to functionally test this set of sequences, we have used a rapid in vivo assay system using zebrafish embryos that allows tissue-specific enhancer activity to be identified. Functional data is presented for highly conserved non-coding sequences associated with four unrelated developmental regulators (SOX21, PAX6, HLXB9, and SHH), in order to demonstrate the suitability of this screen to a wide range of genes and expression patterns. Of 25 sequence elements tested around these four genes, 23 show significant enhancer activity in one or more tissues. We have identified a set of non-coding sequences that are highly conserved throughout vertebrates. They are found in clusters across the human genome, principally around genes that are implicated in the regulation of development, including many transcription factors. These highly conserved non-coding sequences are likely to form part of the genomic circuitry that uniquely defines vertebrate development.

Animals↗

A Fugu-Human Genome Synteny Viewer: web software for graphical display and annotation reports of synteny between Fugu genomic sequence and human genes.

A web server has been developed to access annotation and graphical reports of synteny and gene order between the Fugu genome and human genes. In this system, the assembled Fugu genomic sequences (also known as scaffolds) are annotated. The annotations for each Fugu scaffold are computed, stored and made publicly available. The annotations describe matches to human homologous genes. For each significant human gene match on the Fugu scaffold, the corresponding human chromosome map and measures of the significance of each match are given. The web-based server provides public access to these annotations and graphical displays of the results. The user is provided with a selection of views including a chromosome-colour-coded image and a table containing the details of the matches. The Fugu-Human Genome Synteny Viewer has been tested by comparing results with examples from a paper that includes a study of transcription factors, Fos and Jun encoding regions. The Fugu-human genome synteny views are available for each Fugu scaffold through the clonesearch web page located at the Fugu Genomics website (http://fugu.rfcgr.mrc.ac.uk/).

Animals↗

Molecular characterisation of the SAND protein family: a study based on comparative genomics, structural bioinformatics and phylogeny.

The activities of vertebrate lysosomes are critical to many essential cellular processes. The yeast vacuole is analogous to the mammalian lysosome and is used as a tool to gain insights into vesicle mediated vacuolar/lysosome transport. The protein SAND, which does not contain a SAND domain (PFAM accession number PF01342), has recently been shown to function at the tethering/docking stage of vacuole fusion as a critical component of the vacuole SNARE complex. In this publication we have identified SAND in diverse eukaryotes, from single celled organisms such as the yeasts to complex multi-cellular chordates such as mammals. We have demonstrated subfamily divisions in the SAND proteins and show that in vertebrates, a duplication event gave rise to two SAND sequences. This duplication appears to have occurred during early vertebrate evolution and conceivably with the evolution of lysosomes. Using bioinformatics we predict a secondary structure, solvent accessibility profile and protein fold for the SAND proteins and determine conserved sequence motifs, present in all SAND proteins and those that are specific to subsets. A comprehensive evaluation of yeast and human functional studies in conjunction with our in silico analysis has identified potential roles for some of these motifs.

Amino Acid Sequence↗

Fugu ESTs: new resources for transcription analysis and genome annotation.

The draft Fugu rubripes genome was released in 2002, at which time relatively few cDNAs were available to aid in the annotation of genes. The data presented here describe the sequencing and analysis of 24,398 expressed sequence tags (ESTs) generated from 15 different adult and juvenile Fugu tissues, 74% of which matched protein database entries. Analysis of the EST data compared with the Fugu genome data predicts that approximately 10,116 gene tags have been generated, covering almost one-third of Fugu predicted genes. This represents a remarkable economy of effort. Comparison with the Washington University zebrafish EST assemblies indicates strong conservation within fish species, but significant differences remain. This potentially represents divergence of sequence in the 5' terminal exons and UTRs between these two fish species, although clearly, complete EST data sets are not available for either species. This project provides new Fugu resources, and the analysis adds significant weight to the argument that EST programs remain an essential resource for genome exploitation and annotation. This is particularly timely with the increasing availability of draft genome sequence from different organisms and the mounting emphasis on gene function and regulation.

Animals↗

Theatre: A software tool for detailed comparative analysis and visualization of genomic sequence.

Theatre is a web-based computing system designed for the comparative analysis of genomic sequences, especially with respect to motifs likely to be involved in the regulation of gene expression. Theatre is an interface to commonly used sequence analysis tools and biological sequence databases to determine or predict the positions of coding regions, repetitive sequences and transcription factor binding sites in families of DNA sequences. The information is displayed in a manner that can be easily understood and can reveal patterns that might not otherwise have been noticed. In addition to web-based output, Theatre can produce publication quality colour hardcopies showing predicted features in aligned genomic sequences. A case study using the p53 promoter region of four mammalian species and two fish species is described. Unlike the mammalian sequences the promoter regions in fish have not been previously predicted or characterized and we report the differences in the p53 promoter region of four mammals and that predicted for two fish species. Theatre can be accessed at http://www.hgmp.mrc.ac.uk/Registered/Webapp/theatre/.

Animals↗

AP1 genes in Fugu indicate a divergent transcriptional control to that of mammals.

The draft genomic sequence of the Japanese puffer fish, Fugu rubripes, has now been announced. This is the first complete sequence of a teleost fish and the second available vertebrate sequence, the first being that of human. For the first time, whole-genome comparisons between two vertebrates can be undertaken. Early analysis has suggested that there may be surprising differences in gene regulation between human and fish. In mammals, a gene commonly has several functions, and this may not always be the case in fish. Many gene families comprise more members in fish than they do in mammals, possibly because each fish gene has evolved an individual function. Complexities of gene regulation in mammals has hampered studies of all biological processes from cell proliferation to cell death. Determining the activities of the AP1 transcription factor proteins has been non-trivial. The AP1 complex typically comprises two proteins, a Jun (c-Jun, JunB, and JunD) and a Fos (c-Fos, FosB, Fra1, and Fra2). These proteins can form both homodimers and heterodimers among-themselves and can interact with additional proteins; thus, dissecting their individual roles has been difficult. We have determined that Fugu has more Jun and Fos genes than mammals, and if each proves to have a separate function, then addressing the roles of the individual AP1 proteins in Fugu may be simpler than in human.

Amino Acid Sequence↗

Bioinformatics methods to predict protein structure and function. A practical approach.

Protein structure prediction by using bioinformatics can involve sequence similarity searches, multiple sequence alignments, identification and characterization of domains, secondary structure prediction, solvent accessibility prediction, automatic protein fold recognition, constructing three-dimensional models to atomic detail, and model validation. Not all protein structure prediction projects involve the use of all these techniques. A central part of a typical protein structure prediction is the identification of a suitable structural target from which to extrapolate three-dimensional information for a query sequence. The way in which this is done defines three types of projects. The first involves the use of standard and well-understood techniques. If a structural template remains elusive, a second approach using nontrivial methods is required. If a target fold cannot be reliably identified because inconsistent results have been obtained from nontrivial data analyses, the project falls into the third type of project and will be virtually impossible to complete with any degree of reliability. In this article, a set of protocols to predict protein structure from sequence is presented and distinctions among the three types of project are given. These methods, if used appropriately, can provide valuable indicators of protein structure and function.

Algorithms↗

Whole-genome shotgun assembly and analysis of the genome of Fugu rubripes.

The compact genome of Fugu rubripes has been sequenced to over 95% coverage, and more than 80% of the assembly is in multigene-sized scaffolds. In this 365-megabase vertebrate genome, repetitive DNA accounts for less than one-sixth of the sequence, and gene loci occupy about one-third of the genome. As with the human genome, gene loci are not evenly distributed, but are clustered into sparse and dense regions. Some "giant" genes were observed that had average coding sequence sizes but were spread over genomic lengths significantly larger than those of their human orthologs. Although three-quarters of predicted human proteins have a strong match to Fugu, approximately a quarter of the human proteins had highly diverged from or had no pufferfish homologs, highlighting the extent of protein evolution in the 450 million years since teleosts and mammals diverged. Conserved linkages between Fugu and human genes indicate the preservation of chromosomal segments from the common vertebrate ancestor, but with considerable scrambling of gene order.

Animals↗

Fugu orthologues of human major histocompatibility complex genes: a genome survey.

The major histocompatibility complex (MHC) region in fish has been subjected to piecemeal analysis centering on the in-depth characterization of single genes. The emphasis has been on those genes proven to be involved in the immune response such as the class I and class II antigen presenting genes and the complement genes. The Fugu genome data presents the opportunity to examine the short-range linkage of potentially all the human MHC orthologues and examine conserved synteny with the human and, to a more limited extent, zebrafish genomes. Analysis confirms the existence of a limited MHC locus in Fugu comprising the MHC class Ia genes and associated class II region genes involved in class I antigen presentation. Identification of additional human MHC orthologues indicates the completely dispersed nature of this region in fish, with a maximum of six MHC genes maintained within close proximity in any one contig. The majority of the other genes are present in the genome data as either singletons or pairs. Comparison with zebrafish substantiates previously observed linkages between class III region orthologues and hints at an ancient conserved class III region.

Animals↗