Search PubMed⌕ Search

Biomedical subjects

Richard R Copley

Publications and source records attributed to Richard R Copley.

At least 19 recordsLinked to original sources

Deuterostome phylogeny reveals monophyletic chordates and the new phylum Xenoturbellida.

Deuterostomes comprise vertebrates, the related invertebrate chordates (tunicates and cephalochordates) and three other invertebrate taxa: hemichordates, echinoderms and Xenoturbella. The relationships between invertebrate and vertebrate deuterostomes are clearly important for understanding our own distant origins. Recent phylogenetic studies of chordate classes and a sea urchin have indicated that urochordates might be the closest invertebrate sister group of vertebrates, rather than cephalochordates, as traditionally believed. More remarkable is the suggestion that cephalochordates are closer to echinoderms than to vertebrates and urochordates, meaning that chordates are paraphyletic. To study the relationships among all deuterostome groups, we have assembled an alignment of more than 35,000 homologous amino acids, including new data from a hemichordate, starfish and Xenoturbella. We have also sequenced the mitochondrial genome of Xenoturbella. We support the clades Olfactores (urochordates and vertebrates) and Ambulacraria (hemichordates and echinoderms). Analyses using our new data, however, do not support a cephalochordate and echinoderm grouping and we conclude that chordates are monophyletic. Finally, nuclear and mitochondrial data place Xenoturbella as the sister group of the two ambulacrarian phyla. As such, Xenoturbella is shown to be an independent phylum, Xenoturbellida, bringing the number of living deuterostome phyla to four.

Animals↗

SMART 5: domains in the context of genomes and networks.

The Simple Modular Architecture Research Tool (SMART) is an online resource (http://smart.embl.de/) used for protein domain identification and the analysis of protein domain architectures. Many new features were implemented to make SMART more accessible to scientists from different fields. The new 'Genomic' mode in SMART makes it easy to analyze domain architectures in completely sequenced genomes. Domain annotation has been updated with a detailed taxonomic breakdown and a prediction of the catalytic activity for 50 SMART domains is now available, based on the presence of essential amino acids. Furthermore, intrinsically disordered protein regions can be identified and displayed. The network context is now displayed in the results page for more than 350 000 proteins, enabling easy analyses of domain interactions.

Catalysis↗

A high-resolution single nucleotide polymorphism genetic map of the mouse genome.

High-resolution genetic maps are required for mapping complex traits and for the study of recombination. We report the highest density genetic map yet created for any organism, except humans. Using more than 10,000 single nucleotide polymorphisms evenly spaced across the mouse genome, we have constructed genetic maps for both outbred and inbred mice, and separately for males and females. Recombination rates are highly correlated in outbred and inbred mice, but show relatively low correlation between males and females. Differences between male and female recombination maps and the sequence features associated with recombination are strikingly similar to those observed in humans. Genetic maps are available from http://gscan.well.ox.ac.uk/#genetic_map and as supporting information to this publication.

Animals↗

The EH1 motif in metazoan transcription factors.

BACKGROUND: The Engrailed Homology 1 (EH1) motif is a small region, believed to have evolved convergently in homeobox and forkhead containing proteins, that interacts with the Drosophila protein groucho (C. elegans unc-37, Human Transducin-like Enhancers of Split). The small size of the motif makes its reliable identification by computational means difficult. I have systematically searched the predicted proteomes of Drosophila, C. elegans and human for further instances of the motif. RESULTS: Using motif identification methods and database searching techniques, I delimit which homeobox and forkhead domain containing proteins also have likely EH1 motifs. I show that despite low database search scores, there is a significant association of the motif with transcription factor function. I further show that likely EH1 motifs are found in combination with T-Box, Zinc Finger and Doublesex domains as well as discussing other plausible candidate associations. I identify strong candidate EH1 motifs in basal metazoan phyla. CONCLUSION: Candidate EH1 motifs exist in combination with a variety of transcription factor domains, suggesting that these proteins have repressor functions. The distribution of the EH1 motif is suggestive of convergent evolution, although in many cases, the motif has been conserved throughout bilaterian orthologs. Groucho mediated repression was established prior to the evolution of bilateria.

Amino Acid Motifs↗

Missense mutation in sterile alpha motif of novel protein SamCystin is associated with polycystic kidney disease in (cy/+) rat.

Autosomal dominant polycystic kidney disease (PKD) is the most common genetic disease that leads to kidney failure in humans. In addition to the known causative genes PKD1 and PKD2, there are mutations that result in cystic changes in the kidney, such as nephronophthisis, autosomal recessive polycystic kidney disease, or medullary cystic kidney disease. Recent efforts to improve the understanding of renal cystogenesis have been greatly enhanced by studies in rodent models of PKD. Genetic studies in the (cy/+) rat showed that PKD spontaneously develops as a consequence of a mutation in a gene different from the rat orthologs of PKD1 and PKD2 or other genes that are known to be involved in human cystic kidney diseases. This article reports the positional cloning and mutation analysis of the rat PKD gene, which revealed a C to T transition that replaces an arginine by a tryptophan at amino acid 823 in the protein sequence. It was determined that Pkdr1 is specifically expressed in renal proximal tubules and encodes a novel protein, SamCystin, that contains ankyrin repeats and a sterile alpha motif. The characterization of this protein, which does not share structural homologies with known polycystins, may give new insights into the pathophysiology of renal cyst development in patients.

Animals↗

Variation in structural location and amino acid conservation of functional sites in protein domain families.

BACKGROUND: The functional sites of a protein present important information for determining its cellular function and are fundamental in drug design. Accordingly, accurate methods for the prediction of functional sites are of immense value. Most available methods are based on a set of homologous sequences and structural or evolutionary information, and assume that functional sites are more conserved than the average. In the analysis presented here, we have investigated the conservation of location and type of amino acids at functional sites, and compared the behaviour of functional sites between different protein domains. RESULTS: Functional sites were extracted from experimentally determined structural complexes from the Protein Data Bank harbouring a conserved protein domain from the SMART database. In general, functional (i.e. interacting) sites whose location is more highly conserved are also more conserved in their type of amino acid. However, even highly conserved functional sites can present a wide spectrum of amino acids. The degree of conservation strongly depends on the function of the protein domain and ranges from highly conserved in location and amino acid to very variable. Differentiation by binding partner shows that ion binding sites tend to be more conserved than functional sites binding peptides or nucleotides. CONCLUSION: The results gained by this analysis will help improve the accuracy of functional site prediction and facilitate the characterization of unknown protein sequences.

Amino Acid Sequence↗

A RING-type ubiquitin ligase family member required to repress follicular helper T cells and autoimmunity.

Despite the sequencing of the human and mouse genomes, few genetic mechanisms for protecting against autoimmune disease are currently known. Here we systematically screen the mouse genome for autoimmune regulators to isolate a mouse strain, sanroque, with severe autoimmune disease resulting from a single recessive defect in a previously unknown mechanism for repressing antibody responses to self. The sanroque mutation acts within mature T cells to cause formation of excessive numbers of follicular helper T cells and germinal centres. The mutation disrupts a repressor of ICOS, an essential co-stimulatory receptor for follicular T cells, and results in excessive production of the cytokine interleukin-21. sanroque mice fail to repress diabetes-causing T cells, and develop high titres of autoantibodies and a pattern of pathology consistent with lupus. The causative mutation is in a gene of previously unknown function, roquin (Rc3h1), which encodes a highly conserved member of the RING-type ubiquitin ligase protein family. The Roquin protein is distinguished by the presence of a CCCH zinc-finger found in RNA-binding proteins, and localization to cytosolic RNA granules implicated in regulating messenger RNA translation and stability.

Amino Acid Sequence↗

Animal phylogeny: fatal attraction.

MPhylogenetic analyses of hundreds of genes from model animals have placed flies closer to vertebrates than to nematodes; recent work suggests this may be due to an artefact known as long branch attraction.

Animals↗

Genetic dissection of a behavioral quantitative trait locus shows that Rgs2 modulates anxiety in mice.

Here we present a strategy to determine the genetic basis of variance in complex phenotypes that arise from natural, as opposed to induced, genetic variation in mice. We show that a commercially available strain of outbred mice, MF1, can be treated as an ultrafine mosaic of standard inbred strains and accordingly used to dissect a known quantitative trait locus influencing anxiety. We also show that this locus can be subdivided into three regions, one of which contains Rgs2, which encodes a regulator of G protein signaling. We then use quantitative complementation to show that Rgs2 is a quantitative trait gene. This combined genetic and functional approach should be applicable to the analysis of any quantitative trait.

Animals↗

SMART 4.0: towards genomic data integration.

SMART (Simple Modular Architecture Research Tool) is a web tool (http://smart.embl.de/) for the identification and annotation of protein domains, and provides a platform for the comparative study of complex domain architectures in genes and proteins. The January 2004 release of SMART contains 685 protein domains. New developments in SMART are centred on the integration of data from completed metazoan genomes. SMART now uses predicted proteins from complete genomes in its source sequence databases, and integrates these with predictions of orthology. New visualization tools have been developed to allow analysis of gene intron-exon structure within the context of protein domain structure, and to align these displays to provide schematic comparisons of orthologous genes, or multiple transcripts from the same gene. Other improvements include the ability to query SMART by Gene Ontology terms, improved structure database searching and batch retrieval of multiple entries.

Algorithms↗

Evolutionary convergence of alternative splicing in ion channels.

In Drosophila melanogaster and humans, members of three different ion-channel gene families share tandem exon duplications, which are alternatively spliced. In this article, I demonstrate that the duplication events that give rise to these mutually exclusive exons are unlikely to be ancestral but have probably occurred independently in different lineages. These events provide remarkable examples of evolutionary convergence in alternative splicing. The result has important implications for the analysis of regulation of alternative splicing using comparative genomics and our understanding of molecular evolution.

Alternative Splicing↗

Analysis of the human VPS13 gene family.

The gene mutated in chorea-acanthocytosis (CHAC; approved gene symbol VPS13A) encodes chorein, a protein similar to yeast Vps13p. We detected several similar putative human proteins by BLAST analysis of chorein. We characterized the structure of three new genes encoding these CHAC-similar proteins, located on chromosomes 1p36, 8q22, and 15q21. The most similar gene in yeast to all four human genes is Vps13, and therefore the human genes were named VPS13A (CHAC, 9q21), VPS13B (8q22), VPS13C (15q21), and VPS13D (1p36). VPS13B has recently been reported as COH1, altered in Cohen syndrome. For each gene, we describe several alternative splicing variants; at least two transcripts per gene are major forms. The expression pattern of these genes is ubiquitous, with some tissue-specific differences between several transcript variants. Protein sequence comparisons suggest that intramolecular duplications have played an important role in the evolution of this gene family.

Alternative Splicing↗

Occurrence and consequences of coding sequence insertions and deletions in Mammalian genomes.

Nucleotide insertion and deletion (indel) events, together with substitutions, represent the major mutational processes of gene evolution. Through the alignment of 8148 orthologous genes from human, mouse, and rat, we have identified 1743 indel events within rodent protein-coding sequences. Using human as an out-group, we reconstructed the mutational event underlying each of these indels. Overall, we found an excess of deletions over insertions, particularly for the rat lineage (70% excess). Sequence slippage accounts for at least 52% of insertions and 38% of deletions. We have also evaluated the selective tolerance of identifiable protein structures to indels. Transmembrane domains are the least, and low complexity regions, the most tolerant. Mapping of indels onto known protein structures demonstrated that structural cores are markedly less tolerant to indels than are loop regions. There is a specific enrichment of CpG dinucleotides in close proximity to insertion events, and both insertions and deletions are more common in higher G+C content sequences.

Amino Acid Sequence↗

The InterPro Database, 2003 brings increased coverage and new features.

InterPro, an integrated documentation resource of protein families, domains and functional sites, was created in 1999 as a means of amalgamating the major protein signature databases into one comprehensive resource. PROSITE, Pfam, PRINTS, ProDom, SMART and TIGRFAMs have been manually integrated and curated and are available in InterPro for text- and sequence-based searching. The results are provided in a single format that rationalises the results that would be obtained by searching the member databases individually. The latest release of InterPro contains 5629 entries describing 4280 families, 1239 domains, 95 repeats and 15 post-translational modifications. Currently, the combined signatures in InterPro cover more than 74% of all proteins in SWISS-PROT and TrEMBL, an increase of nearly 15% since the inception of InterPro. New features of the database include improved searching capabilities and enhanced graphical user interfaces for visualisation of the data. The database is available via a webserver (http://www.ebi.ac.uk/interpro) and anonymous FTP (ftp://ftp.ebi.ac.uk/pub/databases/interpro).

Animals↗

Eukaryotic domain evolution inferred from genome comparisons.

Comparative analyses of eukaryotic genomes are providing insights into the mode and tempo of domain family evolution. Gene duplication, the source of family expansion, far exceeds the rate of emergence of domains from non-coding sequence, and the rate of recruitment of domains into novel architectures. Domain families that appear to be restricted to certain lineages are likely to be the result of gene duplication, coupled with rapid sequence diversification. If such families are evidence of past adaptation, then their functions must relate to the underlying mechanism of selection: competition among organisms.

Animals↗

Initial sequencing and comparative analysis of the mouse genome.

The sequence of the mouse genome is a key informational tool for understanding the contents of the human genome and a key experimental tool for biomedical research. Here, we report the results of an international collaboration to produce a high-quality draft sequence of the mouse genome. We also present an initial comparative analysis of the mouse and human genomes, describing some of the insights that can be gleaned from the two sequences. We discuss topics including the analysis of the evolutionary forces shaping the size, structure and sequence of the genomes; the conservation of large-scale synteny across most of the genomes; the much lower extent of sequence orthology covering less than half of the genomes; the proportions of the genomes under selection; the number of protein-coding genes; the expansion of gene families related to reproduction and immunity; the evolution of proteins; and the identification of intraspecies polymorphism.

Animals↗