Search PubMed⌕ Search

Biomedical subjects

Chris P Ponting

Publications and source records attributed to Chris P Ponting.

At least 19 recordsLinked to original sources

Semiparametric efficient estimation of small genetic effects in large-scale population cohorts.

Population genetics seeks to quantify DNA variant associations with traits or diseases, as well as interactions among variants and with environmental factors. Computing millions of estimates in large cohorts in which small effect sizes and tight confidence intervals are expected, necessitates minimizing model-misspecification bias to increase power and control false discoveries. We present TarGene, a unified statistical workflow for the semi-parametric efficient and double robust estimation of genetic effects including $ k $-point interactions among categorical variables in the presence of confounding and weak population dependence. $ k $-point interactions, or Average Interaction Effects (AIEs), are a direct generalization of the usual average treatment effect (ATE). We estimate genetic effects with cross-validated and/or weighted versions of Targeted Minimum Loss-based Estimators (TMLE) and One-Step Estimators (OSE). The effect of dependence among data units on variance estimates is corrected by using sieve plateau variance estimators based on genetic relatedness across the units. We present extensive realistic simulations to demonstrate power, coverage, and control of type I error. Our motivating application is the targeted estimation of genetic effects on trait, including two-point and higher-order gene-gene and gene-environment interactions, in large-scale genomic databases such as UK Biobank and All of Us. All cross-validated and/or weighted TMLE and OSE for the AIE $ k $-point interaction, as well as ATEs, conditional ATEs and functions thereof, are implemented in the general purpose Julia package TMLE.jl. For high-throughput applications in population genomics, we provide the open-source Nextflow pipeline and software TarGene which integrates seamlessly with modern high-performance and cloud computing platforms.

Humans↗

THoR: a tool for domain discovery and curation of multiple alignments.

We describe a tool, THoR, that automatically creates and curates multiple sequence alignments representing protein domains. This exploits both PSI-BLAST and HMMER algorithms and provides an accurate and comprehensive alignment for any domain family. The entire process is designed for use via a web-browser, with simple links and cross-references to relevant information, to assist the assessment of biological significance. THoR has been benchmarked for accuracy using the SMART and pufferfish genome databases.

Algorithms↗

Comparison of the genomes of human and mouse lays the foundation of genome zoology.

The extensive similarities between the genomes of human and model organisms are the foundation of much of modern biology, with model organism experimentation permitting valuable insights into biological function and the aetiology of human disease. In contrast, differences among genomes have received less attention. Yet these can be expected to govern the physiological and morphological distinctions apparent among species, especially if such differences are the result of evolutionary adaptation. A recent comparison of the draft sequences of mouse and human genomes has shed light on the selective forces that have predominated in their recent evolutionary histories. In particular, mouse-specific clusters of homologues associated with roles in reproduction, immunity and host defence appear to be under diversifying positive selective pressure, as indicated by high ratios of non-synonymous to synonymous substitution rates. These clusters are also frequently punctuated by homologous pseudogenes. They thus have experienced numerous gene death, as well as gene birth, events. These regions appear, therefore, to have borne the brunt of adaptive evolution that underlies physiological and behavioural innovation in mice. We predict that the availability of numerous animal genomes will give rise to a new field of genome zoology in which differences in animal physiology and ethology are illuminated by the study of genomic sequence variations.

Animals↗

Comparison of mouse and human genomes followed by experimental verification yields an estimated 1,019 additional genes.

A primary motivation for sequencing the mouse genome was to accelerate the discovery of mammalian genes by using sequence conservation between mouse and human to identify coding exons. Achieving this goal proved challenging because of the large proportion of the mouse and human genomes that is apparently conserved but apparently does not code for protein. We developed a two-stage procedure that exploits the mouse and human genome sequences to produce a set of genes with a much higher rate of experimental verification than previously reported prediction methods. RT-PCR amplification and direct sequencing applied to an initial sample of mouse predictions that do not overlap previously known genes verified the regions flanking one intron in 139 predictions, with verification rates reaching 76%. On average, the confirmed predictions show more restricted expression patterns than the mouse orthologs of known human genes, and two-thirds lack homologs in fish genomes, demonstrating the sensitivity of this dual-genome approach to hard-to-find genes. We verified 112 previously unknown homologs of known proteins, including two homeobox proteins relevant to developmental biology, an aquaporin, and a homolog of dystrophin. We estimate that transcription and splicing can be verified for >1,000 gene predictions identified by this method that do not overlap known genes. This is likely to constitute a significant fraction of the previously unknown, multiexon mammalian genes.

Amino Acid Sequence↗

The InterPro Database, 2003 brings increased coverage and new features.

InterPro, an integrated documentation resource of protein families, domains and functional sites, was created in 1999 as a means of amalgamating the major protein signature databases into one comprehensive resource. PROSITE, Pfam, PRINTS, ProDom, SMART and TIGRFAMs have been manually integrated and curated and are available in InterPro for text- and sequence-based searching. The results are provided in a single format that rationalises the results that would be obtained by searching the member databases individually. The latest release of InterPro contains 5629 entries describing 4280 families, 1239 domains, 95 repeats and 15 post-translational modifications. Currently, the combined signatures in InterPro cover more than 74% of all proteins in SWISS-PROT and TrEMBL, an increase of nearly 15% since the inception of InterPro. New features of the database include improved searching capabilities and enhanced graphical user interfaces for visualisation of the data. The database is available via a webserver (http://www.ebi.ac.uk/interpro) and anonymous FTP (ftp://ftp.ebi.ac.uk/pub/databases/interpro).

Animals↗

The Tudor domain 'Royal Family': Tudor, plant Agenet, Chromo, PWWP and MBT domains.

We have identified a family of 'Agenet' domains that are plant-specific homologs of Tudor domains. This finding has been extended, using a combination of sequence- and structure-dependent approaches, to show that the three beta-stranded core regions of Tudor, PWWP, chromatin-binding (Chromo) and MBT domains are homologous because they originate from a common ancestor. In addition, we have revealed pairs of tandem repeats in the fragile X mental retardation protein (FMRP) family that are also members of this Tudor domain 'Royal Family'.

Amino Acid Sequence↗

UBA domain containing proteins in fission yeast.

The ubiquitin-proteasome pathway for intracellular proteolysis is involved in a series of cellular and molecular functions, including the degradation of bulk proteins, cell cycle control, DNA repair, antigen presentation, vesicle transport and the regulation of signal transudation pathways and transcription. Considering this variety of cell biological processes, it is puzzling that until recently only very few proteins were known to possess the ability to interact specifically with ubiquitin chains. However, several ubiquitin binding proteins have now been identified and the binding domains have been characterised on both the functional and structural levels. One example of a widespread ubiquitin binding module is the ubiquitin associated (UBA) domain. Here, we discuss the approximately 15 UBA domain containing proteins encoded in the relatively small genome of the fission yeast Schizosaccharomyces pombe. The proteins display remarkable differences in their domain organisation, indicating that these potential ubiquitin binding proteins are involved in various cell activities.

Amino Acid Sequence↗

Positional cloning of a quantitative trait locus on chromosome 13q14 that influences immunoglobulin E levels and asthma.

Atopic or immunoglobulin E (IgE)-mediated diseases include the common disorders of asthma, atopic dermatitis and allergic rhinitis. Chromosome 13q14 shows consistent linkage to atopy and the total serum IgE concentration. We previously identified association between total serum IgE levels and a novel 13q14 microsatellite (USAT24G1; ref. 7) and have now localized the underlying quantitative-trait locus (QTL) in a comprehensive single-nucleotide polymorphism (SNP) map. We found replicated association to IgE levels that was attributed to several alleles in a single gene, PHF11. We also found association with these variants to severe clinical asthma. The gene product (PHF11) contains two PHD zinc fingers and probably regulates transcription. Distinctive splice variants were expressed in immune tissues and cells.

Adult↗

Novel domains and orthologues of eukaryotic transcription elongation factors.

The passage of RNA polymerase II across eukaryotic genes is impeded by the nucleosome, an octamer of histones H2A, H2B, H3 and H4 dimers. More than a dozen factors in the yeast Saccharomyces cerevisiae are known to facilitate transcription elongation through chromatin. In order to better understand the evolution and function of these factors, their sequences have been compared with known protein, EST and DNA sequences. Elongator subcomplex components Elp4p and Elp6p are shown to be homologues of ATPases, yet with substitutions of amino acids critical for ATP hydrolysis, and novel orthologues of Elp5p are detectable in human, and other animal, sequences. The yeast CP complex is shown to contain a likely inactive homologue of M24 family metalloproteases in Spt16p/Cdc68p and a 2-fold repeat in Pob3p, the orthologue of mammalian SSRP1. Archaeal DNA-directed RNA polymerase subunit E" is shown to be the orthologue of eukaryotic Spt4p, and Spt5p and prokaryotic NusG are shown to contain a novel 'NGN' domain. Spt6p is found to contain a domain homologous to the YqgF family of RNases, although this domain may also lack catalytic activity. These findings imply that much of the transcription elongation machinery of eukaryotes has been acquired subsequent to their divergence from prokaryotes.

Amino Acid Motifs↗

Identification of a novel family of presenilin homologues.

Presenilin 1 and presenilin 2 are polytopic membrane proteins, whose genes are mutated in some individuals with Alzheimer's disease. Presenilins have been shown to influence limited proteolysis of amyloid beta protein precursor (APP), Notch and ErbB4, and have been proposed to be gamma-secretases that perform the terminal cleavage of APP. In this model, two conserved and apparently intramembranous aspartic acids participate in catalysis. Highly sequence-similar presenilin homologues are known in plants, invertebrates and vertebrates. In this work, we have used a combination of different sequence database search methods to identify a new family of proteins homologous to presenilins. Members of this family, which we term presenilin homologues (PSH), have significant sequence similarities to presenilins and also possess two conserved aspartic acid residues within adjacent predicted transmembrane segments. The PSH family is found throughout the eukaryotes, in fungi as well as plants and animals, and in archaea. Five PSHs are detectable in the human genome, of which three possess "protease-associated" domains that are consistent with the proposed protease function of PSs. Based on these findings, we propose that PSs and PSHs represent different sub-branches of a larger family of polytopic membrane-associated aspartyl proteases.

Alzheimer Disease↗

Recent improvements to the SMART domain-based sequence annotation resource.

SMART (Simple Modular Architecture Research Tool, http://smart.embl-heidelberg.de) is a web-based resource used for the annotation of protein domains and the analysis of domain architectures, with particular emphasis on mobile eukaryotic domains. Extensive annotation for each domain family is available, providing information relating to function, subcellular localization, phyletic distribution and tertiary structure. The January 2002 release has added more than 200 hand-curated domain models. This brings the total to over 600 domain families that are widely represented among nuclear, signalling and extracellular proteins. Annotation now includes links to the Online Mendelian Inheritance in Man (OMIM) database in cases where a human disease is associated with one or more mutations in a particular domain. We have implemented new analysis methods and updated others. New advanced queries provide direct access to the SMART relational database using SQL. This database now contains information on intrinsic sequence features such as transmembrane regions, coiled-coils, signal peptides and internal repeats. SMART output can now be easily included in users' documents. A SMART mirror has been created at http://smart.ox.ac.uk.

Animals↗

TRAM, LAG1 and CLN8: members of a novel family of lipid-sensing domains?

A family of membrane-associated proteins related to yeast Lag1p and mammalian TRAM has been identified. The family includes the protein product of CLN8, a gene mutated in progressive epilepsy with mental retardation. Mouse CLN8 is also mutated in the mnd/mnd mouse, a model for neuronal ceroid lipofuscinoses. The identification of these homologues has potential implications for our understanding of ceramide synthesis, lipid regulation and protein translocation in the endoplasmic reticulum.

Amino Acid Sequence↗

InterPro: an integrated documentation resource for protein families, domains and functional sites.

The exponential increase in the submission of nucleotide sequences to the nucleotide sequence database by genome sequencing centres has resulted in a need for rapid, automatic methods for classification of the resulting protein sequences. There are several signature and sequence cluster-based methods for protein classification, each resource having distinct areas of optimum application owing to the differences in the underlying analysis methods. In recognition of this, InterPro was developed as an integrated documentation resource for protein families, domains and functional sites, to rationalise the complementary efforts of the individual protein signature database projects. The member databases - PRINTS, PROSITE, Pfam, ProDom, SMART and TIGRFAMs - form the InterPro core. Related signatures from each member database are unified into single InterPro entries. Each InterPro entry includes a unique accession number, functional descriptions and literature references, and links are made back to the relevant member database(s). Release 4.0 of InterPro (November 2001) contains 4,691 entries, representing 3,532 families, 1,068 domains, 74 repeats and 15 sites of post-translational modification (PTMs) encoded by different regular expressions, profiles, fingerprints and hidden Markov models (HMMs). Each InterPro entry lists all the matches against SWISS-PROT and TrEMBL (2,141,621 InterPro hits from 586,124 SWISS-PROT and TrEMBL protein sequences). The database is freely accessible for text- and sequence-based searches.

Algorithms↗

Systematic identification of novel protein domain families associated with nuclear functions.

A systematic computational analysis of protein sequences containing known nuclear domains led to the identification of 28 novel domain families. This represents a 26% increase in the starting set of 107 known nuclear domain families used for the analysis. Most of the novel domains are present in all major eukaryotic lineages, but 3 are species specific. For about 500 of the 1200 proteins that contain these new domains, nuclear localization could be inferred, and for 700, additional features could be predicted. For example, we identified a new domain, likely to have a role downstream of the unfolded protein response; a nematode-specific signalling domain; and a widespread domain, likely to be a noncatalytic homolog of ubiquitin-conjugating enzymes.

Amidohydrolases↗

Predicting protein cellular localization using a domain projection method.

We investigate the co-occurrence of domain families in eukaryotic proteins to predict protein cellular localization. Approximately half (300) of SMART domains form a "small-world network", linked by no more than seven degrees of separation. Projection of the domains onto two-dimensional space reveals three clusters that correspond to cellular compartments containing secreted, cytoplasmic, and nuclear proteins. The projection method takes into account the existence of "bridging" domains, that is, instances where two domains might not occur with each other but frequently co-occur with a third domain; in such circumstances the domains are neighbors in the projection. While the majority of domains are specific to a compartment ("locale"), and hence may be used to localize any protein that contains such a domain, a small subset of domains either are present in multiple locales or occur in transmembrane proteins. Comparison with previously annotated proteins shows that SMART domain data used with this approach can predict, with 92% accuracy, the localizations of 23% of eukaryotic proteins. The coverage and accuracy will increase with improvements in domain database coverage. This method is complementary to approaches that use amino-acid composition or identify sorting sequences; these methods may be combined to further enhance prediction accuracy.

Animals↗