Search PubMed⌕ Search

Biomedical subjects

Paul Kellam

Publications and source records attributed to Paul Kellam.

16 recordsLinked to original sources

HIV-phyloTSI: subtype-independent estimation of time since HIV-1 infection for cross-sectional measures of population incidence using deep sequence data.

BACKGROUND: Estimating the time since HIV infection (TSI) at population level is essential for tracking changes in the global HIV epidemic. Most methods for determining TSI give a binary classification of infections as recent or non-recent within a window of several months, and cannot assess the cumulative impact of an intervention. RESULTS: We developed a Random Forest Regression model, HIV-phyloTSI, which combines measures of within-host diversity and divergence to generate continuous TSI estimates directly from viral deep-sequencing data, with no need for additional variables. HIV-phyloTSI provides a continuous measure of TSI up to 9 years, with a mean absolute error of less than 12 months overall and less than 5 months for infections with a TSI of up to a year. It performs equally well for all major HIV subtypes based on data from African and European cohorts. CONCLUSIONS: We demonstrate how HIV-phyloTSI can be used for incidence estimates on a population level.

HIV Infections↗

Assessment of automated genotyping protocols as tools for surveillance of HIV-1 genetic diversity.

BACKGROUND: The routine use of drug resistance testing provides an abundant source of HIV-1 sequence data. However, it is not clear how reliable standard genotyping of these sequences is for describing HIV-1 genetic variation and for detecting novel genetic variants and epidemiological trends. OBJECTIVES: To compare assignment of HIV-1 resistance test sequences to reference strains across commonly used genotyping protocols. METHODS: Subtype assignments were compared across three standard genotyping protocols for 10 537 resistance test sequences, representing approximately one-fifth of all reported infections in the United Kingdom. Sequences that were inconsistently genotyped across methods, or that were unassigned by at least one method, were examined for evidence of recombination using sliding-window-based approaches. RESULTS: Although agreement across methods was high for subtypes B, C and H, it was generally much lower (< 50%) for other subtypes. Disagreement between methods typically involved closely related, but epidemiologically distinct, groups or involved a significant proportion ( approximately 12%) of divergent sequences in which analysis revealed widespread evidence of recombination and a remarkable diversity of unusual recombinant forms. CONCLUSIONS: With frequent long-distance transfer of viral strains and widespread recombination between them, genetic and epidemiological relationships within HIV-1 are becoming increasingly complex. Current methods of subtype assignment vary in their ability to identify novel genetic variants and to distinguish epidemiologically distinct strains. Capturing meaningful epidemiological information from resistance test data will require a critical understanding of the methodologies used in order to appreciate the possible sources of error and misclassification.

Databases, Genetic↗

Infectogenomics: insights from the host genome into infectious diseases.

Five years into the human postgenomic era, we are gaining considerable knowledge about host-pathogen interactions through host genomes. This "infectogenomics" approach should yield further insights into both diagnostic and therapeutic advances, as well as normal cellular function.

Acquired Immunodeficiency Syndrome↗

Attacking pathogens through their hosts.

Through understanding the intricacies of host-pathogen interactions, it is now possible to inhibit the growth of microbes, especially viruses, by targeting host-cell proteins and functions. This new antimicrobial strategy has proved effective in the laboratory and in the clinic, and it has great potential for the future.

Animals↗

Genotyping Hepatitis B virus from whole- and sub-genomic fragments using position-specific scoring matrices in HBV STAR.

Hepatitis B virus (HBV) genomes have been classified into eight genotypes based on phylogenetic analysis of sequence variation. Identifying and tracking the movement of HBV genotypes is important in terms of both monitoring infection rates and predicting disease and treatment. An HBV genotyping tool has been developed that compares query sequences with position-specific scoring matrices representing the eight HBV genotypes. This tool (hbv star) is rapid, robust and accurate and assigns genotype based on a statistically defined scoring model. hbv star confidently assigned 90% of 590 full-length HBV genomes to an HBV genotype (Z score >2.0). Thirty-two of the residual 48 sequences were identified as non-human primate viruses and 16 sequences were identified as recombinant or putative recombinants. Receiver-Operated Characteristic (ROC) analysis was used to compare the accuracy of genotype prediction using basal core promoter sequences and surface and core genes with the accuracy achieved by using full-length sequences. A web interface to hbv star is available at http://www.vgb.ucl.ac.uk/starn.shtml.

Animals↗

Robust Selection of Predictive Genes via a Simple Classifier.

Identifying genes that direct the mechanism of a disease from expression data is extremely useful in understanding how that mechanism works. This in turn may lead to better diagnoses and potentially could lead to a cure for that disease. This task becomes extremely challenging when the data are characterised by only a small number of samples and a high number of dimensions, as is often the case with gene expression data. Motivated by this challenge, we present a general framework that focuses on simplicity and data perturbation. These are the keys for robust identification of the most predictive features in such data. Within this framework, we propose a simple selective naive Bayes classifier discovered using a global search technique, and combine it with data perturbation to increase its robustness for small sample sizes. An extensive validation of the method was carried out using two applied datasets from the field of microarrays and a simulated dataset, all confounded by small sample sizes and high dimensionality. The method has been shown to be capable of selecting genes known to be associated with prostate cancer and viral infections.

Artificial Intelligence↗

Consensus clustering and functional interpretation of gene-expression data.

Microarray analysis using clustering algorithms can suffer from lack of inter-method consistency in assigning related gene-expression profiles to clusters. Obtaining a consensus set of clusters from a number of clustering methods should improve confidence in gene-expression analysis. Here we introduce consensus clustering, which provides such an advantage. When coupled with a statistically based gene functional analysis, our method allowed the identification of novel genes regulated by NFkappaB and the unfolded protein response in certain B-cell lymphomas.

Cluster Analysis↗

Development of a novel human immunodeficiency virus type 1 subtyping tool, Subtype Analyzer (STAR): analysis of subtype distribution in London.

We have developed a high throughput computational tool for assigning subtype to HIV-1, based solely on protease and reverse transcriptase (PR-RT) amino acid sequence, generated routinely for clinical assessment of genotypic drug resistance. Subtype-specific profiles were created by generation of position-specific scoring matrices (PSSMs) from multiple amino acids alignments of HIV-1 sequence data from GenBank, phylogenetically divided into subtypes A, AG, B, C, D, F/K, G, H, and J and the separate groups N and O. Query sequences of unknown subtype are aligned with these profiles and a score is derived by comparing each amino acid position in the unknown sequence to the normalized frequency distribution of amino acids at the corresponding positions in the subtype alignments. The highest score is used to assign subtype to the query sequence. Leave one out cross-validation analysis showed the Subtype Analyzer (STAR) was 99% accurate in subtype assignation. STAR can be updated with additional subtype-specific sequence data from sequence databases. STAR was used to classify HIV-1 PR-RT sequences from 843 HIV-1 clinical isolates submitted for drug resistance profiling in London. Within this dataset 26.9% of sequences were classified by STAR as non-B subtypes.

Algorithms↗

Poxvirus genomes: a phylogenetic analysis.

The evolutionary relationships of 26 sequenced members of the poxvirus family have been investigated by comparing their genome organization and gene content and by using DNA and protein sequences for phylogenetic analyses. The central region of the genome of chordopoxviruses (ChPVs) is highly conserved in gene content and arrangement, except for some gene inversions in Fowlpox virus (FPV) and species-specific gene insertions in FPV and Molluscum contagiosum virus (MCV). In the central region 90 genes are conserved in all ChPVs, but no gene from near the termini is conserved throughout the subfamily. Inclusion of two entomopoxvirus (EnPV) sequences reduces the number of conserved genes to 49. The EnPVs are divergent from ChPVs and between themselves. Relationships between ChPV genera were evaluated by comparing the genome size, number of unique genes, gene arrangement and phylogenetic analyses of protein sequences. Overall, genus Avipoxvirus is the most divergent. The next most divergent ChPV genus is Molluscipoxvirus, whose sole member, MCV, infects only man. The Suipoxvirus, Capripoxvirus, Leporipoxvirus and Yatapoxvirus genera cluster together, with Suipoxvirus and Capripoxvirus sharing a common ancestor, and are distinct from the genus Orthopoxvirus (OPV). Within the OPV genus, Monkeypox virus, Ectromelia virus and Cowpox virus strain Brighton Red (BR) do not group closely with any other OPV, Variola virus and Camelpox virus form a subgroup, and Vaccinia virus is most closely related to CPV-GRI-90. This suggests that CPV-BR and GRI-90 should be separate species.

Amino Acid Sequence↗

Kaposi's sarcoma-associated herpesvirus-infected primary effusion lymphoma has a plasma cell gene expression profile.

Kaposi's sarcoma-associated herpesvirus is associated with three human tumors: Kaposi's sarcoma, and the B cell lymphomas, plasmablastic lymphoma associated with multicentric Castleman's disease, and primary effusion lymphoma (PEL). Epstein-Barr virus, the closest human relative of Kaposi's sarcoma-associated herpesvirus, mimics host B cell signaling pathways to direct B cell development toward a memory B cell phenotype. Epstein-Barr virus-associated B cell tumors are presumed to arise as a consequence of this virus-mediated B cell activation. The stage of B cell development represented by PEL, how this stage relates to tumor pathology, and how this information may be used to treat the disease are largely unknown. In this study we used gene expression profiling to order a range of B cell tumors by stage of development. PEL gene expression closely resembles that of malignant plasma cells, including the low expression of mature B cell genes. The unfolded protein response is partially activated in PEL, but is fully activated in plasma cell tumors, linking endoplasmic reticulum stress to plasma cell development through XBP-1. PEL cells can be defined by the overexpression of genes involved in inflammation, cell adhesion, and invasion, which may be responsible for their presentation in body cavities. Similar to malignant plasma cells, all PEL samples tested express the vitamin D receptor and are sensitive to the vitamin D analogue drug EB 1089 (Seocalcitol).

Calcitriol↗

Rabbit endogenous retrovirus-H encodes a functional protease.

Recent studies have revealed that 'human retrovirus-5' sequences found in human samples belong to a rabbit endogenous retrovirus family named RERV-H. A part of the gag-pro region of the RERV-H genome was amplified by PCR from DNA in human samples and several forms of RERV-H protease were expressed in bacteria. The RERV-H protease was able to cleave itself from a precursor protein and was also able to cleave the RERV-H Gag polyprotein precursor in vitro whereas a form of the protease with a mutation engineered into the active site was inactive. Potential N- and C-terminal autocleavage sites were characterized. The RERV-H protease was sensitive to pepstatin A, showing it to be an aspartic protease. Moreover, it was strongly inhibited by PYVPheStaAMT, a pseudopeptide inhibitor specific for Mason-Pfizer monkey virus and avian myeloblastosis-associated virus. A structural model of the RERV-H protease was constructed that, together with the activity data, confirms that this is a retroviral aspartic protease.

Amino Acid Sequence↗

Viral bioinformatics: computational views of host and pathogen.

Wherever cellular life occurs, viruses are also found. As a result, complex organism and cellular antiviral responses co-evolve with virally encoded countermeasures. Since viruses co-opt or interfere with specific cellular pathways during their replication, knowledge of viral genome sequences has helped fundamental understanding of host biology. During viral infection, shifts in the balance between host and viral biological processes result in acute or chronic viral disease pathology accompanied with either active viral replication, viral containment/persistence or viral clearance. Studying host-virus interactions at the level of single gene effects, however, fails to produce a global systems-level understanding. This should now be achievable in the context of complete host and pathogen genome sequences. New experimental methods and computational approaches are rapidly developing, allowing global views of dynamic viral and cellular molecular mechanisms. Systems level virology using DNA microarrays and specific viral data resources will reveal the detailed cellular context in which viruses replicate, highlighting common and distinct antiviral mechanisms, the effect of different host cell gene expression programs, and the response of cells to similar or diverse virus types. Ultimately, microbiology and immunology will tend towards a systems-level view of how host and pathogen interact.

Amino Acid Motifs↗

PFDB: a generic protein family database integrating the CATH domain structure database with sequence based protein family resources.

MOTIVATION: The PFDB (Protein Family Database) is a new database designed to integrate protein family-related data with relevant functional and genomic data. It currently manages biological data for three projects-the CATH protein domain database (Orengo et al., 1997; Pearl et al., 2001), the VIDA virus domains database (Albà et al., 2001) and the Gene3D database (Buchan et al., 2001). The PFDB has been designed to accommodate protein families identified by a variety of sequence based or structure based protocols and provides a generic resource for biological research by enabling mapping between different protein families and diverse biochemical and genetic data, including complete genomes. RESULTS: A characteristic feature of the PFDB is that it has a number of meta-level entities (for example aggregation, collection and inclusion) represented as base tables in the final design. The explicit representation of relationships at the meta-level has a number of advantages, including flexibility-both in terms of the range of queries that can be formulated and the ability to integrate new biological entities within the existing design. A potential drawback with this approach-poor performance caused by the number of joins across meta-level tables-is avoided by implementing the PFDB with materialized views using the mature relational database technology of Oracle 8i. The resultant database is both fast and flexible. This paper presents the principles on which the database has been designed and implemented, and describes the current status of the database and query facilities supported.

Database Management Systems↗

Identification of new herpesvirus gene homologs in the human genome.

Viruses are intracellular parasites that use many cellular pathways during their replication. Large DNA viruses, such as herpesviruses, have captured a repertoire of cellular genes to block or mimic host immune responses, apoptosis regulation, and cell-cycle control mechanisms. We have conducted a systematic search for all homologs of herpesvirus proteins in the human genome using position-specific scoring matrices representing herpesvirus protein sequence domains, and pair-wise sequence comparisons. The analysis shows that approximately 13% of the herpesvirus proteins have clear sequence similarity to products of the human genome. Different human herpesviruses vary in their numbers of human homologs, indicating distinct rates of gene acquisition in different lineages. Our analysis has identified new families of herpesvirus/human homologs from viruses including human herpesvirus 5 (human cytomegalovirus; HCMV) and human herpesvirus 8 (Kaposi's sarcoma-associated herpesvirus; KSHV), which may play important roles in host-virus interactions.

Amino Acid Sequence↗

Gammaherpesvirus lytic gene expression as characterized by DNA array.

Gammaherpesviruses are associated with a number of diseases including lymphomas and other malignancies. Murine gammaherpesvirus 68 (MHV-68) constitutes the most amenable animal model for this family of pathogens. However experimental characterization of gammaherpesvirus gene expression, at either the protein or RNA level, lags behind that of other, better-studied alpha- and beta-herpesviruses. We have developed a cDNA array to globally characterize MHV-68 gene expression profiles, thus providing an experimental supplement to a genome that is chiefly annotated by homology. Viral genes started to be transcribed as early as 3 h postinfection (p.i.), and this was followed by a rapid escalation of gene expression that could be seen at 5 h p.i. Individual genes showed their own transcription profiles, and most genes were still being expressed at 18 h p.i. Open reading frames (ORFs) M3 (chemokine-binding protein), 52, and M9 (capsid protein) were particularly noticeable due to their very high levels of expression. Hierarchical cluster analysis of transcription profiles revealed four main groups of genes and allowed functional predictions to be made by comparing expression profiles of uncharacterized genes to those of genes of known function. Each gene was also categorized according to kinetic class by blocking de novo protein synthesis and viral DNA replication in vitro. One gene, ORF 73, was found to be expressed with alpha-kinetics, 30 genes were found to be expressed with beta-kinetics, and 42 genes were found to be expressed with gamma-kinetics. This fundamental characterization furthers the development of this model and provides an experimental basis for continued investigation of gammaherpesvirus pathology.

Animals↗

Virus bioinformatics: databases and recent applications.

Bioinformatics is now used as an umbrella term for almost all aspects of computational biology. Bioinformatics research will have an impact on all of biology, and virology is not immune from these research methods. Although virology has been slower to embrace bioinformatics this is now changing, particularly in the areas of viral sequences databasing and the systematic identification of viral and host homologous proteins. Here we will review some of these recent advances focusing mainly on the herpesvirus.

Computational Biology↗