Search PubMed⌕ Search

Biomedical subjects

Jacob de Vlieg

Publications and source records attributed to Jacob de Vlieg.

6 recordsLinked to original sources

Testing statistical significance scores of sequence comparison methods with structure similarity.

BACKGROUND: In the past years the Smith-Waterman sequence comparison algorithm has gained popularity due to improved implementations and rapidly increasing computing power. However, the quality and sensitivity of a database search is not only determined by the algorithm but also by the statistical significance testing for an alignment. The e-value is the most commonly used statistical validation method for sequence database searching. The CluSTr database and the Protein World database have been created using an alternative statistical significance test: a Z-score based on Monte-Carlo statistics. Several papers have described the superiority of the Z-score as compared to the e-value, using simulated data. We were interested if this could be validated when applied to existing, evolutionary related protein sequences. RESULTS: All experiments are performed on the ASTRAL SCOP database. The Smith-Waterman sequence comparison algorithm with both e-value and Z-score statistics is evaluated, using ROC, CVE and AP measures. The BLAST and FASTA algorithms are used as reference. We find that two out of three Smith-Waterman implementations with e-value are better at predicting structural similarities between proteins than the Smith-Waterman implementation with Z-score. SSEARCH especially has very high scores. CONCLUSION: The compute intensive Z-score does not have a clear advantage over the e-value. The Smith-Waterman implementations give generally better results than their heuristic counterparts. We recommend using the SSEARCH algorithm combined with e-values for pairwise sequence comparisons.

Base Sequence↗

The use of in vitro peptide binding profiles and in silico ligand-receptor interaction profiles to describe ligand-induced conformations of the retinoid X receptor alpha ligand-binding domain.

It is hypothesized that different ligand-induced conformational changes can explain the different interactions of nuclear receptors with regulatory proteins, resulting in specific biological activities. Understanding the mechanism of how ligands regulate cofactor interaction facilitates drug design. To investigate these ligand-induced conformational changes at the surface of proteins, we performed a time-resolved fluorescence resonance energy transfer assay with 52 different cofactor peptides measuring the ligand-induced cofactor recruitment to the retinoid X receptor-alpha (RXRalpha) in the presence of 11 compounds. Simultaneously we analyzed the binding modes of these compounds by molecular docking. An automated method converted the complex three-dimensional data of ligand-protein interactions into two-dimensional fingerprints, the so-called ligand-receptor interaction profiles. For a subset of compounds the conformational changes at the surface, as measured by peptide recruitment, correlate well with the calculated binding modes, suggesting that clustering of ligand-receptor interaction profiles is a very useful tool to discriminate compounds that may induce different conformations and possibly different effects in a cellular environment. In addition, we successfully combined ligand-receptor interaction profiles and peptide recruitment data to reveal structural elements that are possibly involved in the ligand-induced conformations. Interestingly, we could predict a possible binding mode of LG100754, a homodimer antagonist that showed no effect on peptide recruitment. Finally, the extensive analysis of the peptide recruitment profiles provided novel insight in the potential cellular effect of the compound; for the first time, we showed that in addition to the induction of coactivator peptide binding, all well-known RXRalpha agonists also induce binding of corepressor peptides to RXRalpha.

Amino Acid Sequence↗

PhyloPat: phylogenetic pattern analysis of eukaryotic genes.

BACKGROUND: Phylogenetic patterns show the presence or absence of certain genes or proteins in a set of species. They can also be used to determine sets of genes or proteins that occur only in certain evolutionary branches. Phylogenetic patterns analysis has routinely been applied to protein databases such as COG and OrthoMCL, but not upon gene databases. Here we present a tool named PhyloPat which allows the complete Ensembl gene database to be queried using phylogenetic patterns. DESCRIPTION: PhyloPat is an easy-to-use webserver, which can be used to query the orthologies of all complete genomes within the EnsMart database using phylogenetic patterns. This enables the determination of sets of genes that occur only in certain evolutionary branches or even single species. We found in total 446,825 genes and 3,164,088 orthologous relationships within the EnsMart v40 database. We used a single linkage clustering algorithm to create 147,922 phylogenetic lineages, using every one of the orthologies provided by Ensembl. PhyloPat provides the possibility of querying with either binary phylogenetic patterns (created by checkboxes) or regular expressions. Specific branches of a phylogenetic tree of the 21 included species can be selected to create a branch-specific phylogenetic pattern. Users can also input a list of Ensembl or EMBL IDs to check which phylogenetic lineage any gene belongs to. The output can be saved in HTML, Excel or plain text format for further analysis. A link to the FatiGO web interface has been incorporated in the HTML output, creating easy access to functional information. Finally, lists of omnipresent, polypresent and oligopresent genes have been included. CONCLUSION: PhyloPat is the first tool to combine complete genome information with phylogenetic pattern querying. Since we used the orthologies generated by the accurate pipeline of Ensembl, the obtained phylogenetic lineages are reliable. The completeness and reliability of these phylogenetic lineages will further increase with the addition of newly found orthologous relationships within each new Ensembl release.

Algorithms↗

Benchmarking ortholog identification methods using functional genomics data.

BACKGROUND: The transfer of functional annotations from model organism proteins to human proteins is one of the main applications of comparative genomics. Various methods are used to analyze cross-species orthologous relationships according to an operational definition of orthology. Often the definition of orthology is incorrectly interpreted as a prediction of proteins that are functionally equivalent across species, while in fact it only defines the existence of a common ancestor for a gene in different species. However, it has been demonstrated that orthologs often reveal significant functional similarity. Therefore, the quality of the orthology prediction is an important factor in the transfer of functional annotations (and other related information). To identify protein pairs with the highest possible functional similarity, it is important to qualify ortholog identification methods. RESULTS: To measure the similarity in function of proteins from different species we used functional genomics data, such as expression data and protein interaction data. We tested several of the most popular ortholog identification methods. In general, we observed a sensitivity/selectivity trade-off: the functional similarity scores per orthologous pair of sequences become higher when the number of proteins included in the ortholog groups decreases. CONCLUSION: By combining the sensitivity and the selectivity into an overall score, we show that the InParanoid program is the best ortholog identification method in terms of identifying functionally equivalent proteins.

Algorithms↗

The nuclear receptor ligand-binding domain: a family-based structure analysis.

Nuclear receptors (NRs) are ligand-dependent transcription factors that play a central role in various physiological processes. The pharmaceutical industry has great interest in this gene-family for the discovery of novel or improved drugs for treatment of, for example, cancer, infertility, or diabetes. The usage of three-dimensional coordinates of protein structures to analyse and predict interactions with ligands is an important aspect of this process. All NR ligand-binding domains have a similar fold, which allows for comparison of the structures of their three main functional sites: the ligand-binding pocket, the cofactor-binding groove, and the dimerization interface. We performed an analysis of nearly one hundred NR ligand-binding domain structures, and identified the functionally important residues. The combined knowledge about the shape of the binding sites and the residues involved in the binding is important for drug design in two ways. First, knowledge about the location of residues that interact with a ligand in all crystal structures or in certain subfamilies assists in the design and docking of drugs. Second, similarities and differences in the residue types of the most frequent ligand- and cofactor-binding residues provide insight about potential cross-reactivity of ligands or cofactors.

Binding Sites↗

A family-based approach reveals the function of residues in the nuclear receptor ligand-binding domain.

Literature studies, 3D structure data, and a series of sequence analysis techniques were combined to reveal important residues in the structure and function of the ligand-binding domain of nuclear hormone receptors. A structure-based multiple sequence alignment allowed for the seamless combination of data from many different studies on different receptors into one single functional model. It was recently shown that a combined analysis of sequence entropy and variability can divide residues in five classes; (1) the main function or active site, (2) support for the main function, (3) signal transduction, (4) modulator or ligand binding and (5) the rest. Mutation data extracted from the literature and intermolecular contacts observed in nuclear receptor structures were analyzed in view of this classification and showed that the main function or active site residues of the nuclear receptor ligand-binding domain are involved in cofactor recruitment. Furthermore, the sequence entropy-variability analysis identified the presence of signal transduction residues that are located between the ligand, cofactor and dimer sites, suggesting communication between these regulatory binding sites. Experimental and computational results agreed well for most residues for which mutation data and intermolecular contact data were available. This allows us to predict the role of the residues for which no functional data is available yet. This study illustrates the power of family-based approaches towards the analysis of protein function, and it points out the problems and possibilities presented by the massive amounts of data that are becoming available in the "omics era". The results shed light on the nuclear receptor family that is involved in processes ranging from cancer to infertility, and that is one of the more important targets in the pharmaceutical industry.

Amino Acids↗