Search PubMed⌕ Search

Biomedical subjects

Igor V Tetko

Publications and source records attributed to Igor V Tetko.

At least 19 recordsLinked to original sources

Spatiotemporal expression control correlates with intragenic scaffold matrix attachment regions (S/MARs) in Arabidopsis thaliana.

Scaffold/matrix attachment regions (S/MARs) are essential for structural organization of the chromatin within the nucleus and serve as anchors of chromatin loop domains. A significant fraction of genes in Arabidopsis thaliana contains intragenic S/MAR elements and a significant correlation of S/MAR presence and overall expression strength has been demonstrated. In this study, we undertook a genome scale analysis of expression level and spatiotemporal expression differences in correlation with the presence or absence of genic S/MAR elements. We demonstrate that genes containing intragenic S/MARs are prone to pronounced spatiotemporal expression regulation. This characteristic is found to be even more pronounced for transcription factor genes. Our observations illustrate the importance of S/MARs in transcriptional regulation and the role of chromatin structural characteristics for gene regulation. Our findings open new perspectives for the understanding of tissue- and organ-specific regulation of gene expression.

Arabidopsis↗

A systematic approach to infer biological relevance and biases of gene network structures.

The development of high-throughput technologies has generated the need for bioinformatics approaches to assess the biological relevance of gene networks. Although several tools have been proposed for analysing the enrichment of functional categories in a set of genes, none of them is suitable for evaluating the biological relevance of the gene network. We propose a procedure and develop a web-based resource (BIOREL) to estimate the functional bias (biological relevance) of any given genetic network by integrating different sources of biological information. The weights of the edges in the network may be either binary or continuous. These essential features make our web tool unique among many similar services. BIOREL provides standardized estimations of the network biases extracted from independent data. By the analyses of real data we demonstrate that the potential application of BIOREL ranges from various benchmarking purposes to systematic analysis of the network biology.

Computational Biology↗

The Mouse Functional Genome Database (MfunGD): functional annotation of proteins in the light of their cellular context.

MfunGD (http://mips.gsf.de/genre/proj/mfungd/) provides a resource for annotated mouse proteins and their occurrence in protein networks. Manual annotation concentrates on proteins which are found to interact physically with other proteins. Accordingly, manually curated information from a protein-protein interaction database (MPPI) and a database of mammalian protein complexes is interconnected with MfunGD. Protein function annotation is performed using the Functional Catalogue (FunCat) annotation scheme which is widely used for the analysis of protein networks. The dataset is also supplemented with information about the literature that was used in the annotation process as well as links to the SIMAP Fasta database, the Pedant protein analysis system and cross-references to external resources. Proteins that so far were not manually inspected are annotated automatically by a graphical probabilistic model and/or superparamagnetic clustering. The database is continuously expanding to include the rapidly growing amount of functional information about gene products from mouse. MfunGD is implemented in GenRE, a J2EE-based component-oriented multi-tier architecture following the separation of concern principle.

Animals↗

In silico approaches to prediction of aqueous and DMSO solubility of drug-like compounds: trends, problems and solutions.

The solubility of drugs and drug-like compounds has been the subject of extensive studies aimed at finding a way to predict solubility from molecular structure. The aqueous solubility of a drug is an important factor that influences its absorption, distribution and elimination in the body. Poor aqueous solubility often causes a drug to appear inactive and may cause other biological problems. Compound solubility in DMSO represents another serious problem in early stages of drug discovery. An appreciation of the factors affecting a compound's DMSO solubility could help in predicting the storage conditions and appropriateness of compounds for primary bioscreening programs. In silico procedures for estimation of water and DMSO solubility represent extremely useful tools for the drug discovery practitioners. In this review, we provide a critical discussion of in silico models for the prediction of DMSO and water solubility of drug-like compounds used for virtual screening. We describe the main tendencies in the field, "booming" approaches and unsolved problems. A critical analysis of the accuracy and applicability of methods is provided.

Biological Availability↗

Computing chemistry on the web.

The development of on-line software tools is changing the way we traditionally perform our analysis in drug design, but will chemoinformatics be forever behind bioinformatics in this development?

Computer Simulation↗

Surrogate data--a secure way to share corporate data.

The privacy of chemical structure is of paramount importance for the industrial sector, in particular for the pharmaceutical industry. At the same time, companies handle large amounts of physico-chemical and biological data that could be shared in order to improve our molecular understanding of pharmacokinetic and toxicological properties, which could lead to improved predictivity and shorten the development time for drugs, in particular in the early phases of drug discovery. The current study provides some theoretical limits on the information required to produce reverse engineering of molecules from generated descriptors and demonstrates that the information content of molecules can be as low as less than one bit per atom. Thus theoretically just one descriptor can be used to completely disclose the molecular structure. Instead of sharing descriptors, we propose to share surrogate data. The sharing of surrogate data is nothing else but sharing of reliably predicted molecules. The use of surrogate data can provide the same information as the original set. We consider the practical application of this idea to predict lipophilicity of chemical compounds and we demonstrate that surrogate and real (original) data provides similar prediction ability. Thus, our proposed strategy makes it possible not only to share descriptors, but also complete collections of surrogate molecules without the danger of disclosing the underlying molecular structures.

Chemical Engineering↗

Eclair--a web service for unravelling species origin of sequences sampled from mixed host interfaces.

The identification of the genes that participate at the biological interface of two species remains critical to our understanding of the mechanisms of disease resistance, disease susceptibility and symbiosis. The sequencing of complementary DNA (cDNA) libraries prepared from the biological interface between two organisms provides an inexpensive way to identify the novel genes that may be expressed as a cause or consequence of compatible or incompatible interactions. Sequence classification and annotation of species origin typically use an orthology-based approach and require access to large portions of either genome, or a close relative. Novel species- or clade-specific sequences may have no counterpart within existing databases and remain ambiguous features. Here we present a web-service, Eclair, which utilizes support vector machines for the classification of the origin of expressed sequence tags stemming from mixed host cDNA libraries. In addition to providing an interface for the classification of sequences, users are presented with the opportunity to train a model to suit their preferred species pair. Eclair is freely available at http://eclair.btk.fi.

Artificial Intelligence↗

Super paramagnetic clustering of protein sequences.

BACKGROUND: Detection of sequence homologues represents a challenging task that is important for the discovery of protein families and the reliable application of automatic annotation methods. The presence of domains in protein families of diverse function, inhomogeneity and different sizes of protein families create considerable difficulties for the application of published clustering methods. RESULTS: Our work analyses the Super Paramagnetic Clustering (SPC) and its extension, global SPC (gSPC) algorithm. These algorithms cluster input data based on a method that is analogous to the treatment of an inhomogeneous ferromagnet in physics. For the SwissProt and SCOP databases we show that the gSPC improves the specificity and sensitivity of clustering over the original SPC and Markov Cluster algorithm (TRIBE-MCL) up to 30%. The three algorithms provided similar results for the MIPS FunCat 1.3 annotation of four bacterial genomes, Bacillus subtilis, Helicobacter pylori, Listeria innocua and Listeria monocytogenes. However, the gSPC covered about 12% more sequences compared to the other methods. The SPC algorithm was programmed in house using C++ and it is available at http://mips.gsf.de/proj/spc. The FunCat annotation is available at http://mips.gsf.de. CONCLUSION: The gSPC calculated to a higher accuracy or covered a larger number of sequences than the TRIBE-MCL algorithm. Thus it is a useful approach for automatic detection of protein families and unsupervised annotation of full genomes.

Algorithms↗

MIPS bacterial genomes functional annotation benchmark dataset.

MOTIVATION: Any development of new methods for automatic functional annotation of proteins according to their sequences requires high-quality data (as benchmark) as well as tedious preparatory work to generate sequence parameters required as input data for the machine learning methods. Different program settings and incompatible protocols make a comparison of the analyzed methods difficult. RESULTS: The MIPS Bacterial Functional Annotation Benchmark dataset (MIPS-BFAB) is a new, high-quality resource comprising four bacterial genomes manually annotated according to the MIPS functional catalogue (FunCat). These resources include precalculated sequence parameters, such as sequence similarity scores, InterPro domain composition and other parameters that could be used to develop and benchmark methods for functional annotation of bacterial protein sequences. These data are provided in XML format and can be used by scientists who are not necessarily experts in genome annotation. AVAILABILITY: BFAB is available at http://mips.gsf.de/proj/bfab

Bacterial Proteins↗

Virtual computational chemistry laboratory--design and description.

Internet technology offers an excellent opportunity for the development of tools by the cooperative effort of various groups and institutions. We have developed a multi-platform software system, Virtual Computational Chemistry Laboratory, http://www.vcclab.org, allowing the computational chemist to perform a comprehensive series of molecular indices/properties calculations and data analysis. The implemented software is based on a three-tier architecture that is one of the standard technologies to provide client-server services on the Internet. The developed software includes several popular programs, including the indices generation program, DRAGON, a 3D structure generator, CORINA, a program to predict lipophilicity and aqueous solubility of chemicals, ALOGPS and others. All these programs are running at the host institutes located in five countries over Europe. In this article we review the main features and statistics of the developed system that can be used as a prototype for academic and industry models.

Computer Simulation↗

Gene selection from microarray data for cancer classification--a machine learning approach.

A DNA microarray can track the expression levels of thousands of genes simultaneously. Previous research has demonstrated that this technology can be useful in the classification of cancers. Cancer microarray data normally contains a small number of samples which have a large number of gene expression levels as features. To select relevant genes involved in different types of cancer remains a challenge. In order to extract useful gene information from cancer microarray data and reduce dimensionality, feature selection algorithms were systematically investigated in this study. Using a correlation-based feature selector combined with machine learning algorithms such as decision trees, naïve Bayes and support vector machines, we show that classification performance at least as good as published results can be obtained on acute leukemia and diffuse large B-cell lymphoma microarray data sets. We also demonstrate that a combined use of different classification and feature selection approaches makes it possible to select relevant genes with high confidence. This is also the first paper which discusses both computational and biological evidence for the involvement of zyxin in leukaemogenesis.

Algorithms↗

Exploiting scale-free information from expression data for cancer classification.

Most studies concerning expression data analyses usually exploit information on the variability of gene intensity across samples. This information is sensitive to initial data processing, which affects the final conclusions. However expression data contains scale-free information, which is directly comparable between different samples. We propose to use the pairwise ratio of gene expression values rather than their absolute intensities for a classification of expression data. This information is stable to data processing and thus more attractive for classification analyses. In proposed schema of data analyses only information on relative gene expression levels in each sample is exploited. Testing on publicly available datasets leads to superior classification results.

Breast Neoplasms↗

Support vector machines for separation of mixed plant-pathogen EST collections based on codon usage.

MOTIVATION: Discovery of host and pathogen genes expressed at the plant-pathogen interface often requires the construction of mixed libraries that contain sequences from both genomes. Sequence identification requires high-throughput and reliable classification of genome origin. When using single-pass cDNA sequences difficulties arise from the short sequence length, the lack of sufficient taxonomically relevant sequence data in public databases and ambiguous sequence homology between plant and pathogen genes. RESULTS: A novel method is described, which is independent of the availability of homologous genes and relies on subtle differences in codon usage between plant and fungal genes. We used support vector machines (SVMs) to identify the probable origin of sequences. SVMs were compared to several other machine learning techniques and to a probabilistic algorithm (PF-IND) for expressed sequence tag (EST) classification also based on codon bias differences. Our software (Eclat) has achieved a classification accuracy of 93.1% on a test set of 3217 EST sequences from Hordeum vulgare and Blumeria graminis, which is a significant improvement compared to PF-IND (prediction accuracy of 81.2% on the same test set). EST sequences with at least 50 nt of coding sequence can be classified using Eclat with high confidence. Eclat allows training of classifiers for any host-pathogen combination for which there are sufficient classified training sequences. AVAILABILITY: Eclat is freely available on the Internet (http://mips.gsf.de/proj/est) or on request as a standalone version. CONTACT: friedel@informatik.uni-muenchen.de.

Algorithms↗

Application of ALOGPS 2.1 to predict log D distribution coefficient for Pfizer proprietary compounds.

Evaluation of the ALOGPS, ACD Labs LogD, and PALLAS PrologD suites to calculate the log D distribution coefficient resulted in high root-mean-squared error (RMSE) of 1.0-1.5 log for two in-house Pfizer's log D data sets of 17,861 and 640 compounds. Inaccuracy in log P prediction was the limiting factor for the overall log D estimation by these algorithms. The self-learning feature of the ALOGPS (LIBRARY mode) remarkably improved the accuracy in log D prediction, and an rmse of 0.64-0.65 was calculated for both data sets.

1-Octanol↗

A web portal for classification of expression data using maximal margin linear programming.

The Maximal Margin (MAMA) linear programming classification algorithm has recently been proposed and tested for cancer classification based on expression data. It demonstrated sound performance on publicly available expression datasets. We developed a web interface to allow potential users easy access to the MAMA classification tool. Basic and advanced options provide flexibility in exploitation. The input data format is the same as that used in most publicly available datasets. This makes the web resource particularly convenient for non-expert machine learning users working in the field of expression data analysis.

Algorithms↗

Optimization models for cancer classification: extracting gene interaction information from microarray expression data.

MOTIVATION: Microarray data appear particularly useful to investigate mechanisms in cancer biology and represent one of the most powerful tools to uncover the genetic mechanisms causing loss of cell cycle control. Recently, several different methods to employ microarray data as a diagnostic tool in cancer classification have been proposed. These procedures take changes in the expression of particular genes into account but do not consider disruptions in certain gene interactions caused by the tumor. It is probable that some genes participating in tumor development do not change their expression level dramatically. Thus, they cannot be detected by simple classification approaches used previously. For these reasons, a classification procedure exploiting information related to changes in gene interactions is needed. RESULTS: We propose a MAximal MArgin Linear Programming (MAMA) method for the classification of tumor samples based on microarray data. This procedure detects groups of genes and constructs models (features) that strongly correlate with particular tumor types. The detected features include genes whose functional relations are changed for particular cancer types. The proposed method was tested on two publicly available datasets and demonstrated a prediction ability superior to previously employed classification schemes. AVAILABILITY: The MAMA system was developed using the linear programming system LINDO http://www.lindo.com. A Perl script that specifies the optimization problem for this software is available upon request from the authors.

Algorithms↗

Application of ALOGPS to predict 1-octanol/water distribution coefficients, logP, and logD, of AstraZeneca in-house database.

The ALOGPS 2.1 was developed to predict 1-octanol/water partition coefficients, logP, and aqueous solubility of neutral compounds. An exclusive feature of this program is its ability to incorporate new user-provided data by means of self-learning properties of Associative Neural Networks. Using this feature, it calculated a similar performance, RMSE = 0.7 and mean average error 0.5, for 2569 neutral logP, and 8122 pH-dependent logD(7.4), distribution coefficients from the AstraZeneca "in-house" database. The high performance of the program for the logD(7.4) prediction looks surprising, because this property also depends on ionization constants pKa. Therefore, logD(7.4) is considered to be more difficult to predict than its neutral analog. We explain and illustrate this result and, moreover, discuss a possible application of the approach to calculate other pharmacokinetic and biological activities of chemicals important for drug development.

1-Octanol↗

An unsupervised automatic method for sorting neuronal spike waveforms in awake and freely moving animals.

The present study introduces an approach to automatic classification of extracellularly recorded action potentials of neurons. The classification of spike waveform is considered a pattern recognition problem of special segments of signal that correspond to the appearance of spikes. The spikes generated by one neuron should be recognized as members of the same class. The spike waveforms are described by the nonlinear oscillating model as an ordinary differential equation with perturbation, thus characterizing the signal distortions in both amplitude and phase. It is shown that the use of local variables reduces the problem of spike recognition to the separation of a mixture of normal distributions in the transformed feature space. We have developed an unsupervised iteration-learning algorithm that estimates the number of classes and their centers according to the distance between spike trajectories in phase space. This algorithm scans the learning set to evaluate spike trajectories with maximal probability density in their neighborhood. Following the learning, the procedure of minimal distance is used to perform spike recognition. Estimation of trajectories in phase space requires calculation of the first- and second-order derivatives, and integral operators with piecewise polynomial kernels were used. This provided the computational efficiency of the developed approach for real-time application as required by recordings in behaving animals and in human neurosurgical operations. The new method of spike sorting was tested on simulated and real data and performed better than other approaches currently used in neurophysiology.

Action Potentials↗