Search PubMed⌕ Search

Biomedical subjects

Pierre Baldi

Publications and source records attributed to Pierre Baldi.

32 records · Page 2Linked to original sources

Three-stage prediction of protein beta-sheets by neural networks, alignments and graph algorithms.

MOTIVATION: Protein beta-sheets play a fundamental role in protein structure, function, evolution and bioengineering. Accurate prediction and assembly of protein beta-sheets, however, remains challenging because protein beta-sheets require formation of hydrogen bonds between linearly distant residues. Previous approaches for predicting beta-sheet topological features, such as beta-strand alignments, in general have not exploited the global covariation and constraints characteristic of beta-sheet architectures. RESULTS: We propose a modular approach to the problem of predicting/assembling protein beta-sheets in a chain by integrating both local and global constraints in three steps. The first step uses recursive neural networks to predict pairing probabilities for all pairs of interstrand beta-residues from profile, secondary structure and solvent accessibility information. The second step applies dynamic programming techniques to these probabilities to derive binding pseudoenergies and optimal alignments between all pairs of beta-strands. Finally, the third step uses graph matching algorithms to predict the beta-sheet architecture of the protein by optimizing the global pseudoenergy while enforcing strong global beta-strand pairing constraints. The approach is evaluated using cross-validation methods on a large non-homologous dataset and yields significant improvements over previous methods. AVAILABILITY: http://www.igb.uci.edu/servers/psss.html.

Algorithms↗

Kernels for small molecules and the prediction of mutagenicity, toxicity and anti-cancer activity.

MOTIVATION: Small molecules play a fundamental role in organic chemistry and biology. They can be used to probe biological systems and to discover new drugs and other useful compounds. As increasing numbers of large datasets of small molecules become available, it is necessary to develop computational methods that can deal with molecules of variable size and structure and predict their physical, chemical and biological properties. RESULTS: Here we develop several new classes of kernels for small molecules using their 1D, 2D and 3D representations. In 1D, we consider string kernels based on SMILES strings. In 2D, we introduce several similarity kernels based on conventional or generalized fingerprints. Generalized fingerprints are derived by counting in different ways subpaths contained in the graph of bonds, using depth-first searches. In 3D, we consider similarity measures between histograms of pairwise distances between atom classes. These kernels can be computed efficiently and are applied to problems of classification and prediction of mutagenicity, toxicity and anti-cancer activity on three publicly available datasets. The results derived using cross-validation methods are state-of-the-art. Tradeoffs between various kernels are briefly discussed. AVAILABILITY: Datasets available from http://www.igb.uci.edu/servers/servers.html

Animals↗

ICBS: a database of interactions between protein chains mediated by beta-sheet formation.

MOTIVATION: Interchain beta-sheet (ICBS) interactions occur widely in protein quaternary structures, interactions between proteins and protein aggregation. These interactions play a central role in many biological processes and in diseases ranging from AIDS and cancer to anthrax and Alzheimer's. RESULTS: We have created a comprehensive database of ICBS interactions that is updated on a weekly basis and allows entries to be sorted and searched by relevance and other criteria through a simple Web interface. We derive a simple ICBS index to quantify the relative contributions of the beta-ladders in the overall interchain interaction and compute first- and second-order statistics regarding amino acid composition and pairing at different relative positions in the beta-strands. Analysis of the database reveals a 15.8% prevalence of significant ICBS interactions, the majority of which involve the formation of antiparallel beta-sheets and many of which involve the formation of dimers and oligomers. The frequencies of amino acids in ICBS interfaces are similar to those in intrachain beta-sheet interfaces. A full range of non-covalent interactions between side chains complement the hydrogen-bonding interactions between the main chains. Polar amino acids pair preferentially with polar amino acids and non-polar amino acids pair preferentially with non-polar amino acids among antiparallel (i, j) pairs. We anticipate that the statistics and insights gained from the database will guide the development of agents that control interchain beta-sheet interactions and that the database will help identify new protein interactions and targets for these agents. AVAILABILITY: The database is available at: http://www.igb.uci.edu/servers/icbs/

Binding Sites↗

Structural proteomics of the poxvirus family.

Recent concerns over the potential use of variola virus-commonly known as smallpox-and other orthopox viruses as weapons of bioterrorism have increased research efforts towards creating new antiviral drugs and safer more effective vaccines. Here we introduce a new resource for structural information of poxvirus proteins: the poxvirus proteomics database (PPDB). In the PPDB, we leverage recently developed bioinformatics structure prediction tools on a genomic scale and provide results in a publicly accessible format. The current version of the system contains both experimentally determined and predicted information about protein structural features, such as secondary structure and relative solvent accessibility, as well as tertiary structure and homology information. The system is automated to read the primary sequences from the database, produce the new information for each sequence, and update the database monthly and as new tools are incorporated. The PPDB contains detailed information on the open reading frames (ORFs) in the Copenhagen strain of the vaccinia virus genome. The contents of the PPDB can be accessed through a simple web interface. Inclusion of additional poxvirus genomes in the PPDB is in progress. The PPDB has an upward scalable informatics infrastructure that can readily be applied to viral, bacterial, as well as eukaryotic genomes.

Antiviral Agents↗

Global gene expression profiling in Escherichia coli K12. The effects of oxygen availability and FNR.

The work presented here is a first step toward a long term goal of systems biology, the complete elucidation of the gene regulatory networks of a living organism. To this end, we have employed DNA microarray technology to identify genes involved in the regulatory networks that facilitate the transition of Escherichia coli cells from an aerobic to an anaerobic growth state. We also report the identification of a subset of these genes that are regulated by a global regulatory protein for anaerobic metabolism, FNR. Analysis of these data demonstrated that the expression of over one-third of the genes expressed during growth under aerobic conditions are altered when E. coli cells transition to an anaerobic growth state, and that the expression of 712 (49%) of these genes are either directly or indirectly modulated by FNR. The results presented here also suggest interactions between the FNR and the leucine-responsive regulatory protein (Lrp) regulatory networks. Because computational methods to analyze and interpret high dimensional DNA microarray data are still at an early stage, and because basic issues of data analysis are still being sorted out, much of the emphasis of this work is directed toward the development of methods to identify differentially expressed genes with a high level of confidence. In particular, we describe an approach for identifying gene expression patterns (clusters) obtained from multiple perturbation experiments based on a subset of genes that exhibit high probability for differential expression values.

Cell Division↗

LineUp: statistical detection of chromosomal homology with application to plant comparative genomics.

The identification of homologous regions between chromosomes forms the basis for studies of genome organization, comparative genomics, and evolutionary genomics. Identification of these regions can be based on either synteny or colinearity, but there are few methods to test statistically for significant evidence of homology. In the present study, we improve a preexisting method that used colinearity as the basis for statistical tests. Improvements include computational efficiency and a relaxation of the colinearity assumption. Two algorithms perform the method: FullPermutation, which searches exhaustively for runs of markers, and FastRuns, which trades faster run times for exhaustive searches. The algorithms described here are available in the LineUp package (http://www.igb.uci.edu/ approximately baldig/lineup). We explore the performance of both algorithms on simulated data and also on genetic map data from maize (Zea mays ssp. mays). The method has reasonable power to detect a homologous region; for example, in >90% of simulations, both algorithms detect a homologous region of 10 markers buried in a random background, even when the homologous regions have diverged by numerous inversion events. The methods were applied to four maize molecular maps. All maps indicate that the maize genome contains extensive regions of genomic duplication and multiplication. Nonetheless, maps differ substantially in the location of homologous regions, probably reflecting the incomplete nature of genetic map data. The variation among maps has important implications for evolutionary inference from genetic map data.

Chromosome Mapping↗

Differential analysis of DNA microarray gene expression data.

Here, we review briefly the sources of experimental and biological variance that affect the interpretation of high-dimensional DNA microarray experiments. We discuss methods using a regularized t-test based on a Bayesian statistical framework that allow the identification of differentially regulated genes with a higher level of confidence than a simple t-test when only a few experimental replicates are available. We also describe a computational method for calculating the global false-positive and false-negative levels inherent in a DNA microarray data set. This method provides a probability of differential expression for each gene based on experiment-wide false-positive and -negative levels driven by experimental error and biological variance.

Analysis of Variance↗

Global gene expression profiling in Escherichia coli K12. The effects of leucine-responsive regulatory protein.

Leucine-responsive regulatory protein (Lrp) is a global regulatory protein that affects the expression of multiple genes and operons in bacteria. Although the physiological purpose of Lrp-mediated gene regulation remains unclear, it has been suggested that it functions to coordinate cellular metabolism with the nutritional state of the environment. The results of gene expression profiles between otherwise isogenic lrp(+) and lrp(-) strains of Escherichia coli support this suggestion. The newly discovered Lrp-regulated genes reported here are involved either in small molecule or macromolecule synthesis or degradation, or in small molecule transport and environmental stress responses. Although many of these regulatory effects are direct, others are indirect consequences of Lrp-mediated changes in the expression levels of other global regulatory proteins. Because computational methods to analyze and interpret high dimensional DNA microarray data are still an early stage, much of the emphasis of this work is directed toward the development of methods to identify differentially expressed genes with a high level of confidence. In particular, we describe a Bayesian statistical framework for a posterior estimate of the standard deviation of gene measurements based on a limited number of replications. We also describe an algorithm to compute a posterior estimate of differential expression for each gene based on the experiment-wide global false positive and false negative level for a DNA microarray data set. This allows the experimenter to compute posterior probabilities of differential expression for each individual differential gene expression measurement.

Base Sequence↗

Prediction of coordination number and relative solvent accessibility in proteins.

Knowing the coordination number and relative solvent accessibility of all the residues in a protein is crucial for deriving constraints useful in modeling protein folding and protein structure and in scoring remote homology searches. We develop ensembles of bidirectional recurrent neural network architectures to improve the state of the art in both contact and accessibility prediction, leveraging a large corpus of curated data together with evolutionary information. The ensembles are used to discriminate between two different states of residue contacts or relative solvent accessibility, higher or lower than a threshold determined by the average value of the residue distribution or the accessibility cutoff. For coordination numbers, the ensemble achieves performances ranging within 70.6-73.9% depending on the radius adopted to discriminate contacts (6A-12A). These performances represent gains of 16-20% over the baseline statistical predictor, always assigning an amino acid to the largest class, and are 4-7% better than any previous method. A combination of different radius predictors further improves performance. For accessibility thresholds in the relevant 15-30% range, the ensemble consistently achieves a performance above 77%, which is 10-16% above the baseline prediction and better than other existing predictors, by up to several percentage points. For both problems, we quantify the improvement due to evolutionary information in the form of PSI-BLAST-generated profiles over BLAST profiles. The prediction programs are implemented in the form of two web servers, CONpro and ACCpro, available at http://promoter.ics.uci.edu/BRNN-PRED/.

Amino Acids↗

Improving the prediction of protein secondary structure in three and eight classes using recurrent neural networks and profiles.

Secondary structure predictions are increasingly becoming the workhorse for several methods aiming at predicting protein structure and function. Here we use ensembles of bidirectional recurrent neural network architectures, PSI-BLAST-derived profiles, and a large nonredundant training set to derive two new predictors: (a) the second version of the SSpro program for secondary structure classification into three categories and (b) the first version of the SSpro8 program for secondary structure classification into the eight classes produced by the DSSP program. We describe the results of three different test sets on which SSpro achieved a sustained performance of about 78% correct prediction. We report confusion matrices, compare PSI-BLAST to BLAST-derived profiles, and assess the corresponding performance improvements. SSpro and SSpro8 are implemented as web servers, available together with other structural feature predictors at: http://promoter.ics.uci.edu/BRNN-PRED/.

Algorithms↗

Distribution patterns of over-represented k-mers in non-coding yeast DNA.

MOTIVATION: Over-represented k-mers in genomic DNA regions are often of particular biological interest. For example, over-represented k-mers in co-regulated families of genes are associated with the DNA binding sites of transcription factors. To measure over-representation, we introduce a statistical background model based on single-mismatches, and apply it to the pooled 500 bp ORF Upstream Regions (USRs) of yeast. More importantly, we investigate the context and spatial distribution of over-represented k-mers in yeast USRs. RESULTS: Single and double-stranded spatial distributions of most over-represented k-mers are highly non-random, and predominantly cluster into a small number of classes that are robust with respect to over-representation measures. Specifically, we show that the three most common distribution patterns can be related to DNA structure, function, and evolution and correspond to: (a) homologous ORF clusters associated with sharply localized distributions; (b) regulatory elements associated with a symmetric broad hill-shaped distribution in the 50-200 bp USR; and (c) runs of As, Ts, and ATs associated with a broad hill-shaped distribution also in the 50-200 bp USR, with extreme structural properties. Analysis of over-representation, homology, localization, and DNA structure are essential components of a general data-mining approach to finding biologically important k-mers in raw genomic DNA and understanding the 'lexicon' of regulatory regions.

Amino Acid Motifs↗

Why are complementary DNA strands symmetric?

MOTIVATION: Over sufficiently long windows, complementary strands of DNA tend to have the same base composition. A few reports have indicated that this first-order parity rule extends at higher orders to oligonucleotide composition, at least in some organisms or taxa. However, the scientific literature falls short of providing a comprehensive study of reverse-complement symmetry at multiple orders and across the kingdom of life. It also lacks a characterization of this symmetry and a convincing explanation or clarification of its origin. RESULTS: We develop methods to measure and characterize symmetry at multiple orders, and analyze a wide set of genomes, encompassing single- and double-stranded RNA and DNA viruses, bacteria, archae, mitochondria, and eukaryota. We quantify symmetry at orders 1 to 9 for contiguous sequences and pools of coding and non-coding upstream regions, compare the observed symmetry levels to those predicted by simple statistical models, and factor out the effect of lower-order distributions. We establish the universality and variability range of first-order strand symmetry, as well as of its higher-order extensions, and demonstrate the existence of genuine high-order symmetric constraints. We show that ubiquitous reverse-complement symmetry does not result from a single cause, such as point mutation or recombination, but rather emerges from the combined effects of a wide spectrum of mechanisms operating at multiple orders and length scales.

Base Composition↗

Functional census of mutation sequence spaces: the example of p53 cancer rescue mutants.

Many biomedical problems relate to mutant functional properties across a sequence space of interest, e.g., flu, cancer, and HIV. Detailed knowledge of mutant properties and function improves medical treatment and prevention. A functional census of p53 cancer rescue mutants would aid the search for cancer treatments from p53 mutant rescue. We devised a general methodology for conducting a functional census of a mutation sequence space by choosing informative mutants early. The methodology was tested in a double-blind predictive test on the functional rescue property of 71 novel putative p53 cancer rescue mutants iteratively predicted in sets of three (24 iterations). The first double-blind 15-point moving accuracy was 47 percent and the last was 86 percent; r = 0.01 before an epiphanic 16th iteration and r = 0.92 afterward. Useful mutants were chosen early (overall r = 0.80). Code and data are freely available (http://www.igb.uci.edu/research/research.html, corresponding authors: R.H.L. for computation and R.K.B. for biology).

Artificial Intelligence↗