Search PubMed⌕ Search

Biomedical subjects

Pierre Baldi

Publications and source records attributed to Pierre Baldi.

At least 19 recordsLinked to original sources

Cilia in the brain display region-dependent oscillations of length and orientation.

In this study, we conducted high-throughput spatiotemporal analysis of primary cilia length and orientation across 22 mouse brain regions. We developed automated image analysis algorithms, which enabled us to examine over 10 million individual cilia, generating the largest spatiotemporal atlas of cilia. We found that cilia length and orientation display substantial variations across different brain regions and exhibit fluctuations over a 24-h period, with region-specific peaks during light-dark phases. Our analysis revealed unique orientation patterns of cilia, suggesting that cilia orientation within the brain is not random but follows specific patterns. Using BioCycle, we identified rhythmic fluctuations in cilia length across five brain regions: the nucleus accumbens core, somatosensory cortex, and the dorsomedial, ventromedial, and arcuate hypothalamic nuclei. Our findings present novel insights into the brain cilia dynamics, and highlight the need for further investigation into cilia's role in the brain's response to environmental changes and regulation of oscillatory physiological processes.

Animals↗

An enhanced MITOMAP with a global mtDNA mutational phylogeny.

The MITOMAP (http://www.mitomap.org) data system for the human mitochondrial genome has been greatly enhanced by the addition of a navigable mutational mitochondrial DNA (mtDNA) phylogenetic tree of approximately 3000 mtDNA coding region sequences plus expanded pathogenic mutation tables and a nuclear-mtDNA pseudogene (NUMT) data base. The phylogeny reconstructs the entire mutational history of the human mtDNA, thus defining the mtDNA haplogroups and differentiating ancient from recent mtDNA mutations. Pathogenic mutations are classified by both genotype and phenotype, and the NUMT sequences permits detection of spurious inclusion of pseudogene variants during mutation analysis. These additions position MITOMAP for the implementation of our automated mtDNA sequence analysis system, Mitomaster.

DNA, Mitochondrial↗

Identification of humoral immune responses in protein microarrays using DNA microarray data analysis techniques.

MOTIVATION: We present a study of antigen expression signals from a newly developed high-throughput protein microarray technique. These signals are a measure of antibody-antigen binding activity and provide a basis for understanding humoral immune responses to various infectious agents and supporting vaccine and diagnostic development. RESULTS: We investigate the characteristics of these expression profiles and show that noise models, normalization, variance estimation and differential expression analysis techniques developed in the context of DNA microarray analysis can be adapted and applied to these protein arrays. Using a high-dimensional dataset containing measurements of expression profiles of antibody reactivity against each protein (295 antigens and 9 controls) in 42 malaria (Plasmodium falciparum) protein arrays derived from 22 donors with various clinical presentations of malaria, we present a methodology for the analysis and identification of significantly expressed antigens targeted by immune responses for individual sera, groups of sera and across stages of infection. We also conduct a short study highlighting the top immunoreactive antigens where we identify three novel high priority antigens for future evaluation. AVAILABILITY: All software programs (in R) used for the analysis described in this paper are freely available for academic purposes at www.igb.uci.edu/servers/servers.html.

Algorithms↗

A machine learning information retrieval approach to protein fold recognition.

MOTIVATION: Recognizing proteins that have similar tertiary structure is the key step of template-based protein structure prediction methods. Traditionally, a variety of alignment methods are used to identify similar folds, based on sequence similarity and sequence-structure compatibility. Although these methods are complementary, their integration has not been thoroughly exploited. Statistical machine learning methods provide tools for integrating multiple features, but so far these methods have been used primarily for protein and fold classification, rather than addressing the retrieval problem of fold recognition-finding a proper template for a given query protein. RESULTS: Here we present a two-stage machine learning, information retrieval, approach to fold recognition. First, we use alignment methods to derive pairwise similarity features for query-template protein pairs. We also use global profile-profile alignments in combination with predicted secondary structure, relative solvent accessibility, contact map and beta-strand pairing to extract pairwise structural compatibility features. Second, we apply support vector machines to these features to predict the structural relevance (i.e. in the same fold or not) of the query-template pairs. For each query, the continuous relevance scores are used to rank the templates. The FOLDpro approach is modular, scalable and effective. Compared with 11 other fold recognition methods, FOLDpro yields the best results in almost all standard categories on a comprehensive benchmark dataset. Using predictions of the top-ranked template, the sensitivity is approximately 85, 56, and 27% at the family, superfamily and fold levels respectively. Using the 5 top-ranked templates, the sensitivity increases to 90, 70, and 48%.

Algorithms↗

Large-scale prediction of disulphide bridges using kernel methods, two-dimensional recursive neural networks, and weighted graph matching.

The formation of disulphide bridges between cysteines plays an important role in protein folding, structure, function, and evolution. Here, we develop new methods for predicting disulphide bridges in proteins. We first build a large curated data set of proteins containing disulphide bridges to extract relevant statistics. We then use kernel methods to predict whether a given protein chain contains intrachain disulphide bridges or not, and recursive neural networks to predict the bonding probabilities of each pair of cysteines in the chain. These probabilities in turn lead to an accurate estimation of the total number of disulphide bridges and to a weighted graph matching problem that can be addressed efficiently to infer the global disulphide bridge connectivity pattern. This approach can be applied both in situations where the bonded state of each cysteine is known, or in ab initio mode where the state is unknown. Furthermore, it can easily cope with chains containing an arbitrary number of disulphide bridges, overcoming one of the major limitations of previous approaches. It can classify individual cysteine residues as bonded or nonbonded with 87% specificity and 89% sensitivity. The estimate for the total number of bridges in each chain is correct 71% of the times, and within one from the true value over 94% of the times. The prediction of the overall disulphide connectivity pattern is exact in about 51% of the chains. In addition to using profiles in the input to leverage evolutionary information, including true (but not predicted) secondary structure and solvent accessibility information yields small but noticeable improvements. Finally, once the system is trained, predictions can be computed rapidly on a proteomic or protein-engineering scale. The disulphide bridge prediction server (DIpro), software, and datasets are available through www.igb.uci.edu/servers/psss.html.

Amino Acid Sequence↗

Prediction of protein stability changes for single-site mutations using support vector machines.

Accurate prediction of protein stability changes resulting from single amino acid mutations is important for understanding protein structures and designing new proteins. We use support vector machines to predict protein stability changes for single amino acid mutations leveraging both sequence and structural information. We evaluate our approach using cross-validation methods on a large dataset of single amino acid mutations. When only the sign of the stability changes is considered, the predictive method achieves 84% accuracy-a significant improvement over previously published results. Moreover, the experimental results show that the prediction accuracy obtained using sequence alone is close to the accuracy obtained using tertiary structure information. Because our method can accurately predict protein stability changes using primary sequence information only, it is applicable to many situations where the tertiary structure is unknown, overcoming a major limitation of previous methods which require tertiary information. The web server for predictions of protein stability changes upon mutations (MUpro), software, and datasets are available at http://www.igb.uci.edu/servers/servers.html.

Binding Sites↗

Structure-based inhibitor design of AccD5, an essential acyl-CoA carboxylase carboxyltransferase domain of Mycobacterium tuberculosis.

Mycolic acids and multimethyl-branched fatty acids are found uniquely in the cell envelope of pathogenic mycobacteria. These unusually long fatty acids are essential for the survival, virulence, and antibiotic resistance of Mycobacterium tuberculosis. Acyl-CoA carboxylases (ACCases) commit acyl-CoAs to the biosynthesis of these unique fatty acids. Unlike other organisms such as Escherichia coli or humans that have only one or two ACCases, M. tuberculosis contains six ACCase carboxyltransferase domains, AccD1-6, whose specific roles in the pathogen are not well defined. Previous studies indicate that AccD4, AccD5, and AccD6 are important for cell envelope lipid biosynthesis and that its disruption leads to pathogen death. We have determined the 2.9-Angstroms crystal structure of AccD5, whose sequence, structure, and active site are highly conserved with respect to the carboxyltransferase domain of the Streptomyces coelicolor propionyl-CoA carboxylase. Contrary to the previous proposal that AccD4-5 accept long-chain acyl-CoAs as their substrates, both crystal structure and kinetic assay indicate that AccD5 prefers propionyl-CoA as its substrate and produces methylmalonyl-CoA, the substrate for the biosyntheses of multimethyl-branched fatty acids such as mycocerosic, phthioceranic, hydroxyphthioceranic, mycosanoic, and mycolipenic acids. Extensive in silico screening of National Cancer Institute compounds and the University of California, Irvine, ChemDB database resulted in the identification of one inhibitor with a K(i) of 13.1 microM. Our results pave the way toward understanding the biological roles of key ACCases that commit acyl-CoAs to the biosynthesis of cell envelope fatty acids, in addition to providing a target for structure-based development of antituberculosis therapeutics.

Antitubercular Agents↗

A tandem affinity tag for two-step purification under fully denaturing conditions: application in ubiquitin profiling and protein complex identification combined with in vivocross-linking.

Tandem affinity strategies reach exceptional protein purification grades and have considerably improved the outcome of mass spectrometry-based proteomic experiments. However, current tandem affinity tags are incompatible with two-step purification under fully denaturing conditions. Such stringent purification conditions are desirable for mass spectrometric analyses of protein modifications as they result in maximal preservation of posttranslational modifications. Here we describe the histidine-biotin (HB) tag, a new tandem affinity tag for two-step purification under denaturing conditions. The HB tag consists of a hexahistidine tag and a bacterially derived in vivo biotinylation signal peptide that induces efficient biotin attachment to the HB tag in yeast and mammalian cells. HB-tagged proteins can be sequentially purified under fully denaturing conditions, such as 8 m urea, by Ni(2+) chelate chromatography and binding to streptavidin resins. The stringent separation conditions compatible with the HB tag prevent loss of protein modifications, and the high purification grade achieved by the tandem affinity strategy facilitates mass spectrometric analysis of posttranslational modifications. Ubiquitination is a particularly sensitive protein modification that is rapidly lost during purification under native conditions due to ubiquitin hydrolase activity. The HB tag is ideal to study ubiquitination because the denaturing conditions inhibit hydrolase activity, and the tandem affinity strategy greatly reduces nonspecific background. We tested the HB tag in proteome-wide ubiquitin profiling experiments in yeast and identified a number of known ubiquitinated proteins as well as so far unidentified candidate ubiquitination targets. In addition, the stringent purification conditions compatible with the HB tag allow effective mass spectrometric identification of in vivo cross-linked protein complexes, thereby expanding proteomic analyses to the description of weakly or transiently associated protein complexes.

Affinity Labels↗

Modular DAG-RNN architectures for assembling coarse protein structures.

We develop and test machine learning methods for the prediction of coarse 3D protein structures, where a protein is represented by a set of rigid rods associated with its secondary structure elements (alpha-helices and beta-strands). First, we employ cascades of recursive neural networks derived from graphical models to predict the relative placements of segments. These are represented as discretized distance and angle maps, and the discretization levels are statistically inferred from a large and curated dataset. Coarse 3D folds of proteins are then assembled starting from topological information predicted in the first stage. Reconstruction is carried out by minimizing a cost function taking the form of a purely geometrical potential. We show that the proposed architecture outperforms simpler alternatives and can accurately predict binary and multiclass coarse maps. The reconstruction procedure proves to be fast and often leads to topologically correct coarse structures that could be exploited as a starting point for various protein modeling strategies. The fully integrated rod-shaped protein builder (predictor of contact maps + reconstruction algorithm) can be accessed at http://distill.ucd.ie/.

Algorithms↗

Global landscape of recent inferred Darwinian selection for Homo sapiens.

By using the 1.6 million single-nucleotide polymorphism (SNP) genotype data set from Perlegen Sciences [Hinds, D. A., Stuve, L. L., Nilsen, G. B., Halperin, E., Eskin, E., Ballinger, D. G., Frazer, K. A. & Cox, D. R. (2005) Science 307, 1072-1079], a probabilistic search for the landscape exhibited by positive Darwinian selection was conducted. By sorting each high-frequency allele by homozygosity, we search for the expected decay of adjacent SNP linkage disequilibrium (LD) at recently selected alleles, eliminating the need for inferring haplotype. We designate this approach the LD decay (LDD) test. By these criteria, 1.6% of Perlegen SNPs were found to exhibit the genetic architecture of selection. These results were confirmed on an independently generated data set of 1.0 million SNP genotypes (International Human Haplotype Map Phase I freeze). Simulation studies indicate that the LDD test, at the megabase scale used, effectively distinguishes selection from other causes of extensive LD, such as inversions, population bottlenecks, and admixture. The approximately 1,800 genes identified by the LDD test were clustered according to Gene Ontology (GO) categories. Based on overrepresentation analysis, several predominant biological themes are common in these selected alleles, including host-pathogen interactions, reproduction, DNA metabolism/cell cycle, protein metabolism, and neuronal function.

Alleles↗

ChemDB: a public database of small molecules and related chemoinformatics resources.

MOTIVATION: The development of chemoinformatics has been hampered by the lack of large, publicly available, comprehensive repositories of molecules, in particular of small molecules. Small molecules play a fundamental role in organic chemistry and biology. They can be used as combinatorial building blocks for chemical synthesis, as molecular probes in chemical genomics and systems biology, and for the screening and discovery of new drugs and other useful compounds. RESULTS: We describe ChemDB, a public database of small molecules available on the Web. ChemDB is built using the digital catalogs of over a hundred vendors and other public sources and is annotated with information derived from these sources as well as from computational methods, such as predicted solubility and three-dimensional structure. It supports multiple molecular formats and is periodically updated, automatically whenever possible. The current version of the database contains approximately 4.1 million commercially available compounds and 8.2 million counting isomers. The database includes a user-friendly graphical interface, chemical reactions capabilities, as well as unique search capabilities. AVAILABILITY: Database and datasets are available on http://cdb.ics.uci.edu.

Access to Information↗

On the relationship between deterministic and probabilistic directed Graphical models: from Bayesian networks to recursive neural networks.

Machine learning methods that can handle variable-size structured data such as sequences and graphs include Bayesian networks (BNs) and Recursive Neural Networks (RNNs). In both classes of models, the data is modeled using a set of observed and hidden variables associated with the nodes of a directed acyclic graph. In BNs, the conditional relationships between parent and child variables are probabilistic, whereas in RNNs they are deterministic and parameterized by neural networks. Here, we study the formal relationship between both classes of models and show that when the source nodes variables are observed, RNNs can be viewed as limits, both in distribution and probability, of BNs with local conditional distributions that have vanishing covariance matrices and converge to delta functions. Conditions for uniform convergence are also given together with an analysis of the behavior and exactness of Belief Propagation (BP) in 'deterministic' BNs. Implications for the design of mixed architectures and the corresponding inference algorithms are briefly discussed.

Bayes Theorem↗

Graph kernels for chemical informatics.

Increased availability of large repositories of chemical compounds is creating new challenges and opportunities for the application of machine learning methods to problems in computational chemistry and chemical informatics. Because chemical compounds are often represented by the graph of their covalent bonds, machine learning methods in this domain must be capable of processing graphical structures with variable size. Here, we first briefly review the literature on graph kernels and then introduce three new kernels (Tanimoto, MinMax, Hybrid) based on the idea of molecular fingerprints and counting labeled paths of depth up to d using depth-first search from each possible vertex. The kernels are applied to three classification problems to predict mutagenicity, toxicity, and anti-cancer activity on three publicly available data sets. The kernels achieve performances at least comparable, and most often superior, to those previously reported in the literature reaching accuracies of 91.5% on the Mutag dataset, 65-67% on the PTC (Predictive Toxicology Challenge) dataset, and 72% on the NCI (National Cancer Institute) dataset. Properties and tradeoffs of these kernels, as well as other proposed kernels that leverage 1D or 3D representations of molecules, are briefly discussed.

Anticarcinogenic Agents↗

The absence of favorable aromatic interactions between beta-sheet peptides.

This paper asks whether interactions between phenylalanine (Phe) residues of the non-hydrogen-bonded cross-strand pairs of antiparallel beta-sheets are important and finds that they are not. Peptides 1a-d [o-BuO-C6H4CO-AA1-Orn(i-PrCO-Hao)-Phe-Ile-AA5-NHMe: 1a AA1, AA5 = Phe; 1b AA1, AA5 = Cha (cyclohexylalanine); 1c AA1 = Phe, AA5 = Cha; 1d AA1 = Cha, AA5 = Phe] provide a sensitive system for probing interactions between phenylalanine residues. These peptides form beta-sheet homodimers in organic solvents. When the homodimers of different peptides are mixed, they equilibrate to form heterodimers, as well as homodimers. The position of the equilibrium reflects the propensity of the first (AA1) and fifth (AA5) amino acids to interact within the non-hydrogen-bonded cross-strand pairs of beta-sheets. Mixing peptides 1a-d in all six possible binary combinations provides a measure of the relative propensities of Phe and Cha to pair. Analysis by 1H NMR spectroscopy of the equilibrium constants in CDCl3 solution reveals no significant preference for the formation of Phe-Phe pairs. The equilibria in all six experiments are essentially statistical (K approximately 4), and no (<0.1 kcal/mol) preference is seen for any pairing combination. A survey of Phe-Phe pairs in the Interchain beta-Sheet Database (http://www.igb.uci.edu/servers/icbs/) corroborates that little significant contact occurs between the aromatic rings in the non-hydrogen-bonded cross-strand pairs of antiparallel beta-sheets at the interface between polypeptide chains. Even though contacts between aromatic rings are favorable when they are of suitable geometry, the energetic price of achieving suitable geometries appears to offset the energetic benefits of such contacts in the current model system, as well as in proteins.

Amino Acids, Aromatic↗

Retroviruses and yeast retrotransposons use overlapping sets of host genes.

A collection of 4457 Saccharomyces cerevisiae mutants deleted for nonessential genes was screened for mutants with increased or decreased mobilization of the gypsylike retroelement Ty3. Of these, 64 exhibited increased and 66 decreased Ty3 transposition compared with the parental strain. Genes identified in this screen were grouped according to function by using GOnet software developed as part of this study. Gene clusters were related to chromatin and transcript elongation, translation and cytoplasmic RNA processing, vesicular trafficking, nuclear transport, and DNA maintenance. Sixty-six of the mutants were tested for Ty3 proteins and cDNA. Ty3 cDNA and transposition were increased in mutants affected in nuclear pore biogenesis and in a subset of mutants lacking proteins that interact physically or genetically with a replication clamp loader. Our results suggest that nuclear entry is linked mechanistically to Ty3 cDNA synthesis but that host replication factors antagonize Ty3 replication. Some of the factors we identified have been previously shown to affect Ty1 transposition and others to affect retroviral budding. Host factors, such as these, shared by distantly related Ty retroelements and retroviruses are novel candidates for antiviral targets.

Blotting, Southern↗

Global gene expression profiling in Escherichia coli K12: effects of oxygen availability and ArcA.

The ArcAB two-component system of Escherichia coli regulates the aerobic/anaerobic expression of genes that encode respiratory proteins whose synthesis is coordinated during aerobic/anaerobic cell growth. A genomic study of E. coli was undertaken to identify other potential targets of oxygen and ArcA regulation. A group of 175 genes generated from this study and our previous study on oxygen regulation (Salmon, K., Hung, S. P., Mekjian, K., Baldi, P., Hatfield, G. W., and Gunsalus, R. P. (2003) J. Biol. Chem. 278, 29837-29855), called our gold standard gene set, have p values <0.00013 and a posterior probability of differential expression value of 0.99. These 175 genes clustered into eight expression patterns and represent genes involved in a large number of cell processes, including small molecule biosynthesis, macromolecular synthesis, and aerobic/anaerobic respiration and fermentation. In addition, 119 of these 175 genes were also identified in our previous study of the fnr allele. A MEME/weight matrix method was used to identify a new putative ArcA-binding site for all genes of the E. coli genome. 16 new sites were identified upstream of genes in our gold standard set. The strict statistical analyses that we have performed on our data allow us to predict that 1139 genes in the E. coli genome are regulated either directly or indirectly by the ArcA protein with a 99% confidence level.

Alleles↗

Profiling the humoral immune response to infection by using proteome microarrays: high-throughput vaccine and diagnostic antigen discovery.

Despite the increasing availability of genome sequences from many human pathogens, the production of complete proteomes remains at a bottleneck. To address this need, a high-throughput PCR recombination cloning and expression platform has been developed that allows hundreds of genes to be batch-processed by using ordinary laboratory procedures without robotics. The method relies on high-throughput amplification of each predicted ORF by using gene specific primers, followed by in vivo homologous recombination into a T7 expression vector. The proteins are expressed in an Escherichia coli-based cell-free in vitro transcription/translation system, and the crude reactions containing expressed proteins are printed directly onto nitrocellulose microarrays without purification. The protein microarrays are useful for determining the complete antigen-specific humoral immune-response profile from vaccinated or infected humans and animals. The system was verified by cloning, expressing, and printing a vaccinia virus proteome consisting of 185 individual viral proteins. The chips were used to determine Ab profiles in serum from vaccinia virus-immunized humans, primates, and mice. Human serum has high titers of anti-E. coli Abs that require blocking to unmask vaccinia-specific responses. Naive humans exhibit reactivity against a subset of 13 antigens that were not associated with vaccinia immunization. Naive mice and primates lacked this background reactivity. The specific profiles between the three species differed, although a common subset of antigens was reactive after vaccinia immunization. These results verify this platform as a rapid way to comprehensively scan humoral immunity from vaccinated or infected humans and animals.

Animals↗

MITOMAP: a human mitochondrial genome database--2004 update.

MITOMAP (http://www.MITOMAP.org), a database for the human mitochondrial genome, has grown rapidly in data content over the past several years as interest in the role of mitochondrial DNA (mtDNA) variation in human origins, forensics, degenerative diseases, cancer and aging has increased dramatically. To accommodate this information explosion, MITOMAP has implemented a new relational database and an improved search engine, and all programs have been rewritten. System administrative changes have been made to improve security and efficiency, and to make MITOMAP compatible with a new automatic mtDNA sequence analyzer known as Mitomaster.

DNA, Mitochondrial↗