Search PubMed⌕ Search

Biomedical subjects

Jean-Loup Faulon

Publications and source records attributed to Jean-Loup Faulon.

13 recordsLinked to original sources

dAMN: a genome-scale neural-mechanistic hybrid model to predict bacterial growth dynamics.

SUMMARY: This study presents dAMN, a genome-scale neural-mechanistic hybrid model that combines neural networks with dynamic flux balance analysis to predict bacterial growth dynamics across diverse nutrient environments. Using a residual network architecture, dAMN predicts reaction fluxes and lag-phase parameters from initial medium composition, then integrates these predictions under stoichiometric constraints derived from genome-scale metabolic models. Trained on Escherichia coli and Pseudomonas putida growth datasets across combinatorial media, dAMN accurately forecasts temporal growth dynamics and generalizes to unseen media conditions, with mean R² ≥ 0.9. The model also reproduces biologically relevant behaviors including substrate depletion, acetate overflow, and diauxic shifts, while explicitly modeling lag phases usually absent from standard dFBA. AVAILABILITY AND IMPLEMENTATION: The dAMN software, associated models, and datasets are available at https://github.com/brsynth/dAMN-main-release and via Zenodo DOI: 10.5281/zenodo.17908125.

Escherichia coli↗

Prediction of beta-strand packing interactions using the signature product.

The prediction of beta-sheet topology requires the consideration of long-range interactions between beta-strands that are not necessarily consecutive in sequence. Since these interactions are difficult to simulate using ab initio methods, we propose a supplementary method able to assign beta-sheet topology using only sequence information. We envision using the results of our method to reduce the three-dimensional search space of ab initio methods. Our method is based on the signature molecular descriptor, which has been used previously to predict protein-protein interactions successfully, and to develop quantitative structure-activity relationships for small organic drugs and peptide inhibitors. Here, we show how the signature descriptor can be used in a Support Vector Machine to predict whether or not two beta-strands will pack adjacently within a protein. We then show how these predictions can be used to order beta-strands within beta-sheets. Using the entire PDB database with ten-fold cross-validation, we have achieved 74.0% accuracy in packing prediction and 75.6% accuracy in the prediction of edge strands. For the case of beta-strand ordering, we are able to predict the correct ordering accurately for 51.3% of the beta-sheets. Furthermore, using a simple confidence metric, we can determine those sheets for which accurate predictions can be obtained. For the top 25% highest confidence predictions, we are able to achieve 95.7% accuracy in beta-strand ordering. [Figure: see text].

Amino Acid Sequence↗

Reverse engineering chemical structures from molecular descriptors: how many solutions?

Physical, chemical and biological properties are the ultimate information of interest for chemical compounds. Molecular descriptors that map structural information to activities and properties are obvious candidates for information sharing. In this paper, we consider the feasibility of using molecular descriptors to safely exchange chemical information in such a way that the original chemical structures cannot be reverse engineered. To investigate the safety of sharing such descriptors, we compute the degeneracy (the number of structure matching a descriptor value) of several 2D descriptors, and use various methods to search for and reverse engineer structures. We examine degeneracy in the entire chemical space taking descriptors values from the alkane isomer series and the PubChem database. We further use a stochastic search to retrieve structures matching specific topological index values. Finally, we investigate the safety of exchanging of fragmental descriptors using deterministic enumeration.

Chemical Engineering↗

A deterministic algorithm for constrained enumeration of transmembrane protein folds.

A deterministic algorithm for enumeration of transmembrane protein folds is presented. Using a set of sparse pairwise atomic distance constraints (such as those obtained from chemical cross-linking, FRET, or dipolar EPR experiments), the algorithm performs an exhaustive search of secondary structure element packing conformations distributed throughout the entire conformational space. The end result is a set of distinct protein conformations, which can be scored and refined as part of a process designed for computational elucidation of transmembrane protein structures.

Algorithms↗

Optimal bundling of transmembrane helices using sparse distance constraints.

We present a two-step approach to modeling the transmembrane spanning helical bundles of integral membrane proteins using only sparse distance constraints, such as those derived from chemical cross-linking, dipolar EPR and FRET experiments. In Step 1, using an algorithm, we developed, the conformational space of membrane protein folds matching a set of distance constraints is explored to provide initial structures for local conformational searches. In Step 2, these structures refined against a custom penalty function that incorporates both measures derived from statistical analysis of solved membrane protein structures and distance constraints obtained from experiments. We begin by describing the statistical analysis of the solved membrane protein structures from which the theoretical portion of the penalty function was derived. We then describe the penalty function, and, using a set of six test cases, demonstrate that it is capable of distinguishing helical bundles that are close to the native bundle from those that are far from the native bundle. Finally, using a set of only 27 distance constraints extracted from the literature, we show that our method successfully recovers the structure of dark-adapted rhodopsin to within 3.2 A of the crystal structure.

Animals↗

Predicting protein-protein interactions using signature products.

MOTIVATION: Proteome-wide prediction of protein-protein interaction is a difficult and important problem in biology. Although there have been recent advances in both experimental and computational methods for predicting protein-protein interactions, we are only beginning to see a confluence of these techniques. In this paper, we describe a very general, high-throughput method for predicting protein-protein interactions. Our method combines a sequence-based description of proteins with experimental information that can be gathered from any type of protein-protein interaction screen. The method uses a novel description of interacting proteins by extending the signature descriptor, which has demonstrated success in predicting peptide/protein binding interactions for individual proteins. This descriptor is extended to protein pairs by taking signature products. The signature product is implemented within a support vector machine classifier as a kernel function. RESULTS: We have applied our method to publicly available yeast, Helicobacter pylori, human and mouse datasets. We used the yeast and H.pylori datasets to verify the predictive ability of our method, achieving from 70 to 80% accuracy rates using 10-fold cross-validation. We used the human and mouse datasets to demonstrate that our method is capable of cross-species prediction. Finally, we reused the yeast dataset to explore the ability of our algorithm to predict domains. CONTACT: smartin@sandia.gov

Algorithms↗

The signature molecular descriptor. 3. Inverse-quantitative structure-activity relationship of ICAM-1 inhibitory peptides.

We present a methodology for solving the inverse-quantitative structure-activity relationship (QSAR) problem using the molecular descriptor called signature. This methodology is detailed in four parts. First, we create a QSAR equation that correlates the occurrence of a signature to the activity values using a stepwise multilinear regression technique. Second, we construct constraint equations, specifically the graphicality and consistency equations, which facilitate the reconstruction of the solution compounds directly from the signatures. Third, we solve the set of constraint equations, which are both linear and Diophantine in nature. Last, we reconstruct and enumerate the solution molecules and calculate their activity values from the QSAR equation. We apply this inverse-QSAR method to a small set of LFA-1/ICAM-1 peptide inhibitors to assist in the search and design of more-potent inhibitory compounds. Many novel inhibitors were predicted, a number of which are predicted to be more potent than the strongest inhibitor in the training set. Two of the more potent inhibitors were synthesized and tested in-vivo, confirming them to be the strongest inhibiting peptides to date. Some of these compounds can be recycled to train a new QSAR and develop a more focused library of lead compounds.

Drug Design↗

Exploring the conformational space of membrane protein folds matching distance constraints.

Herein we present a computational technique for generating helix-membrane protein folds matching a predefined set of distance constraints, such as those obtained from NMR NOE, chemical cross-linking, dipolar EPR, and FRET experiments. The purpose of the technique is to provide initial structures for local conformational searches based on either energetic considerations or ad-hoc scoring criteria. In order to properly screen the conformational space, the technique generates an exhaustive list of conformations within a specified root-mean-square deviation (RMSD) where the helices are positioned in order to match the provided distances. Our results indicate that the number of structures decreases exponentially as the number of distances increases, and increases exponentially as the errors associated with the distances increases. We also found the number of solutions to be smaller when all the distances share one helix in common, compared to the case where the distances connect helices in a daisy-chain manner. We found that for 7 helices, at least 15 distances with errors up to 8 A are needed to produce a number of solutions that is not too large to be processed by local search refinement procedures. Finally, without energetic considerations, our enumeration technique retrieved the transmembrane domains of Bacteriorhodopsin (PDB entry1c3w), Halorhodopsin (1e12), Rhodopsin (1f88), Aquaporin-1 (1fqy), Glycerol uptake facilitator protein (1fx8), Sensory Rhodopsin (1jgj), and a subunit of Fumarate reductase flavoprotein (1qlaC) with Calpha level RMSDs of 3.0 A, 2.3 A, 3.2 A, 4.6 A, 6.0 A, 3.7 A, and 4.4 A, respectively.

Membrane Proteins↗

Developing a methodology for an inverse quantitative structure-activity relationship using the signature molecular descriptor.

The concept of signature as a molecular descriptor is introduced and various topological indices used in quantitative structure-activity relationships (QSARs) are expressed as functions of the new descriptor. The effectiveness of signature versus commonly used descriptors in QSAR analysis is demonstrated by correlating the activities of 121 HIV-1 protease inhibitors. Our approach to the inverse-QSAR problem consists of first finding the optimum sets of descriptor values best matching a target activity and then generating a focused library of candidate structures from the solution set of descriptor values. Both steps are facilitated by the use of signature.

Algorithms↗

The signature molecular descriptor. 1. Using extended valence sequences in QSAR and QSPR studies.

We present a new descriptor named signature based on extended valence sequence. The signature of an atom is a canonical representation of the atom's environment up to a predefined height h. The signature of a molecule is a vector of occurrence numbers of atomic signatures. Two QSAR and QSPR models based on signature are compared with models obtained using popular molecular 2D descriptors taken from a commercially available software (Molconn-Z). One set contains the inhibition concentration at 50% for 121 HIV-1 protease inhibitors, while the second set contains 12865 octanol/water partitioning coefficients (Log P). For both data sets, the models created by signature performed comparable to those from the commercially available descriptors in both correlating the data and in predicting test set values not used in the parametrization. While probing signature's QSAR and QSPR performances, we demonstrates that for any given molecule of diameter D, there is a molecular signature of height h </= D+1, from which any 2D descriptor can be computed. As a consequence of this finding any QSAR or QSPR involving 2D descriptors can be replaced with a relationship involving occurrence number of atomic signatures.

Algorithms↗

The signature molecular descriptor. 2. Enumerating molecules from their extended valence sequences.

We present a new algorithm that enumerates molecular structures matching a predefined extended valence sequence or signature. The algorithm can construct molecular structures composed of about 50 non-hydrogen atoms in CPU seconds time scale. The algorithm is run to produce all molecular structures matching the binding affinities (IC(50)) of some HIV-1 protease inhibitors. The algorithm is also used to compute the degeneracy, or the number of molecular structures, corresponding to a given signature. Signature degeneracy is systematically studied for varying signature heights on four molecular series, alkanes, alcohols, fullerene-type structures, and peptides. Signature degeneracy is compared with similar results obtained with popular topological indices (TIs). As a general rule, we find that signature degeneracy decreases as the signature height increases. We also find that alkanes, alcohols, and fullerene-type structures comprising n non-hydrogen atoms are uniquely characterized by signatures of height n/4, while peptides up to 4000 amino acids can be singled out with signatures of heights as small as 2 and 3.

Algorithms↗

The signature molecular descriptor. 4. Canonizing molecules using extended valence sequences.

We present a new algorithm to canonize molecular graphs using the signature molecular descriptor introduced in the previous papers of this series. While developed specifically for molecular structures, the algorithm can be used for any graph and is not limited to acyclic graphs, planar graphs, bounded valence, or bounded genus graphs, for which polynomial time algorithms exist. The algorithm is tested with benzenoid hydrocarbons and a database of 126,705 organic compounds. The algorithm's performances are compared against Brendan Mc Kay's Nauty algorithm, which is believed to be the fastest graph canonization algorithm for general graphs, with five series of graphs each comprising up to 30,000 vertices: 2D meshes (pericondensed benzenoids), 3D cages (fullerenes and nanotubes), 3D meshes (crystal lattices), 4D cages, and power law graphs (protein and gene networks). The algorithm can be downloaded as an open source code at http://www.cs.sandia.gov/ approximately jfaulon/QSAR.

Journal Article↗

Designing novel polymers with targeted properties using the signature molecular descriptor.

A method for solving the inverse quantitative structure-property relationship (QSPR) problem is presented which facilitates the design of novel polymers with targeted properties. Here, we demonstrate the efficacy of the approach using the targeted design of polymers exhibiting a desired glass transition temperature, heat capacity, and density. We present novel QSPRs based on the signature molecular descriptor capable of predicting glass transition temperature, heat capacity, density, molar volume, and cohesive energies of linear homopolymers with cross-validation squared correlation coefficients ranging between 0.81 and 0.95. Using these QSPRs, we show how the inverse problem can be solved to design poly(N-methyl hexamethylene sebacamide) despite the fact that the polymer was used not used in the training of this model.

Algorithms↗