Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Knowledge base”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Data and knowledge based experimental design for fermentation process optimization.

A novel method for the sequential experimental design in order to optimize fed-batch fermentations was applied to a hyaluronidase fermentation by Streptococcus agalactiae. A Lambda-optimal design was introduced to minimize the model parameter estimation error and to maximize the performance of the fermentation process. The method employs hybrid models that contain mechanistic, fuzzy and neural network components.

Journal Article↗

CORA--a knowledge-based system for the analysis of case-control studies.

Carrying out a statistical analysis, the researcher is concerned with the problem of choosing an appropriate statistical technique from a large number of competing methods. Most common statistical software offer different methods for analysing the data without giving any support regarding the adequacy of a method for a particular data set. This paper outlines the main features of the computer system CORA which provides a statistical analysis of stratified contingency tables and additionally supports the researcher at the different steps of this analysis. Here, the support given by the system consists of two different aspects. On the one hand, the help system of CORA contains general information on the implemented statistical methods which can be obtained on request. On the other hand, an advice tool recommends an adequate statistical method which depends on the actual empirical case-control data to be analysed. To build up the advice tool, a set of rules being discovered by machine learning from simulation studies is integrated into the system CORA.

Case-Control Studies↗

Knowledge-based modeling of a legume lectin and docking of the carbohydrate ligand: the Ulex europaeus lectin I and its interaction with fucose.

Ulex europaeus isolectin I is specific for fucose-containing oligosaccharide such as H type 2 trisaccharide alpha-L-Fuc (1-->2) beta-D-Gal (1-->4) beta-D-GlcNAc. Several legume lectins have been crystallized and modeled, but no structural data are available concerning such fucose-binding lectin. The three-dimensional structure of Ulex europaeus isolectin I has been constructed using seven legume lectins for which high-resolution crystal structures were available. Some conserved water molecules, as well as the structural cations, were taken into account for building the model. In the predicted binding site, the most probable locations of the secondary hydroxyl groups were determined using the GRID method. Several possible orientations could be determined for a fucose residue. All of the four possible conformations compatible with energy calculations display several hydrogen bonds with Asp-87 and Ser-132 and a stacking interaction with Tyr-220 and Phe-136. In two orientations, the O-3 and O-4 hydroxyl groups of fucose are the most buried ones, whereas two other, the O-2 and O-3 hydroxyl groups are at the bottom of the site. Possible docking modes are also studied by analysis of the hydrophobic and hydrophilic surfaces for both the ligand and the protein. The SCORE method allows for a quantitative evaluation of the complementarity of these surfaces, on the basis of molecular lipophilicity calculations. The predictions presented here are compared with known biochemical data.

Artificial Intelligence↗

A model for single and multiple knowledge based networks.

The inherent black-box nature of neural networks is an important drawback with respect to the problem of explanation of neural network responses. Although several articles have tackled the problem of rule extraction from a single neural network, just a few papers have investigated rule extraction from several combined neural networks. In this article we describe how to translate symbolic rules into the Discretized Interpretable Multi-Layer Perceptron (DIMLP) and how to extract rules from one or several combined neural networks. Our approach consists of characterizing discriminant hyperplane frontiers. Unordered rules are extracted in polynomial time with respect to the size of the problem and the size of the network. Moreover, the degree of matching between extracted rules and neural network responses is 100% on training examples. We applied single DIMLP networks to 17 data sets related to medical diagnosis and medical prognosis problems. Results based on 10-fold cross-validation showed that the DIMLP model was on average as accurate as standard multi-layer perceptrons (MLP). Furthermore, DIMLP networks were significantly more accurate than CN2 on eight problems, whereas only on one problem CN2 was better than DIMLP. Finally, a non-Hodgkin lymphoma diagnosis problem based on classification of electrophoresis gels was defined. It turned out that ensembles of DIMLP networks were significantly more accurate than CN2 (96.1% +/- 1.4 versus 82.7% +/- 4.0). Finally, symbolic rules revealed the presence of five important spots for the discrimination of the class of Lymphocyte Leukemia/Chronic Lymphoid Leukemia (Lc/LLc), and the class of Centrocytic Lymphoma (Cc).

Algorithms↗

Knowledge-based design of bimodular and trimodular polyketide synthases based on domain and module swaps: a route to simple statin analogues.

BACKGROUND: Polyketides are structurally diverse natural products that have a range of medically useful activities. Nonaromatic bacterial polyketides are synthesised on modular polyketide synthase (PKS) multienzymes, in which each cycle of chain extension requires a different 'module' of enzymatic activities. Attempts to design and construct modular PKSs that synthesise specified novel polyketides provide a particularly stringent test of our understanding of PKS structure and function. RESULTS: We have constructed bimodular and trimodular PKSs based on DEBS1-TE, a derivative of the erythromycin PKS that contains only modules 1 and 2 and a thioesterase (TE), by substituting multiple domains with appropriate counterparts derived from the rapamycin PKS. Hybrid PKSs were obtained that synthesised the predicted target triketide lactones, which are simple analogues of cholesterol-lowering statins. In constructing intermodular fusions, whether between modules in the same or in different proteins, it was found advantageous to preserve intact the acyl carrier protein-ketosynthase (ACP-KS) didomain that spans the junction between successive modules. CONCLUSIONS: Relatively simple considerations govern the construction of functional hybrid PKSs. Fusion sites should be chosen either in the surface-accessible linker regions between enzymatic domains, as previously revealed, or just inside the conserved margins of domains. The interaction of an ACP domain with the adjacent KS domain, whether on the same polyketide or not, is of particular importance, both through conservation of appropriate protein-protein interactions, and through optimising molecular recognition of the altered polyketide chain in the key transfer of the acyl chain from the ACP of one module to the KS of the downstream module.

Amino Acid Sequence↗

Knowledge-based structure prediction of MHC class I bound peptides: a study of 23 complexes.

BACKGROUND: The binding of T-cell antigenic peptides to MHC molecules is a prerequisite for their immunogenicity. The ability to identify binding peptides based on the protein sequence is of great importance to the rational design of peptide vaccines. As the requirements for peptide binding cannot be fully explained by the peptide sequence per se, structural considerations should be taken into account and are expected to improve predictive algorithms. The first step in such an algorithm requires accurate and fast modeling of the peptide structure in the MHC-binding groove. RESULTS: We have used 23 solved peptide-MHC class I complexes as a source of structural information in the development of a modeling algorithm. The peptide backbones and MHC structures were used as the templates for prediction. Sidechain conformations were built based on a rotamer library, using the 'dead end elimination' approach. A simple energy function selects the favorable combination of rotamers for a given sequence. It further selects the correct backbone structure from a limited library. The influence of different parameters on the prediction quality was assessed. With a specific rotamer library that incorporates information from the peptide sidechains in the solved complexes, the algorithm correctly identifies 85% (92%) of all (buried) sidechains and selects the correct backbones. Under cross-validation, 70% (78%) of all (buried) residues are correctly predicted and most of all backbones. The interaction between peptide sidechains has a negligible effect on the prediction quality. CONCLUSIONS: The structure of the peptide sidechains follows from the interactions with the MHC and the peptide backbone, as the prediction is hardly influenced by sidechain interactions. The proposed methodology was able to select the correct backbone from a limited set. The impairment in performance under cross-validation suggests that, currently, the specific rotamer library is not satisfactorily representative. The predictions might improve with an increase in the data.

Animals↗

Text-based knowledge discovery: search and mining of life-sciences documents.

Text literature is playing an increasingly important role in biomedical discovery. The challenge is to manage the increasing volume, complexity and specialization of knowledge expressed in this literature. Although information retrieval or text searching is useful, it is not sufficient to find specific facts and relations. Information extraction methods are evolving to extract automatically specific, fine-grained terms corresponding to the names of entities referred to in the text, and the relationships that connect these terms. Information extraction is, in turn, a means to an end, and knowledge discovery methods are evolving for the discovery of still more-complex structures and connections among facts. These methods provide an interpretive context for understanding the meaning of biological data.

Biological Science Disciplines↗

Validation of a knowledge based reminder system for diagnostic test ordering in general practice.

We describe the validation of a real-time automated reminder system that assists General Practitioners (GP) in appropriate test ordering. We compared the comments of human experts with the comments of the reminder system using a retrospective random selection of 253 request forms. A panel of three expert physicians judged the requested tests independently based on their interpretations of the practice guidelines. The majority assessment of the physicians was compared with the assessment of the reminder system. In case the system's output differed from the majority assessment the written practice guidelines were consulted. On average 1.75 reminders were produced per form. In total 32 of the 442 given reminders (7%) were given incorrectly. The amount of information and the level of detail (the specificity of the terms) in which the GP describes the patients' medical status are crucial for the reminder system to react correctly.

Adult↗

Improving the sensitivity and dynamic range of reagentless fluorescent immunosensors by knowledge-based design.

The variable fragment (Fv) of an antibody can be transformed into a reagentless fluorescent biosensor by mutating a residue into a cysteine in the neighborhood of the paratope (antigen-binding site) and then coupling an environment-sensitive fluorophore, e.g., N-((2-(iodoacetoxy)ethyl)-N-methyl)amino-7-nitrobenz-2-oxa-1,3-diazole (IANBD ester), to the mutant cysteine. For some residues, named operational, the formation of the conjugate does not affect the affinity of the Fv fragment for the antigen, and the binding of the antigen generates a measurable variation in the fluorescence intensity of the conjugate. We tested if this signal variation could be increased by coupling several molecules of fluorophores to the same molecule of Fv. Seven operational residues have been previously identified in the single-chain Fv (scFv) of monoclonal antibody D1.3 (mAbD1.3), directed against lysozyme. Ten double mutants of scFvD1.3, involving these residues, were constructed and coupled to the IANBD ester. The fluorescence of the double conjugates revealed a transfer of resonance energy between the two identical fluorescent groups. This homotranfer could be more important in the free state of the conjugate than in its antigen-bound state and increase its sensitivity for the detection of the antigen by up to 2.9-fold. A poorly sensitive conjugate could be improved by coupling a second molecule of fluorophore to residues located far from the paratope. Mutations altering the affinity of scFvD1.3 for lysozyme were introduced into one of its fluorescent conjugates. Using a mixture of three mutant derivatives of this unique conjugate, we could titrate lysozyme with precision in a concentration range encompassing 3 orders of magnitude.

Binding Sites↗

A knowledge-based approach in designing combinatorial or medicinal chemistry libraries for drug discovery. 1. A qualitative and quantitative characterization of known drug databases.

The discovery of various protein/receptor targets from genomic research is expanding rapidly. Along with the automation of organic synthesis and biochemical screening, this is bringing a major change in the whole field of drug discovery research. In the traditional drug discovery process, the industry tests compounds in the thousands. With automated synthesis, the number of compounds to be tested could be in the millions. This two-dimensional expansion will lead to a major demand for resources, unless the chemical libraries are made wisely. The objective of this work is to provide both quantitative and qualitative characterization of known drugs which will help to generate "drug-like" libraries. In this work we analyzed the Comprehensive Medicinal Chemistry (CMC) database and seven different subsets belonging to different classes of drug molecules. These include some central nervous system active drugs and cardiovascular, cancer, inflammation, and infection disease states. A quantitative characterization based on computed physicochemical property profiles such as log P, molar refractivity, molecular weight, and number of atoms as well as a qualitative characterization based on the occurrence of functional groups and important substructures are developed here. For the CMC database, the qualifying range (covering more than 80% of the compounds) of the calculated log P is between -0.4 and 5.6, with an average value of 2.52. For molecular weight, the qualifying range is between 160 and 480, with an average value of 357. For molar refractivity, the qualifying range is between 40 and 130, with an average value of 97. For the total number of atoms, the qualifying range is between 20 and 70, with an average value of 48. Benzene is by far the most abundant substructure in this drug database, slightly more abundant than all the heterocyclic rings combined. Nonaromatic heterocyclic rings are twice as abundant as the aromatic heterocycles. Tertiary aliphatic amines, alcoholic OH and carboxamides are the most abundant functional groups in the drug database. The effective range of physicochemical properties presented here can be used in the design of drug-like combinatorial libraries as well as in developing a more efficient corporate medicinal chemistry library.

Algorithms↗

Automatic generation of knowledge base from infrared spectral database for substructure recognition

This paper presents a new methodology of chemical substructure recognition by interpretation of an infrared spectrum. The approach in spectrum interpretation is based on the determination of functional groups, which may be present or absent in compounds whose structure is unknown. The process of searching for spectrum-substructure correlation is realized by application of a statistical algorithm. In this method, correlations are generalized and condensed into a set of interpretation rules which are applied to the interpretation of an unknown compound's spectrum in order to predict whether the respective substructures are present or absent in the unknown molecule.

Journal Article↗

Simple knowledge-based descriptors to predict protein-ligand interactions. methodology and validation.

A new type of shape descriptor is proposed to describe the spatial orientation for non-covalent interactions. It is built from simple, anisotropic Gaussian contributions that are parameterised by 10 adjustable values. The descriptors have been used to fit propensity distributions derived from scatter data stored in the IsoStar database. This database holds composite pictures of possible interaction geometries between a common central group and various interacting moieties, as extracted from small-molecule crystal structures. These distributions can be related to probabilities for the occurrence of certain interaction geometries among different functional groups. A fitting procedure is described that generates the descriptors in a fully automated way. For this purpose, we apply a similarity index that is tailored to the problem, the Split Hodgkin Index. It accounts for the similarity in regions of either high or low propensity in a separate way. Although dependent on the division into these two subregions, the index is robust and performs better than the regular Hodgkin index. The reliability and coverage of the fitted descriptors was assessed using SuperStar. SuperStar usually operates on the raw IsoStar data to calculate propensity distributions, e.g., for a binding site in a protein. For our purpose we modified the code to have it operate on our descriptors instead. This resulted in a substantial reduction in calculation time (factor of five to eight) compared to the original implementation. A validation procedure was performed on a set of 130 protein-ligand complexes, using four representative interacting probes to map the properties of the various binding sites: ammonium nitrogen, alcohol oxygen, carbonyl oxygen, and methyl carbon. The predicted 'hot spots' for the binding of these probes were compared to the actual arrangement of ligand atoms in experimentally determined protein-ligand complexes. Results indicate that the version of SuperStar that applies to our descriptors is capable to predict the above-mentioned atom types in ligands correctly with success rates of 59% and 74%, respectively, for all ligand atoms (regardless of their solvent accessibility), and a subset of solvent-inaccessible ones. If not only exact atom-type matches are counted, but also those that identify ligand atoms of similar physicochemical properties, the prediction rates rise to 75% and 89%. These rates are close to those obtained by the original SuperStar method (being 67% and 82%, respectively, for the prediction of exact matching atom types, and 81% and 91% in the case of predicting similar atom types).

Hydrogen Bonding↗