Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,351 records · Page 75Linked to original sources

Web and database software for identification of intact proteins using "top down" mass spectrometry.

For the identification and characterization of proteins harboring posttranslational modifications (PTMs), a "top down" strategy using mass spectrometry has been forwarded recently but languishes without tailored software widely available. We describe a Web-based software and database suite called ProSight PTM constructed for large-scale proteome projects involving direct fragmentation of intact protein ions. Four main components of ProSight PTM are a database retrieval algorithm (Retriever), MySQL protein databases, a file/data manager, and a project tracker. Retriever performs probability-based identifications from absolute fragment ion masses, automatically compiled sequence tags, or a combination of the two, with graphical rendering and browsing of the results. The database structure allows known and putative protein forms to be searched, with prior or predicted PTM knowledge used during each search. Initial functionality is illustrated with a 36-kDa yeast protein identified from a processed cell extract after automated data acquisition using a quadrupole-FT hybrid mass spectrometer. A +142-Da delta(m) on glyceraldehyde-3-phosphate dehydrogenase was automatically localized between Asp90 and Asp192, consistent with its two cystine residues (149 and 153) alkylated by acrylamide (+71 Da each) during the gel-based sample preparation. ProSight PTM is the first search engine and Web environment for identification of intact proteins (https://prosightptm.scs.uiuc.edu/).

Amino Acid Sequence↗

Discovering known and unanticipated protein modifications using MS/MS database searching.

We present an MS/MS database search algorithm with the following novel features: (1) a novel protein database structure containing extensive preindexing and (2) zone modification searching, which enables the rapid discovery of protein modifications of known (i.e., user-specified) and unanticipated delta masses. All of these features are implemented in Interrogator, the search engine that runs behind the Pro ID, Pro ICAT, and Pro QUANT software products. Speed benchmarks demonstrate that our modification-tolerant database search algorithm is 100-fold faster than traditional database search algorithms when used for comprehensive searches for a broad variety of modification species. The ability to rapidly search for a large variety of known as well as unanticipated modifications allows a significantly greater percentage of MS/MS scans to be identified. We demonstrate this with an example in which, out of a total of 473 identified MS/MS scans, 315 of these scans correspond to unmodified peptides, while 158 scans correspond to a wide variety of modified peptides. In addition, we provide specific examples where the ability to search for unanticipated modifications allows the scientist to discover: unexpected modifications that have biological significance; amino acid mutations; salt-adducted peptides in a sample that has nominally been desalted; peptides arising from nontryptic cleavage in a sample that has nominally been digested using trypsin; other unintended consequences of sample handling procedures.

Algorithms↗

DBDigger: reorganized proteomic database identification that improves flexibility and speed.

Database search identification algorithms, such as Sequest and Mascot, constitute powerful enablers for proteomic tandem mass spectrometry. We introduce DBDigger, an algorithm that reorganizes the database identification process to remove a problematic bottleneck. Typically such algorithms determine which candidate sequences can be compared to each spectrum. Instead, DBDigger determines which spectra can be compared to each candidate sequence, enabling the software to generate candidate sequences only once for each HPLC separation rather than for each spectrum. This reorganization also reduces the number of times a spectrum must be predicted for a particular candidate sequence and charge state. As a result, DBDigger can accelerate some database searches by more than an order of magnitude. In addition, the software offers features to reduce the performance degradation introduced by posttranslational modification (PTM) searching. DBDigger allows researchers to specify the sequence context in which each PTM is possible. In the case of CNBr digests, for example, modified methionine residues can be limited to occur only at the C-termini of peptides. Use of "context-dependent" PTM searching reduces the performance penalty relative to traditional PTM searching. We characterize the performance possible with DBDigger, showcasing MASPIC, a new statistical scorer. We describe the implementation of these innovations in the hope that other researchers will employ them for rapid and highly flexible proteomic database search.

Algorithms↗

Identifying proteins using matrix-assisted laser desorption/ionization in-source fragmentation data combined with database searching.

Metastable ion decay in matrix-assisted laser desorption/ionization (MALDI) has become a routine method for obtaining primary structures of peptides. Significant fragmentation occurs in the MALDI ion source and can be observed via delayed ion extraction TOF-MS. In-source decay (ISD) can provide C- and N-terminal primary sequence data for even moderate-sized peptides (< 5000 Da). The unique cn series fragmentation that occurs in ISD has been exploited to obtain partial C-terminal sequences for proteins as large as human apotransferrin (75 kDa). Two approaches for combining this ISD MALDI-generated partial sequence information with protein database searching techniques are presented. In one approach, cyanogen bromide is used to cleave relatively large peptide fragments from a sample of human apotransferrin. One of the larger cleavage products (6034.84 Da) was isolated by HPLC and subjected to ISD MALDI analysis. An easily identified cn fragment ion series allowed two noncontiguous segments of the peptide's sequence to be determined (about 55% of the total sequence). This partial sequence information was used to search protein and oligonucleotide sequence databases. In addition to uniquely identifying human apotransferrin in a protein sequence database, an example of the use of this ISD MALDI-determined partial sequence information to search expressed sequence tag databases is presented. Such searches have the potential for rapidly identifying new genes that code for target proteins. An alternate approach for obtaining partial sequence information on proteins is also demonstrated that utilizes ISD MALDI fragmentation of the intact protein to generate partial sequence information. This approach is shown to generate about 5-7% of a protein's sequence, usually near the C-terminus of the protein. Examples of the ISD MALDI fragmentation data obtained from intact (reduced) human apotransferrin and intact (nonreduced) bovine serum albumin (66 kDa) proteins are presented.

Animals↗

Protein identification with a single accurate mass of a cysteine-containing peptide and constrained database searching.

A method for rapid and unambiguous identification of proteins by sequence database searching using the accurate mass of a single peptide and specific sequence constraints is described. Peptide masses were measured using electrospray ionization-Fourier transform ion cyclotron resonance mass spectrometry to an accuracy of 1 ppm. The presence of a cysteine residue within a peptide sequence was used as a database searching constraint to reduce the number of potential database hits. Cysteine-containing peptides were detected within a mixture of peptides by incorporating chlorine into a general alkylating reagent specific for cysteine residues. Secondary search constraints included the specificity of the protease used for protein digestion and the molecular mass of the protein estimated by gel electrophoresis. The natural isotopic distribution of chlorine encoded the cysteine-containing peptide with a distinctive isotopic pattern that allowed automatic screening of mass spectra. The method is demonstrated for a peptide standard and unknown proteins from a yeast lysate using all 6118 possible yeast open reading frames as a database. As judged by calculation of codon bias, low-abundance proteins were identified from the yeast lysate using this new method but not by traditional methods such as tandem mass spectrometry via data-dependent acquisition or mass mapping.

Amino Acid Sequence↗

A knowledge-based approach in designing combinatorial or medicinal chemistry libraries for drug discovery. 1. A qualitative and quantitative characterization of known drug databases.

The discovery of various protein/receptor targets from genomic research is expanding rapidly. Along with the automation of organic synthesis and biochemical screening, this is bringing a major change in the whole field of drug discovery research. In the traditional drug discovery process, the industry tests compounds in the thousands. With automated synthesis, the number of compounds to be tested could be in the millions. This two-dimensional expansion will lead to a major demand for resources, unless the chemical libraries are made wisely. The objective of this work is to provide both quantitative and qualitative characterization of known drugs which will help to generate "drug-like" libraries. In this work we analyzed the Comprehensive Medicinal Chemistry (CMC) database and seven different subsets belonging to different classes of drug molecules. These include some central nervous system active drugs and cardiovascular, cancer, inflammation, and infection disease states. A quantitative characterization based on computed physicochemical property profiles such as log P, molar refractivity, molecular weight, and number of atoms as well as a qualitative characterization based on the occurrence of functional groups and important substructures are developed here. For the CMC database, the qualifying range (covering more than 80% of the compounds) of the calculated log P is between -0.4 and 5.6, with an average value of 2.52. For molecular weight, the qualifying range is between 160 and 480, with an average value of 357. For molar refractivity, the qualifying range is between 40 and 130, with an average value of 97. For the total number of atoms, the qualifying range is between 20 and 70, with an average value of 48. Benzene is by far the most abundant substructure in this drug database, slightly more abundant than all the heterocyclic rings combined. Nonaromatic heterocyclic rings are twice as abundant as the aromatic heterocycles. Tertiary aliphatic amines, alcoholic OH and carboxamides are the most abundant functional groups in the drug database. The effective range of physicochemical properties presented here can be used in the design of drug-like combinatorial libraries as well as in developing a more efficient corporate medicinal chemistry library.

Algorithms↗

The PHARMSEARCH database.

PHARMSEARCH, a database produced by the French Patent and Trademark Office (INPI), covers pharmaceutical patents issued by the Europeans, French, and United States patent offices from November 1986 onward. PHARMSEARCH is composed of MPHARM, a structure file searchable using Markush DARC software, and PHARM, the companion bibliographic file. Markush structures claimed in the patent documents are entered into the database as variable generic structures. Specific structures are also included in the database, when they are not part of a Markush structure in the patent document. Chemical index terms describe all moieties of the structure. Indexing also describes the therapeutic activities and preparation processes for the compounds. The indexing policies used in the production of this database are described.

Abstracting and Indexing↗

A database of lipid phase transition temperatures and enthalpy changes.

The systematic study of the mesomorphic phase properties of synthetic and biologically derived lipids began some 30 years ago. In the past decade, interest in this area has grown enormously. As a result, there exists a wealth of information on lipid phase behavior, but unfortunately, these data have, until now, been scattered throughout the literature in a variety of books, proceedings, and journals. The data have recently been compiled in a centralized database with a view to providing ready access to the same and to the appropriate literature. The compilation facilitates review of what has thus far been accomplished and highlights what remains to be done in this active research area. As such, it represents a convenient summary of the existing data which, when evaluated, will enable us to identify where deficits exist in the data, to reveal the fundamental physicochemical principles upon which lipid phase behavior is based, and to understand more completely lipid phase relations in biological, reconstituted, and formulated systems. The compilation consists of a tabulation of all known mesomorphic and polymorphic phase transition temperatures and enthalpy changes for synthetic and biologically derived lipids in the dry and in the partially and fully hydrated states. Also included is the effect on these thermodynamic values of pH, and of salt and metal ion concentration and other additives such as proteins, drugs, etc. The methods used in making the measurements and the experimental conditions are reported. Bibliographic information includes complete literature referencing and list of authors. As of this writing, the database is current through June 1990 and contains 9500 records. Each record contains 28 fields. Here, we describe how the database originated, its scope and contents, data abstraction procedures, and issues relating to mesophase and lipid nomenclature, data analysis, and evaluation, and database maintenance and distribution.

Databases, Factual↗

National Cancer Institute Drug Information System 3D database.

A searcheable database of three-dimensional structures has been developed from the chemistry database of the NCI Drug Information System (DIS), a file of about 450,000 primarily organic compounds which have been tested by NCI for anticancer activity. The DIS database is very similar in size and content to the proprietary databases used in the pharmaceutical industry; its development began in the 1950s; and this history led to a number of problems in the generation of 3D structures.

Antineoplastic Agents↗

A web-based 3D-database pharmacophore searching tool for drug discovery.

Three-Dimensional (3D) structural database pharmacophore searching has become a very effective approach for discovery of novel lead compounds in drug discovery. Although several commercial programs are available, these commercial programs are primarily used as a stand alone and require a local database. In recent years, the Internet has become the main medium of choice for multiuser application program distribution. Herein, we describe our development of a Web-based 3D-database pharmacophore-searching tool based on the server-client Web architecture. Both rigid and conformationally flexible searching methods are implemented. Our results show that for a typical three-center rigid pharmacophore search, the run time for searching 50 000 compounds is less than three minutes, and for four-center pharmacophore searching, the run time is less than 10 minutes on a desktop computer. For a flexible 3D-pharmacophore search, the run time for searching 50 000 compounds generally takes between one and several hours. The search results are comparable to those obtained using a commercial program. We expect that this Web-based tool will be very useful for scientists who are interested in 3D-database pharmacophore searching via the Internet.

Database Management Systems↗

Recursive median partitioning for virtual screening of large databases.

Recently, we have introduced the median partitioning (MP) method for diversity selection and compound classification. The MP approach utilizes property descriptors with continuous value ranges, transforms these descriptors into a binary classification scheme by determining their medians in source databases, and divides database molecules in subsequent steps into populations above or below these medians. Having previously demonstrated the usefulness of MP for the classification of molecules according to biological activity, we have now gone a step further and extended the methodology for application in virtual screening. In these calculations, a series of bait molecules having desired activity is added to large compound databases, and subsequent iterations or recursions are carried out to reduce the number of candidate molecules until a small number of compounds are found in partitions enriched with bait molecules. For each recursion step, descriptor combinations are identified that copartition as many active molecules as possible. Descriptor selection is facilitated by application of a genetic algorithm (GA). The recursive MP approach (RMP) has been applied to five diverse biological activity classes in virtual screening of a database consisting of approximately 1.34 million molecules to which different types of active compounds were added. RMP analysis produced hit rates of up to 21%, dependent on the biological activity class, and led to an average approximately 3600-fold improvement over random selection for the activity classes that were used as test cases.

Algorithms↗

ZINC--a free database of commercially available compounds for virtual screening.

A critical barrier to entry into structure-based virtual screening is the lack of a suitable, easy to access database of purchasable compounds. We have therefore prepared a library of 727,842 molecules, each with 3D structure, using catalogs of compounds from vendors (the size of this library continues to grow). The molecules have been assigned biologically relevant protonation states and are annotated with properties such as molecular weight, calculated LogP, and number of rotatable bonds. Each molecule in the library contains vendor and purchasing information and is ready for docking using a number of popular docking programs. Within certain limits, the molecules are prepared in multiple protonation states and multiple tautomeric forms. In one format, multiple conformations are available for the molecules. This database is available for free download (http://zinc.docking.org) in several common file formats including SMILES, mol2, 3D SDF, and DOCK flexibase format. A Web-based query tool incorporating a molecular drawing interface enables the database to be searched and browsed and subsets to be created. Users can process their own molecules by uploading them to a server. Our hope is that this database will bring virtual screening libraries to a wide community of structural biologists and medicinal chemists.

Databases, Factual↗

A searching and reporting system for relational databases using a graph-based metadata representation.

Relational databases are the current standard for storing and retrieving data in the pharmaceutical and biotech industries. However, retrieving data from a relational database requires specialized knowledge of the database schema and of the SQL query language. At Anadys, we have developed an easy-to-use system for searching and reporting data in a relational database to support our drug discovery project teams. This system is fast and flexible and allows users to access all data without having to write SQL queries. This paper presents the hierarchical, graph-based metadata representation and SQL-construction methods that, together, are the basis of this system's capabilities.

Computer Simulation↗

Prediction of biological targets for compounds using multiple-category Bayesian models trained on chemogenomics databases.

Target identification is a critical step following the discovery of small molecules that elicit a biological phenotype. The present work seeks to provide an in silico correlate of experimental target fishing technologies in order to rapidly fish out potential targets for compounds on the basis of chemical structure alone. A multiple-category Laplacian-modified naïve Bayesian model was trained on extended-connectivity fingerprints of compounds from 964 target classes in the WOMBAT (World Of Molecular BioAcTivity) chemogenomics database. The model was employed to predict the top three most likely protein targets for all MDDR (MDL Drug Database Report) database compounds. On average, the correct target was found 77% of the time for compounds from 10 MDDR activity classes with known targets. For MDDR compounds annotated with only therapeutic or generic activities such as "antineoplastic", "kinase inhibitor", or "anti-inflammatory", the model was able to systematically deconvolute the generic activities to specific targets associated with the therapeutic effect. Examples of successful deconvolution are given, demonstrating the usefulness of the tool for improving knowledge in chemogenomics databases and for predicting new targets for orphan compounds.

Bayes Theorem↗

StructSorter: a method for continuously updating a comprehensive protein structure alignment database.

Advances in protein crystallography and homology modeling techniques are producing vast amounts of high resolution protein structure data at ever increasing rates. As such, the ability to quickly and easily extract structural similarities is a key tool in discovering important functional relationships. We report on an approach for creating and maintaining a database of pairwise structure alignments for a comprehensive database comprising the PDB and homology models for the human and select pathogen genomes. Our approach consists of a novel, multistage method for determining pairwise structural similarity coupled with an efficient clustering protocol that approximates a full NxN assessment in a fraction of the time. Since biologists are commonly interested in recently released structures, and the homology models built from them, an automatically updating database of structural alignments has great value. Our approach yields a querying system that allows scientists to retrieve databank-wide protein structure similarities as easily as retrieving protein sequence similarities via BLAST or PSI-BLAST. Basic, noncommercial access to the database can be requested at https://tip.eidogen-sertanty.com/.

Databases, Protein↗

Mining the chemical quarry with joint chemical probes: an application of latent semantic structure indexing (LaSSI) and TOPOSIM (Dice) to chemical database mining.

In this study we use a novel similarity search technique called latent semantic structure indexing (LaSSI) with joint chemical probes as queries to mine the MDL drug data report database. LaSSI is based on latent semantic indexing developed for searching textual databases. We use atom pair and topological torsion descriptors in our calculations. The results obtained with LaSSI are compared with another in-house similarity search technique TOPOSIM. The results from the similarity searches using joint chemical probes are significantly better than searches using single chemical probes for both LaSSI and TOPOSIM. The selected molecules are closely related in activity to their queries and are ranked among the top 300 scoring molecules of the 82 860 entries in the database. Our implementation of LaSSI is very fast and efficient in finding active compounds. The results also show that LaSSI consistently retrieves more diverse chemical structures representative of the joint chemical probes in comparison to TOPOSIM. The use of multimolecule topological probes to identify compounds complements the use of searching databases with 3D pharmacophore hypotheses.

Angiotensin-Converting Enzyme Inhibitors↗

Protein-based virtual screening of chemical databases. 1. Evaluation of different docking/scoring combinations.

Three different database docking programs (Dock, FlexX, Gold) have been used in combination with seven scoring functions (Chemscore, Dock, FlexX, Fresno, Gold, Pmf, Score) to assess the accuracy of virtual screening methods against two protein targets (thymidine kinase, estrogen receptor) of known three-dimensional structure. For both targets, it was generally possible to discriminate about 7 out of 10 true hits from a random database of 990 ligands. The use of consensus lists common to two or three scoring functions clearly enhances hit rates among the top 5% scorers from 10% (single scoring) to 25-40% (double scoring) and up to 65-70% (triple scoring). However, in all tested cases, no clear relationships could be found between docking and ranking accuracies. Moreover, predicting the absolute binding free energy of true hits was not possible whatever docking accuracy was achieved and scoring function used. As the best docking/consensus scoring combination varies with the selected target and the physicochemistry of target-ligand interactions, we propose a two-step protocol for screening large databases: (i) screening of a reduced dataset containing a few known ligands for deriving the optimal docking/consensus scoring scheme, (ii) applying the latter parameters to the screening of the entire database.

Algorithms↗

Docking and database screening reveal new classes of Plasmodium falciparum dihydrofolate reductase inhibitors.

Plasmodium falciparum dihydrofolate reductase (PfDHFR) is an important target for antimalarial chemotherapy. Unfortunately, the emergence of resistant parasites has significantly reduced the efficiency of classical antifolate drugs such as cycloguanil and pyrimethamine. In this study, an approach toward molecular docking of the structures contained in the Available Chemicals Directory (ACD) database to search for novel inhibitors of PfDHFR is described. Instead of docking the whole ACD database, specific 3D pharmacophores were used to reduce the number of molecules in the database by excluding a priori molecules lacking essential requisites for the interaction with the enzyme and potentially unable to bind to resistant mutant PfDHFRs. The molecules in the resulting "focused" database were then evaluated with regard to their fit into the PfDHFR active site. Twelve new compounds whose structures are completely unrelated to known antifolates were identified and found to inhibit, at the micromolar level, the wild-type and resistant mutant PfDHFRs harboring A16V, S108T, A16V + S108T, C59R + S108N + I164L, and N51I + C59R + S108N + I164L mutations. Depending on the functional groups interacting with key active site residues of the enzyme, these inhibitors were classified as N-hydroxyamidine, hydrazine, urea, and thiourea derivatives. The structures of the complexes of the most active inhibitors, as refined by molecular mechanics and molecular dynamics, provided insight into how these inhibitors bind to the enzyme and suggested prospects for these novel derivatives as potential leads for antimalarial development.

Animals↗