Search PubMedSearch

SEARCH · Search PubMed

Results for “database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

UTAB: a computer database on residues of xenobiotic organic chemicals and heavy metals in plants.

The UTAB Database contains information concerned with the uptake/accumulation, translocation, adhesion, and biotransformation of both xenobiotic organic chemicals and heavy metals by vascular plants. UTAB can be used to estimate the accumulation of chemicals in vegetation and their subsequent movement through the food chain. The database contains actual data from papers in the published literature dating from 1926 for organic chemicals and from 1976 for heavy metals. At present the database is comprised of more than 37,000 records pertaining to 900 different organic chemicals, 21 heavy metals, and over 350 plant species. Each record contains information on a single combination of species, chemical, and dose. Other information includes the application and destination sites, amount accumulated, rates of uptake or translocation, products and sites of biotransformation, experimental condition parameters, and the source paper. Thus, the database can be used to quickly obtain specific data pertaining to a chemical, plant species, mine spoil, etc. or it can be used for the comparative analysis of a set of data pertaining to groups of chemicals and plants.

Databases, Bibliographic

GBRAP: A Comprehensive Database and Tool for Exploring Genomic Diversity Across All Domains of Life.

Evolutionary studies require extensive examination of genomic information across all domains of life. Despite the availability of a large number of genomes through GenBank, the effective visualization or comparison of the information they contain is challenging due to many reasons, including their size. We introduce genome-based retrieval and analysis parser, a comprehensive software tool to analyze genome files, and an online database housing an extensive collection of carefully curated, high-quality genome statistics for all the organisms available in the RefSeq database of National Center for Biotechnology Information. Users can either directly search, or select from precategorized groups, the organisms of their choice and retrieve data, and the output is generated as tables containing more than 200 columns of useful genomic information (base counts, GC content, Shannon entropy, codon usage, etc.) separately calculated for different genomic elements (e.g. coding sequences, introns, transfer RNA, ribosomal RNA, noncoding RNA, etc.). The data are independently displayed (if applicable) for each chromosomal, mitochondrial, plastid, or plasmid sequence. All the data can be visualized on the database or downloaded as comma-separated value or Excel files. The genome-based retrieval and analysis parser database is free to access without any registration and is publicly available at http://tacclab.org/gbrap/.

Software

PlantPan: A comprehensive multi-species plant pan-genome database.

The pan-genome represents the complete genomic diversity of specific species, serving as a valuable resource for studying species evolution, crop domestication, and guiding crop breeding and improvement. While there are several single-species-specific plant pan-genome databases, the availability of multi-species pan-genome databases is limited. Additionally, variations in methods and data types used for plant pan-genome analysis across different databases hinder the comparison and integration of pan-genome information from various projects at multi-species or single-species levels. To tackle this challenge, we introduce PlantPan, a comprehensive database housing the results of pan-genome analysis for 195 genomes from 11 plant species. PlantPan aims to provide extensive information, including gene-centric and sequence-centric pan-genome information, graph-based pan-genome, pan-genome openness profiles, gene functions and its variation characteristics, homologous genes, and gene clusters across different species. Statistically, PlantPan incorporates 9 163 011 genes, 694 191 gene clusters, 526 973 370 genome variations, and 1 616 089 non-redundant genome variation groups at the species level, 33 455,098 genome synteny, and 177 827 non-redundant genome synteny groups at the species level. Regarding functional genes, PlantPan contains 5 222 720 genes related to transcription factors, 395 247 literature-reported resistance genes, 455 748 predicted microbial/disease resistance genes, and 1 612 112 genes related to molecular pathways. In summary, PlantPan is a vital platform for advancing the application of pan-genomes in molecular breeding for crops and evolutionary research for plants.

Genome, Plant

Structure-activity relations: maximizing the usefulness of mutagenicity and carcinogenicity databases.

The most important criteria for the development and analysis of databases for elucidating the structural bases of toxicological activity include the integrity of the databases with respect to uniformity of the experimental protocol and interpretation of the test results and inclusion of chemicals representing different chemical classes and differing mechanisms of action. Within these criteria, it is demonstrated that when the chemicals are chosen at random, the larger the database, the better the predictivity of chemicals not included in the learning set. It is shown however, that when chemicals are selected on the basis of structural features, that a learning set of approximately 180 chemicals is as informative as a database consisting of 800 chemicals chosen at random.

Animals

Methanog: a specialized database on methanogenic bacteria.

A specialized, interdisciplinary database on various types of related information on methanogenic bacteria is described. Derived from other sequence databases etc., this database collects information from many sources, including unpublished work from research laboratories working in this field, and makes them accessible from a single source, to interested scientists, free of cost. It is presently held in eight 48 T.P.I. floppy disks and can be run on any IBM PC under DOS 3.0 or above, making this database of particular interest to researchers with limited resources and on-line search/access facilities.

Amino Acid Sequence

The MRC-5 human embryonal lung fibroblast two-dimensional gel cellular protein database: quantitative identification of polypeptides whose relative abundance differs between quiescent, proliferating and SV40 transformed cells.

A new version of the MRC-5 two-dimensional gel cellular protein database (Celis et al., Electrophoresis 1989, 10, 76-115) is presented. Gels were scanned with a Molecular Dynamics laser scanner and processed by the PDQUEST II software. A total of 1895 [35S]methionine-labeled cellular polypeptides (1323 with isoelectric focusing and 572 with nonequilibrium pH gradient electrophoresis) are recorded in this database, containing quantitative and qualitative data on the relative abundance of cellular proteins synthesized by quiescent, proliferating and SV40 transformed MRC-5 fibroblasts. Of the 592 proteins quantitated so far, the levels of 138 were up- or down-regulated (51 and 87, respectively) by two times or more in the transformed cells as compared to their normal proliferating counterparts, while only 14 behaved similarly in quiescent cells. Seven MRC-5 SV40 proteins, including plastin and two interferon-induced proteins, were not detected in the master MRC-5 images. The identity of 36 of the transformation-sensitive proteins whose levels are up or down regulated by two times or more was determined and additional information can be transferred from the master transformed human epithelial amnion cells (AMA) database (Celis et al., Electrophoresis 1990, 11, 989-1071) for those polypeptides of known and unknown identity that have been matched to AMA polypeptides. As more information is gathered in this and other laboratories, including data on oncogene proteins and transcription factors, this comprehensive database will outline an integrated picture of the expression levels and properties of the thousands of protein components of organelles, pathways and cytoskeletal systems that may be directly or indirectly involved in properties associated with the transformed state.

Cell Transformation, Viral

A database of protein structure families with common folding motifs.

The availability of fast and robust algorithms for protein structure comparison provides an opportunity to produce a database of three-dimensional comparisons, called families of structurally similar proteins (FSSP). The database currently contains an extended structural family for each of 154 representative (below 30% sequence identity) protein chains. Each data set contains: the search structure; all its relatives with 70-30% sequence identity, aligned structurally; and all other proteins from the representative set that contain substructures significantly similar to the search structure. Very close relatives (above 70% sequence identity) rarely have significant structural differences and are excluded. The alignments of remote relatives are the result of pairwise all-against-all structural comparisons in the set of 154 representative protein chains. The comparisons were carried out with each of three novel automatic algorithms that cover different aspects of protein structure similarity. The user of the database has the choice between strict rigid-body comparisons and comparisons that take into account interdomain motion or geometrical distortions; and, between comparisons that require strictly sequential ordering of segments and comparisons, which allow altered topology of loop connections or chain reversals. The data sets report the structurally equivalent residues in the form of a multiple alignment and as a list of matching fragments to facilitate inspection by three-dimensional graphics. If substructures are ignored, the result is a database of structure alignments of full-length proteins, including those in the twilight zone of sequence similarity.(ABSTRACT TRUNCATED AT 250 WORDS)

Algorithms

A relational database for sequence-specific protein NMR data.

A protein NMR database has been designed and is being implemented. The database is intended to contain solution NMR results from proteins and peptides (larger than 12 residues). A relational database format has been chosen that indexes data by: primary journal citation, molecular species, sequence-related and atom-specific assignments, and experimental conditions. At present, all data are entered from the primary refereed literature. Examples are given of sample queries to the database. Possible distribution formats are discussed.

Animals

A database of lipid phase transition temperatures and enthalpy changes.

The systematic study of the mesomorphic phase properties of synthetic and biologically derived lipids began some 30 years ago. In the past decade, interest in this area has grown enormously. As a result, there exists a wealth of information on lipid phase behavior, but unfortunately these data have, until now, been scattered throughout the literature in a variety of books, proceedings and journals. The data have recently been compiled in a centralized database with a view to providing ready access to same and to the appropriate literature. The compilation facilitates review of what has thus far been accomplished and highlights what remains to be done in this active research area. As such, it represents a convenient summary of the existing data which, when evaluated, will enable us to identify where deficits exist in the data, to reveal the fundamental physicochemical principles upon which lipid phase behavior is based and to understand more completely lipid phase relations in biological, reconstituted and formulated systems. The compilation consists of a tabulation of all known mesomorphic and polymorphic phase transition temperatures and enthalpy changes for synthetic and biologically-derived lipids in the dry and in the partially and fully hydrated states. Also included is the effect on these thermodynamic values of pH, and of salt and metal ion concentration and other additives such as proteins, drugs, etc. The methods used in making the measurements and the experimental conditions are reported. Bibliographic information includes complete literature referencing and list of authors. As of this writing, the database is current through June, 1990 and contains in excess of 9500 records. Each record contains 28 fields. Here, we describe how the database originated, its scope and contents, data abstraction procedures, and issues relating to mesophase and lipid nomenclature, data analysis and evaluation, and database maintenance and distribution.

Databases, Factual

The PDQ (Physician Data Query), the cancer database, in oncological clinical practice.

The above illustrates the fact that a physician interested in consulting the PDQ database must dedicate a certain amount of time to an analytical review of the database. It is difficult to determine how much time is required to acquire a sufficient level of control because there are many variables affecting the learning time: experience in using computerized systems, cultural background, personal inclination, etc. However, a certain amount of caution and humility should be exercised whenever a physician approaches a database of this type for the first time, in order to avoid the mistake of dangerously underestimating the nature of the problem. On the other hand, the physician's specific competence and professionalism will not be questioned at all, since they are fundamental to obtain productive search results. If, indeed, the above discussion focussed heavily on the most closely documental aspect of the problem, it should not be forgotten that the contents of the database can be fully understood only by experts who are used to encountering certain terms and procedures on a daily basis. In fact, when a physician turns to a documentation center for a PDQ research, the physician's assistance is always requested in order pair clinical and documental competence. It is this second skill that the physician must acquire to become totally independent.

Databases, Factual

An adjuvant database for preclinical evaluation of vaccines and immunotherapeutics.

Adjuvants are immunostimulators used to enhance vaccine efficacy against infectious diseases. However, current methods for evaluating their efficacy and safety are limited, hindering large-scale screening. To address this, we developed a prototype Adjuvant Database (ADB) containing transcriptome data, generated using the same protocols as the widely used Open TG-GATEs (OTG) toxicogenomics database, covering 25 adjuvants across multiple species, organs, time points, and doses. This enabled cross-database integration of ADB and OTG. Transcriptomic patterns successfully distinguished each adjuvant regardless of organs or species. Using both databases, we built machine learning models to predict adjuvanticity and hepatotoxicity. Notably, we identified colchicine's adjuvant activity and FK565's liver toxicity through data-driven analysis. Overall, ADB combined with OTG offers a framework for transcriptomics-based, data-driven screening of adjuvant candidates.

Animals

Unique signatures of highly constrained genes across publicly available genomic databases.

PURPOSE: Publicly available genomic databases are critical in understanding human genetic variation. They also provide unique insights into patterns of genetic constraints and their relationship with human disease. METHODS: We utilized one of the largest publicly available databases, Genome Aggregate Database, to determine genes that are highly constrained for only loss-of-function, only missense, and both loss-of-function/missense variants. We identified their unique signatures and explored their causal relationship with human diseases. Those genes were also evaluated for chromosomal location, tissue-level expression, Gene Ontology analysis, and gene family categorization using multiple publicly available databases. RESULTS: We identified unique patterns of inheritance, protein size, and enrichment in distinct molecular pathways for those constrained genes associated with human disease. In addition, we identified genes that are currently not known to cause human disease, which may be excellent gene discovery candidates. CONCLUSION: We elucidate biological pathways of highly constrained genes that expand our understanding of critical cellular proteins. The findings can also advance research in rare diseases.

Humans

The Human Communication Research Centre dialogue database.

The HCRC dialogue database consists of over 700 transcribed and coded dialogues from pairs of speakers aged from seven to fourteen. The speakers are recorded while tackling co-operative problem-solving tasks and the same pairs of speakers are recorded over two years tackling 10 different versions of our two tasks. In addition there are over 200 dialogues recorded between pairs of undergraduate speakers engaged on versions of the same tasks. Access to the database, and to its accompanying custom-built search software, is available electronically over the JANET system by contacting liz@psy.glasgow.ac.uk, from whom further information about the database and a user's guide to the database can be obtained.

Adolescent

The PHARMSEARCH database.

PHARMSEARCH, a database produced by the French Patent and Trademark Office (INPI), covers pharmaceutical patents issued by the Europeans, French, and United States patent offices from November 1986 onward. PHARMSEARCH is composed of MPHARM, a structure file searchable using Markush DARC software, and PHARM, the companion bibliographic file. Markush structures claimed in the patent documents are entered into the database as variable generic structures. Specific structures are also included in the database, when they are not part of a Markush structure in the patent document. Chemical index terms describe all moieties of the structure. Indexing also describes the therapeutic activities and preparation processes for the compounds. The indexing policies used in the production of this database are described.

Abstracting and Indexing

A database of lipid phase transition temperatures and enthalpy changes.

The systematic study of the mesomorphic phase properties of synthetic and biologically derived lipids began some 30 years ago. In the past decade, interest in this area has grown enormously. As a result, there exists a wealth of information on lipid phase behavior, but unfortunately, these data have, until now, been scattered throughout the literature in a variety of books, proceedings, and journals. The data have recently been compiled in a centralized database with a view to providing ready access to the same and to the appropriate literature. The compilation facilitates review of what has thus far been accomplished and highlights what remains to be done in this active research area. As such, it represents a convenient summary of the existing data which, when evaluated, will enable us to identify where deficits exist in the data, to reveal the fundamental physicochemical principles upon which lipid phase behavior is based, and to understand more completely lipid phase relations in biological, reconstituted, and formulated systems. The compilation consists of a tabulation of all known mesomorphic and polymorphic phase transition temperatures and enthalpy changes for synthetic and biologically derived lipids in the dry and in the partially and fully hydrated states. Also included is the effect on these thermodynamic values of pH, and of salt and metal ion concentration and other additives such as proteins, drugs, etc. The methods used in making the measurements and the experimental conditions are reported. Bibliographic information includes complete literature referencing and list of authors. As of this writing, the database is current through June 1990 and contains 9500 records. Each record contains 28 fields. Here, we describe how the database originated, its scope and contents, data abstraction procedures, and issues relating to mesophase and lipid nomenclature, data analysis, and evaluation, and database maintenance and distribution.

Databases, Factual

Role of exposure databases in epidemiology.

At present, exposure databases record data primarily for regulatory purposes; they have not focused on serving the needs of epidemiologists or public health. However, the modification of exposure databases could facilitate their use in epidemiology. Characteristics necessary to enhance the use of all databases include easy access by users; documentation of methods, sampling bias, error, and inconsistences; widespread coverage in time and space; and methods and measures for estimating exposure of individuals as well as populations. Also needed are exposure scenarios and models to estimate exposures for geographic areas and time intervals not currently sampled. Multidisciplinary teams are needed to examine current databases, to review strategies for improving data collection, and to suggest and help implement appropriate changes. A long-term goal is to develop and validate data from exposure scenarios and models using data on the relationship of exposure to doses measured in humans.

Databases, Factual

The European ST-T database: standard for evaluating systems for the analysis of ST-T changes in ambulatory electrocardiography.

The project for the development of the European ST-T annotated Database originated from a 'Concerted Action' on Ambulatory Monitoring, set up by the European Community in 1985. The goal was to prototype an ECG database for assessing the quality of ambulatory ECG monitoring (AECG) systems. After the 'concerted action', the development of the full database was coordinated by the Institute of Clinical Physiology of the National Research Council (CNR) in Pisa and the Thoraxcenter of Erasmus University in Rotterdam. Thirteen research groups from eight countries provided AECG tapes and annotated beat by beat the selected 2-channel records, each 2 h in duration. ST segment (ST) and T-wave (T) changes were identified and their onset, offset and peak beats annotated in addition to QRSs, beat types, rhythm and signal quality changes. In 1989, the European Society of Cardiology sponsored the remainder of the project. Recently the 90 records were completed and stored on CD-ROM. The records include 372 ST and 423 T changes. In cooperation with the Biomedical Engineering Centre of MIT (developers of the MIT-BIH arrhythmia database), the annotation scheme was revised to be consistent with both MIT-BIH and American Heart Association formats.

Algorithms

Similarity graphing and enzyme-reaction database: methods to detect sequence regions of importance for recognition of chemical structures.

We developed a new method which searches sequence segments responsible for the recognition of a given chemical structure. These segments are detected as those locally conserved among a sequence to be analyzed (target sequence) and a set of sequences (reference sequences). Reference sequences are the sequences of functionally related proteins, ligands of which contain a common chemical substructure in their molecular structures. 'Similarity graphing' cuts target sequences into segments, aligns them with reference sequence pairwise, calculates the degree of similarity for each alignment, and shows graphically cumulative similarity values on target sequence. Any locally conserved regions, short or long in length and weak or strong in similarity, are detected at their optimal conditions by adjusting three parameters. The 'enzyme-reaction database' contains chemical structures and their related enzymes. When a chemical substructure is input into the database, sequences of the enzymes related to the input substructure are systematically searched from the NBRF sequence database and output as reference sequences. Examples of analysis using similarity graphing in combination with the enzyme-reaction database showed a great potentiality in the systematic analysis of the relationships between sequences and molecular recognitions for protein engineering.

Algorithms