Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Accelerating approximate subsequence search on large protein sequence databases.

Bioinformatics has become an active research area in recent years. The amount of mapped sequences doubles every fourteen months. BLAST has been widely employed for retrieving sequences which has similar portion(s) to a given sequence. However, BLAST has to scan the entire database every time when a query is issued. This can be very time consuming especially when the database is large. In this paper, we study the problem on how to build a persistent index structure for protein sequences to support approximate match. The suffix tree has been proposed as a solution to index sequence database and has been deployed on organizing DNA sequences (Hunt et al. 2001). Unfortunately, it suffers from the problem of "memory bottleneck" that prevents it from being applied efficiently to a large database. The performance even degrades further for protein database due to a larger fanout at each node. Here, we employ an indexing structure, called BASS-tree, to support approximate match in sublinear time on a large protein database. We call this indexing method as sequence approximate match (SAM) index method. The search of approximate matches can be properly directed to the portion in the database with a high potential of matching quickly. It has been demonstrated in our experiments that the potential performance improvement is in an order of magnitude over alternative methods such as the BLAST algorithm and the suffix tree.

Algorithms↗

Guidelines and recommendations for content, structure, and deployment of mutation databases.

These Guidelines recognize the need for annotated online mutation databases documenting allelic variation (both pathogenic and phenotype modifying, and also neutral polymorphic); the databases will be both generalized (genomic) and specialized (locus specific), and a seamless integration of the two types is intended. Each requires a Document (its "biography"). Different mutation databases will have different content and structure, but a minimum core of content in a shared syntax is a necessity; the core includes: (1) a unique identifier of the allele; (2) the source/report of the data; (3) context of the allele; and (4) the allele itself (the description). The allele description should be validated. There is no single correct way to design a mutation database. The uses to which databases are put dictate the design. Software and deployment together recognize the different needs of specialized and generalized databases, while making them mutually compatible through shared content and the appropriate search facilities. A set of eight Recommendations completes these Guidelines for Content, Design, and Deployment of Mutation Databases.

Alleles↗

Quantitative exploration of the REF52 protein database: cluster analysis reveals the major protein expression profiles in responses to growth regulation, serum stimulation, and viral transformation.

Quantitative protein databases reveal the response of cells to experimental variables, such as exposure to growth factors or transfection with a transforming gene. The nature of the response depends on the type of cell and its internal state at the time of the stimulus. By constructing a protein database to study a given cell line, we can better understand the differentiated state of the cell, the growth regulatory mechanisms it employs, the particular mechanisms it uses to cope with its environment, and the ways these mechanisms may have been compromised through mutation or transformation. The REF52 database is a quantitative database designed to study growth control and transformation in a well-defined family of normal and transformed rat cell lines. The database, which has been described and analyzed elsewhere (J. I. Garrels and B. R. Franza, J. Biol. Chem. 1989, 264, 5283-5298 and J. I. Garrels and B. R. Franza, J. Biol. Chem. 1989, 264, 5299-5312) is further explored here using cluster analysis. This method reveals the most common protein expression profiles for each series of two-dimensional gels without requiring any prior hypothesis or queries on the part of the investigator. This study reveals, for each experiment, large and small clusters of protein expression profiles, most of which have readily apparent biological meaning. For example, large clusters of proteins induced or repressed during growth to confluence have been revealed, and several clusters of transformation-sensitive proteins reveal differential effects of transformation by DNA- and RNA-tumor viruses. This analysis extends our earlier quantitative explorations of the REF52 protein database and helps to show how such a database can be used to provide context and guidance for molecular studies of regulation in a given cell system.

Algorithms↗

Hepatitis C databases, principles and utility to researchers.

Part of the effort to develop hepatitis C-specific drugs a nd vaccines is the study of genetic variability of allpublicly available HCV sequences. Three HCV databases are currently available to aid this effort and to provide additional insight into the basic biology, immunology, and evolution of the virus. The Japanese HCV database (http://s2as02.genes.nig.ac.jp) gives access to a genomic mapping of sequences as well as their phylogenetic relationships. The European HCV database (http://euhcvdb.ibcp.fr) offers access to a computer-annotated set of sequences and molecular models of HCV proteins and focuses on protein sequence, structure and function analysis. The HCV database at the Los Alamos National Laboratory in the United States (http://hcv.lanl.gov) provides access to a manually annotated sequence database and a database of immunological epitopes which contains concise descriptions of experimental results. In this paper, we briefly describe each of these databases and their associated websites and tools, and give some examples of their use in furthering HCV research.

Biomedical Research↗

The IARC TP53 database: new online mutation analysis and recommendations to users.

Mutations in the tumor suppressor gene TP53 are frequent in most human cancers. Comparison of the mutation patterns in different cancers may reveal clues on the natural history of the disease. Over the past 10 years, several databases of TP53 mutations have been developed. The most extensive of these databases is maintained and developed at the International Agency for Research on Cancer. The database compiles all mutations (somatic and inherited), as well as polymorphisms, that have been reported in the published literature since 1989. The IARC TP53 mutation dataset is the largest dataset available on the variations of any human gene. The database is available at www.iarc.fr/P53/. In this paper, we describe recent developments of the database. These developments include restructuring of the database, which is now patient-centered, with more detailed annotations on the patient (carcinogen exposure, virus infection, genetic background). In addition, a new on-line application to retrieve somatic mutation data and analyze mutation patterns is now available. We also discuss limitations on the use of the database and provide recommendations to users.

Computational Biology↗

SNP databases and pharmacogenetics: great start, but a long way to go.

With the recent publication of the human genome project there has been an explosion of data available for pharmacogenetic research. Web-based databases containing information on single nucleotide polymorphisms (SNPs) are readily accessible to researchers, but there has been little comment on their utility. We used seven major international databases to identify SNPs in 74 genes involved in drug pathways. Very little overlap was seen among the databases, with only eight out of a putative 893 SNPs ( approximately 1%) common to the most commonly used databases. Problems with false positives, secondary to a high degree of homology in gene families, were also observed. These studies suggest researchers limiting their studies to one database would miss a great deal of information. Effort to update compilation databases, such as HGVbase, GeneSNP, PharmGKB, and HOWDY, and the aggressive removal of false positives from all databases is required if these resources are to facilitate the intended growth in pharmacogenetics research.

Amino Acid Sequence↗

The human FOXL2 mutation database.

Blepharophimosis-ptosis-epicanthus inversus syndrome (BPES; MIM# 110100) is an autosomal dominant genetic condition in which an eyelid malformation is associated (type I) or not associated (type II) with premature ovarian failure (POF). In 2001, mutations in the FOXL2 gene, encoding a forkhead transcription factor, were shown to cause both BPES type I and II. Since then, a number of reports have appeared that describe intragenic FOXL2 mutations in BPES patients. In addition, a few FOXL2 variants have been reported in isolated POF patients and XX males. Previously, our group has described a large number of FOXL2 mutations, thereby demonstrating the existence of two mutational hotspots in FOXL2, intra- and interfamilial phenotypic variability in BPES families, and genotype-phenotype correlations for a number of mutations in BPES patients. Here we describe a locus-specific Human FOXL2 Mutation Database (http://medgen.ugent.be/foxl2/), created using the MuStaR software. Our database contains general information about the FOXL2 gene, as well as details about 135 intragenic mutations and variants of FOXL2, obtained from published papers, abstracts of meetings, and from unpublished data produced by our group. Not included in the current version of the database are variants residing outside the coding region of FOXL2 and molecular cytogenetic rearrangements of the FOXL2 locus. The Human FOXL2 Mutation Database was created to provide a unique publicly available online resource of information about human FOXL2 mutations/variants associated with BPES and POF. It allows remote users to submit new mutations to the database and to query the database using a web form. It will facilitate evaluation of the pathogenicity of a particular mutation, as it contains data about disease-causing mutations and polymorphisms in BPES and isolated POF patients, and a link to the known FOXL2 orthologs. Moreover, it will allow us to establish more accurate genotype-phenotype correlations, since clinical information is contained in the database.

Alleles↗

The IMGT/HLA and IPD databases.

The IMGT/HLA database (www.ebi.ac.uk/imgt/hla) has provided a centralized repository for the sequences of the alleles named by the WHO Nomenclature Committee for Factors of the HLA System since 1998. Since its initial release, the database has rapidly grown in size and is recognized as the primary source of information for the study of sequences of the human major histocompatibility complex. The Immuno Polymorphism Database (IPD; www.ebi.ac.uk/ipd) is a set of specialist databases related to the study of polymorphic genes in the immune system. The IPD currently consists of four databases: IPD-KIR contains the allelic sequences of killer-cell immunoglobulin-like receptors; IPD-MHC is a database of sequences of the major histocompatibility complex of different species; IPD-HPA contains alloantigens expressed only on platelets (human platelet antigens or HPA); and IPD-ESTDAB provides access to the European Searchable Tumour Cell-Line Database, a cell bank of immunologically characterized melanoma cell lines.

Alleles↗

Allergen sequence databases.

A number of specialized databases have been developed to facilitate studies of human allergens. These include molecular databases focused on protein sequences and structures, informational databases focused on clinical, biochemical and epidemiological data related to protein allergens, a database on allergen nomenclature, and other knowledge bases or informational websites that are peripherally-related to research on allergens. Examples of each type of databases are listed and described briefly in this review. Database construction and maintenance and their impact on database quality and usefulness are also discussed.

Allergens↗

The ConSurf-HSSP database: the mapping of evolutionary conservation among homologs onto PDB structures.

The HSSP (Homology-Derived Secondary Structure of Proteins) database provides multiple sequence alignments (MSAs) for proteins of known three-dimensional (3D) structure in the Protein Data Bank (PDB). The database also contains an estimate of the degree of evolutionary conservation at each amino acid position. This estimate, which is based on the relative entropy, correlates with the functional importance of the position; evolutionarily conserved positions (i.e., positions with limited variability and low entropy) are occasionally important to maintain the 3D structure and biological function(s) of the protein. We recently developed the Rate4Site algorithm for scoring amino acid conservation based on their calculated evolutionary rate. This algorithm takes into account the phylogenetic relationships between the homologs and the stochastic nature of the evolutionary process. Here we present the ConSurf-HSSP database of Rate4Site estimates of the evolutionary rates of the amino acid positions, calculated using HSSP's MSAs. The database provides precalculated evolutionary rates for nearly all of the PDB. These rates are projected, using a color code, onto the protein structure, and can be viewed online using the ConSurf server interface. To exemplify the database, we analyzed in detail the conservation pattern obtained for pyruvate kinase and compared the results with those observed using the relative entropy scores of the HSSP database. It is reassuring to know that the main functional region of the enzyme is detectable using both conservation scores. Interestingly, the ConSurf-HSSP calculations mapped additional functionally important regions, which are moderately conserved and were overlooked by the original HSSP estimate. The ConSurf-HSSP database is available online (http://consurf-hssp.tau.ac.il).

Algorithms↗

Database of homology-derived protein structures and the structural meaning of sequence alignment.

The database of known protein three-dimensional structures can be significantly increased by the use of sequence homology, based on the following observations. (1) The database of known sequences, currently at more than 12,000 proteins, is two orders of magnitude larger than the database of known structures. (2) The currently most powerful method of predicting protein structures is model building by homology. (3) Structural homology can be inferred from the level of sequence similarity. (4) The threshold of sequence similarity sufficient for structural homology depends strongly on the length of the alignment. Here, we first quantify the relation between sequence similarity, structure similarity, and alignment length by an exhaustive survey of alignments between proteins of known structure and report a homology threshold curve as a function of alignment length. We then produce a database of homology-derived secondary structure of proteins (HSSP) by aligning to each protein of known structure all sequences deemed homologous on the basis of the threshold curve. For each known protein structure, the derived database contains the aligned sequences, secondary structure, sequence variability, and sequence profile. Tertiary structures of the aligned sequences are implied, but not modeled explicitly. The database effectively increases the number of known protein structures by a factor of five to more than 1800. The results may be useful in assessing the structural significance of matches in sequence database searches, in deriving preferences and patterns for structure prediction, in elucidating the structural role of conserved residues, and in modeling three-dimensional detail by homology.

Amino Acid Sequence↗

HomologyPlot: searching for homology to a family of proteins using a database of unique conserved patterns.

A new database of conserved amino acid residues is derived from the multiple sequence alignment of over 84 families of protein sequences that have been reported in the literature. This database contains sequences of conserved hydrophobic core patterns which are probably important for structure and function, since they are conserved for most sequences in that family. This database differs from other single-motif or signature databases reported previously, since it contains multiple patterns for each family. The new database is used to align a new sequence with the conserved regions of a family. This is analogous to reports in the literature where multiple sequence alignments are used to improve a sequence alignment. A program called HomologyPlot (suitable for IBM or compatible computers) uses this database to find homology of a new sequence to a family of protein sequences. There are several advantages to using multiple patterns. First, the program correctly identifies a new sequence as a member of a known family. Second, the search of the entire database is rapid and requires less than one minute. This is similar to performing a multiple sequence alignment of a new sequence to all of the known protein family sequences. Third, the alignment of a new sequence to family members is reliable and can reproduce the alignment of conserved regions already described in the literature. The speed and efficiency of this method is enhanced, since there is no need to score for insertions or deletions as is done in the more commonly used sequence alignment methods. In this method only the patterns are aligned. HomologyPlot also provides general information on each family, as well as a listing of patterns in a family.

Amino Acid Sequence↗

The effect of a multiple literature database search--a numerical evaluation in the domain of Japanese life science.

In literature database searching, we show that it is necessary to use plural databases for a more improved search. We also compare the results of a single database search with that of multiple database search in the domain of Japanese life sciences. We searched the MEDLINE and EMBASE using the same search terms. There were some differences in the results, owing to differences in the journals and recording methods. We herein show some of the differences in the journals contained in both databases. Furthermore, we show the differences in the number of papers derived from the same journal. Next, as an example of a practical search, we selected some universities in Japan, searched both databases regarding papers published from these universities and then merged the results by hand. According to our results, only 63% of all papers were common to both databases.

Biological Science Disciplines↗

An analysis of concordance among hospital databases and physician records.

BACKGROUND: Hospital databases contain vital demographic patient information, which is increasingly being used as a basis to dictate care. It is hypothesized that the validity of data administratively generated from such sources is suboptimal, especially for rare subspecialties. The authors examined three databases to determine their concordance in an academic orthopaedic oncology subspecialty practice. METHODS: A 2-year retrospective review was performed on three databases searching for seven fundamental variables: additions/deletions; identification number; birthdate; procedure date; admit/discharge date; procedure code; and diagnostic code. Two university-maintained hospital databases (medical records and physician billing) were compared to the surgeon's personal handwritten daily log, which served as the "gold standard." RESULTS: All seven variables were in agreement with the physician's log in only 60% of the medical records and 61% of the physician billing patient entries (n = 564). On more detailed statistical analysis using chi(2), cross tabulations, and the K statistic for interobserver agreement, it was determined that poor concordance exists among the databases. CONCLUSION: Surgeons delivering quartenary care should maintain his or her own database because the hospital's information often differs on one or more important variables. Further investigation into the accuracy of hospital databases regarding commonly practiced medical disciplines appears warranted.

Academic Medical Centers↗

A human friendly reporting and database system for brain PET analysis.

We have developed a human friendly reporting and database system for clinical brain PET (Positron Emission Tomography) scans, which enables statistical data analysis on qualitative information obtained from image interpretation. Our system consists of a Brain PET Data (Input) Tool and Report Writing Tool. In the Brain PET Data Tool, findings and interpretations are input by selecting menu icons in a window panel instead of writing a free text. This method of input enables on-line data entry into and update of the database by means of pre-defined consistent words, which facilitates statistical data analysis. The Report Writing Tool generates a one page report of natural English sentences semi-automatically by using the above input information and the patient information obtained from our PET center's main database. It also has a keyword selection function from the report text so that we can save a set of keywords on the database for further analysis. By means of this system, we can store the data related to patient information and visual interpretation of the PET examination while writing clinical reports in daily work. The database files in our system can be accessed by means of commercially available databases. We have used the 4th Dimension database that runs on a Macintosh computer and analyzed 95 cases of 18F-FDG brain PET studies. The results showed high specificity of parietal hypometabolism for Alzheimer's patients.

Alzheimer Disease↗

Chemical databases evaluated by order theoretical tools.

Data on environmental chemicals are urgently needed to comply with the future chemicals policy in the European Union. The availability of data on parameters and chemicals can be evaluated by chemometrical and environmetrical methods. Different mathematical and statistical methods are taken into account in this paper. The emphasis is set on a new, discrete mathematical method called METEOR (method of evaluation by order theory). Application of the Hasse diagram technique (HDT) of the complete data-matrix comprising 12 objects (databases) x 27 attributes (parameters + chemicals) reveals that ECOTOX (ECO), environmental fate database (EFD) and extoxnet (EXT)--also called multi-database databases--are best. Most single databases which are specialised are found in a minimal position in the Hasse diagram; these are biocatalysis/biodegradation database (BID), pesticide database (PES) and UmweltInfo (UMW). The aggregation of environmental parameters and chemicals (equal weight) leads to a slimmer data-matrix on the attribute side. However, no significant differences are found in the "best" and "worst" objects. The whole approach indicates a rather bad situation in terms of the availability of data on existing chemicals and hence an alarming signal concerning the new and existing chemicals policies of the EEC.

Biodegradation, Environmental↗

Antidepressant drug prescribing in Italy, 2000: analysis of a general practice database.

OBJECTIVE: Databases of subjects receiving antidepressants provide evidence on the use of drugs in typical patients and settings under real-world conditions. This study analysed a general practice database to estimate the prevalence of antidepressant drug use, describe the use of these compounds by gender and age and estimate the prevalence of occasional versus non-occasional users. METHODS: The general practice database of Chivasso, a city near Turin in Piedmont, was analysed. The database includes all community (i.e. outside hospitals) prescriptions reimbursed by the National Health System in the population living in the study area. From the database, the total number of units of antidepressant drugs prescribed over a 6-month period was extracted. Using the general practice patient code, all records were converted into a sample of patients receiving one or more prescriptions of one or more antidepressants. RESULTS: During the 6 months surveyed, 12,930 antidepressant prescriptions were dispensed to 3751 patients, resulting in a prevalence of use of 19 patients per 1000 inhabitants (confidence interval 18.3, 19.5). The prevalence of use progressively increased with age and was more than double in females than males (female/male ratio 2.16). Paroxetine was the most prescribed compound, followed by amitriptyline and fluoxetine. However, in older subjects, the top two antidepressants were trazodone and amitriptyline. Nearly one-fourth of all dispensed antidepressants were prescribed on one occasion only; occasional users were slightly younger than non-occasional users. CONCLUSIONS: In Italy, databases have been used to monitor the prescription of medicines, but they have always provided aggregate data on drug sales and consumption. In this study, a sample of typical patients receiving antidepressants under real-world conditions was analysed to help clarify what happens in clinical practice. Databases of patients receiving antidepressants should be adopted to suggest public health priorities and generate original research hypotheses to be formally tested with experimental studies.

Adolescent↗

Practical considerations in the management of large multiinstitutional databases.

Large multiinstitutional databases are excellent sources of information that provide clinically useful insight into the practice of cardiac surgery. Fully informed subscribers should be aware of the practical concerns associated with the management and interpretation of database results. During development of The Society of Thoracic Surgeons National Database, three such areas have become particularly important: the database population, the database quality, and the significance of results. Appreciation of the real and philosophical problems associated with these issues will allow for greater appreciation of the intricacies of the database and will enhance the users' ability to interpret information gained from the database.

Bias↗