Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,693 records · Page 94Linked to original sources

Protein structural domains: analysis of the 3Dee domains database.

The 3Dee database of domain definitions was developed as a comprehensive collection of domain definitions for all three-dimensional structures in the Protein Data Bank (PDB). The database includes definitions for complex, multiple-segment and multiple-chain domains as well as simple sequential domains, organized in a structural hierarchy. Two different snapshots of the 3Dee database were analyzed at September 1996 and November 1999. For the November 1999 release, 7,995 PDB entries contained 13,767 protein chains and gave rise to 18,896 domains. The domain sequences clustered into 1,715 domain sequence families, which were further clustered into a conservative 1,199 domain structure families (families with similar folds). The proportion of different domain structure families per domain sequence family increases from 84% for domains 1-100 residues long to 100% for domains greater than 600 residues. This is in keeping with the idea that longer chains will have more alternative folds available to them. Of the representative domains from the domain sequence families, 49% are in the range of 51-150 residues, whereas 64% of the representative chains over 200 residues have more than 1 domain. Of the representative chains, 8.5% are part of multichain domains. The largest multichain domain in the database has 14 chains and 1,400 residues, whereas the largest single-chain domain has 907 residues. The largest number of domains found in a protein is 13. The analysis shows that over the history of the PDB, new domain folds have been discovered at a slower rate than by random selection of all known folds. Between 1992 and 1997, a constant 1 in 11 new domains deposited in the PDB has shown no sequence similarity to a previously known domain sequence family, and only 1 in 15 new domain structures has had a fold that has not been seen previously. A comparison of the September 1996 release of 3Dee to the Structural Classification of Proteins (SCOP) showed that the domain definitions agreed for 80% of the representative protein chains. However, 3Dee provided explicit domain boundaries for more proteins. 3Dee is accessible on the World Wide Web at http://barton.ebi.ac.uk/servers/3Dee.html.

Computational Biology↗

Identification of novel glutathione transferases and polymorphic variants by expressed sequence tag database analysis.

The human expressed sequence tag (EST) database can be searched by different sequence alignment strategies to identify new members of gene families and allelic variants. To illustrate the value of database analysis for gene discovery, we have focused on the glutathione S-transferase (GST) super family, an approach that has led to the identification of the Zeta class. The Zeta class GSTs catalyze the glutathione-dependent biotransformation of alpha-haloacids and the isomerization of maleylacetoacetic acid to fumarylacetoacetic acid, an essential step in the catabolism of tyrosine. Allelic variants of the GST Z1 and GST A2 genes have also been identified by EST database analysis. One GST Z1 variant (GST Z1A) has significantly higher activity with dichloroacetic acid as a substrate than other GST Z1 isoforms. This variant may be important in the clinical treatment of lactic acidosis where dichloroacetic acid is prescribed. Our experience with the application of EST database searching methods suggests that it may be productively applied to other gene families of pharmacogenetic interest.

Amino Acid Sequence↗

[Liver damage and nonsteroidal anti-inflammatory drugs: case non-case study in the French Pharmacovigilance Database].

This study investigates the relationship between exposure to non-steroidal anti-inflammatory drugs (NSAIDs) and liver injuries using the French Pharmacovigilance Database. We use the case/non-case methodology, where 'cases' were reports of the reactions of interest (liver injuries as recorded in the database according to the WHO-ART classification including cytolytic and cholestatic hepatitis, acute hepatitis, liver enzyme elevations). 'Non-cases' were all reports of reactions other than these being studied. Amineptine and acetaminophen were used as 'positive controls'. Among the 42,913 adverse drug reactions recorded in the database between January 1995 and December 1997, 5708 (13 per cent) were liver injuries. In comparison with other drugs in the database, liver injuries were inversely associated with exposure to NSAIDs, whatever the class of the drugs (OR 0.3 [0.3-0.4]). In contrast, liver injuries were significantly related to acetaminophen (OR 2.1 [1.9-2.3]), and amineptine (OR 14.0 [10.5-18.7]). Naproxen and diclofenac were associated with a higher frequency of liver injuries, respectively 15.7 per cent and 11.5 per cent. The risk associated with NSAIDs alone significantly decreased when the analysis was performed after exclusion of hepatotoxic drugs associated with NSAIDs (except for naproxen). The present results show the low frequency of liver damage associated with NSAIDs. The main factor appears to be concomitant exposure to other hepatotoxic drugs.

Anti-Inflammatory Agents↗

Migration of legacy mumps applications to relational database servers.

An extended implementation of the Mumps language is described that facilitates vendor neutral migration of legacy Mumps applications to SQL-based relational database servers. Implemented as a compiler, this system translates Mumps programs to operating system independent, standard C code for subsequent compilation to fully stand-alone, binary executables. Added built-in functions and support modules extend the native hierarchical Mumps database with access to industry standard, networked, relational database management servers (RDBMS) thus freeing Mumps applications from dependence upon vendor specific, proprietary, unstandardized database models. Unlike Mumps systems that have added captive, proprietary RDMBS access, the programs generated by this development environment can be used with any RDBMS system that supports common network access protocols. Additional features include a built-in web server interface and the ability to interoperate directly with programs and functions written in other languages.

Computer Communication Networks↗

Predicting risk of prostate specific antigen recurrence after radical prostatectomy with the Center for Prostate Disease Research and Cancer of the Prostate Strategic Urologic Research Endeavor databases.

PURPOSE: Biostatistical models to predict stage or outcome in patients with clinically localized prostate cancer with pretreatment prostate specific antigen (PSA), Gleason sum on biopsy or prostatectomy specimen, clinical or pathological stage and other variables, including ethnicity, have been developed. However, to date models have relied on small subsets from academic centers or military populations that may not be representative. Our study validates and updates a model published previously with the Cancer of the Prostate Strategic Urologic Research Endeavor (CaPSURE, UCSF, Urology Outcomes Research Group and TAP Pharmaceutical Products, Inc.), a large multicenter, community based prostate cancer database and Center for Prostate Disease Research (CPDR), a large military database. MATERIALS AND METHODS: We validated a biostatistical model that includes pretreatment PSA, highest Gleason sum on prostatectomy specimen, prostatectomy organ confinement status and ethnicity, including white and black patients. We then revised it with the Cox regression analysis of the combined 503 PSA era surgical cases from the CPDR prospective cancer database and 1,012 from the CaPSURE prostate cancer outcomes database. RESULTS: The original equation with 3 risk groups stratified CaPSURE cases into distinct categories with 7-year disease-free survival rates of 72%, 42.1% and 27.6% for low, intermediate and high risk men, respectively. Parameter estimates obtained from a Cox regression analysis provided a revised model equation that calculated the relative risk of recurrence as: exponent (exp)[(0.54 x Race) + (0.05 x sigmoidal transformation of PSA [PSA(ST)]) + (0.23 x Postop Gleason) + (0.69 x Pathologic stage). The relative risk of recurrence, as calculated by the aforementioned equation, was used to stratify the cases into 4 risk groups. Very low-4.7 or less, low-4.7 to 7.1, high-7.1 to 16.7 and very high-greater than 16.7, and patients at risk had 7-year disease-free survival rates of 85.4%, 66.0%, 50.6% and 21.3%, respectively. CONCLUSIONS: With a broad cohort of community based, academic and military cases, we developed an equation that stratifies men into 4 discrete risk groups of recurrence after radical prostatectomy and confirmed use of a prior 3 risk group model. Although the variables of ethnicity, pretreatment PSA, highest Gleason sum on prostatectomy specimen and organ confinement status on surgical pathology upon which the model is based are easily obtained, more refined modeling with additional variables are needed to improve prediction of intermediate risk in individuals.

Databases, Factual↗

Comparison of three databases with a decision tree approach in the medical field of acute appendicitis.

Decision trees have been successfully used for years in many medical decision making applications. Transparent representation of acquired knowledge and fast algorithms made decision trees one of the most often used symbolic machine learning approaches. This paper concentrates on the problem of separating acute appendicitis, which is a special problem of acute abdominal pain from other diseases that cause acute abdominal pain by use of an decision tree approach. Early and accurate diagnosing of acute appendicitis is still a difficult and challenging problem in everyday clinical routine. An important factor in the error rate is poor discrimination between acute appendicitis and other diseases that cause acute abdominal pain. This error rate is still high, despite considerable improvements in history-taking and clinical examination, computer-aided decision-support and special investigation, such as ultrasound. We investigated three different large databases with cases of acute abdominal pain to complete this task as successful as possible. The results show that the size of the database does not necessary directly influence the success of the decision tree built on it. Surprisingly we got the best results from the decision trees built on the smallest and the biggest database, where the database with medium size (relative to the other two) was not so successful. Despite that we were able to produce decision tree classifiers that were capable of producing correct decisions on test data sets with accuracy up to 84%, sensitivity to acute appendicitis up to 90%, and specificity up to 80% on the same test set.

Abdomen, Acute↗

[Single nucleotide polymorphisms(SNPs)and SNP databases].

Along the rapid development of human genome sequencing project, the variation data of human DNA sequence has become more and more useful not only for studying the origin, evolution and the mechanisms of maintenance of genetic variability in human populations, but also for detection of genetic association in complex disease such as diabetes, obesity, hypertension, Alzheimer's disease, etc. In recent two years, the databases such as dbSNP, CGAP, HGBASE, JST and Go!Poly etc. to collect and exploit data of genomic polymorphisms mainly single nucleotide polymorphisms (SNPs) have been respectively established in the United States, European countries, Japan and China. This overview summarized the development and applications of those SNP databases and also discussed some issues regarding the potential improvement of accuracy of SNP data collected. China has one fifth population in the world. Therefore, development of the SNP database for Chinese populations is of importance in developing complete SNP databases of human genome and may also stimulate the further development of biomedical research and production in China.

Databases, Nucleic Acid↗

Rubella vaccine and arthritic adverse reactions: an analysis of the Vaccine Adverse Events Reporting System (VAERS) database from 1991 through 1998.

OBJECTIVE: The United States Academy of Sciences, Institute of Medicine (IOM) reported in 1991 that the evidence indicates a causal relationship between the currently used rubella vaccine and acute and chronic arthritis. The purpose of this study was to analyze the associated arthritic reactions reported following rubella immunization from 1991 through 1998 to the Vaccine Adverse Events Reporting System (VAERS) database. METHODS: A certified copy of the VAERS database was obtained from the CDC. Microsoft Access was used to analyze the database. RESULTS: The results show that rubella vaccine is associated with a number of arthritic reactions reported to the VAERS database. CONCLUSION: Adult female patients need to make informed decisions on whether or not rubella vaccination is right for them. Doctors and patients must together make an informed consent decision about the risk verses the benefit to the patient in their particular life situation. Additionally, those patients who have had an adverse reaction to rubella vaccination should be informed that they may seek compensation under the no-fault Vaccine Compensation Act, which is administered by the US Claims Court.

Adult↗

Linkage of the Canadian Study of Health and Aging to provincial administrative health care databases in Nova Scotia.

The Canadian Study of Health and Aging (CSHA) was a cohort study that included 528 Nova Scotian community-dwelling participants. Linkage of CSHA and provincial Medical Services Insurance (MSI) data enabled examination of health care utilization in this subsample. This article discusses methodological and ethical issues of database linkage and explores variation in the use of health services by demographic variables and health status. Utilization over 24 months following baseline was extracted from MSI's physician claims, hospital discharge abstracts, and Pharmacare claims databases. Twenty-nine subjects refused consent for access to their MSI file; health card numbers for three others could not be retrieved. A significant difference in healthcare use by age and self-rated health was revealed. Linkage of population-based data with provincial administrative health care databases has the potential to guide health care planning and resource allocation. This process must include steps to ensure protection of confidentiality. Standard practices for linkage consent and routine follow-up should be adopted. The Canadian Study of Health and Aging (CSHA) began in 1991-92 to explore dementia, frailty, and adverse health outcomes (Canadian Study of Health and Aging Working Group, 1994). The original CSHA proposal included linkage to provincial administrative health care databases by the individual CSHA study centers to enhance information on health care utilization and outcomes of study participants. In Nova Scotia, the Medical Services Insurance (MSI) administration, which drew the sampling frame for the original CSHA, did not retain the list of corresponding health card numbers. Furthermore, consent for this access was not asked of participants at the time of the first interview. The objectives of this study reported here were to examine the feasibility and ethical considerations of linking data from the CSHA to MSI utilization data, and to explore variation in health services use by demographic and health status characteristics in the Nova Scotia community cohort.

Aged↗

[Construction of a proteomic map database].

The first proteomic map database of China will be published by Bioinformation Center and Research Center of Proteomics, Shanghai Institutes for Biological Sciences, the Chinese Academy of Sciences.The database consists of physical layer, link layer and interfacial layer. As Java technique was introduced to construct the database system, the database does not depend on special working platform. With the search instruments users can easily look through the proteomic map and get concrete information of proteins they are interested in.

China↗

[Construction and application of colorectal polyp database].

OBJECTIVE: To study the construction and application of computerized database of colorectal polyp in the clinical management and research of this disease. METHOD: A colorectal polyp database and its management system was constructed on the basis of Microsoft Access 2000. Clinical, endoscopic and pathological data, which went through standardized and elemental processing, of 2 627 cases (4 850 records) of colorectal polyp collected from 1990 to 2000 in Nanfang Hospital was entered into this database. RESULTS: Using this new database, the information on the population and age distribution, location and clinical features of colorectal polyps were obtained. Comparative study of the clinical and pathological findings in the cases, evaluation of the therapeutic effects, statistical review of the identification of the polyp and its canceration in the previous years as well as the analysis of other relevant factors were successfully accomplished, which greatly facilitated the follow-up study of some chosen cases that may be of clinical significance. CONCLUSIONS: Applications of modern informatics and computer technology greatly facilitates case management and clinical research of colorectal polyps, and standardized and elemental processing of the clinical data offers a new possibility for easy case information management.

Colonic Polyps↗

[Use of a health insurance company database for study of theoretical exposure to hypolipidemic agents].

OBJECTIVE: The objective of the investigation was to analyze theoretical exposures to hypolipidaemics in patients treated chronically with these drugs, using the database of the health insurance company. INVESTIGATED GROUP: From the database (with information on age, sex of the insured person, the number of packages and the type of hypolipidemic and year of issue of the prescription) of subjects insured at the Employees Health Insurance Skoda Mladá Boleslav comprising some 100,000 insured subjects in 1994-2000. Patients with long-term (more than one year) hypolipidaemic treatment were selected in years from 1995 to 1999. The group increased every year. In 1995 it comprised 668 cases in 1999, 2396 subjects. METHOD: The consumption of hypolipidaemics was expressed in defined daily doses (DDD). The authors investigated the ratio of chronically treated patients and the proportion of the following groups of patients according to their annual consumption in 1995-1999: group of of drug "vacation" (0 DDD) and the group with a low (< 121.7 DDD), medium (< 243.3 & > 121.7 DDD) and optimal (> 243.3 DDD) consumption of hypolipidaemics and their relationship to sex and age. For statistical ealuation software SPSS 10.1 was used. RESULTS: In the course of the investigation among the insured subjects the statin consumption increased 76 times and the consumption of fibrates 5 times. The ratio of consumption of resin derivatives and of nicotinic acid was negligible. The size of the group of subjects treated with hypolipidaemics for longer than one year increased from 0.8% in 1995 to 2.2% of the database. The average age increased from 55 to 59 years. The ratio of seniors (> or = 65 years) increased in the course of the investigation and reached 33% in 1999 of all members of the investigated group. The mean annual consumption of hypolipidaemics increased significantly as compared with 1995 and the interannual increase as compared with the previous year was statistically significant in 1997 and 1999. In 1999 it was 237 DDD/per consumer. A lower consumption was recorded in women and in seniors. Drug "vacations" were recorded in 6% of the insured subjects of the group and the frequency did not change significantly in the course of the investigation and no relationship with age and sex was found. A low exposure according to DDD was found in 20%, medium exposure in about 40% and optimal exposure in only one third of the subjects of the investigated group. CONCLUSION: The authors developed a method which makes it possible, when individual data of the health insurance company are available, to investigate the theoretical exposure to hypolipidaemics in insured subjects treated on a long-term basis with these drugs. The authors provided evidence that analysis of the database of the health insurance company can provide certain signals for further pharmacoepidemiological research and for application in a defined medical discipline. Some of the insured subjects are exposed to smaller doses than theoretically assumed. It is necessary to extend the investigation so that the results will better reflect the population of patients and prescribing physicians. Complete evaluation of cases with a low exposure from the aspect of morbidity and drug compliance will be also essential.

Aged↗

Rapid assignment of nucleotide sequence data to allele types for multi-locus sequence analysis (MLSA) of bacteria using an adapted database and modified alignment program.

A novel database and modified alignment program is described which provides a fast and accurate procedure for assigning nucleotide sequences to allele types for multi-locus sequence analysis (MLSA). The database has between 40 and 160 alleles per organism including Neisseria meningitidis, Streptococcus pneumoniae, Staphylococcus aureus and Haemophilus influenzae. The database directly compares the query nucleotide sequence against all alleles within the database and this system reduces the time taken for the analysis of nucleotide sequence data and assignment of alleles for subsequent sequence analysis.

Alleles↗

Generating a mortality model from a pediatric ICU (PICU) database utilizing knowledge discovery.

Current models for predicting outcomes are limited by biases inherent in a priori hypothesis generation. Knowledge discovery algorithms generate models directly from databases, minimizing such limitations. Our objective was to generate a mortality model from a PICU database utilizing knowledge discovery techniques. The database contained 5067 records with 192 clinically relevant variables. It was randomly split into training (75%) and validation (25%) groups. We used decision tree induction to generate a mortality model from the training data, and validated its performance on the validation data. The original PRISM algorithm was used for comparison. The decision tree model contained 25 variables and predicted 53/88 deaths; 29 correctly (Sens:33%, Spec:98%, PPV:54%). PRISM predicted 27/88 deaths correctly (Sens:30%, Spec:98%, PPV:51%). Performance difference between models was not significant. We conclude that knowledge discovery algorithms can generate a mortality model from a PICU database, helping establish validity of such tools in the clinical medical domain.

Algorithms↗

Automated concept matching between laboratory databases.

To address the problem of semantic inconsistencies between medical databases, semantic network representations can be utilized to automate the matching of medical concepts between the databases. The performance of automated concept matching was tested by creating semantic network representations for two laboratory databases, one from a pediatric hospital and the other from an oncology institute. The matching algorithms identified all equivalent concepts that were present in both databases, and did not leave any equivalent concepts unmatched. By automatically identifying semantically equivalent concepts, the Medical Information Acquisition and Transmission Enabler (MEDIATE) facilitates data exchange between heterogeneous systems because no pre-negotiation is required. Consequently, system scalability and stability is improved.

Algorithms↗

[Construction and application of Access database of colorectal carcinoma cases].

OBJECTIVE: To construct a database using Access software in which clinical information of colorectal cancer cases can be loaded, to facilitate relevant large-sample clinical studies. METHODS: A retrospective study was conducted in 1 374 cases of colorectal carcinoma with surgical treatment between 1975 to 1999 in Nanfang Hospital. According to the National Standards for Pathological Study of Colorectal Carcinoma, an Access2000 database consisting of 1 145 pathologically confirmed colorectal carcinoma cases was established, designated as The Specialized Access Database of Colorectal Carcinoma. RESULTS AND CONCLUSION: The database system has been successfully constructed and operates smoothly, which possesses powerful capacity for information processing of colorectal cancer cases.

Adolescent↗

Limitations of electronic databases: a caution.

OBJECTIVE: The purpose of this study was to assess the completeness and accuracy of information from two electronic datasets, one of which is voluntarily submitted, the other submitted by mandate. METHODS: Emergency department (ED) data have been voluntarily submitted by several hospitals to the Kentucky Emergency Medical Services Information System (KEMSIS). Similar information on all patients admitted to the hospital has been submitted by mandate to the state under the Uniform Billing Act (UB92). UB92 data for patients with at least one diagnosis code > or = 800 were available. The KEMSIS and UB 92 data for one hospital were compared to those manually abstracted from the ED log and medical records for completeness and accuracy. RESULTS: There were 316 patients listed on the ED log that were subsequently admitted to the hospital. The KEMSIS database contained 266 (84%) of these records, but only 91 (34%) were classified as having been admitted. Of those correctly classified as admitted, only 25 (27%, or 9% of the total 266) were correctly classified as to the hospital of admission (directly to hospital or transferred to another facility). Discharge diagnoses in the KEMSIS database and hospital records were concordant in 240 (90%) of the patients, even for those misclassified as to disposition. There were 37 patients listed in the ED log admitted during the study period with at least one discharge diagnosis field > or = 800. Eight patients were transferred to another institution, making the total population available for study period 29. Only eight (28%) of these patients were included in the UB92 database. The diagnosis codes were concordant between the UB 92 data and ED log in all cases. CONCLUSIONS: There is significant misclassification and/or omission in electronic databases. This is true regardless of whether data is reported voluntarily or by mandate. Electronic data must be independently validated before they are used for policy or research purposes.

Databases, Factual↗

DAVID: Database for Annotation, Visualization, and Integrated Discovery.

BACKGROUND: Functional annotation of differentially expressed genes is a necessary and critical step in the analysis of microarray data. The distributed nature of biological knowledge frequently requires researchers to navigate through numerous web-accessible databases gathering information one gene at a time. A more judicious approach is to provide query-based access to an integrated database that disseminates biologically rich information across large datasets and displays graphic summaries of functional information. RESULTS: Database for Annotation, Visualization, and Integrated Discovery (DAVID; http://www.david.niaid.nih.gov) addresses this need via four web-based analysis modules: 1) Annotation Tool - rapidly appends descriptive data from several public databases to lists of genes; 2) GoCharts - assigns genes to Gene Ontology functional categories based on user selected classifications and term specificity level; 3) KeggCharts - assigns genes to KEGG metabolic processes and enables users to view genes in the context of biochemical pathway maps; and 4) DomainCharts - groups genes according to PFAM conserved protein domains. CONCLUSIONS: Analysis results and graphical displays remain dynamically linked to primary data and external data repositories, thereby furnishing in-depth as well as broad-based data coverage. The functionality provided by DAVID accelerates the analysis of genome-scale datasets by facilitating the transition from data collection to biological meaning.

Computational Biology↗