Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Statistical models for protein validation using tandem mass spectral data and protein amino acid sequence databases.

The purpose of this work is to develop and verify statistical models for protein identification using peptide identifications derived from the results of tandem mass spectral database searches. Recently we have presented a probabilistic model for peptide identification that uses hypergeometric distribution to approximate fragment ion matches of database peptide sequences to experimental tandem mass spectra. Here we apply statistical models to the database search results to validate protein identifications. For this we formulate the protein identification problem in terms of two independent models, two-hypothesis binomial and multinomial models, which use the hypergeometric probabilities and cross-correlation scores, respectively. Each database search result is assumed to be a probabilistic event. The Bernoulli event has two outcomes: a protein is either identified or not. The probability of identifying a protein at each Bernoulli event is determined from relative length of the protein in the database (the null hypothesis) or the hypergeometric probability scores of the protein's peptides (the alternative hypothesis). We then calculate the binomial probability that the protein will be observed a certain number of times (number of database matches to its peptides) given the size of the data set (number of spectra) and the probability of protein identification at each Bernoulli event. The ratio of the probabilities from these two hypotheses (maximum likelihood ratio) is used as a test statistic to discriminate between true and false identifications. The significance and confidence levels of protein identifications are calculated from the model distributions. The multinomial model combines the database search results and generates an observed frequency distribution of cross-correlation scores (grouped into bins) between experimental spectra and identified amino acid sequences. The frequency distribution is used to generate p-value probabilities of each score bin. The probabilities are then normalized with respect to score bins to generate normalized probabilities of all score bins. A protein identification probability is the multinomial probability of observing the given set of peptide scores. To reduce the effect of random matches, we employ a marginalized multinomial model for small values of cross-correlation scores. We demonstrate that the combination of the two independent methods provides a useful tool for protein identification from results of database search using tandem mass spectra. A receiver operating characteristic curve demonstrates the sensitivity and accuracy level of the approach. The shortcomings of the models are related to the cases when protein assignment is based on unusual peptide fragmentation patterns that dominate over the model encoded in the peptide identification process. We have implemented the approach in a program called PROT_PROBE.

Amino Acid Sequence↗

Clinical databases and data protection: are they compatible?

In the current climate of clinical governance and audit, and in the setting of an active academic unit, an effective clinical database is an invaluable tool. In this article, we will present our neurovascular database, discuss the issues related to setting up the ideal clinical database, discuss the problems related to accurate data input and review the legal requirements of data protection. The success of a clinical database is reflected by the completeness of the data, the accessibility of the information and how useful it has proven to be. After 4 years of experimentation we currently use a database designed on Microsoft Access. The form is a single page. Junior medical staff input the information as medical staff have been found to be the most reliable personnel for data input in terms of accuracy. However, time is generally in short supply amongst this group. For our purposes, the ideal database is one that is simple, that can be used to flag up cases, rather than provide all of the information and ensures a complete dataset. The arrival of the UK 1998 Data Protection Act has put many clinical databases and registries in jeopardy, and introduced further bureaucracy to research. We discuss the Act and its interpretation by the General Medical Council, Medical Research Council, British Medical Association, Department of Health and our own trust with respect to databases and research.

Cerebrovascular Disorders↗

EUROPOEM, a predictive occupational exposure database for registration purposes of pesticides.

For registration of agricultural pesticides, the risks for humans, animals, and the environment must be determined. The risk assessment is based on an appraisal of the levels of exposure and the hazards of the active substance(s) in the plant protection product, that is, the agricultural pesticide. Funded by the European Commission (AIR3 CT93-1370), the EUROPOEM database has been developed by a group of experts, representing governments, industry, and academia. The currently available exposure database reflects exposure to operators (mixer/loaders and applicators). The EUROPOEM approach is based on a harmonized protocol for conduct of field studies of operator exposure (presently published as an Organization for Economic Cooperation and Development [OECD] Guidance Document) and a tiered approach to exposure and risk assessment. The database is constructed from exposure data obtained in representative field studies. These field studies are considered according to criteria reflecting the quality of documentation, study design, adequate methodology, number of replicates, and QA/QC elements, for use of the inhalation and dermal exposure data. The resulting exposure data were combined according to comparable use scenarios. From the resulting databases typical surrogate potential exposure values have been obtained, which are determined by their use for either acute or chronic health effects, and the size of the database. For large databases (over 50-100 data points), from many different field studies (10 or more), the 75th percentile is taken if the exposure is considered leading to chronic effects. For smaller databases, a more conservative 90th percentile is taken as surrogate value, or none at all for very small databases (15-20 or less data points from 3 or less different field studies). The choice for the 75th percentile is based on the assumed or observed lognormal distribution of the exposure data, as being the most relevant typical value for long-term effects, since the 75th percentile of log-normal distributions is nominally very similar to a calculated arithmetic mean (AM). The AM, as such however, is irrelevant for log-normal distributions.

Agriculture↗

The thyrotropin receptor mutation database: update 2003.

In 1999 we have created a TSHR mutation database compiling TSHR mutations with their basic characteristics and associated clinical conditions (www.uni-leipzig.de/innere/tshr). Since then, more than 2887 users from 36 countries have logged into the TSHR mutation database and have contributed several valuable suggestions for further improvement of the database. We now present an updated and extended version of the TSHR database to which several novel features have been introduced: 1. detailed functional characteristics on all 65 mutations (43 activating and 22 inactivating mutations) reported to date, 2. 40 pedigrees with detailed information on molecular aspects, clinical courses and treatment options in patients with gain-of-function and loss-of-function germline TSHR mutations, 3. a first compilation of site-directed mutagenesis studies, 4. references with Medline links, 5. a user friendly search tool for specific database searches, user-specific database output and 6. an administrator tool for the submission of novel TSHR mutations. The TSHR mutation database is installed as one of the locus specific HUGO mutation databases. It is listed under index TSHR 603372 (http://ariel.ucs.unimelb.edu.au/~cotton/glsdbq.htm) and can be accessed via www.uni-leipzig.de/innere/tshr.

Databases, Factual↗

Efficiency of 22 online databases in the search for physicochemical, toxicological and ecotoxicological information on chemicals.

The objective of this study was to evaluate the efficiency of 22 free online databases that could be used for an exhaustive search of physicochemical, toxicological and/or ecotoxicological information about various chemicals. Twenty-two databases with free access on the Internet were referenced. We then selected 27 major physicochemical, toxicological and ecotoxicological criteria and 14 compounds belonging to seven different chemical classes which were used to interrogate all the databases. Two indices were successively calculated to evaluate the efficiency with taking or not taking account of their specialization. More than 50% of the 22 databases 'knew' all of the 14 chemicals, but the quantity of information provided is very different from one to the other and most are poorly documented. Two categories clearly appear with specialized and non-specialized databases. The HSDB database is the most efficient general database to be searched first, because it is well documented for most of the 27 criteria. However, some specialized databases (i.e. EXTOXNET, SOLVEDB, etc.) must be searched secondarily to find additional information.

Databases, Bibliographic↗

Analysis and comparison of metabolic pathway databases.

Enormous amounts of data result from genome sequencing projects and new experimental methods. Within this tremendous amount of genomic data 30-40 per cent of the genes being identified in an organism remain unknown in terms of their biological function. As a consequence of this lack of information the overall schema of all the biological functions occurring in a specific organism cannot be properly represented. To understand the functional properties of the genomic data more experimental data must be collected. A pathway database is an effort to handle the current knowledge of biochemical pathways and in addition can be used for interpretation of sequence data. Some of the existing pathway databases can be interpreted as detailed functional annotations of genomes because they are tightly integrated with genomic information. However, experimental data are often lacking in these databases. This paper summarises a list of pathway databases and some of their corresponding biological databases, and also focuses on information about the content and the structure of these databases, the organisation of the data and the reliability of stored information from a biological point of view. Moreover, information about the representation of the pathway data and tools to work with the data are given. Advantages and disadvantages of the analysed databases are pointed out, and an overview to biological scientists on how to use these pathway databases is given.

Animals↗

RSDB: representative protein sequence databases have high information content.

MOTIVATION: Biological sequence databases are highly redundant for two main reasons: 1. various databanks keep redundant sequences with many identical and nearly identical sequences 2. natural sequences often have high sequence identities due to gene duplication. We wanted to know how many sequences can be removed before the databases start losing homology information. Can a database of sequences with mutual sequence identity of 50% or less provide us with the same amount of biological information as the original full database? RESULTS: Comparisons of nine representative sequence databases (RSDB) derived from full protein databanks showed that the information content of sequence databases is not linearly proportional to its size. An RSDB reduced to mutual sequence identity of around 50% (RSDB50) was equivalent to the original full database in terms of the effectiveness of homology searching. It was a third of the full database size which resulted in a six times faster iterative profile searching. The RSDBs are produced at different granularity for efficient homology searching. AVAILABILITY: All the RSDB files generated and the full analysis results are available through internet: ftp://ftp.ebi.ac. uk/pub/contrib/jong/RSDB/http://cyrah.e bi.ac.uk:1111/Proj/Bio/RSDB

Algorithms↗

The PIR-International databases.

PIR-International is an association of macromolecular sequence data collection centers dedicated to fostering international cooperation as an essential element in the development of scientific databases. PIR-International is most noted for the Protein Sequence Database. This database originated in the early 1960's with the pioneering work of the late Margaret Dayhoff as a research tool for the study of protein evolution and intersequence relationships; it is maintained as a scientific resource, organized by biological concepts, using sequence homology as a guiding principle. PIR-International also maintains a number of other genomic, protein sequence, and sequence-related databases. The databases of PIR-International are made widely available. This paper briefly describes the architecture of the Protein Sequence Database, a number of other PIR-International databases, and mechanisms for providing access to and for distribution of these databases.

Amino Acid Sequence↗

MitBASE : a comprehensive and integrated mitochondrial DNA database. The present status.

MitBASE is an integrated and comprehensive database of mitochondrial DNA data which collects, under a single interface, databases for Plant, Vertebrate, Invertebrate, Human, Protist and Fungal mtDNA and a Pilot database on nuclear genes involved in mitochondrial biogenesis in Saccharomyces cerevisiae. MitBASE reports all available information from different organisms and from intraspecies variants and mutants. Data have been drawn from the primary databases and from the literature; value adding information has been structured, e.g., editing information on protist mtDNA genomes, pathological information for human mtDNA variants, etc. The different databases, some of which are structured using commercial packages (Microsoft Access, File Maker Pro) while others use a flat-file format, have been integrated under ORACLE. Ad hoc retrieval systems have been devised for some of the above listed databases keeping into account their peculiarities. The database is resident at the EBI and is available at the following site: http://www3.ebi.ac.uk/Research/Mitbase/mitbas e.pl. The impact of this project is intended for both basic and applied research. The study of mitochondrial genetic diseases and mitochondrial DNA intraspecies diversity are key topics in several biotechnological fields. The database has been funded within the EU Biotechnology programme.

Animals↗

The Diatom EST Database.

The Diatom EST database provides integrated access to expressed sequence tag (EST) data from two eukaryotic microalgae of the class Bacillariophyceae, Phaeodactylum tricornutum and Thalassiosira pseudonana. The database currently contains sequences of close to 30,000 ESTs organized into PtDB, the P.tricornutum EST database, and TpDB, the T.pseudonana EST database. The EST sequences were clustered and assembled into a non-redundant set for each organism, and these non-redundant sequences were then subjected to automated annotation using similarity searches against protein and domain databases. EST sequences, clusters of contiguous sequences, their annotation and analysis with reference to the publicly available databases, and a codon usage table derived from a subset of sequences from PtDB and TpDB can all be accessed in the Diatom EST Database. The underlying RDBMS enables queries over the raw and annotated EST data and retrieval of information through a user-friendly web interface, with options to perform keyword and BLAST searches. The EST data can also be retrieved based on Pfam domains, Cluster of Orthologous Groups (COG) and Gene Ontologies (GO) assigned to them by similarity searches. The Database is available at http://avesthagen.sznbowler.com.

DNA, Algal↗

The Molecular Biology Database Collection: 2007 update.

The NAR online Molecular Biology Database Collection is a public resource that contains links to the databases described in this issue of Nucleic Acids Research, previous NAR database issues, as well as a selection of other molecular biology databases that are freely available on the web and might be useful to the molecular biologist. The 2007 update includes 968 databases, 110 more than the previous one. Many databases that have been described in earlier issues of NAR come with updated summaries, which reflect recent progress and, in some instances, an expanded scope of these databases. The complete database list and summaries are available online on the Nucleic Acids Research web site http://nar.oxfordjournals.org/.

Databases, Genetic↗

FINDbase: a relational database recording frequencies of genetic defects leading to inherited disorders worldwide.

Frequency of INherited Disorders database (FINDbase) (http://www.findbase.org) is a relational database, derived from the ETHNOS software, recording frequencies of causative mutations leading to inherited disorders worldwide. Database records include the population and ethnic group, the disorder name and the related gene, accompanied by links to any corresponding locus-specific mutation database, to the respective Online Mendelian Inheritance in Man entries and the mutation together with its frequency in that population. The initial information is derived from the published literature, locus-specific databases and genetic disease consortia. FINDbase offers a user-friendly query interface, providing instant access to the list and frequencies of the different mutations. Query outputs can be either in a table or graphical format, accompanied by reference(s) on the data source. Registered users from three different groups, namely administrator, national coordinator and curator, are responsible for database curation and/or data entry/correction online via a password-protected interface. Databaseaccess is free of charge and there are no registration requirements for data querying. FINDbase provides a simple, web-based system for population-based mutation data collection and retrieval and can serve not only as a valuable online tool for molecular genetic testing of inherited disorders but also as a non-profit model for sustainable database funding, in the form of a 'database-journal'.

Databases, Genetic↗

Searching bibliographic databases for literature on chronic disease and work participation.

BACKGROUND: The work participation of people with chronic diseases is a growing concern within the field of occupational medicine. Information on this topic is dispersed across a variety of data sources, making it difficult for health professionals to find relevant studies for literature reviews and guidelines. AIM: The goal of this project was to identify bibliographic databases and search terms that could be most useful for retrieving relevant studies on this topic. METHODS: Five broad questions regarding work participation and chronic disease were formulated, focusing on angina pectoris, depression, diabetes mellitus, hearing impairment and rheumatoid arthritis. A search strategy for retrieving information on these questions was developed and run in five bibliographic databases: Medline, EMBASE, PsycINFO, Cinahl and OSHROM. Relevant publications were selected from the search results. The utility of the selected databases and search terms was evaluated by analysing the number of relevant publications that were retrieved. RESULTS: The number of relevant publications retrieved from each database varied. Most (84%) of the relevant publications that were retrieved from each database were unique to that source. For each database, specific search terms for the concept of 'work' were useful for retrieving relevant publications. CONCLUSION: Medline, EMBASE and PsycINFO are useful databases for quick searches. Useful search terms for the concept of 'work' are work capacity, work disability, vocational rehabilitation, occupational health, sick leave, absenteeism, return to work, retirement, employment status and work status. For comprehensive searches, we recommend additional searches in Cinahl and OSHROM, adapting the search terms to specific databases.

Chronic Disease↗

Clinical information for research; the use of general practice databases.

General practice computers have been widely used in the United Kingdom for the last 10 years and there are over 30 different systems currently available. The commercially available databases are based on two of the most widely used systems--VAMP Medical and Meditel. These databases provide both longitudinal and cross-sectional data on between 1.8 and 4 million patients. Despite their availability only limited use has been made of them for epidemiological and health service research purposes. They are a unique source of population-based information and deserve to be better recognized. The advantages of general practice databases include the fact that they are population based with excellent prescribing data linked to diagnosis, age and gender. The problems are that their primary purpose is patient care and the database population is constantly changing, as well as the usual problems of bias and confounding that occur in any observational studies. The barriers to the use of general practice databases include the cost of access, the size of the databases and that they are not structured in a way that easily allows analysis. Proper utilization of these databases requires powerful computers, staff proficient in writing computer programs to facilitate analysis and epidemiologists skilled in their use. If these structural problems are overcome then the databases are an invaluable source of data for epidemiological studies.

Bias↗

Comparing checklists and databases with physicians' ratings as measures of students' history and physical-examination skills.

PURPOSE: To compare two methods of rating students' performances on history and physical examination: (1) by using checklists completed by standardized patients (SPs) and databases completed by students, and (2) by using ratings of students by three physicians for each SP-student encounter. METHOD: Four cases were chosen for the study, and 30 students were examined per case. The students were all in their fourth year at the Southern Illinois University School of Medicine in the spring of 1991. Two of the cases had both checklists and databases, and the remaining two had databases only. Each SP-student encounter was videotaped and was viewed independently by three physicians unfamiliar with the contents of the checklists and databases. The physicians' pooled ratings were then compared with the checklist and database scores. Uncorrected and corrected correlations were obtained, with the generalizability coefficient used as the index of reliability. RESULTS: Interrater generalizability of physicians' ratings was very good, ranging from .65 to .93 for overall ratings. Generalizability of physicians' ratings pooled across the four cases was .85. Checklist scores tended to correlate higher with physicians' ratings than did database scores: across the cases, correlation coefficients between physicians' ratings and checklist scores and database scores were .65 and .39, respectively. CONCLUSION: The checklist scores correlated strongly with the physicians' ratings of history and physical-examination skills, providing some evidence of validity for their use. The checklist scores correlated much better with the physicians' ratings than did the database scores. Possible explanations for this finding are discussed.

Clinical Clerkship↗

Identification of neonatal hearing impairment: experimental protocol and database management.

OBJECTIVE: The purposes of this article are to describe the overall protocol for the Identification of Neonatal Hearing Impairment (INHI) project and to describe the management of the data collected as part of this project. A well-defined protocol and database management techniques were needed to ensure that data were 1) collected accurately and in the same way across sites; 2) maintained in a database that could be used to provide feedback to individual sites regarding enrollment and the extent to which the protocol was complete on individual subjects; and 3) available to answer project questions. This article describes techniques that were used to meet these needs. DESIGN: This study was a prospective, randomized study that was designed to evaluate auditory brain stem responses, transient evoked otoacoustic emissions, and distortion product otoacoustic emissions as hearing-screening tools, and to relate neonatal test findings to hearing status, defined by visual reinforcement audiometry at 8 to 12 mo of age. Measures of middle-ear function also were obtained at some sites as part of the neonatal test battery. In addition, other clinical and demographic data were gathered to determine the extent to which factors, other than auditory status, influenced test behavior. Three groups were evaluated: neonatal intensive care unit (NICU) infants (those who spent 3 or more days in a NICU), well babies with risk factors for hearing loss, and well babies without risk factors. Six centers participated in the trial. The testers for the project included audiologists, technicians, audiology graduate students, and medical research staff. The same computerized neonatal test program was applied at each center. This program generated the neonatal test database automatically. Clinical and demographic data were collected by means of concise data collection forms and were entered into a database at each site. After the neonatal test, subjects from the NICU and at-risk well babies were evaluated with visual reinforcement audiometry starting at 8 to 12 mo of age. All data were electronically transmitted to the core site where they were merged into one overall database. This database was exercised to provide feedback and to identify discrepancies throughout the course of the study. In its final form, it served as the database on which all analyses were performed. RESULTS AND CONCLUSION: The protocol was a departure from typical hearing screening procedures in terms of 1) its regimented application of three screening measures; 2) the detailed information that was obtained regarding subject clinical and demographic factors; and 3) its application of the same procedures across six centers having diverse geographic location and subject demographics. A learning curve for successfully executing the study protocols was observed. Throughout the study, monthly reports were generated to monitor subject enrollment, check for data completeness, and to perform data integrity checks. In combination with monthly data reports and checks that occurred throughout the progression of the study, miscellaneous data audits were performed to check accuracy of neonatal testing programs and to cross-check information entered in the clinical and demographic database. The data management techniques used in this project helped to ensure the quality of the data collection process and also allowed for detailed analyses once data were collected. This was particularly important because it enabled us to evaluate not only the performance of individual measures as screening tools, but also permitted an evaluation of the influence of other variables on screening test results.

Acoustic Stimulation↗

The Albion Street Centre database, Sydney, Australia.

The Albion Street Centre was established in 1985 as an HIV testing and early management center. More than 22,000 people have been screened for HIV and other blood-borne infections at the Centre, and approximately 3,600 people with HIV/AIDS have been managed there. Approximately 1,600 patients with various stages of HIV disease are currently managed at the Centre by a staff of 60 health care professionals and about 1,000 volunteers. The Albion Street Centre's computer database began recording selected demographic, epidemiologic, clinical, and laboratory characteristics when the first patient presented in 1985. Since then, the complexity and utilization of the database has increased in parallel with improvement in the understanding of the natural history and pathogenesis of HIV infection. Over 100 peer-reviewed publications and presentations have been produced from the database and 45 clinical trials have used the database to identify potential subjects. All data are de-identified and are protected by multiple password codes. Approximately 700 variables are collected from each HIV-positive patient at the initial visit to the Centre and up to 200 variables are added at each subsequent routine clinic visit. The variables collected include the following: standard epidemiologic characteristics; transmission and behavioral parameters, clinical signs and symptoms; laboratory test results; treatments; nutritional history; body composition parameters; psychological assessment results; and management history, including neuropsychological testing. The overall number and characteristics of patients recorded in the database are reported monthly, and are used to plan services, for prevention and educational programs, and as an indicator of the effectiveness of campaigns to encourage HIV-positive people to attend clinics for early management. When these patients have been identified they are invited to participate in the study. Individual patient records are identified and accessed if they meet certain criteria for flagging. For example, patients who have lost more than 5% of their maximal weight are flagged and referred to the dietician for assessment. Further uses for the database are to identify cohorts of patients who are seroconverters and to follow their natural history-the Centre has over 250 patients for whom a documented HIV-positive test has been obtained within 12 months of a documented HIV-negative test; to investigate clinical observations that have been associated with particular drug therapy, e.g., investigation of the reported association between the use of valacyclovir and the thrombotic thrombocytopenic purpura/hemolytic uremic syndrome (TTP/HUS)-like complex showed patients with terminal-stage AIDS demonstrated this syndrome independently of their therapy and probably as a consequence of multiorgan failure; and to document the relationship between nutritional intervention and survival, for which use of the database enabled an historical cohort that matched the cases under investigation to be selected. In conclusion, the database is a dynamic and integral part of the assessment, management, and research program of the Albion Street Centre, where it is used by all professional staff.

Adolescent↗

The potential of UK clinical databases in enhancing paediatric medication research.

The research potential of many UK clinical databases is not being realized. A recent report published by the Royal College of Paediatrics & Child Health stated that there is a need to build research capacity and support in the area of paediatric pharmacology, with specific emphasis on the use of clinical databases. This article presents the databases available in the UK for medication research and gives some examples of paediatric studies conducted. The databases discussed include the Prescription Pricing Authority database, the General Practice Research Database, IMS Health databases (Medical Data Index, MIDAS Prescribing Insights, Disease-Analyser-Mediplus) and the Yellow Card Scheme. Other databases such as the Medicines Monitoring Unit (MEMO) and the Scottish Primary Care Computer System also have research potential in paediatric pharmacoepidemiology, but their population sizes are relatively small.

Adolescent↗