Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

The use of databases to manage fertility.

Dairy farming now needs more records to be kept for quality assurance as well as for management. Herd fertility management is best brought about through the use of computerised records for each animal that integrate fertility, health and production. The development of dairy information systems over the last 25 years has allowed the creation of databases that give rise to "standards" of performance and "interference levels". These databases are of limited use for research unless the coding system has a structure and definition that works across herds. There is an increasing need to incorporate carefully coded disease records into these databases as there is increasing concern about welfare, zoonoses, assurance and the environment. Rules can be determined for satisfactory fertility so interference at an early stage is cost-effective. Integrated indices have been developed (using databases) that incorporate the costs of wastage caused by poor fertility, thus highlighting the priorities for management. Databases are best operated near to the farm, either in the veterinarian's office or on-line in the farm office. Databases can be made into expert systems that deliver high standards of fertility management. A checklist is included that can be followed to analyse the causes of poor fertility in a dairy herd.

Animals↗

Image matching algorithms for breech face marks and firing pins in a database of spent cartridge cases of firearms.

On the market several systems exist for collecting spent ammunition data for forensic investigation. These databases store images of cartridge cases and the marks on them. Image matching is used to create hit lists that show which marks on a cartridge case are most similar to another cartridge case. The research in this paper is focused on the different methods of feature selection and pattern recognition that can be used for optimizing the results of image matching. The images are acquired by side light images for the breech face marks and by ring light for the firing pin impression. For these images a standard way of digitizing the images used. For the side light images and ring light images this means that the user has to position the cartridge case in the same position according to a protocol. The positioning is important for the sidelight, since the image that is obtained of a striation mark depends heavily on the angle of incidence of the light. In practice, it appears that the user positions the cartridge case with +/-10 degrees accuracy. We tested our algorithms using 49 cartridge cases of 19 different firearms, where the examiner determined that they were shot with the same firearm. For testing, these images were mixed with a database consisting of approximately 4900 images that were available from the Drugfire database of different calibers.In cases where the registration and the light conditions among those matching pairs was good, a simple computation of the standard deviation of the subtracted gray levels, delivered the best-matched images. For images that were rotated and shifted, we have implemented a "brute force" way of registration. The images are translated and rotated until the minimum of the standard deviation of the difference is found. This method did not result in all relevant matches in the top position. This is caused by the effect that shadows and highlights are compared in intensity. Since the angle of incidence of the light will give a different intensity profile, this method is not optimal. For this reason a preprocessing of the images was required. It appeared that the third scale of the "à trous" wavelet transform gives the best results in combination with brute force. Matching the contents of the images is less sensitive to the variation of the lighting. The problem with the brute force method is however that the time for calculation for 49 cartridge cases to compare between them, takes over 1 month of computing time on a Pentium II-computer with 333MHz. For this reason a faster approach is implemented: correlation in log polar coordinates. This gave similar results as the brute force calculation, however it was computed in 24h for a complete database with 4900 images.A fast pre-selection method based on signatures is carried out that is based on the Kanade Lucas Tomasi (KLT) equation. The positions of the points computed with this method are compared. In this way, 11 of the 49 images were in the top position in combination with the third scale of the à trous equation. It depends however on the light conditions and the prominence of the marks if correct matches are found in the top ranked position. All images were retrieved in the top 5% of the database. This method takes only a few minutes for the complete database if, and can be optimized for comparison in seconds if the location of points are stored in files. For further improvement, it is useful to have the refinement in which the user selects the areas that are relevant on the cartridge case for their marks. This is necessary if this cartridge case is damaged and other marks that are not from the firearm appear on it.

Algorithms↗

A brief history of the formation of DNA databases in forensic science within Europe.

The introduction of DNA analysis to forensic science brought with it a number of choices for analysis, not all of which were compatible. As laboratories throughout Europe were eager to use the new technology different systems became routine in different laboratories and consequently, there was no basis for the exchange of results. A period of co-operation then started in which a nucleus of forensic scientists agreed on an uniform system. This collaboration spread to incorporate most of the established forensic science laboratories in Europe and continued through two major changes in the technology. At each step agreement was reached on which systems to use. From the beginning it was realised that DNA databases would provide the criminal justice systems with an efficient way of crime solving and consequently some local databases were created. It was not until the introduction of the amplification technology linked to the analysis of short tandem repeats that a sufficiently sensitive and robust system was available for the formation of efficient and effective DNA databases. Comprehensive legislation enacted in the UK in 1995 enabled forensic scientists to set up the first national DNA database which would hold both personal DNA profiles together with results obtained from crime scenes. Other countries quickly followed but in some the legislation has severely restricted the amount and type of data which can be retained and, therefore, effectiveness of the databases is limited. The widespread use of commercially produced multiplex kits has produced a situation in which nearly all European laboratories are using compatible systems and there is, therefore, the potential for the introduction of a pan-European DNA database. However, the exchange of results between countries is hampered by the various legislations which currently exist.

DNA Fingerprinting↗

Impact of different definitions on estimates of accuracy of the diagnosis data in a clinical database.

Computerized medical databases are increasingly used for research. The influence of different definitions of the accuracy of matching on the estimated accuracy of diagnosis data was assessed in a database of visits to a public pediatric clinic. Differences between definitions involved 1) unit of analysis, 2) number of diagnoses required to match per visit, and/or 3) whether database contents are required to match the medical record or medical record contents are required to be matched in the database. Overall, 90% of diagnoses in the database (391/435) were accurately coded relative to the medical record. Alternatively, 77% of diagnoses listed in the medical record (391/506) were accurately coded in the database. When individual visits were used as the unit of analysis, estimates of accuracy using six definitions ranged from 65% to 92%. The most appropriate definition to use for estimating accuracy of diagnosis data likely depends on the purpose of the study. Use of two or more such definitions may enhance portrayal of the accuracy of diagnosis data.

Algorithms↗

A new way of building a database of EEG findings.

Whereas computer-based electroencephalography (EEG) is widely applied, the EEG interpretations are usually not stored in a way that favours exploitation of modern computer technology. This paper reports an EEG description system facilitating categorization of EEG data in a computerized database. The system interactively communicates with the digital EEG system and also with the general patient administrative system. The main new quality of this system is the methods for data input and automatic data retrieval from several systems, rather than the establishment of a database of EEG data itself. The EEGs are visually analysed and categorized. Manually marked EEG events are automatically transferred to the database and such events as well as defined electrode positions within these epochs are directly linked to their corresponding descriptions. The database is updated without demand for filling in the events in the database in a second operation. Thereby, the EEG interpreter builds the database while analysing the EEG. This system provides an improved accessibility of EEG data for clinical, normative, educational and scientific use.

Brain↗

Method to correlate tandem mass spectra of modified peptides to amino acid sequences in the protein database.

A method to correlate uninterpreted tandem mass spectra of modified peptides, produced under low-energy (10-50 eV) collision conditions, with amino acid sequences in a protein database has been developed. The fragmentation patterns observed in the tandem mass spectra of peptides containing covalent modifications is used to directly search and fit linear amino acid sequences in the database. Specific information relevant to sites of modification is not contained in the character-based sequence information of the databases. The search method considers each putative modification site as both modified and unmodified in one pass through the database and simultaneously considers up to three different sites of modification. The search method will identify the correct sequence if the tandem mass spectrum did not represent a modified peptide. This approach is demonstrated with peptides containing modifications such as S-carboxymethylated cysteine, oxidized methionine, phosphoserine, phosphothreonine, or phosphotyrosine. In addition, a scanning approach is used in which neutral loss scans are used to initiate the acquisition of product ion MS/MS spectra of doubly charged phosphorylated peptides during a single chromatographic run for data analysis with the database-searching algorithm. The approach described in this paper provides a convenient method to match the nascent tandem mass spectra of modified peptides to sequences in a protein database and thereby identify previously unknown sites of modification.

Algorithms↗

Suitability of molecular descriptors for database mining. A comparative analysis.

Database mining methods rely on the molecular descriptors used to characterize a structural database. In the present investigation, five different types of descriptors (log P, UNITY fingerprints, ISIS keys, VolSurf, and GRIND) are applied to characterize various databases (n = 1007, 100, and 229) comprising drugs almost exclusively. The validity of the descriptors is comparatively analyzed via principal component analysis and its hierarchical variant, consensus principal component analysis. Both pharmacodynamic and pharmacokinetic aspects of database mining are treated. For pharmacodynamic aspects, clustering behavior achieved with the different descriptors is tested on the chemically homogeneous beta-blockers, benzodiazepines, and penicillins and on the chemically more diverse class I antiarrhythmics. The following ranking is observed: UNITY fingerprints > ISIS keys and GRIND > VolSurf > log P. Regarding information content, the CPCA superweight plot indicates similarity between fingerprints and ISIS keys as well as between VolSurf and log P, while GRIND differs from all the remaining descriptors. Solubility data and blood/brain barrier penetrating behavior serve as test cases for pharmacokinetic aspects. Comparison of the descriptors applied to these data reveals that VolSurf has the most realistic and consistent behavior, GRIND shows intermediate behavior, while UNITY fingerprints and ISIS keys are not well suited for pharmacokinetic profiling. From this comparative analysis, we conclude that VolSurf descriptors exhibit particular advantages in treating pharmacokinetic aspects; UNITY fingerprints, ISIS keys, and GRIND descriptors are of special value for tackling pharmacodynamic aspects of database mining. The parameter log P is of limited applicability in database mining because of rather poor reliability and lack of completeness of data.

Computing Methodologies↗

Database diversity assessment: new ideas, concepts, and tools.

We present some new ideas for characterizing and comparing large chemical databases. The comparison of the contents of large databases is not trivial since it implies pairwise comparison of hundreds of thousands of compounds. We have developed methods for categorizing compounds into groups or series based on their ring-system content, using precalculated structure-based hashcodes. Two large databases can then be compared by simply comparing their hashcode tables. Furthermore, the number of distinct ring-system combinations can be used as an indicator of database diversity. We also present an independent technique for diversity assessment called the saturation diversity approach. This method is based on picking as many mutually dissimilar compounds as possible from a database or a subset thereof. We show that both methods yield similar results. Since the two methods measure very different properties, this probably says more about the properties of the databases studied than about the methods.

Benzene Derivatives↗

Identification of protein functions from a molecular surface database, eF-site.

A bioinformatics method was developed to identify the protein surface around the functional site and to estimate the biochemical function, using a newly constructed molecular surface database named the eF-site (electrostatic surface of Functional site. Molecular surfaces of protein molecules were computed based on the atom coordinates, and the eF-site database was prepared by adding the physical properties on the constructed molecular surfaces. The electrostatic potential on each molecular surface was individually calculated solving the Poisson-Boltzmann equation numerically for the precise continuum model, and the hydrophobicity information of each residue was also included. The eF-site database is accessed by the internet (http://pi.protein.osaka-u.ac.jp/eF-site/). We have prepared four different databases, eF-site/antibody, eF-site/prosite, eF-site/P-site, and eF-site/ActiveSite, corresponding to the antigen binding sites of antibodies with the same orientations, the molecular surfaces for the individual motifs in PROSITE database, the phosphate binding sites, and the active site surfaces for the representatives of the individual protein family, respectively. An algorithm using the clique detection method as an applied graph theory was developed to search of the eF-site database, so as to recognize and discriminate the characteristic molecular surfaces of the proteins. The method identifies the active site having the similar function to those of the known proteins.

Antibodies↗

Development of a database of functional assessment measures related to work disability.

The development of the Functional Assessment Measures Database is described. The database provides a method to organize and search for measures that are used to assess the functional abilities of people with medical impairments to determine work disability. The project identified 4,200 different measures that are used in the functional assessment of persons with disability across the life span, 812 of which are used to evaluate adults in terms of work disability. The database has 3,033 scales that are found in 633 measures. In the database, each measure is described and is linked to at least one functional assessment construct. The use of the database in the Social Security Administration Redesign Project is described. Other possible uses for the database are presented.

Databases, Factual↗

Allergen databases.

Allergies represent a significant medical and industrial problem. Molecular and clinical data on allergens are growing exponentially and in this article we have reviewed nine specialized allergen databases and identified data sources related to protein allergens contained in general purpose molecular databases. An analysis of allergens contained in public databases indicates a high level of redundancy of entries and a relatively low coverage of allergens by individual databases. From this analysis we identify current database needs for allergy research and, in particular, highlight the need for a centralized reference allergen database.

Allergens↗

Population of the HLA ligand database.

We have established an HLA ligand database to provide scientists and clinicians with access to Major Histocompatibility Complex (MHC) class I and II motif and ligand data. The HLA Ligand Database is available on the world wide web at http://hlaligand.ouhsc.edu and contains ligands that have been published in peer-reviewed journals. HLA peptide datasets prove useful in several areas: ligands are important as targets for various immune responses while algorithms built upon ligand datasets allow identification of new peptides without time-consuming experimental procedures. A review of the HLA class I ligands in the database identifies strengths and deficiencies in the database and, therefore, the utility of the dataset for identifying new peptides. For instance, 212 HLA-A phenotypes exist of which 23 have a motif determined and 43 have peptides characterized. In terms of number of ligands, HLA-A*0201 has 258 characterized ligands, A*1101 has 25 peptides, while the remaining two-thirds of the HLA-A phenotypes have less than 10 associated peptide sequences. Characterization of ligands and motifs remains roughly the same at the HLA-B locus while the peptides of the HLA-C locus tend to be less characterized. These data show that 74% of HLA class I molecules do not have ligands represented in the database and thus algorithms based on the dataset could not predict ligands for a majority of the US population. Building upon this dataset and knowledge of HLA allelic frequencies, it is possible to plan a systematic expansion of the HLA class I ligand database to better identify ligands useful throughout the population.

Databases, Protein↗

Overview of large database analysis in renal transplantation.

The discipline of renal transplantation has been fortunate in having one of the largest and most complete databases of any field of medical inquiry. Many analyses have been performed utilizing these databases and a wide array of views regarding the role and limitations of these analyses exists. In this manuscript, we hope to present the merits and limitations of large database analysis in renal transplantation. In addition, we will attempt to go over the major databases' structures and the critical issues of coding, verification and clinical awareness that help maintain the integrity of this type of analysis. Other fundamental issues that will be covered include the relationship of single center and randomized prospective studies with database analyses, statistical considerations, and hidden selection bias inherent in registry analysis. We hope to present both a guide to investigators who are first embarking on large-scale database analysis and also to present a fair view of the utility and limitations of a potentially very useful research tool in renal transplantation.

Databases, Factual↗

Validity of self-reported energy intake in lean and obese young women, using two nutrient databases, compared with total energy expenditure assessed by doubly labeled water.

OBJECTIVE: To compare self-reported total energy intake (TEI) estimated using two databases with total energy expenditure (TEE) measured by doubly labeled water in physically active lean and sedentary obese young women, and to compare reporting accuracy between the two subject groups. DESIGN: A cross-sectional study in which dietary intakes of women trained in diet-recording procedures were analyzed using the Minnesota Nutrition Data System (NDS; versions 2.4/6A/21, 2.6/6A/23 and 2.6/8.A/23) and Nutritionist III (N3; version 7.0) software. Reporting accuracy was determined by comparison of average TEI assessed by an 8 day estimated diet record with average TEE for the same period. RESULTS: Reported TEI differed from TEE for both groups irrespective of nutrient database (P<0.01). Measured TEE was 11.10+/-2.54 and 11.96+/-1.21 MJ for lean and obese subjects, respectively. Reported TEI, using either database, did not differ between groups. For lean women, TEI calculated by NDS was 7.66+/-1.73 MJ and by N3 was 8.44+/-1.59 MJ. Corresponding TEI for obese women were 7.46+/-2.17 MJ from NDS and 7.34+/-2.27 MJ from N3. Lean women under-reported by 23% (N3) and 30% (NDS), and obese women under-reported by 39% (N3) and 38% (NDS). Regardless of database, lean women reported higher carbohydrate intakes, and obese women reported higher total fat and individual fatty acid intakes. Higher energy intakes from mono- and polyunsaturated fatty acids were estimated by NDS than by N3 in both groups of women (P< or =0.05). CONCLUSIONS: Both physically active lean and sedentary obese women under-reported TEI regardless of database, although the magnitude of under-reporting may be influenced by the database for the lean women. SPONSORSHIP: USDA Hatch Project award (ARZT-136528-H-23-111) to LB Houtkooper and WH Howell.

Adolescent↗

Use of the UK General Practice Research Database for pharmacoepidemiology.

The last decade has seen a surge in the use of computerized health care data for pharmacoepidemiology. Of all European databases, the General Practice Research Database (GPRD) in the UK, has been the most widely used for pharmacoepidemiological research. Since 1994, this database has belonged to the UK Department of Health, and is maintained by the Office of National Statistics (ONS). Currently, around 1500 general practitioners with a population coverage in excess of 3 million, systematically provide their computerized medical data anonymously to ONS. Validation studies of the GPRD have documented the recording of medical data into general practitioners' computers to be near to complete. The GPRD collects truly population-based data, has a size that makes it possible to follow-up large cohorts of users of specific drugs, and includes both outpatient and inpatient clinical information. The access to original medical records is excellent. Desirable improvements to the GPRD would be additional computerized information on certain variables and linkage to other health care databases. Most published studies to date have been in the area of drug safety. The General Practice Research Database has proved that valuable data can be collected in a general practice setting. The full potential of this rich computerized database has yet to come. This experience should serve to encourage others to develop similar population-based data in other countries.

Databases, Factual↗

A scientific relational database combined with a report generator for endoscopy in networks: EndoNet.

BACKGROUND AND STUDY AIMS: The flexibility required in academic endoscopy units is not provided by the available database systems. In a project involving substantial cooperation between endoscopists and computer scientists, we have developed an adaptable database, combined with a report generator embedded in the hospital's intranet. PATIENTS AND METHODS: Six workstations in different areas of the hospital were clustered with a UNIX operating system to implement multi-user capability and access control. A relational database was used to design an application appropriate to the specific needs of the endoscopy unit in a teaching hospital engaged in scientific research. Both the terminology used in standardized endoscopy nomenclature and a free text block facility were included. A graphical user interface was developed to assemble pertinent data, generate the reports, and supervise the database. RESULTS: A total of 4936 examinations including 2988 patients were entered consecutively during continuous routine operation of the system. Complete report generation required five minutes (median; range 1-9 minutes). Both structured items and free text were used in all the reports. Querying of the database was possible, concerning matters such as the need for repeated endoscopic therapy in acute gastrointestinal bleeding (4%), the search for Helicobacter pylori in appropriate patients (64%), the rate of accidental pancreatic duct visualization in endoscopic retrograde cholangiography (24%), and links between examinations and active trials (2%). Indicating improved report quality, the number and the diameter of esophageal varices in patients with varices were more frequently reported with the new report system than with previous typed reports (P<0.001). An anonymous questionnaire revealed that the readability of the computer-generated reports was better than that of the previous typewritten reports (P=0.01). CONCLUSIONS: This report describes the creation of a database application and a report generator meeting the needs of scientific and routine use, and the successful application of this system in an academic endoscopy unit.

Computer Communication Networks↗

Using the NTP database to assess the value of rodent carcinogenicity studies for determining human cancer risk.

The large database of carcinogenicity results generated by the National Toxicology Program (NTP) provides a unique opportunity to critically evaluate important scientific issues such as (1) the frequency of positive outcomes, (2) the interspecies correlation in carcinogenic response between rats and mice, (3) the correlation between body weight and tumor incidence, (4) estimates of the false-positive and false-negative rates, and (5) the frequency of decreasing tumor incidences. Such database evaluations enable us to better understand the value and limitations of rodent carcinogenicity studies for determining human cancer risk. However, as the NTP database becomes increasingly accessible to the general scientific community, there is also increased opportunity for misuse of the database. This article reexamines and updates previous database evaluations, presents four scientific principles that should be employed by anyone attempting to use this database, and illustrates how failure to apply these principles can lead to misleading results.

Animals↗

Access to databases in complementary medicine.

Access to medical databases is a keystone for obtaining up-to-date and complete information for physicians. In the last few years, the rapid growth of the World Wide Web has given rise to an information revolution, enabling health care providers to gain access (often free) to an expanding volume of information that was previously inaccessible. Search engines and online databases assist the search for health information. In this article we examine the biomedical databases of primary interest in the field of alternative and complementary medicine, dividing them into Web accessible and nonaccessible databases and emphasizing the freely available ones. A further classification is major biomedical bibliographic databases specific to complementary medicine, and dedicated therapy or modality-specific databases.

Complementary Therapies↗