Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Scriptable access to the Caenorhabditis elegans genome sequence and other ACEDB databases.

Much of the world's genomic data are available to the community through networked databases that are accessed via Web interfaces. Although this paradigm provides browse-level access and has greatly facilitated linking between databases, it does not provide any convenient mechanism for programmatically fetching and integrating data from diverse databases. We have created a library and an application programming interface (API) named AcePerl that provides simple, direct access to ACEDB databases from the Perl programming language. With this library, programmers and computer-savvy biologists can write software to pose complex queries on local and remote ACEDB databases, retrieve the data, integrate the results, and move data objects from one database to another. In addition, a set of Web scripts running on top of AcePerl provides Web-based browsing of any local or remote ACEDB database. AcePerl and the AceBrowser Web browser run on Unix systems and are available under a license that allows for unrestricted use and redistribution. Both packages can be downloaded from URL. A Microsoft Windows port of AcePerl is in the planning stages.

Animals↗

Expanding vaginal microbiome pangenomes via a custom MIDAS database reveals Lactobacillus crispatus accessory genes associated with cervical dysplasia.

The vaginal microbiome plays a central role in reproductive health. Vaginal microbiome dysbiosis is associated with many adverse reproductive health outcomes, but most studies have focused on associations at the species level. The potential contribution of intraspecies microbial variation, especially gene content differences across bacterial strains, remains underexplored in reproductive health contexts. The Metagenomic Intra-Species Diversity Analysis (MIDAS) framework enables such analyses, but depends on comprehensive reference databases. We constructed a MIDAS-compatible pangenome database from over 18,000 genomes in the Vaginal Microbiome Genome Collection (VMGC). Compared to the Genome Taxonomy Database (GTDB)-derived reference, the VMGC-derived database expanded the pangenomes of prevalent vaginal species, better capturing vaginal-specific intraspecies diversity. Applying this database to vaginal samples from a cervical dysplasia cohort, we identified 13 Lactobacillus crispatus accessory genes significantly associated with cervical dysplasia, including a HicAB toxin-antitoxin system, three transcriptional regulators, and three phage-derived genes. These findings highlight the utility of body site-specific reference resources and shotgun metagenomic sequencing for uncovering intraspecies microbial variation relevant to reproductive health.IMPORTANCEThe vaginal microbiome plays a critical role in reproductive health, and different bacteria from the same species can carry different genes that influence how the strains interact with the host and other microbes. These strain-level differences are often overlooked when microbiomes are analyzed only at the species level. Existing genomic reference databases are heavily biased toward gut and environmental bacteria, leaving the genetic diversity of vaginal microbes understudied. We built a specialized reference database from over 18,000 vaginal bacterial genomes that better reflects this diversity. We then applied this resource to quantify gene-level variation in vaginal samples from a cervical dysplasia cohort. Focusing on Lactobacillus crispatus, a prevalent and often beneficial vaginal species, we identified 13 genes that were more common in women with cervical dysplasia than in controls. This work demonstrates that body site-specific genomic resources are essential for uncovering strain-level bacterial differences relevant to reproductive health.

Lactobacillus crispatus↗

Nutrient-related analysis of pathway/genome databases.

We present an algorithm that solves two related problems in the analysis of metabolic networks stored within a pathway/genome database. (1) The Forward Propagation Problem: given a set of nutrients that are inputs to the metabolic network, what compounds will be produced by the metabolic network? (2) The Backtracking Problem: given the results of a forward propagation, and given a set of essential compounds that are not produced as a result of the forward propagation, what precursors must be supplied to produce those essential compounds? A program based on this algorithm is applied to the EcoCyc database, which is a pathway/genome database for E. coli that consists of annotated genomes and the metabolic reactions and pathways associated with the known gene products. The inputs to the program are a description of the metabolic network of an organism (EcoCyc), a set of nutrients corresponding to a known minimal growth medium, and a list of essential compounds to be produced. The program "fires" the microorganism's metabolism contained in the database and predicts all synthesized and nonsynthesized essential compounds, along with the missing precursors required to produce the latter. When applied to the EcoCyc database, the program identifies a number of missing precursors that indicate incomplete regions of the database. Thus the program results can be used to evaluate existing pathway databases like EcoCyc.

Algorithms↗

Carcinogenicity evaluations and ongoing studies: the IARC databases.

Many thousands of chemicals are produced industrially and many more occur naturally. Information on the toxicology of these chemicals is often minimal or absent. The International Agency for Research on Cancer (IARC) has published evaluations of the carcinogenic risk to humans of over 700 chemicals, groups of chemicals, and complex mixtures as a regular series of monographs. A database has been created containing summaries of all the relevant epidemiological, animal carcinogenicity, and other relevant biological data for each chemical or mixture evaluated. Additional databases have been created for ongoing epidemiological studies of cancer in humans and for long-term carcinogenicity studies in rodents, as well as a database containing information on genotoxic and related effects of chemicals. Some of these databases have been published in print form. IARC now plans to publish them electronically, together with other databases, in the form of a CDROM (compact disk, read-only memory). The objective will be to make the entire IARC database of cancer information as widely available as possible in an integrated format conducive to efficient and combined exploitation of all the component databases.

Animals↗

Flicker image comparison of 2-D gel images for putative protein identification using the 2DWG meta-database.

With the availability of two-dimensional (2-D) gel electrophoresis databases that have many characterized proteins, it may be possible to compare a researcher's gel images with those in relevant databases. This may lead to the putative identification of unknown protein spots in a researcher's gel with those characterized in a given database, saving the researcher time and money by suggesting monoclonal antibodies to try in confirming these identifications. We have developed two tools to help with this comparison: (1) Flicker, http:/(/)www.lecb.ncifcrf.gov/flicker/, a Java applet program running in the researcher's Web browser, to visually compare their gels against gels on the Internet; and (2) the 2DWG meta-database, http:/(/)www.lecb.ncifcrf.gov/2dwgDB /, a searchable database of locations of 2-D electrophoretic gel images found on the Internet. Recent additions to Flicker allow users to click on a protein spot in a gel that is linked to a federated 2D gel database, such as SWISS-2DPAGE, and have it retrieve a report from that Web database for that protein.

Data Display↗

The role of insurance claims databases in drug therapy outcomes research.

The use of insurance claims databases in drug therapy outcomes research holds great promise as a cost-effective alternative to post-marketing clinical trials. Claims databases uniquely capture information about episodes of care across healthcare services and settings. They also facilitate the examination of drug therapy effects on cohorts of patients and specific patient subpopulations. However, there are limitations to the use of insurance claims databases including incomplete diagnostic and provider identification data. The characteristics of the population included in the insurance plan, the plan benefit design, and the variables of the database itself can influence the research results. Given the current concerns regarding the completeness of insurance claims databases, and the validity of their data, outcomes research usually requires original data to validate claims data or to obtain additional information. Improvements to claims databases such as standardisation of claims information reporting, addition of pertinent clinical and economic variables, and inclusion of information relative to patient severity of illness, quality of life, and satisfaction with provided care will enhance the benefit of such databases for outcomes research.

Clinical Trials as Topic↗

The extraction of quality-of-care clinical indicators from State health department administrative databases.

OBJECTIVE: To assess whether three proposed quality-of-care indicators (unplanned readmissions, hospital-acquired bacteraemia, and postoperative wound infection) can be accurately identified from State health department databases. DESIGN: Algorithms were applied to State health department databases to maximise the identification of individuals potentially positive for each indicator. Records of these patients were then examined to determine the percentage of cases that met the precise indicator definitions. SETTING: 10 public, acute-care hospitals from Victoria, South Australia and New South Wales. Data from the 1994-95 and 1995-96 financial years were collected. PARTICIPANTS: Individuals 18 years of age or older who were identified from State health department administrative databases as potentially meeting the indicator criteria. MAIN OUTCOME MEASURES: The proportion of screened cases that met the precise indicator definitions, and the elements of the indicator definitions which could not be extracted from the administrative databases. RESULTS: The proportions of cases confirmed by medical record review to be positive for the indicator events were 76.3% for unplanned readmissions within 28 days, 20% for hospital-acquired bacteraemia, 43.5% for wound infections after clean surgery, and 34.8% for wound infections after contaminated surgery. The clinical elements of each indicator definition were not easily extracted from the administrative databases. CONCLUSIONS: The three proposed clinical indicators could not be extracted from current State health department databases without an extensive process of secondary medical record review. If administrative databases are to be used for assessing quality of care, more systematic recording of data is needed.

Algorithms↗

Identifying diabetes mellitus or heart disease among health maintenance organization members: sensitivity, specificity, predictive value, and cost of survey and database methods.

We conducted a study of the sensitivity, specificity, positive predictive value, and cost of two methods of identifying diagnosed diabetes mellitus or heart disease among members of a health maintenance organization (HMO). Among 3186 adult HMO members who were attending one primary care clinic, 2326 were reached for a telephone survey (efficiency = 0.73). Among these members, 1991 answered standardized questions to ascertain whether they had diabetes or heart disease (corrected response rate = 0.85). Linkage was then made to computerized diagnostic databases. By means of both a database method and a survey method, the 1976 members with complete data for analysis were classified as having or not having diabetes or heart disease. When results with the two methods disagreed, charts were reviewed to confirm the presence or absence of diabetes or heart disease. Diabetes was identified among 4.7% of adult members, and heart disease was identified among 3.7%. Identification of diabetes differed between the database method and the survey method (sensitivity 0.91 vs 0.98, specificity 0.99 vs 0.99, positive predictive value 0.94 vs 0.83). Identification of heart attach history was similar for the database method and the survey method (sensitivity 0.89 vs 0.95, specificity 0.99 vs 0.99, positive predictive value 0.79 vs 0.81). The cost of obtaining data was $13.50 per member for the survey method and $0.30 per member for the database method. Database methods or survey methods of identifying selected chronic diseases among HMO members may be acceptable for various purposes, but database identification methods appear to be less expensive and provide information on a higher proportion of HMO members than do survey methods. Accurate identification of chronic diseases among patients supports clinic-level measures for clinical improvement, research, and accountability.

Adult↗

Designing an international industrial hygiene database of exposures among workers in the asphalt industry.

OBJECTIVES: The objective of this project was to construct a database of exposure measurements which would be used to retrospectively assess the intensity of various exposures in an epidemiological study of cancer risk among asphalt workers. METHODS: The database was developed as a stand-alone Microsoft Access 2.0 application, which could work in each of the national centres. Exposure data included in the database comprised measurements of exposure levels, plus supplementary information on production characteristics which was analogous to that used to describe companies enrolled in the study. RESULTS AND DISCUSSION: The database has been successfully implemented in eight countries, demonstrating the flexibility and data security features adequate to the task. The database allowed retrieval and consistent coding of 38 data sets of which 34 have never been described in peer-reviewed scientific literature. We were able to collect most of the data intended. As of February 1999 the database consisted of 2007 sets of measurements from persons or locations. The measurements appeared to be free from any obvious bias. CONCLUSIONS: The methodology embodied in the creation of the database can be usefully employed to develop exposure assessment tools in epidemiological studies.

Data Collection↗

[Surveillance of communicable diseases using a computer database of reported cases].

Epidemiology services during the surveillance of communicable diseases collects of different sorts of data, which are used for an analysis of epidemiologic situation. Those data are the starting point for timeline control and preventive activities. Data processing of notified communicable diseases cases provides information on types of diseases, number of cases, time and place of their occurrence. Manual data processing, used till 1993, was slow, unreliable and considerably decreased the efficiency of epidemiology service activities. In this paper we have set the hypothesis that is possible to form a computerized database with the following aims: to form user friendly computerized database model for those without knowledge in using computers: to get output spread sheets with information needed for epidemiologic situation analyses at any time. Database was developed in 1993 and has been used as source of the information in epidemiologic diagnosis process. The significant accuracy, reliability, timelines, and shortening of the time of data processing was achieved. The database can also serve as the initial component for designing an epidemiologic services information network in Belgrade county. In designing such a network it is necessary to form the additional databases of isolated infectious agents and their drug resistance, database of health status of persons under surveillance and database of environmental and sanitary condition in children and youth facilities.

Communicable Disease Control↗

The Southern Alberta Renal Program database: a prototype for patient management and research initiatives.

The Southern Alberta Renal Program (SARP) database was developed to respond to an urgent need for local information on clinical outcomes, laboratory information, and health care costs, and to enable our local renal program to monitor the implementation of established clinical practice guidelines. The database captures detailed demographic, clinical, and laboratory information and is unique by also capturing comorbidity, health-related quality of life and costing information for patients with end-stage renal disease (ESRD) in southern Alberta, storing the information in one common database. By collecting information on patient comorbidity, health outcomes and costs, the SARP database has enabled many quality assurance initiatives as well as research opportunities for projects involving patients with ESRD. Due to the availability of links with other available local clinical and administrative databases, information is collected with a minimal need for manual data entry. This type of database is a method by which health programs could improve the quality of patient care. Programs caring for patients with chronic medical conditions such as ESRD should examine how computer databases could assist in clinical care and improve the efficiency with which that care is delivered to their patients.

Acute Kidney Injury↗

[A new database system for radiological reports].

We have designed and developed a new database system to facilitate automatic feedback of the content of radiology reports to radiologists. The prototype of this database system has been implemented in the RGSS-IDJ, a developmental computer system that applies artificial intelligence methods to a reporting system. This prototype system was constructed to test the feasibility of overcoming the limitations of conventional database systems. The new database system is based on our semantic model for radiology reports and is able to treat data with unnormalized relations. Operations specific to our database system include the ability to acquire information about a set of reports that contains any semantic expression included in the lexicon and the ability to obtain the expressions that belong to a set of several semantic expressions in the reports. Thus, our new database system will offer a more powerful tool for analyzing the content of reports than conventional database systems.

Databases, Bibliographic↗

A virtual repository approach to clinical and utilization studies: application in mammography as alternative to a national database.

A national mammography database was proposed, based on a centralized architecture for collecting, monitoring, and auditing mammography data. We have developed an alternative architecture relying on Internet-based distributed queries to heterogeneous databases. This architecture creates a "virtual repository", or a federated database which is constructed dynamically, for each query and makes use of data available in legacy systems. It allows the construction of custom-tailored databases at individual sites that can serve the dual purposes of providing data (a) to researchers through a common mammography repository and (b) to clinicians and administrators at participating institutions. We implemented this architecture in a prototype system at the Brigham and Women's Hospital to show its feasibility. Common queries are translated dynamically into database-specific queries, and the results are aggregated for immediate display or download by the user. Data reside in two different databases and consist of structured mammography reports, coded per BIRADS Standardized Mammography Lexicon, as well as pathology results. We prospectively collected data on 213 patients, and showed that our system can perform distributed queries effectively. We also implemented graphical exploratory analysis tools to allow visualization of results. Our findings indicate that the architecture is not only feasible, but also flexible and scaleable, constituting a good alternative to a national mammography database.

Computer Communication Networks↗

[The organization of the database and data flow in mass screening for cervical cancer].

Mass screening, because of very many potential patients, requires storing and processing a great deal of medical and population information. That is why it should be supported not only by human resources but by computer techniques as well. The example of a computer science application in medicine is Populations Database System (PDB) which was designed and implemented in the Department of Institute of Mother and Child in Białystok. The aim of this work is to evaluate PDB System's effectiveness in mass screening for cervical cancer. Population database contains several standard database files (DBF) and indexes. All the data is organized as a relational database. Every data relationship is at least in 1NF (first normal form). Functional dependency holds for the structures of database. Because of great variety of stored data it was essential to design how to enter information and how to combine database files to avoid redundancy. It has particular importance for the special functions of system, for example printing and sending individual invitation for an examination. In addition the system can realize all standard database functions and some statistical analysis. Special attention was paid to the problem of data security which is particularly important for medical information. Thanks to PDB system we could realize mass and active screening for cervical cancer in Białystok. Without computer techniques it would be impossible to store, process and interpret so much data.

Databases as Topic↗

DBGET/LinkDB: an integrated database retrieval system.

The integrated database retrieval system DBGET/LinkDB is the backbone of the Japanese GenomeNet service. DBGET is used to search and extract entries from a wide range of molecular biology databases, while LinkDB is used to search and compute links between entries in different databases. DBGET/LinkDB is designed to be a network distributed database system with an open architecture, which is suitable for incorporating local databases or establishing a specialized server environment. It also has an advantage of simple architecture allowing rapid daily updates of all the major databases. The WWW version of DBGET/LinkDB at GenomeNet is integrated with other search tools, such as BLAST, FASTA and MOTIF, and with local helper applications, such as RasMol. In addition to factual links between database entries, LinkDB is being extended to included similarity links and biological links toward computerization of logical reasoning processes.

Databases, Factual↗

Proclass protein family database: new version with motif alignments.

ProClass is a protein family database which organizes non-redundant sequence entries into families defined collectively by the ProSite patterns and PIR superfamilies. The database consists of about 100,000 entries, more than half of which are classified in about 3,000 families. The new version includes links to various protein family/domain and structural class databases and contains gapped motif alignments for all ProSite patterns. The motif sequences are retrieved from both SwissProt and PIR-international databases, including numerous new members detected by our GeneFIND family identification system. The motif collection represents a 50% increase from those catalogued in ProSite. The ProClass database can be used to maximize family information retrieval, help organize protein sequence databases, and support full-scale genomic annotation. The database and its query program are freely available for on-line record retrieval and direct file transfer from our WWW server at http:/(/)diana.uthct.edu/proclass.html+ ++.

Amino Acid Sequence↗

Large scale database scrubbing using object oriented software components.

Now that case managers, quality improvement teams, and researchers use medical databases extensively, the ability to share and disseminate such databases while maintaining patient confidentiality is paramount. A process called scrubbing addresses this problem by removing personally identifying information while keeping the integrity of the medical information intact. Scrubbing entire databases, containing multiple tables, requires that the implicit relationships between data elements in different tables of the database be maintained. To address this issue we developed DBScrub, a Java program that interfaces with any JDBC compliant database and scrubs the database while maintaining the implicit relationships within it. DBScrub uses a small number of highly configurable object-oriented software components to carry out the scrubbing. We describe the structure of these software components and how they maintain the implicit relationships within the database.

Confidentiality↗

Management of severe hypokalemia in hospitalized patients: a study of quality of care based on computerized databases.

BACKGROUND: While administrative databases are used to assess general indicators of quality of care, a detailed audit of the process of clinical care usually requires review of hospital medical records. OBJECTIVE: To evaluate the feasibility of assessing the management of severe hypokalemia using computerized administrative and laboratory databases. METHODS: The study included all patients hospitalized in 1997 who experienced serum potassium levels of less than 3.0 mmol/L at Hadassah University Hospital, Jerusalem, Israel, a tertiary care center. Using the computerized databases, we measured the following: (1) whether a subsequent serum potassium test was performed, (2) time to the subsequent test and to normalization of the serum potassium level, (3) achievement of normokalemia, and (4) in-hospital mortality. In a random subsample of 100 patients, these measures were compared with the blinded assessment of the quality of medical management of hypokalemia, as determined from medical records, using predetermined criteria for adequate management. RESULTS: The computerized databases revealed that severe hypokalemia occurred in 866 patients (2.6% of the yearly hospitalizations): 55 patients (6.4%) had no subsequent serum potassium levels measured, and 260 (30.0%) were discharged from the hospital with a subnormal potassium level. The mean time to a subsequent test was 20 hours, and to normokalemia, 50 hours; both intervals varied by department. In-hospital mortality was 20.4%, or 10-fold that of the entire hospitalized population. A review of hospital medical records revealed inadequate clinical management of hypokalemia in 24%, which was associated with nonperformance of a subsequent test (likelihood ratio, 8.4), failure to normalize the serum potassium level (likelihood ratio, 4.2), discharge from the hospital with a subnormal potassium level (likelihood ratio, 2.1), and in-hospital death (likelihood ratio, 2.5), all of which could be determined by the computerized databases. CONCLUSIONS: The computerized laboratory database is useful in ascertaining the prevalence of severe hypokalemia and in assessing shortcomings in its management. Databases can be used to derive valid and efficient measures of the quality of the clinical management of electrolyte disorders.

Clinical Laboratory Information Systems↗