Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “data integration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

International INtegrated Database for the Evaluation of severe sePsis and drotrecogin alfa (activated) THerapy: component trials and statistical methods for INDEPTH.

OBJECTIVES: To better understand the effects of drotrecogin alfa (activated) (DrotAA) in severe sepsis patients, and the natural progression of severe sepsis, by creating a database of severe sepsis patients using the appropriate statistical analysis methods to integrate data from various trials. PATIENTS AND METHODS: Patient-level data from five severe sepsis trials, conducted by the same sponsor (Eli Lilly and Company, Indianapolis, IN, USA), were combined in an integrated database. Patients from various studies were included and received either DrotAA at 24 microg/kg/h for 96 hours (n = 3228) or placebo (n = 1231), in addition to standard supportive care. The following adjustments to the analyses were made to allow for the combined, and thus non-randomized, nature of the data: (1) differences in observed outcomes between studies were investigated to assess the extent of study-to-study variation before combining study-level data across trials for statistical analysis; (2) random study effects were included in models for patient-level data to capture potential extraneous study-to-study variation; and (3) propensity scores were computed and included as covariates in models for patient-level data to adjust for the nonrandomized nature of the data. RESULTS: Baseline characteristics were similar across the studies, supporting the combination of study-level data across trials. Comparing aggregate event rates between the two treatment arms yielded a relative risk for mortality (DrotAA versus placebo) of 0.79 (95% confidence interval [CI] 0.71-0.88), p < 0.0001. For patient-level analyses, after adjustment for 13 independent variables and random study effects, the odds ratio for mortality in the DrotAA versus placebo patients was 0.71 (95% CI 0.59-0.86), p = 0.0003. With adjustment for 13 independent variables and propensity score, the odds ratio was 0.79 (95% CI 0.67-0.93), p = 0.006. Limitations of this integrated database include the modest total number of the trials in the database and the fact that only one component trial in the database contributed data from both placebo and DrotAA-treated patients. SUMMARY: A robust severe sepsis database was developed which will be suitable for future studies on the progression of severe sepsis and the mechanism of action of DrotAA. Initial analysis of data from INDEPTH provides additional evidence that treatment of severe sepsis patients with DrotAA is associated with a sustained survival advantage throughout 28-day follow-up.

Aged↗

Evolution of web services in bioinformatics.

Bioinformaticians have developed large collections of tools to make sense of the rapidly growing pool of molecular biological data. Biological systems tend to be complex and in order to understand them, it is often necessary to link many data sets and use more than one tool. Therefore, bioinformaticians have experimented with several strategies to try to integrate data sets and tools. Owing to the lack of standards for data sets and the interfaces of the tools this is not a trivial task. Over the past few years building services with web-based interfaces has become a popular way of sharing the data and tools that have resulted from many bioinformatics projects. This paper discusses the interoperability problem and how web services are being used to try to solve it, resulting in the evolution of tools with web interfaces from HTML/web form-based tools not suited for automatic workflow generation to a dynamic network of XML-based web services that can easily be used to create pipelines.

Computational Biology↗

A computerized maintenance management system's requirements for standard operating procedures.

From this review of the 6 aspects of opportunity for inconsistency to corrupt or skew the reliability of data, it becomes apparent why members of management must provide the standards of operation and use within the CMMS for their employees. The possibility of poor data integrity due to any one of these aspects may not be severe; however, the severity is compounded and inevitable when different aspects are combined. Responding to information collected through the CMMS can be effective only if the data are reliable. With SOPs, management has provided their personnel with the necessary tools to ensure department-wide consistency. Management cannot afford to allow any one [table: see text] individual to apply personal interpretations of the importance and requirements in their approach to using the CMMS. If this is permitted, the loss of integrity due to one individual's judgment grows rapidly when data are analyzed at the departmental level. Standard operating procedures go beyond creating a "how to" for the CMMS; they provide the critical elements for collecting responsible and reliable data.

Biomedical Engineering↗

Spatial registration of multichannel multi-subject fNIRS data to MNI space without MRI.

The registration of functional brain data to the common brain space offers great advantages for inter-modal data integration and sharing. However, this is difficult to achieve in functional near-infrared spectroscopy (fNIRS) because fNIRS data are primary obtained from the head surface and lack structural information of the measured brain. Therefore, in our previous articles, we presented a method for probabilistic registration of fNIRS data to the standard Montreal Neurological Institute (MNI) template through international 10-20 system without using the subject's magnetic resonance image (MRI). In the current study, we demonstrate our method with a new statistical model to facilitate group studies and provide information on different components of variability. We adopt an analysis similar to the single-factor one-way classification analysis of variance based on random effects model to examine the variability involved in our improvised method of probabilistic registration of fNIRS data. We tested this method by registering head surface data of twelve subjects to seventeen reference MRI data sets and found that the standard deviation in probabilistic registration thus performed for given head surface points is approximately within the range of 4.7 to 7.0 mm. This means that, if the spatial registration error is within an acceptable tolerance limit, it is possible to perform multi-subject fNIRS analysis to make inference at the population level and to provide information on positional variability in the population, even when subjects' MRIs are not available. In essence, the current method enables the multi-subject fNIRS data to be presented in the MNI space with clear description of associated positional variability. Such data presentation on a common platform, will not only strengthen the validity of the population analysis of fNIRS studies, but will also facilitate both intra- and inter-modal data sharing among the neuroimaging community.

Adult↗

Quantitative quality control in microarray experiments and the application in data filtering, normalization and false positive rate prediction.

Data preprocessing including proper normalization and adequate quality control before complex data mining is crucial for studies using the cDNA microarray technology. We have developed a simple procedure that integrates data filtering and normalization with quantitative quality control of microarray experiments. Previously we have shown that data variability in a microarray experiment can be very well captured by a quality score q(com) that is defined for every spot, and the ratio distribution depends on q(com). Utilizing this knowledge, our data-filtering scheme allows the investigator to decide on the filtering stringency according to desired data variability, and our normalization procedure corrects the q(com)-dependent dye biases in terms of both the location and the spread of the ratio distribution. In addition, we propose a statistical model for false positive rate determination based on the design and the quality of a microarray experiment. The model predicts that a lower limit of 0.5 for the replicate concordance rate is needed in order to be certain of true positives. Our work demonstrates the importance and advantages of having a quantitative quality control scheme for microarrays.

Algorithms↗

GenBank.

GenBank (R) is a comprehensive sequence database that contains publicly available DNA sequences for more than 119 000 different organisms, obtained primarily through the submission of sequence data from individual laboratories and batch submissions from large-scale sequencing projects. Most submissions are made using the BankIt (web) or Sequin programs and accession numbers are assigned by GenBank staff upon receipt. Daily data exchange with the EMBL Data Library in the UK and the DNA Data Bank of Japan helps ensure worldwide coverage. GenBank is accessible through NCBI's retrieval system, Entrez, which integrates data from the major DNA and protein sequence databases along with taxonomy, genome, mapping, protein structure and domain information, and the biomedical journal literature via PubMed. BLAST provides sequence similarity searches of GenBank and other sequence databases. Complete bimonthly releases and daily updates of the GenBank database are available by FTP. To access GenBank and its related retrieval and analysis services, go to the NCBI home page at: http://www.ncbi.nlm.nih.gov.

Animals↗

GenBank: update.

GenBank is a comprehensive database that contains publicly available DNA sequences for more than 140 000 named organisms, obtained primarily through submissions from individual laboratories and batch submissions from large-scale sequencing projects. Most submissions are made using the BankIt (web) or Sequin program and accession numbers are assigned by GenBank staff upon receipt. Daily data exchange with the EMBL Data Library in the UK and the DNA Data Bank of Japan helps ensure worldwide coverage. GenBank is accessible through NCBI's retrieval system, Entrez, which integrates data from the major DNA and protein sequence databases along with taxonomy, genome mapping, protein structure and domain information, and the biomedical journal literature via PubMed. BLAST provides sequence similarity searches of GenBank and other sequence databases. Complete bimonthly releases and daily updates of the GenBank database are available by FTP. To access GenBank and its related retrieval and analysis services, go to the NCBI home page at: http://www.ncbi.nlm.nih.gov.

Animals↗

Promoting transparency of long-term environmental decisions: the Hanford Decision Mapping System pilot project.

Nuclear waste cleanup is a challenging and complex problem that requires both scientific analysis and dialogue among a variety of stakeholders. This article describes an effort to develop an online information system that supports this analytic-deliberative dialogue by integrating cleanup information for the Hanford Site, and making it more "transparent." A framework for understanding and evaluating transparency guided system development. Working directly with stakeholders, we identified information needs and developed new ways to organize and present the information so that it would be more transparent to interested parties, with the ultimate aim of fostering greater participation in decision dialogues and processes. The complexity of the information needed for dialogue suggested that several types of communication devices ("information structures") were warranted. Five information structures were developed for the pilot Decision Mapping System (http://nalu.geog.washington.edu/dms). Decision maps hyperlinked decision information to maps of Hanford. Background Information provided context in a narrative format. Decision Paths organized decision process information on a timeline and provided direct hyperlinks to online documentation. The Geographic Library hyperlinked decision documents to maps. Finally, a Discussion Forum allowed users to make comments and view remarks from others. Early lessons from this work suggest that transparency is integral to long-term management, a participatory design process contributed greatly to its perceived success, and better data integration to support decision making is needed. This work has broad implications for risk communicators and risk managers because it speaks to the design of information systems to support "analytic-deliberative" decision processes (i.e., those that rely upon both risk science and public dialogue).

Communication↗

Harnessing the Power of Large Language Models for Drug Discovery: A Systematic Review of Current Applications and Future Directions.

INTRODUCTION: The demand for inventive approaches to drug discovery has increased due to the rising costs, time, and failure rates in pharmaceutical research. Large Language Models (LLMs), with their sophisticated natural language processing and generative capabilities, have become potent instruments that have the potential to revolutionize biomedical research. The function of LLMs in different phases of drug development is methodically examined in this article. METHODS: The PRISMA 2020 principles were adhered to in this systematic study. A thorough search for research published between 2018 and 2025 was done using PubMed, Scopus, Web of Science, and Google Scholar. The search terms "large language model," "transformer," "drug discovery," and important sub-domains (such as "de-novo design" and "ADMET") were merged, and two reviewers independently screened the results. Predetermined inclusion and exclusion criteria were used to filter studies for relevance. 98 studies out of the 1,285 records that were initially retrieved met the requirements for the final qualitative synthesis. RESULTS: 98 studies that demonstrated the use of LLMs in various drug discovery domains were found during the review. These covered molecular generation, genomics, protein-ligand modeling, ADME/T and toxicity profiling, drug-target interaction and DTI prediction, and biomedical text mining. 42 different LLM-based tools were mapped, including BioBERT, SciSpacy, Drug- LLM, DNA-BERT, GPT-4, and ChatGPT. Predictive accuracy, hypothesis creation, target prioritization, and multi-modal data integration all showed notable gains with these techniques. DISCUSSION: By providing scalable, precise, and effective solutions for data-driven drug discovery, LLMs are revolutionizing the pharmaceutical industry. They allow for the creation of hypotheses and individualized insights across multi-modal biological data, and they perform better than conventional approaches in a number of subdomains. Improvements in performance were task-dependent; the most consistent gains occurred for biomedical text mining, disease-genedrug relationship mapping and drug-target interaction prediction tasks. Yet most evidence for clinical applications is still derived from retrospective studies and benchmark datasets, suggesting a higher need for prospective validation. CONCLUSION: There is revolutionary potential in incorporating LLMs into drug discovery processes. Clinical translation and regulatory uptake will depend heavily on collaborative validation, ethical deployment, and standardization as models become more multimodal and interpretable. Before normal use, extensive prospective benchmarking and head-to-head comparisons with established chemoinformatics pipelines are necessary.

De novo design↗

An integrated model for cellular analysis.

We present the MOlecular NETwork (MONET) ontology as a model to integrate data from different networks that govern cell function. To achieve this, different existing ontologies were analyzed and an integrated ontology was built in a way to make it possible to share and reuse knowledge, support interoperability between systems, and also allow the formulation of hypotheses through inferences. By studying the cell as an entity of a myriad of elements and networks of interactions, we aim to offer a means to understand the large-scale characteristics responsible for the behavior of the cell and to enable new biological insights.

Algorithms↗

High-throughput DNA sequencing on a capillary array electrophoresis system.

A capillary array electrophoresis apparatus capable of running and analyzing 48 DNA sequencing samples simultaneously has been constructed. The instrument uses a replaceable sieving buffer and incorporates a convenient method for introducing the buffer into the capillaries. Data from laser-induced fluorescence are collected as four separate images, one for each optical channel. The integrated data analysis software employs an open architecture that allows use of any DNA base-calling algorithm. DNA sequencing runs are completed in approx. 1 hr (approximately 500 bases), and instrument turnaround time between runs is less than 15 min. Overall, the instrument throughput is on the order of 720 templates/day, or 360,000 bases/day.

Animals↗

Web-based image review and data acquisition for multiinstitutional research.

OBJECTIVE: In this article, we describe a user-friendly Web-based interface that allows review of images combined with integrated data collection and entry for use at multiple sites involved in a large multicenter research project. CONCLUSION: The Web-based system that we present uses a commercially available Internet browser and Web platform and allows automated data entry that can be easily uploaded into standard data analysis programs. The system simplifies the complex logistics of using multiple sites and reviewers for radiology research and can preserve human subject confidentiality. We tested the system using a large-scale multicenter cohort study of pelvic fracture-related hemorrhage (the "Evaluating Pelvic Hemorrhage" study). Program testing revealed seamless remote image interpretation and data acquisition.

Biomedical Research↗

Data merging for integrated microarray and proteomic analysis.

The functioning of even a simple biological system is much more complicated than the sum of its genes, proteins and metabolites. A premise of systems biology is that molecular profiling will facilitate the discovery and characterization of important disease pathways. However, as multiple levels of effector pathway regulation appear to be the norm rather than the exception, a significant challenge presented by high-throughput genomics and proteomics technologies is the extraction of the biological implications of complex data. Thus, integration of heterogeneous types of data generated from diverse global technology platforms represents the first challenge in developing the necessary foundational databases needed for predictive modelling of cell and tissue responses. Given the apparent difficulty in defining the correspondence between gene expression and protein abundance measured in several systems to date, how do we make sense of these data and design the next experiment? In this review, we highlight current approaches and challenges associated with integration and analysis of heterogeneous data sets, focusing on global analysis obtained from high-throughput technologies.

Animals↗

Re-analysis of data and its integration.

To understand a biological process it is clear that a single approach will not be sufficient, just like a single measurement on a protein--such as its expression level--does not describe protein function. Using reference sets of proteins as benchmarks different approaches can be scaled and integrated. Here, we demonstrate the power of data re-analysis and integration by applying it in a case study to data from deletion phenotype screens and mRNA expression profiling.

Computational Biology↗

Rhinoplasty perioperative database using a personal digital assistant.

OBJECTIVE: To construct a reliable, accurate, and easy-to-use handheld computer database that facilitates the point-of-care acquisition of perioperative text and image data specific to rhinoplasty. METHODS: A user-modified database (Pendragon Forms [v.3.2]; Pendragon Software Corporation, Libertyville, Ill) and graphic image program (Tealpaint [v.4.87]; Tealpaint Software, San Rafael, Calif) were used to capture text and image data, respectively, on a Palm OS (v.4.11) handheld operating with 8 megabytes of memory. The handheld and desktop databases were maintained secure using PDASecure (v.2.0) and GoldSecure (v.3.0) (Trust Digital LLC, Fairfax, Va). The handheld data were then uploaded to a desktop database of either FileMaker Pro 5.0 (v.1) (FileMaker Inc, Santa Clara, Calif) or Microsoft Access 2000 (Microsoft Corp, Redmond, Wash). DESIGN: Patient data were collected from 15 patients undergoing rhinoplasty in a private practice outpatient ambulatory setting. Data integrity was assessed after 6 months' disk and hard drive storage. RESULTS: The handheld database was able to facilitate data collection and accurately record, transfer, and reliably maintain perioperative rhinoplasty data. Query capability allowed rapid search using a multitude of keyword search terms specific to the operative maneuvers performed in rhinoplasty. CONCLUSIONS: Handheld computer technology provides a method of reliably recording and storing perioperative rhinoplasty information. The handheld computer facilitates the reliable and accurate storage and query of perioperative data, assisting the retrospective review of one's own results and enhancement of surgical skills.

Computers, Handheld↗

Improving data systems about juvenile victimization in the United States.

OBJECTIVE: To suggest improvements to 13 data sets and systems that collect information about juvenile victimization in United States. METHOD: The suggestions were gathered from a variety of sources, including data system users and administrators, as well as a special meeting convened on the topic by the National Consortium on Children, Families and the Law in Washington, DC (December 2000). RESULTS: Key areas of improvement were identified for each of 13 US data systems and possible solutions were identified. CONCLUSIONS: This paper suggests three broad categories of improvements that apply to a number of data systems. First, data systems could expand the coverage of the systems to include more jurisdictions or other segments of the population. Second, in order to be more comprehensive and specific to child victimization, the systems need to create more specific data items, questions, or response categories. Finally, the data systems need to be modified to provide continuity and interrelationships among systems, either by using uniform definitions, or integrating data systems to facilitate the tracking of children across systems.

Adolescent↗

Developmental language learning impairments.

Developmental language learning impairments (LLI) are one of the most prevalent of all developmental disabilities, can occur in children for a wide variety of reasons, and have been shown to co-occur frequently with other developmental social, emotional and behavioral disorders, as well as with academic achievement problems. Research pertaining to developmental LLI of unknown origin, with an emphasis on the continuum between oral and written language impairment, is the focus of this review. Given the complexity of language learning, research has focused on multiple levels of analysis, including linguistic, neuropsychological, genetic, neurobiological, and remediation studies. To date, the vast majority of data on LLI derive from studies focused on a single level of analysis. Although attempts have been made to integrate data across studies and multiple levels of analysis, this has proven to be problematic, given the heterogeneity of the subject populations used to study LLI, as well as the differences in ages, degree of impairment, and types of impairment included in each study. Given that LLI is a complex developmental disability, it is suggested that future research would benefit from taking a multiple levels of analysis approach with the same individuals, incorporating mathematical models designed to analyze dynamically changing complex systems, and studying individual differences in language learning, prospectively and longitudinally, throughout the most dynamic stages of the process.

Brain↗

GenBank.

GenBank is a comprehensive database that contains publicly available DNA sequences for more than 165,000 named organisms, obtained primarily through submissions from individual laboratories and batch submissions from large-scale sequencing projects. Most submissions are made using the web-based BankIt or standalone Sequin programs and accession numbers are assigned by GenBank staff upon receipt. Daily data exchange with the EMBL Data Library in the UK and the DNA Data Bank of Japan helps to ensure worldwide coverage. GenBank is accessible through NCBI's retrieval system, Entrez, which integrates data from the major DNA and protein sequence databases along with taxonomy, genome, mapping, protein structure and domain information, and the biomedical journal literature via PubMed. BLAST provides sequence similarity searches of GenBank and other sequence databases. Complete bimonthly releases and daily updates of the GenBank database are available by FTP. To access GenBank and its related retrieval and analysis services, go to the NCBI Homepage at http://www.ncbi.nlm.nih.gov.

Animals↗