Data management systems: science versus technology?
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Randomized clinical trial is the main source of evidence in medicine. Information from clinical trials, however, is limited because only a limited number of questions can be addressed with this approach and, in particular, the issue of drug effectiveness (in real life) remains unanswered. On the other hand, the literature about drug effects (through animal studies, model of effect, observational studies, etc) is very large, but provides little scientific evidence useful for drug prescription: most studies do not address relevant clinical questions or have methodologic problems. In practice, physicians are overwhelmed by thousands of new articles among which they have to select the few that are useful and valid for their practice. This task is difficult because physicians do not have enough time in their clinical activity and they are not trained for critical appraisal of the literature. Therefore, they are extremely dependent on the diffusion process that selects and synthesizes the information for them. Another important factor that has been shown to influence prescription is patient's opinion, as one of the main objective for physicians is patient's satisfaction. Some physicians also use prescription as a way to terminate a consultation or as a medical act independently from the drug which is prescribed (drug prescription being part of the physician-patient relationship). At least, the absence of real control over drug prescription is considered to be one of the main factors explaining the large discrepancy between numerous prescription guidelines (from consensus conference or expert opinions) and the reality.
Data for four STR loci have been collected from 400 samples taken from complainers and suspects encountered in casework at the Strathclyde Police Forensic Science Laboratory (SPFSL). This paper describes statistical testing which demonstrates that its use will provide operationally robust procedures. Comparisons made with data collected from other British samples confirmed no practical differences between the different frequency distributions. This work provides further confirmation of the reliability of the so-called "product rule' in estimating the frequency of multilocus genotypes in British forensic casework.
Modern research techniques have led to exponential growth in the volume and complexity of scientific data. Consequently, managing these volumes securely and efficiently has become a major challenge. While all research domains face these challenges, life science research is particularly affected because current approaches often rely on a large set of different file formats, with metadata stored in separated databases or spreadsheets. This leads to fragmented datasets, orphaned data, and compromised research reproducibility. Traditional solutions also force researchers to choose between security and accessibility, with encrypted files preventing selective access and indexed formats lacking adequate security for sensitive data. These limitations are particularly problematic in large-scale genomic studies where researchers must decompress multi-gigabyte files to access specific regions, creating computational bottlenecks and inefficient network usage when working with cloud-stored datasets. We introduce Pithos, a next-generation file format specifically designed for scientific data management in distributed cloud environments. The format uses content-defined chunking to enable efficient deduplication across distributed storage systems, thereby reducing storage costs and bandwidth requirements. The append-only structure ensures data immutability and allows for incremental updates without compromising content. Benchmark results show that Pithos outperforms existing solutions in read and write performance, with comparable or improved storage efficiency.
The DNA Data Bank of Japan (DDBJ, http://www.ddbj.nig.ac.jp) has made an effort to collect as much data as possible mainly from Japanese researchers. The increase rates of the data we collected, annotated and released to the public in the past year are 43% for the number of entries and 52% for the number of bases. The increase rates are accelerated even after the human genome was sequenced, because sequencing technology has been remarkably advanced and simplified, and research in life science has been shifted from the gene scale to the genome scale. In addition, we have developed the Genome Information Broker (GIB, http://gib.genes.nig.ac.jp) that now includes more than 50 complete microbial genome and Arabidopsis genome data. We have also developed a database of the human genome, the Human Genomics Studio (HGS, http://studio.nig.ac.jp). HGS provides one with a set of sequences being as continuous as possible in any one of the 24 chromosomes. Both GIB and HGS have been updated incorporating newly available data and retrieval tools.
Emerging data from diverse organisms indicate that we are only at the threshold of our understanding of the genome-wide implications of epigenetics. This relatively new field, entitled epigenomics, will be advanced by the recently completed sequence of the Tetrahymena thermophila macronuclear genome.
Using a scientific measurement without an estimate of its error is like lending money to a stranger. Given the explosion in nucleic acid and protein sequence and structural data, what risks are the scientific and medical communities running in using these databases. Is there an 'ombudsman' who speaks for the users of the data? CODATA, the Committee on Data for Science and Technology of the International Council of Scientific Unions was established to improve the quality, reliability, processing, management, and accessibility of data for science and technology. The CODATA Task Group on Biological Macromolecules has surveyed quality control procedures of archival databanks in molecular biology. Our role is 'to advise, to be consulted, and to warn.' This report describes the kinds and extents of errors that may appear in nucleic acid and protein databases, and presents an agenda for future work to improve the quality of these databases. The results of the survey appear on the webhttp://www.codata.org/codata/tgreports/ tg_reps.html.
OBJECTIVES: The Clinical data repository (CDR) at the University of Virginia Health System is a data warehouse that provides direct access to data for clinical research and effective decision making. We undertook an evaluation of the CDR to understand factors affecting its adoption. DESIGN: We used a theoretical framework that is based on diffusion of innovation theory. Building on validated survey instruments, we developed a questionnaire and conducted interviews of key executive leaders. Fifty-three individuals with logon ids to the CDR completed our questionnaire. Twelve executive leaders were interviewed. MEASUREMENTS: The outcome variables were the initial and continued use of the CDR. Independent variables included attributes suggested by diffusion theory (i.e. relative advantage, complexity), knowledge and skills expected to correlate with computer usage, and the influence of communication channels. RESULTS: Our overall response rate was 82%. We identified characteristics of users associated with the initial decision to use the CDR. Compatibility with an individual's skills and work style was associated strongly with satisfaction and continued use. Secondly, the importance of organizational culture and the need for data was illuminated by management interviews. CONCLUSIONS: We have shown that diffusion of innovation theory can be used to help understand factors contributing to the success of a data warehouse in a healthcare setting. Our results suggest areas for future research and inquiry as the CDR evolves.
The sharing of neuroimagery data offers great benefits to science, however, data owners sharing their data face substantial custodial responsibilities, such as ensuring data sets are correctly interpreted in their new shared context, protecting the identity and privacy of human research participants, and safeguarding the understood order of use. Given choices of sharing widely or not at all, the result will often be no sharing, due to the inability of data owners to control their exposure to the risks associated with data sharing. In this context, data sharing is enabled by providing data owners with well-defined intermediate levels of data visibility, progressing incrementally toward public visibility. In this paper, we define a novel and general data sharing model, Structured Sharing Communities (SSC), meeting this requirement. Arbitrary visibility levels representing collaborative agreements, consortium memberships, research organizations, and other affiliations are structured into a policy space through explicit paths of permissible information flow. Operations enable users and applications to manage the visibility of data and enforce access permissions and restrictions. We show how a policy space can be implemented in realistic neuroinformatic architectures with acceptable assurance of correctness, and briefly describe an open source implementation effort.
PURPOSE/OBJECTIVES: To examine how delays in breast cancer care currently are conceptualized and to introduce philosophical and theoretical tenets of critical realism as an alternative approach. DATA SOURCES: Health and social sciences literature. DATA SYNTHESIS: Diagnostic and treatment delays in breast cancer most frequently are conceptualized as patient, provider, or system related. The approach has limited utility in guiding explanatory analysis because it does not acknowledge the social context in which the delays occur. The philosophical tenets of critical realism and two related theoretical approaches are an alternative. They illustrate how an individual's biologic, social, and material resources may undermine or support structural inequities in access to breast cancer care. CONCLUSIONS: Critical realism provides a useful framework for analysis of links between social inequalities and delays in breast cancer diagnosis and treatment. IMPLICATIONS FOR NURSING: Access to breast cancer care could be better understood and conceptualized by basing future research and theoretical endeavors on a critical realist perspective.
The "premedical syndrome" has been widely discussed but only anecdotally described. To learn whether the syndrome exists in the South Carolina schools and which traits compose it, the authors surveyed faculty members and students of 13 undergraduate colleges in the state. Premedical students were perceived as differing from nonpremedical students in being excessively competitive, academically, overspecialized , overachieving , more highly motivated, more highly self-disciplined, goal-oriented, and proud of their career choice. The perception by students and faculty members of the premedical syndrome may have important effects on the undergraduate curriculum and students' choices of major areas of study. Only 3 percent of the premedical students who responded to the survey were majoring in the liberal arts, and only 9 percent of the nonpremedical students were majoring in the natural sciences. These data suggest that the natural science departments in U.S. colleges may have become training grounds for premedical students to the exclusion of others. Modification of medical school admissions policies may be able to reverse some features of the premedical syndrome and some of its effects.
Explore the source record for details and available documents.
In the use of ANOVA for hypothesis testing in animal science experiments, the assumption of homogeneity of errors often is violated because of scale effects and the nature of the measurements. We demonstrate a method for transforming data so that the assumptions of ANOVA are met (or violated to a lesser degree) and apply it in analysis of data from a physiology experiment. Our study examined whether melatonin implantation would affect progesterone secretion in cycling pony mares. Overall treatment variances were greater in the melatonin-treated group, and several common transformation procedures failed. Application of the Box-Cox transformation algorithm reduced the heterogeneity of error and permitted the assumption of equal variance to be met.