Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Science”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 991 records · Page 55Linked to original sources

Testing mutual independence between two discrete-valued spatial processes: a correction to pearson chi-squared.

A common feature of data collected in environmental and earth sciences is that they typically exhibit spatial autocorrelation. Violating the assumption of independent observations can have dramatic effects on inferences derived from standard statistical methods. In this article, we examine the consequences of spatial autocorrelation on Pearson's chi-squared test of mutual independence between two categorical responses with a general number of classes. Correspondingly, we suggest a simple modification to the standard test statistic that allows for spatial autocorrelation. Our modified statistic is based on a first-order correction factor and thus provides only an approximate test. However, we show by Monte Carlo simulation that this approximation results in satisfactory inferences in several situations of practical interest. The usefulness of the method is displayed through an application to categorical data arising in the study of the relationship between the distribution pattern of plant species and woodland age in a forest in northern Belgium.

Belgium↗

Knowledge integration and reasoning as a function of instruction in a hybrid medical curriculum.

This study investigates the effect of curricular change on knowledge integration and reasoning processes during problem-solving by medical students. The curricular change involved the introduction of problem-based, small group tutorials into a conventional health science curriculum (CC). Students at three levels of training were asked to provide diagnostic explanations of two clinical cases, both before (spontaneous) and after (primed) being exposed to basic science information relevant to the clinical problems. Data were analyzed using techniques of propositional and semantic analysis. Based on theories of instruction and cognition, we expected that the instructional changes would facilitate knowledge integration and influence the reasoning patterns of the students. The results show that students generated fewer inferences and used more information from the basic science text (text-based) to explain the clinical problems. However, they generated a greater number of elaborations during explanations using a mixture of data-driven and hypothesis-driven strategies. The spontaneous and primed problem-solving conditions produced more hypothesis-driven and data-driven strategies, respectively, as would be expected in a hybrid curriculum. We conclude that a) problem-based, small group tutorials facilitate integration of clinical and biomedical knowledge through the use of elaborations and hypothesis-driven strategies, and b) aspects of problem-based learning can be successfully integrated into traditional curricula.

Clinical Competence↗

Tricross : using dot-plots in sequence-id space to detect uncataloged intergenic features.

MOTIVATION: The process of determining the functional sequence content of an organism is confounded by several factors. Large protein coding sequences are relatively easy to find by statistical methods. Smaller proteins however may escape detection due to their size falling below some arbitrary researcher-defined minimum cutoff, or the inability to precisely define a promoter, or translational start (Delcher et al., Nucleic Acids Res., 27, 4636-4641, 1999). Promoter and regulatory sequences themselves are difficult to define due to a significant amount of allowable sequence variation, as well as a probable lack of any completely accurate whole-organismal gene catalogs to date. Finally, certain genes coding functional RNAs may have insufficient structural or sequence constraints to be detectable by normal sequence structure/pattern searching methods (Eddy and Rivas, Bioinformatics, 16, 583-605, 2000). In those cases where there are multiple closely related organisms that have been sequenced, there is additional information that may be used in the investigation of sequence content-that being the possible conserved nature of functional sequences between the organisms. We present a method for the utilization of this conserved information to detect genes and other potentially functional sequences that may be missed by standard ORF-calling, RNA finding, and pattern matching software. The tricross programs produce a multi-way cross comparison of three sets of sequences, determine which are conserved in all three sets, and produce a graphical (Virtual Reality Modelling Language-VRML; (ISO/IEC 14772-1: 1997, VDC), 1997) representation as well as alignments of all sequence triples found. The software can also be applied to a pair of sequence sets, though the noise in the results increases. RESULTS: Tricross has been used to examine the intergenic-sequence content of the three archaeal Pyrococcus genomes to determine the most highly related sequences remaining between the annotated protein and RNA coding sequences. Set to relatively stringent similarity requirements for the search, tricross found 101 intergenic sequences conserved among the three organisms. Interestingly, 29 of these appear to contain members of a family of small RNA molecules (Kiss-Laszlo et al., EMBO J., 17, 797-807, 1998) only recently discovered in the Archaea (Armbruster, OSU, Diss., 1988; Omer et al., Science, 288, 517-522, 2000; Gaspin et al., J. Mol. Biol., 297, 895-906, 2000). While some of the remaining 72 appear to be individual highly conserved promoter sequences, others have no currently known biological significance. Although originally developed to facilitate the examination of intergenic sequences, none of the tricross logic is inherently specific to intergenic sequences. The software can also be applied to gene sequences, and has been used to produce inter-genomic gene order dot-plots for Haemophilus influenzae (Fleischmann et al., Science, 269, 496-512, 1995) versus H.ducreyi (unpublished data), and Neisseria meningiditis Z2491 (serogroup A) (Parkhill et al., Nature, 404, 502-506, 2000) versus Neisseria meningiditis Z58 (serogroup B) (Tettelin et al., Science, 287, 1809-1815, 2000) versus Neisseria gonorrhoeae (Lewis et al., http://micro-gen.ouhsc.edu/, 2000). AVAILABILITY: The tricross software package is available from http://www.biosci.ohio-state.edu/~ray/bioinformatics/tricross.html. CONTACT: ray@biosci.ohio-state.edu; daniels.7@osu.edu; munsonr@pediatrics.ohio-state.edu SUPPLEMENTARY INFORMATION: Additional data from the cross-genomic comparisons examined in the discussion section are linked from http://www.biosci.ohio-state.edu/~ray/bioinformatics/tricross.html.

Base Sequence↗

Methods for defining equity-stratifying variables: a systematic review of validation studies.

BACKGROUND AND OBJECTIVE: Disease burden is often disproportionally higher among those who are socially disadvantaged by factors defined in the PROGRESS-Plus framework (ie, Place of residence, Race/ethnicity/culture/language, Occupation, Gender/sex, Religion, Education, Socioeconomic status, and Social capital, with "Plus" covering features like age and disability). The accuracy and applicability of case definitions to identify these variables from administrative and clinical health data are unknown. We conducted a systematic review to explore how equity-stratifying variables, as categorized by the PROGRESS-Plus framework, have been defined and validated in epidemiologic studies using administrative health, population-level, or electronic health record (EHR) data. METHODS: Medline, EMBASE, CINAHL, Web of Science, and Google Scholar were searched from the inception of the databases to 2024 for validation studies of equity-stratifying variables in adults using administrative health datasets, health registries, or EHR data. Titles and abstracts, followed by relevant full-text articles, were screened in duplicate by two reviewers for eligibility. The data sources utilized, algorithms employed, and their associated performance measures were extracted and synthesized from included studies. Given substantial heterogeneity in study design, equity-stratifying variable definition, and performance metrics, meta-analysis was not possible. RESULTS: Of the 9099 unique citations screened, 188 full texts were reviewed and 116 were included in this review. Most studies were published between 2019 and 2024 (n = 64, 55%) and were validation studies of race/ethnicity definitions that used race/ethnicity codes or surname list algorithms (n = 66, 57%). No studies examined religion. Regarding the reported performance measure estimates, the race/ethnicity/culture/language equity-stratifying variables category had the largest variability across sensitivity, positive predictive value (PPV), and Cohen's Kappa. Occupation validation studies had the lowest variation in sensitivity and PPV. CONCLUSION: Despite an increasing number of publications reporting on the validation of equity-stratifying variables relevant to the PROGRESS-Plus framework, performance measures varied widely across studies. The significant heterogeneity in equity-stratifying variable definitions and methods used to validate them support the need for further rigorous validation of equity-stratifying variables in administrative and clinical health data. PLAIN LANGUAGE SUMMARY: Disease burden is often higher in people who experience financial hardships, lower level of education, discrimination due to race/ethnicity, and unstable housing. These social factors can be considered health equity factors and are important for understanding health inequalities. Health researchers often use large datasets, such as hospital or electronic health records (EHRs), to study these health equity factors. However, it is not clear how accurately these data sources capture information about people's social circumstances and how these factors are defined. In this study, we reviewed existing research to understand how health equity factors have been defined across health data sources and how accurate they are at measuring aspects of health equity and social disadvantage. Of the more than 9000 studies we identified, we included 116 that met our criteria for this systematic review. Most included studies focused on identifying race and ethnicity, often using codes or surname-based methods. We found that the accuracy of these methods varied widely across studies, meaning results may not always be reliable or comparable. Overall, our findings show that there are inconsistencies in how social factors are defined and measured in health data. This makes it difficult to fully understand and address health inequalities using routinely collected health data. More work is needed to develop and validate better quality and more consistent methods for capturing these important social factors.

Humans↗

Bismuth subgallate-epinephrine paste in adenotonsillectomies.

OBJECTIVE: To evaluate the role of bismuth subgallate-epinephrine (BSE) paste as a hemostatic in adenotonsillectomies. DATA SOURCES: MEDLINE (January 1966-October 1999) and Current Contents (January 1997-October 1999) were searched, using bismuth subgallate, adenoidectomy, tonsillectomy, and adenotonsillectomy as search terms. A citation search was performed using Science Citation Index (January 1977-October 1999). DATA SYNTHESIS: Adenotonsillectomies are common procedures; although there are few complications, hemorrhage is a concern. Bismuth subgallate has historically been used as an astringent and hemostatic. An evaluation of studies of bismuth subgallate and BSE paste was conducted. CONCLUSIONS: There is minimal evidence to support this practice, but data suggest that epinephrine may be the active ingredient in BSE paste. BSE paste is inexpensive, poses little risk, and may decrease postoperative bleeding; therefore, it may be a reasonable hemostatic agent.

Clinical Trials as Topic↗

Apolipoprotein E in hyperlipidemia.

PURPOSE: To review DNA analysis of apolipoprotein E used to assess patients with hyperlipidemia. DATA SOURCES AND STUDY SELECTION: 44 basic science studies of molecular analysis; 42 basic science studies of the biochemical, cellular biological, and molecular biological features of apolipoprotein E; and 29 clinical investigational studies, meta-analyses, and case series of patients with mutations in apolipoprotein E. DATA EXTRACTION: Methods of DNA analysis were reviewed, using specific examples in human disease, and the role of apolipoprotein E in normal and disordered lipoprotein metabolism was reviewed. Genetic analysis of apolipoprotein E in populations and particularly in persons with type III hyperlipoproteinemia is reviewed. DATA SYNTHESIS: In the general population, common DNA variants of apolipoprotein E are consistently associated with modest differences in plasma lipids and lipoproteins. Homozygosity for the E2 isoform of apolipoprotein E predisposes some patients to the development of type III hyperlipoproteinemia, a condition that involves an additional genetic or environmental factor for full clinical expression. Rare mutations of apolipoprotein E also cause hyperlipidemia. CONCLUSIONS: DNA variation of apolipoprotein E is one of several genetic and environmental factors that interact in a complex manner to affect plasma lipoproteins. DNA analysis of apolipoprotein E can be used in persons with hyperlipidemia to identify those with type III hyperlipoproteinemia and in relatives of affected persons to identify those who are predisposed.

Apolipoproteins E↗

Surveillance of sexually transmitted infections in New Zealand, 1998.

AIMS: To describe the surveillance and epidemiology of sexually transmitted infections (STIs) in New Zealand. METHODS: Sexual health clinics submitted STI data to the Institute of Environmental Science and Research (ESR). Infection rates were calculated by dividing the number of diagnoses by the number of total clinic visits. Because the denominator used to calculate infection rates changed in 1998, STI rates in 1998 cannot be directly compared with previous years and case numbers were used to identify recent trends. RESULTS: In 1998, genital warts was the most commonly diagnosed STI (4.7%), followed by chlamydia (3.0%) and genital herpes (1.0%). Approximately two-thirds of gonorrhoea, chlamydia and genital warts diagnoses were in people aged less than 25 years. Chlamydia rates were 7.3% in Maori, 7.1% in Pacific Island people, and 2.1% in European. Gonorrhoea rates were 1.6% in Maori, 1.9% in Pacific Island people and 0.2% in European. The number of chlamydia and gonorrhoea cases increased between 1995 and 1998. CONCLUSIONS: The reporting of data by age, sex and ethnicity has allowed a more useful evaluation of the incidence of STIs. The majority of STIs were diagnosed among young New Zealanders, and disproportionately high chlamydia and gonorrhoea infection rates were found among Maori and Pacific Island people.

Adolescent↗

A perspective on current and future uses of alternative models for carcinogenicity testing.

This perspective is based upon the data presented at the International Life Sciences Institute (ILSI), Health and Environmental Sciences Institute Workshop on the Evaluation of Alternative Methods for Carcinogenicity Testing (ILSI Workshop). It is important to understand that all models discussed at the Workshop have limitations and that they are not designed to be employed as stand-alone assays. Although they may have other, appropriate applications. I do not recommend use of the SHE cell assay and the Tg.AC model for the regulatory purposes of a safety assessment. In my view, the neonatal mouse, p53+/-, XPA-/-, XPA-/- and p53+/-, and the rasH2 models can, as a component of an overall assessment, provide information on potential carcinogenicity of a chemical that is appropriate for consideration in a regulatory context. Generally, these models exhibit the ability to detect genotoxic compounds. In most cases these compounds would be detected in a standard battery of genotoxicity tests and, therefore, quite often the use of an alternative is not necessary. Actually, I believe that a bioassay in rats will suffice most of the time, that is, in my view, a routine bioassay in mice is not necessary. Specific circumstances where data obtained from one of the "recommended" alternative models might be helpful are discussed. With regard to lessons for the future, there is a particular need for models that are responsive to chemicals that exhibit a nongenotoxic mode of action. Additionally, new models will continue to be developed and their half-life will likely be substantially shorter than the time required for traditional validation. The development of enhanced paradigms for validation should be a priority so that improved safety assessment decisions can be made more quickly. However, while evaluating and validating such models, it is important to consider the fundamental issues, for example, rational dose selection, evaluation of mode of action in the context of dose-response relationships including the existence of thresholds and secondary mechanisms, and species-to-species extrapolation. The alternatives to carcinogenicity testing project was a very major undertaking. In addition to the valuable information provided, it serves to illustrate the value of cooperation between academia, government, and industry. Furthermore, the involvement of the International Life Sciences Institute as the overall organizing, facilitating umbrella was crucial for the success of the project.

Academies and Institutes↗

The ArrayExpress gene expression database: a software engineering and implementation perspective.

MOTIVATION: The lack of microarray data management systems and databases is still one of the major problems faced by many life sciences laboratories. While developing the public repository for microarray data ArrayExpress we had to find novel solutions to many non-trivial software engineering problems. Our experience will be both relevant and useful for most bioinformaticians involved in developing information systems for a wide range of high-throughput technologies. RESULTS: ArrayExpress has been online since February 2002, growing exponentially to well over 10,000 hybridizations (as of September 2004). It has been demonstrated that our chosen design and implementation works for databases aimed at storage, access and sharing of high-throughput data. AVAILABILITY: The ArrayExpress database is available at http://www.ebi.ac.uk/arrayexpress/. The software is open source. CONTACT: ugis@ebi.ac.uk.

Algorithms↗

Education level and rheumatoid arthritis: evidence from five data centers.

Data on 2,006 patients with rheumatoid arthritis (RA) for 1984 and 1986 were analyzed from 5 American Rheumatism Association Medical Information Systems patient centers to assess covariates of future disability. The dependent variable was the Health Assessment Questionnaire disability index, measured in 1986. Independent variables, measured in 1984, include years of schooling, age, age squared, sex, labor force status, occupation, marital status, race, income, sedimentation rate, log latex, number of tender joints, and duration of illness. A negative association between schooling and the disability index is strongly apparent for men and weaker in women. Results for men persist even after adjustment for occupation and income, but not after additional adjustment for biological variables. The effects of schooling upon progression of RA are complex and interpretation requires simultaneous assessment of a variety of other variables. Causal effects of level of schooling are seen with a social science perspective which considers the biologic data to represent dependent variables yet cannot be inferred from a clinical model which considers the biology of the disease as an independent variable.

Arthritis, Rheumatoid↗

A generalized likelihood ratio test to identify differentially expressed genes from microarray data.

MOTIVATION: Microarray technology emerges as a powerful tool in life science. One major application of microarray technology is to identify differentially expressed genes under various conditions. Currently, the statistical methods to analyze microarray data are generally unsatisfactory, mainly due to the lack of understanding of the distribution and error structure of microarray data. RESULTS: We develop a generalized likelihood ratio (GLR) test based on the two-component model proposed by Rocke and Durbin to identify differentially expressed genes from microarray data. Simulation studies show that the GLR test is more powerful than commonly used methods, like the fold-change method and the two-sample t-test. When applied to microarray data, the GLR test identifies more differentially expressed genes than the t-test, has a lower false discovery rate and shows more consistency over independently repeated experiments. AVAILABILITY: The approach is implemented in software called GLR, which is freely available for downloading at http://www.cc.utah.edu/~jw27c60

Algorithms↗

DNA and the revolutions of molecular evolution, computational biology, and bioinformatics.

The discovery of the structure of DNA was a necessary prerequisite for determining the sequence of DNA molecules. Technological advances have now made it possible to sequence DNA rapidly and has resulted in public databases with over 30 billion nucleotides of known sequence. The analysis of these data has lead to new fields of science and to amazing advances in our understanding of evolution.

Animals↗

How to begin a new topic in mathematics: does it matter to students' performance in mathematics?

The authors use Canadian data from the Third International Mathematics and Science Study to examine six instructional methods that mathematics teachers use to introduce new topics in mathematics on performance of eighth-grade students in six mathematical areas (mathematics as a whole, algebra, data analysis, fraction, geometry, and measurement). Results of multilevel analysis with students nested within schools show that the instructional methods of having the teacher explain the rules and definitions and looking at the textbook while the teacher talks about it had little instructional effects on student performance in any mathematical area. In contrast, the instructional method in which teachers try to solve an example related to the new topic was effective in promoting student performance across all mathematical areas.

Adolescent↗

Estimation of the hereditary risks of exposure to ionizing radiation: history, current status, and emerging perspectives.

This paper provides a brief overview of the advances in the field of the estimation of the genetic risks of exposure of human populations to ionizing radiation from the early 1950's to the present and of the developments that are anticipated in the coming years. The latter are based on the view that the insights gained from human genetics, especially human molecular genetics, will be increasingly applied to address problems in risk estimation. Owing to the paucity of human data on radiation-induced mutations, mouse data on radiation-induced mutations are used to predict the risk of genetic diseases in humans using the doubling dose method. With this method, the risk per unit dose is expressed as a product of three quantities, i.e., P x 1/DD x MC where P is the baseline frequency of genetic diseases, 1/DD (the relative mutation risk per unit dose; DD refers to the doubling dose, i.e., the radiation dose required to produce as many mutations as those that occur spontaneously in a generation) and MC is the disease class-specific mutation component (a measure of the relative increase in disease frequency per unit relative increase in mutation rate). The five important changes that are now introduced in genetic risk estimation include (1) an upward revision of the baseline frequency of Mendelian diseases to 2.4% (from 1.25% used until the early 1990's); (2) a reversion to the conceptual basis for DD calculations used in the 1972 BEIR report of the U.S. National Academy of Sciences, namely, the use of human data on spontaneous mutation rates and mouse data on induced mutation rates (instead of the use of mouse data for both these rates as has been the case from mid-1970's until the early 1990's); (3) the fuller development and use of the MC concept for predicting the responsiveness of Mendelian and multifactorial diseases to increases in mutation rate; (4) the introduction of a new disease-class-specific quantity called the "potential recoverability correction factor" or PRCF in the risk equation to bridge the gap between the rates of induced mutations in mice and the risk of inducible genetic diseases in humans; and (5) the introduction of the concept that multisystem developmental abnormalities are likely to be among the principal phenotypes of radiation induced genetic damage in humans. All these advances now permit, for the first time in 40 y, the estimation of risks for all classes of genetic diseases. For a population exposed to low-LET, chronic or low-dose irradiation, the risks predicted for the first generation progeny are the following (all estimates are per million live born progeny per gray of parental irradiation): autosomal dominant and x-linked diseases, approximately 750 to 1,500 cases; autosomal recessive, nearly zero; chronic multifactorial diseases, approximately 250 to 1,200 cases; and congenital abnormalities, approximately 2000 cases. The total risk per gray is of the order of approximately 3,000 to 4,700 cases, which represent approximately 0.4 to 0.6% of the baseline frequency of these diseases (738,000 per million) in the population. The advances anticipated in the coming years are likely to permit the estimation of genetic risks of radiation with greater precision than is now possible.

Animals↗

Automated smear counting and data processing using a notebook computer in a biomedical research facility.

An automated smear counting and data processing system for a life science laboratory was developed to facilitate routine surveys and eliminate human errors by using a notebook computer. This system was composed of a personal computer, a liquid scintillation counter and a well-type NaI(Tl) scintillation counter. The radioactivity of smear samples was automatically measured by these counters. The personal computer received raw signals from the counters through an interface of RS-232C. The software for the computer evaluated the surface density of each radioisotope and printed out that value along with other items as a report. The software was programmed in Pascal language. This system was successfully applied to routine surveys for contamination in our facility.

Computers↗

On robust partial discriminant analysis as a decision-making tool with clinical and analytical chemical data.

Classification is one of the fundamental goals of science and is basic to the diagnosis of disease. Unfortunately, classifying objects (e.g., patients) on the basis of clinical and/or laboratory experimental observations into various groups can be difficult when the groups overlap or contain outlying points. Recently, Broffitt, Randles, and co-workers proposed a procedure, robust partial discriminant analysis (RPDA) for dealing with such problems, but testing of the procedure was limited to Monte Carlo simulation. In this study, RPDA was applied to real data, in order to compare its effectiveness with ordinary discriminant analysis, as well as to determine if RPDA was a suitable procedure to use to classify chemical compounds on the basis of experimental observations and as a tool in the diagnosis of disease (in particular, multiple sclerosis and thyrotoxicosis), with data based on experimental and clinical observations. The resulting RPDA classifications were an improvement over those obtained from ordinary discriminant analysis.

Aldehydes↗

Service-oriented science.

New information architectures enable new approaches to publishing and accessing valuable data and programs. So-called service-oriented architectures define standard interfaces and protocols that allow developers to encapsulate information tools as services that clients can access without knowledge of, or control over, their internal workings. Thus, tools formerly accessible only to the specialist can be made available to all; previously manual data-processing and analysis tasks can be automated by having services access services. Such service-oriented approaches to science are already being applied successfully, in some cases at substantial scales, but much more effort is required before these approaches are applied routinely across many disciplines. Grid technologies can accelerate the development and adoption of service-oriented science by enabling a separation of concerns between discipline-specific content and domain-independent software and hardware infrastructure.

Algorithms↗

Ethics issues in academic-industry relationships in the life sciences: the continuing debate.

The author reviews in detail the status of academic-industry relationships (AIRs) in the life sciences from both ethical and empirical perspectives, and identifies ethical issues that have been resolved and those that must still be debated. He summarizes by stating that ethical reasoning militates against the involvement of scientists and universities in those AIRs in which a financial conflict of interest on the part of life science investigators may affect the welfare of human subjects and trainees. Even in other types of AIRs, conflicts of interest have effects on professional decision making that could damage the integrity and productivity of life sciences research, especially scientists' withholding of data and their redirecting of research in more commercial directions. These effects could also help undermine public trust in and support of university researchers. Balanced against these worrisome effects are the benefits of AIRs in increasing some investigators' creativity and productivity, in encouraging technology transfer, and thus in promoting economic growth and public health. He concludes that more research is needed on the harms and benefits of AIRs, especially the development of better data on the effects of withholding data, and also on the economic and health benefits of AIRs and public attitudes toward issues of scientific research that involve possible conflicts of interest. More information on these questions would allow policymakers to make more realistic estimates of the gains and losses associated with AIRs. In the meantime, current information suggests that in general the conflicts of interest created by AIRs are real, consequential, but tolerable if managed carefully. Until more is known about the effects of AIRs, it is prudent for universities and faculty to participate at modest levels in such relationships and to monitor them carefully. This article is one of three in this issue of Academic Medicine that deal with issues of conflict of interest in university-industry research relationships. These articles are discussed in an overview that precedes them.

Biological Science Disciplines↗