Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Science”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,009 records · Page 56Linked to original sources

Locus-specific mutation databases: pitfalls and good practice based on the p53 experience.

Between 50,000 and 60,000 mutations have been described in various genes that are associated with a wide variety of diseases. Reporting, storing and analysing these data is an important challenge as such data provide invaluable information for both clinical medicine and basic science. Locus-specific databases have been developed to exploit this huge volume of data. The p53 mutation database is a paradigm, as it constitutes the largest collection of somatic mutations (22,000). However, there are several biases in this database that can lead to serious erroneous interpretations. We describe several rules for mutation database management that could benefit the entire scientific community.

Database Management Systems↗

Muscle strain injuries.

One of the most common injuries seen in the office of the practicing physician is the muscle strain. Until recently, little data were available on the basic science and clinical application of this basic science for the treatment and prevention of muscle strains. Studies in the last 10 years represent action taken on the direction of investigation into muscle strain injuries from the laboratory and clinical fronts. Findings from the laboratory indicate that certain muscles are susceptible to strain injury (muscles that cross multiple joints or have complex architecture). These muscles have a strain threshold for both passive and active injury. Strain injury is not the result of muscle contraction alone, rather, strains are the result of excessive stretch or stretch while the muscle is being activated. When the muscle tears, the damage is localized very near the muscle-tendon junction. After injury, the muscle is weaker and at risk for further injury. The force output of the muscle returns over the following days as the muscle undertakes a predictable progression toward tissue healing. Current imaging studies have been used clinically to document the site of injury to the muscle-tendon junction. The commonly injured muscles have been described and include the hamstring, the rectus femoris, gastrocnemius, and adductor longus muscles. Injuries inconsistent with involvement of a single muscle-tendon junction proved to be at tendinous origins rather than within the muscle belly. Important information has also been provided regarding injuries with poor prognosis, which are potentially repairable surgically, including injuries to the rectus femoris muscle, the hamstring origin, and the abdominal wall. Data important to the management of common muscle injuries have been published. The risks of reinjury have been documented. The early efficacy and potential for long-term risks of nonsteroidal antiinflammatory agents have been shown. New data can also be applied to the field with respect to the beneficial effects of warm-up, temperature, and stretching on the mechanical properties of muscle. These benefits potentially reduce the risks of strain injury to the muscle. Fortunately, many of the factors protecting muscle, such as strength, endurance, and flexibility, are also essential for maximum performance. Future studies should delineate the repair and recovery process emphasizing not only the recovery of function, but also the susceptibility to reinjury during the recovery phase.

Animals↗

Data mining: sophisticated forms of managed care modeling through artificial intelligence.

Data mining is a recent development in computer science that combines artificial intelligence algorithms and relational databases to discover patterns automatically, without the use of traditional statistical methods. Work with data mining tools in health care is in a developmental stage that holds great promise, given the combination of demographic and diagnostic information.

Algorithms↗

The role of migration research in regional science.

"In this paper we try to provide an assessment of the role that migration research has played over the course of the more than 40 years in which regional science has existed as a recognizable, multidisciplinary academic enterprise.... To carry out our analyses we developed a data base of papers published in five leading regional science journals." The authors "attempt to set the regional science contributions in the context of migration research more generally, comparing the results of the journal analysis to a broader sample of migration abstracts published in the Population Index."

Demography↗

[Data processing computer system for evaluation of health effects due to occupational diseases diagnosed in Poland].

General consequences of occupational diseases for both employees and the country's economy have been known for many decades. Nevertheless there were no legal instruments and financial means to carry out studies leading to a comprehensive evaluation of health effects induced by occupational diseases diagnosed in Poland. Only recently, the legal basis has been provided by the revision of the regulations on occupational diseases (Official Journal of the Ministry of Health and Social Welfare no 9, heading 51, 1989), and the financial means allocated according to the Governmental Strategic Programme (SPR-1) for the years 1995-1998. The objectives of the study and the results obtained have been presented in another publication. Here the authors concentrated on some aspects of data collection and their analysis. An essential element in the establishment of data base on health effects of occupational diseases was its integration, namely the number of diagnosed occupational disease had to correspond with the entity in the Register of Occupational Diseases. Such an approach helped to avoid including in the base all metrical data, as information on effects of occupational diseases was collected exclusively for persons with detected and diagnosed occupational disease, which had to be previously placed in the Register of Occupational Diseases. The correctness of data input and the data analysis were carried out outside the Paradox system. The analysis of histograms of characteristics under study and the explanation of departing values were employed to evaluate the freedom from bias, and the statistical package for social sciences (SPSS) was used to analyse the data.

Computer Systems↗

Some aspects of modern population mathematics.

"The purpose of this paper is to survey a number of the technical tools and models that have found use in the study of human and other populations, and to indicate some problems of current interest. These tools and models are varied: integral equations, nonlinear oscillations, differential geometry, dynamical systems, nonlinear operation, bifurcation theory, semigroup theory, martingale theory, Markov processes, diffusion processes, branching processes, ergodic theory, prediction theory and state-space models. A fairly extensive bibliography is provided. Also an Appendix has been added describing the analysis of a classical entomological data set." (summary in FRE)

Demography↗

Educational Reform in Athletic Training: A Policy Analysis.

OBJECTIVE: To apply a policy-analysis framework to the athletic training educational reform policy that will be fully implemented by January 2004. DATA SOURCES: Policy analysis is not a specific science. No one framework exists for conducting all policy analyses. I used literature from the education, policy analysis, and athletic training fields as data sources to provide background and to create a framework from which to conduct the policy analysis. DATA SYNTHESIS: Once the policy-analysis framework was selected, I began data synthesis, using several athletic training sources in support of the findings. The tension among the myriad stakeholders in this policy is clear. Although many see the benefits of accreditation, some experience hardships from the imposed policy. CONCLUSIONS/RECOMMENDATIONS: Of the 4 possible alternatives suggested, following the route currently under implementation (Committee on Accreditation of Health Education Programs accreditation) was the most agreeable solution. The goals as stated by the policy makers are attained by the policy. However, issues within the accreditation process itself need to be addressed. Of the many stakeholders in the reform effort, some will see little gain and have many hardships imposed on them. As the policy is implemented, unintended implications will likely arise, as with any new policy. Thus, I recommend that the National Athletic Trainers' Association develop a system dedicated solely to reducing the hardships faced by many of its members as the policy is implemented.

Journal Article↗

How well do HapMap SNPs capture the untyped SNPs?

BACKGROUND: The recent advancement in human genome sequencing and genotyping has revealed millions of single nucleotide polymorphisms (SNP) which determine the variation among human beings. One of the particular important projects is The International HapMap Project which provides the catalogue of human genetic variation for disease association studies. In this paper, we analyzed the genotype data in HapMap project by using National Institute of Environmental Health Sciences Environmental Genome Project (NIEHS EGP) SNPs. We first determine whether the HapMap data are transferable to the NIEHS data. Then, we study how well the HapMap SNPs capture the untyped SNPs in the region. Finally, we provide general guidelines for determining whether the SNPs chosen from HapMap may be able to capture most of the untyped SNPs. RESULTS: Our analysis shows that HapMap data are not robust enough to capture the untyped variants for most of the human genes. The performance of SNPs for European and Asian samples are marginal in capturing the untyped variants, i.e. approximately 55%. Expectedly, the SNPs from HapMap YRI panel can only capture approximately 30% of the variants. Although the overall performance is low, however, the SNPs for some genes perform very well and are able to capture most of the variants along the gene. This is observed in the European and Asian panel, but not in African panel. Through observation, we concluded that in order to have a well covered SNPs reference panel, the SNPs density and the association among reference SNPs are important to estimate the robustness of the chosen SNPs. CONCLUSION: We have analyzed the coverage of HapMap SNPs using NIEHS EGP data. The results show that HapMap SNPs are transferable to the NIEHS SNPs. However, HapMap SNPs cannot capture some of the untyped SNPs and therefore resequencing may be needed to uncover more SNPs in the missing region.

Asian People↗

A causal model for the effectiveness of internal quality assurance for the health science area.

The purposes of this research were 1) to study the effectiveness of Internal Quality Assurance (IQA) of the Health science area, and 2) to study the factors affecting the effectiveness of the IQA of the Health science area. A causal model has been developed by the researcher comprised of the 6 exogenous latent variables: Attitude towards quality assurance, Teamwork, Staff training, Resource sufficiency, Organizational culture, and Leadership, and the 4 endogenous latent variables, which are the effectiveness of the IQA, Student-centered approach, Decentralized administration, PDCA cycle of work (Plan-Do-Check-Act), and Staff job satisfaction. The research sample consisted of 108 health science faculties derived by stratified random sampling technique. Data were collected by 10 questionnaires having reliability ranging from 0.79 to 0.96. Data analyses were descriptive statistics, and Linear Structure Relationship (LISREL) analysis. The major findings were as follows: 1. The 4 dimensions of effectiveness for the IQA of the Health science areas were significantly higher at the .05 level, after the Health science faculty applied the IQA programme according to the National Education Act of 1999. 2. The causal model of the effectiveness of the IQA was valid and fitted the empirical data. The 6 predictors accounted for 83% of the variance in the effectiveness of IQA. Culture and Leadership were the predictors that significantly accounted for the effectiveness of the IQA.

Health Facility Administration↗

Estimating missing data: an iterative regression approach.

The problem of missing data is common in all fields of science. Various methods of estimating missing values in a dataset exist, such as deletion of cases, insertion of sample mean, and linear regression. Each approach presents problems inherent in the method itself or in the nature of the pattern of missing data. We report a method that (1) is more general in application and (2) provides better estimates than traditional approaches, such as one-step regression. The model is general in that it may be applied to singular matrices, such as small datasets or those that contain dummy or index variables. The strength of the model is that it builds a regression equation iteratively, using a bootstrap method. The precision of the regressed estimates of a variable increases as regressed estimates of the predictor variables improve. We illustrate this method with a set of measurements of European Upper Paleolithic and Mesolithic human postcranial remains, as well as a set of primate anthropometric data. First, simulation tests using the primate data set involved randomly turning 20% of the values to "missing". In each case, the first iteration produced significantly better estimates than other estimating techniques. Second, we applied our method to the incomplete set of human postcranial measurements. MISDAT estimates always perform better than replacement of missing data by means and better than classical multiple regression. As with classical multiple regression, MISDAT performs when squared multiple correlation values approach the reliability of the measurement to be estimated, e.g., above about 0. 8.

Animals↗

Coverage and perceptions of Medical Sciences students towards hepatitis B virus vaccine in Sana'a City, Yemen.

OBJECTIVE: The present study was conducted to estimate vaccination coverage against hepatitis B virus and the perceptions of 1198 medical sciences students in Sana'a City, Yemen. METHODS: Only those who practice clinical training or are in contact with body fluids were included. The students were enrolled in the Faculty of Medicine and Health Sciences, Sana'a University, Republic of Yemen. Data was collected from 1999-2000. Arabic pre-tested questionnaire forms were completed by 840 students at a response rate of 70.6%. RESULTS: The study revealed a reported vaccination rate of 29.5%. The rate among Faculty of Medicine and Health Sciences students was 32.3%, whereas only 21.3% among the students of High Institute of Health Sciences. Students of dentistry attained the highest rate of vaccination (38.8%), while nursing students of the High Institute of Health Sciences achieved the lowest rate (17.1%). Rate of vaccination (46.6%) among female students was significantly higher than male students (22.3%) with a P- value of 0.0001. Medical assistants of the High Institute of Health Sciences scored the best (56%) in terms of knowledge, medical laboratory sciences students achieved the highest (43.6%) in attitude and dentistry students had the highest scores (35.5%) in practices. The mean knowledge of females and males was comparable, however, females achieved higher attitudes and practices. Final stage students attained better attitude scores than the pre-final and intermediate students. CONCLUSION: Vaccination coverage of medical sciences students in Sana'a City, Yemen is low. Knowledge of medical assistants is the best, attitude of medical laboratory sciences students and practices of dental students is the highest. Attitudes and practices of female students are better than that of males.

Health Knowledge, Attitudes, Practice↗

Implications of the human genome for understanding human biology and medicine.

Clinical researchers, practicing physicians, patients, and the general public now live in a world in which the 2.9 billion nucleotide codes of the human genome are available as a resource for scientific discovery. Some of the findings from the sequencing of the human genome were expected, confirming knowledge presaged by many decades of research in both human and comparative genetics. Other findings are unexpected in their scientific and philosophical implications. In either case, the availability of the human genome is likely to have significant implications, first for clinical research and then for the practice of medicine. This article provides our reflections on what the new genomic knowledge might mean for the future of medicine and how the new knowledge relates to what we knew in the era before the availability of the genome sequence. In addition, practicing physicians in many communities are traditionally also ambassadors of science, called on to translate arcane data or the complex ramifications of biology into a language understood by the public at large. This article also may be useful for physicians who serve in this capacity in their communities. We address the following issues: the number of protein-coding genes in the human genome and certain classes of noncoding repeat elements in the genome; features of genome evolution, including large-scale duplications; an overview of the predicted protein set to highlight prominent differences between the human genome and other sequenced eukaryotic genomes; and DNA variation in the human genome. In addition, we show how this information lays the foundations for ongoing and future endeavors that will revolutionize biomedical research and our understanding of human health.

Clinical Medicine↗

Interspecific variation at the Y-linked RPS4Y locus in hominoids: implications for phylogeny.

Within- and between-species variation in restriction endonuclease recognition sites was examined at the Y-linked RPS4Y locus of six hominoid species: human (Homo sapiens), gorilla (Gorilla gorilla), chimpanzee (Pan troglodytes), bonobo (Pan paniscus), orangutan (Pongo pygmaeus), and gibbon (Hylobates lar). RPS4Y is an expressed gene that maps to the non-recombining region of the Y chromosome. An approximately 1,490 base pair fragment of the RPS4Y gene, including all of intron 3, was amplified by PCR from DNA extracted from each of the six species. Forty-seven restriction sites were identified on the six-species composite map derived from double-digest restriction analyses of the amplified fragment. As expected, maximum parsimony analysis indicated that chimpanzee and bonobo are the two most closely related living hominoids. The same analysis suggested that the closest living relative of Homo is Gorilla, not Pan, although support for this relationship was relatively weak. These results disagree with recently published phylogenies based on analyses of mtDNA sequences (Horai et al. [1995] Proc. Natl. Acad. Sci. U.S.A. 88:7401-7404) and the Y-linked ZFY locus (Dorit et al. [1995] Science 268:1183-1185). A combined data set derived from three distinct Y-linked loci-RPS4Y, SRY, and ZFY-was also analyzed. The maximum parsimony topology for the combined data provided only weak support for a shared common ancestor for Homo and Pan subsequent to divergence from the Gorilla lineage. Taken together, the data from the Y chromosome do not provide unequivocal support for any single, dichotomously branching species tree linking Homo, Pan, and Gorilla.

Animals↗

Estimation of maximum increment age in height and weight during adolescence and the effect of World War II.

An attempt was made to estimate the maximum increment age (MIA) in height and weight of Japanese boys and girls during the birth years 1893-1990 through the published data of the Ministry of Education, Science, Sports and Culture in Japan. In cases where the same maximum annual increment occurred in two or three successive age classes in a birth year cohort, a new formula (see Eq. 2) was developed to estimate the MIA. The existing formula for estimating MIA was modified to remove the mathematical deficiency (Eq. 1). Estimated MIA shows an overall declining trend, except in birth year cohorts 1934-1951. The effect of World War II on MIA was investigated by a dummy variable regression model. On average, during the birth years 1934-1951, MIA in height decelerated by 1.35 years in boys and 0.54 year in girls, while MIA in weight decelerated by 0.95 year in boys and 0.78 year in girls. Am. J. Hum. Biol. 12:363-370, 2000. Copyright 2000 Wiley-Liss, Inc.

Journal Article↗