Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Datasets as Topic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Estimates of the number of cancer patients hospitalized in a geographic area using claims data without a unique personal identifier.

OBJECTIVE: In French national claims databases, claims are currently anonymous i.e. not linked to individual patients. In order to improve our estimate of the medical activity related to cancer in one French region, a statistical method was developed to use claims data to assess the number of cancer patients hospitalized in acute care. METHODS: This method used the medical and administrative information available in the claims (i.e. age, primary site, length of stay) to predict an average number of stays per patient, followed by a number of patients. It was based on a two-phase study design using an internal dataset which contained personal identifiers to estimate the model parameters. RESULTS: The predicted number of acute care patients hospitalized in one or several health care centers in one French region was 38,109 with a 95% predictive interval (37,990; 38,228) for the first six months of 2002. A prediction error of 24 per thousand was found. CONCLUSION: We provide a good estimate of the morbidity in acute care hospitals using claims data that is not linked to individual patients. This estimate reflects the medical activity and can be used to anticipate acute care needs.

Adolescent↗

A population needs assessment profile for dementia.

The Tayside Profile for Dementia Planning is an instrument designed to obtain data for population needs assessment and planning. It provides a brief tool to collect a minimum dataset by non-specialists. Third-party informants-informal carers or involved professionals-are used as data sources. The key concept is the use of a descriptive profile rather than a summative score or categorization. The profile consists of a set of needs indicators, information on current service response and demographic and background data. Key levels of dependency are measured by time interval dependency. Validity, reliability, acceptability and usability are satisfactory, with the crucial exception that informal carers and professionals appear to perceive needs differently. Further research is needed to assess which type of informant provides the more useful data.

Caregivers↗

Publishing large proteome datasets: scientific policy meets emerging technologies.

Currently, there are various approaches to proteomic analyses based on either 2D gel or HPLC separation platforms, generating data of different formats, structures and types. Identification of these separated proteins or peptide fragments is typically achieved by mass spectrometry (MS) measurements that use either accurate mass measurements or fragmentation (MS-MS) information. Integrating the information generated from these different platforms is essential if proteomics is to succeed. A further challenge lies in generating standards that can accept the hundreds-of-thousands of mass spectra produced per analysis based on threshold or probability measurements. Finally, peer review and electronic publication processes will be crucial to the dissemination and use of proteomic information. Merging the policy requirements of data-intensive research with information technology will enable scientists to gain real value from global proteomics information.

Chromatography, Liquid↗

Quantification of 99Tcm-HMPAO brain SPET in two series of healthy volunteers using different triple-headed SPET configurations: normal databases and methodological considerations.

We evaluated the methodological issues underlying the assessment of normal confidence intervals, as used in clinically based region-of-interest (ROI) semi-quantification of 99Tcm-HMPAO brain SPET. At two different centres equipped with high-resolution, triple-headed gamma cameras, HMPAO SPET scans were performed on two groups of 24 and 15 healthy volunteers respectively. Together with an operator-defined analysis (ODA), a semi-automated analysis (SAA) was conducted on the normal datasets in one centre. Tests of intra- and inter-observer variability were performed. Repeat scans were performed within 72 h after the first to analyse short-term regional inter-study variations. The overall regional uptake showed significant differences in most regions between both normal datasets. Intra-observer and inter-observer reproducibility were on average within 4% for the ODA, while for the SAA it was less than 1%. Inter-study variations were excellent for both centres, ranging from -4% to +3% for most regions studied. The variability in clinical brain perfusion studies largely depends on the reproducibility of the data analysis technique. A semi-automated approach shows clear advantages over an entirely operator-defined approach. Intra-subject repeat studies show enough stability for use as reliable baseline measurements in the construction of a normal database or to allow activation studies with high sensitivity.

Adult↗

Creating and validating an algorithm to measure AIDS mortality in the adult population using verbal autopsy.

BACKGROUND: Vital registration and cause of death reporting is incomplete in the countries in which the HIV epidemic is most severe. A reliable tool that is independent of HIV status is needed for measuring the frequency of AIDS deaths and ultimately the impact of antiretroviral therapy on mortality. METHODS AND FINDINGS: A verbal autopsy questionnaire was administered to caregivers of 381 adults of known HIV status who died between 1998 and 2003 in Manicaland, eastern Zimbabwe. Individuals who were HIV positive and did not die in an accident or during childbirth (74%; n = 282) were considered to have died of AIDS in the gold standard. Verbal autopsies were randomly allocated to a training dataset (n = 279) to generate classification criteria or a test dataset (n = 102) to verify criteria. A rule-based algorithm created to minimise false positives had a specificity of 66% and a sensitivity of 76%. Eight predictors (weight loss, wasting, jaundice, herpes zoster, presence of abscesses or sores, oral candidiasis, acute respiratory tract infections, and vaginal tumours) were included in the algorithm. In the test dataset of verbal autopsies, 69% of deaths were correctly classified as AIDS/non-AIDS, and it was not necessary to invoke a differential diagnosis of tuberculosis. Presence of any one of these criteria gave a post-test probability of AIDS death of 0.84. CONCLUSIONS: Analysis of verbal autopsy data in this rural Zimbabwean population revealed a distinct pattern of signs and symptoms associated with AIDS mortality. Using these signs and symptoms, demographic surveillance data on AIDS deaths may allow for the estimation of AIDS mortality and even HIV prevalence.

Adolescent↗

Evaluation of a novel infrared range vibration-based descriptor (EVA) for QSAR studies. 1. General application.

A novel molecular descriptor (EVA) based upon calculated infrared range vibrational frequencies is evaluated for use in QSAR studies. The descriptor is invariant to both translation and rotation of the structures concerned. The method was applied to 11 QSAR datasets exhibiting both a range of biological endpoints and various degrees of structural diversity. This study demonstrates that robust QSAR models can be obtained using the EVA descriptor and examines the effect of EVA parameter changes on these models; recommendations are made as to the appropriate choice of parameters. The performance of EVA was found to be comparable in statistical terms to that of CoMFA, despite the fact that EVA does not require the generation of a structural alignment. Models derived using semiempirical (MOPAC AM1 and PM3) and AMBER mechanics calculated normal mode frequencies are compared, with the overall conclusion that the semiempirical methods perform equally well and both outperform the AMBER-based models.

Computer Simulation↗

Comparison of alternative methods for assessing injury severity based on anatomic descriptors.

BACKGROUND: There is mounting confusion as to which anatomic scoring systems can be used to adequately control for trauma case mix when predicting patient survival. METHODS: Several Abbreviated Injury Scale (AIS) and International Classification of Disease Clinical (ICD-9CM)-based methods of scoring severity were compared by using data from the Pennsylvania Trauma Outcome Study. By using a design dataset, the probability of survival was modeled as a function of each score or profile. Resulting coefficients were used to derive expected probabilities in a test dataset; expected and observed probabilities were then compared by using standard measures of discrimination and calibration. RESULTS: The modified Anatomic Profile, Anatomic Profile, and New Injury Severity Score outperformed the International Classification of Disease-based Injury Severity Score. This finding remains true when AIS values are obtained by means of a conversion from International Classification of Disease to AIS. CONCLUSION: Results support the integrity of the AIS and argue for its continued use in research and evaluation. The modified Anatomic Profile, Anatomic Profile, and New Injury Severity Score, however, should be used in preference to the Injury Severity Score as an overall measure of severity.

Humans↗

DMARD use in early rheumatoid arthritis. Lessons from observations in patients with established disease.

The concept of early and aggressive therapy of rheumatoid arthritis (RA) has been well documented in the past years. It includes immediate DMARD institution after diagnosis, the use of the most effective DMARDs, and rapid switching of regimens if a level of disease activity close to remission is not achieved. In this review we briefly explore to what degree this new concept has been implemented in routine clinical care. Based on an observational dataset comprising 3342 DMARD courses, we present evidence of a change in DMARD patterns in newly diagnosed RA patients towards a higher prescription rate of more aggressive drugs like methotrexate (MTX), as well as a decreasing lag time until MTX was instituted in RA patients over the years. One consequence of recent changes in therapeutic strategies is that comparative analyses of formerly versus recently employed DMARDs will be considerably biased in observational studies. By contrast to changes in DMARD usage, survey data show neither a shortening of referral time nor a change in the approach to diagnose early RA. These data indicate a need for more dissemination of the early arthritis concept.

Adult↗

Diverse prognosis in metastatic breast cancer: who should be offered alternative initial therapies?

In an attempt to clarify appropriate treatment options for women with stage IV breast cancer, we studied the survival experience of a large dataset of patients treated on Cancer and Leukemia Group B (CALGB) protocols. The study, restricted to women who had had no prior chemotherapy for metastatic disease, demonstrated a surprisingly poor prognosis, with an estimated median survival of 1.6 years and only 26% alive at 3 years. Analysis of prognostic factors permitted the identification of subsets with even shorter survival, such as women with estrogen receptor negative tumor in more than one metastatic site and prior adjuvant chemotherapy. We feel that an evaluation of intensive investigational treatment approaches, such as trials using autologous bone marrow transplantation, is justified for most stage IV breast cancer patients, in view of their poor prognosis.

Biomarkers, Tumor↗

Neural network based automated algorithm to identify joint locations on hand/wrist radiographs for arthritis assessment.

Arthritis is a significant and costly healthcare problem that requires objective and quantifiable methods to evaluate its progression. Here we describe software that can automatically determine the locations of seven joints in the proximal hand and wrist that demonstrate arthritic changes. These are the five carpometacarpal (CMC1, CMC2, CMC3, CMC4, CMC5), radiocarpal (RC), and the scaphocapitate (SC) joints. The algorithm was based on an artificial neural network (ANN) that was trained using independent sets of digitized hand radiographs and manually identified joint locations. The algorithm used landmarks determined automatically by software developed in our previous work as starting points. Other than requiring user input of the location of nonanatomical structures and the orientation of the hand on the film, the procedure was fully automated. The software was tested on two datasets: 50 digitized hand radiographs from patients participating in a large clinical study, and 60 from subjects participating in arthritis research studies and who had mild to moderate rheumatoid arthritis (RA). It was evaluated by a comparison to joint locations determined by a trained radiologist using manual tracing. The success rate for determining the CMC, RC, and SC joints was 87%-99%, for normal hands and 81%-99% for RA hands. This is a first step in performing an automated computer-aided assessment of wrist joints for arthritis progression. The software provides landmarks that will be used by subsequent image processing routines to analyze each joint individually for structural changes such as erosions and joint space narrowing.

Algorithms↗

Secured distributed service to manage biological data on EGEE grid.

Biological data are most times published and then become public ones. They, then, do not need to be isolated or encrypted. But, in some cases, these data stemed from patients or are analyzed with, for instance, pharmaceutical or agronomics goals. Also in simple ways , these data, before to become public, have to be kept confidential while researchers haven't been able to publish their work or to register them. So they are a lot of cases where the integrity and the confidentiality of biological data have to be protected against unauthorized accesses. But, as these private data are also large datasets, they need high-throughput computing and huge data storage to processed, such as ones produced by complete genome projects. These requirements are enhanced in the context of a Grid such EGEE, where the computing and storage resources are distributed across a large-scale platform. We have developed a secured distributed service to manage biological data on grid: the EncFile encrypted files management system. We have deployed it on the production platform of the EGEE grid project. Thus we provided grid users with a user-friendly component that doesn't require any user privileges. And we have integrated into a bioinformatics grid portal associated to encrypted representative biological resources: world-famous databases and programs.

Computational Biology↗

Long terminal repeat retrotransposons of Mus musculus.

BACKGROUND: Long terminal repeat (LTR) retrotransposons make up a large fraction of the typical mammalian genome. They comprise about 8% of the human genome and approximately 10% of the mouse genome. On account of their abundance, LTR retrotransposons are believed to hold major significance for genome structure and function. Recent advances in genome sequencing of a variety of model organisms has provided an unprecedented opportunity to evaluate better the diversity of LTR retrotransposons resident in eukaryotic genomes. RESULTS: Using a new data-mining program, LTR_STRUC, in conjunction with conventional techniques, we have mined the GenBank mouse (Mus musculus) database and the more complete Ensembl mouse dataset for LTR retrotransposons. We report here that the M. musculus genome contains at least 21 separate families of LTR retrotransposons; 13 of these families are described here for the first time. CONCLUSIONS: All families of mouse LTR retrotransposons are members of the gypsy-like superfamily of retroviral-like elements. Several different families of unrelated non-autonomous elements were identified, suggesting that the evolution of non-autonomy may be a common event. High sequence similarity between several LTR retrotransposons identified in this study and those found in distantly-related species suggests that horizontal transfer has been a significant factor in the evolution of mouse LTR retrotransposons.

Animals↗

YeastHub: a semantic web use case for integrating data in the life sciences domain.

MOTIVATION: As the semantic web technology is maturing and the need for life sciences data integration over the web is growing, it is important to explore how data integration needs can be addressed by the semantic web. The main problem that we face in data integration is a lack of widely-accepted standards for expressing the syntax and semantics of the data. We address this problem by exploring the use of semantic web technologies-including resource description framework (RDF), RDF site summary (RSS), relational-database-to-RDF mapping (D2RQ) and native RDF data repository-to represent, store and query both metadata and data across life sciences datasets. RESULTS: As many biological datasets are presently available in tabular format, we introduce an RDF structure into which they can be converted. Also, we develop a prototype web-based application called YeastHub that demonstrates how a life sciences data warehouse can be built using a native RDF data store (Sesame). This data warehouse allows integration of different types of yeast genome data provided by different resources in different formats including the tabular and RDF formats. Once the data are loaded into the data warehouse, RDF-based queries can be formulated to retrieve and query the data in an integrated fashion. AVAILABILITY: The YeastHub website is accessible via the following URL: http://yeasthub.gersteinlab.org.

Biology↗

Measuring coral reef decline through meta-analyses.

Coral reef ecosystems are in decline worldwide, owing to a variety of anthropogenic and natural causes. One of the most obvious signals of reef degradation is a reduction in live coral cover. Past and current rates of loss of coral are known for many individual reefs; however, until recently, no large-scale estimate was available. In this paper, we show how meta-analysis can be used to integrate existing small-scale estimates of change in coral and macroalgal cover, derived from in situ surveys of reefs, to generate a robust assessment of long-term patterns of large-scale ecological change. Using a large dataset from Caribbean reefs, we examine the possible biases inherent in meta-analytical studies and the sensitivity of the method to patchiness in data availability. Despite the fact that our meta-analysis included studies that used a variety of sampling methods, the regional estimate of change in coral cover we obtained is similar to that generated by a standardized survey programme that was implemented in 1991 in the Caribbean. We argue that for habitat types that are regularly and reasonably well surveyed in the course of ecological or conservation research, meta-analysis offers a cost-effective and rapid method for generating robust estimates of past and current states.

Animals↗

Tuber on a chip: differential gene expression during potato tuber development.

Potato tuber development has proven to be a valuable model system for studying underground sink organ formation. Research on this topic has led to the identification of many genes involved in this complex process and has aided in the unravelling of the mechanisms underlying starch synthesis. However, less attention has been paid to the biochemical pathways of other important metabolites or to the changing metabolic fluxes occurring during potato tuber development. In this paper, we describe the construction of a potato complementary DNA (cDNA) microarray specifically designed for genes involved in processes related to tuber development and tuber quality traits. We present expression profiles of 1315 cDNAs during tuber development where the predominant profiles were strong up- and down-regulation. Gene expression profiles showing transient increases or decreases were less abundantly represented and followed more moderate changes, mainly during tuber initiation. In addition to the confirmation of gene expression patterns during tuber development, many novel differentially expressed genes were identified and are considered as candidate genes for direct involvement in potato tuber development. A detailed analysis of starch metabolism genes provided a unique overview of expression changes during tuber development. Characteristic expression profiles were often clearly different between gene family members. A link between differential gene expression during tuber development and potato tissue specificity is described. This dataset provides a firm basis for the identification of key regulatory genes in a number of metabolic pathways that may provide researchers with new tools to achieve breeding goals for use in industrial applications.

Journal Article↗

Are decisions using cost-utility analyses robust to choice of SF-36/SF-12 preference-based algorithm?

BACKGROUND: Cost utility analysis (CUA) using SF-36/SF-12 data has been facilitated by the development of several preference-based algorithms. The purpose of this study was to illustrate how decision-making could be affected by the choice of preference-based algorithms for the SF-36 and SF-12, and provide some guidance on selecting an appropriate algorithm. METHODS: Two sets of data were used: (1) a clinical trial of adult asthma patients; and (2) a longitudinal study of post-stroke patients. Incremental costs were assumed to be 2000 dollars per year over standard treatment, and QALY gains realized over a 1-year period. Ten published algorithms were identified, denoted by first author: Brazier (SF-36), Brazier (SF-12), Shmueli, Fryback, Lundberg, Nichol, Franks (3 algorithms), and Lawrence. Incremental cost-utility ratios (ICURs) for each algorithm, stated in dollars per quality-adjusted life year (dollars/QALY), were ranked and compared between datasets. RESULTS: In the asthma patients, estimated ICURs ranged from Lawrence's SF-12 algorithm at 30,769 dollars/QALY (95% CI: 26,316 to 36,697) to Brazier's SF-36 algorithm at 63,492 dollars/QALY (95% CI: 48,780 to 83,333). ICURs for the stroke cohort varied slightly more dramatically. The MEPS-based algorithm by Franks et al. provided the lowest ICUR at 27,972 dollars/QALY (95% CI: 20,942 to 41,667). The Fryback and Shmueli algorithms provided ICURs that were greater than 50,000 dollars/QALY and did not have confidence intervals that overlapped with most of the other algorithms. The ICUR-based ranking of algorithms was strongly correlated between the asthma and stroke datasets (r = 0.60). CONCLUSION: SF-36/SF-12 preference-based algorithms produced a wide range of ICURs that could potentially lead to different reimbursement decisions. Brazier's SF-36 and SF-12 algorithms have a strong methodological and theoretical basis and tended to generate relatively higher ICUR estimates, considerations that support a preference for these algorithms over the alternatives. The "second-generation" algorithms developed from scores mapped from other indirect preference-based measures tended to generate lower ICURs that would promote greater adoption of new technology. There remains a need for an SF-36/SF-12 preference-based algorithm based on the US general population that has strong theoretical and methodological foundations.

Adult↗

From terminology to terminology services.

Terminologies have traditionally been considered as static datasets held in books or databases. The GALEN Terminology Server presents a prototype for a new view of terminologies delivered as a set of functions and services provided to other applications. This facilitates their development and integration as part of a strategy for sharing and re-using information and knowledge. The essential features of the Terminology server are the functions which it can perform; questions which it can answer and statements which it can be told. The GALEN Terminology Server supports these operations through a modular architecture and uniform applications programming interface which allows client applications to ignore the internal structure and simply use the Server for terminological, coding, and linguistic functions.

Computer Systems↗

The Prostate, Lung, Colorectal and Ovarian (PLCO) Cancer Screening Trial of the National Cancer Institute: history, organization, and status.

The Prostate, Lung, Colorectal and Ovarian (PLCO) Cancer Screening Trial is enrolling 148,000 men and women ages 55-74 at ten screening centers nationwide with balanced randomization to intervention and control arms. For prostate cancer, men receive a digital rectal examination and a blood test for prostate-specific antigen. For lung cancer, men and women receive a posteroanterior view chest X-ray. For colorectal cancer, men and women undergo a 60-cm flexible sigmoidoscopy. For ovarian cancer, women receive a blood test for the CA125 tumor marker and transvaginal ultrasound. Members of the control arm continue with their usual care. Follow-up in both groups will continue for at least 13 years from randomization to assess health status and cause of death. The primary endpoint is mortality from the four PLCO cancers, which accounts for about 53% of all cancer deaths in men and 41% of cancer deaths in women in the United States each year. Blood specimens are collected from screened participants, buccal cell DNA from controls, and histology slides from cases; these are maintained in a biorepository. Participants complete a baseline questionnaire (covering health status and risk factors) and a dietary questionnaire. More than 12,000 participants were enrolled in the pilot phase (concluded in September 1994). Changes in the eligibility criteria followed. As of April 2000, enrollment exceeded 144,500. Data are scanned into designated on-site computers for uploading by participant identification number to the coordinating center for quality checks, archival storage, and preparation of analysis datasets for use by the National Cancer Institute (NCI). Scientific direction is provided by NCI scientists, trial investigators, external consultants, and an independent data safety and monitoring board. Performance and data quality are monitored via data edits, site visits, random record audits, and teleconferences. The PLCO trial is formally endorsed by the American Cancer Society and has been ranked by the American Urological Association as one of the most important prostate cancer studies being conducted. Special efforts to enroll black participants are cosponsored by the U.S. Centers for Disease Control and Prevention.

Aged↗