Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Files”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,783 records · Page 99Linked to original sources

A web-accessible complete transcriptome of normal human and DMD muscle.

We present an assessment of the complete transcriptome of human skeletal muscle in Duchenne muscular dystrophy patient muscle and non-dystrophic controls (36 RNAs analyzed from ten Duchenne dystrophy and eight controls; approximately 65,000 gene/expressed sequence tag/probe sets queried on U95 five-GeneChip series and MuscleChip). The use of the multiple chip types allowed us to compare results from different probe sets for the same gene: we found excellent concordance between different probe sets on different microarrays. We found 30% of human genes expressed in muscle at detectable levels. Three percent of these showed differential regulation in dystrophin deficiency. Among 1,882 dysregulated probe sets, 1,324 corresponded to characterized genes/proteins (891 non-redundant transcript units), and 588 to expressed sequence tags or predicted genes. Data interpretation was limited to the insulin-like growth factor pathway members, an investigation of possible de-regulation towards a cardiac lineage, and identification of male- and female-specific transcripts. We found transcriptional upregulation of both IGF-I and IGF-II in dystrophic muscle, however the possible beneficial effects of the growth factors appear offset by transcriptional upregulation of inhibitory IGF-binding proteins and regulators (IGFBP-2, -4, -6 and -7; and PRSS11 [IGFBP-5 protease]). We hypothesize that the beneficial effects of IGF-I or IGF-II supplementation in dystrophic muscle may be the result of dose-dependent sequestration of inhibitory IGF-binding proteins. We also focused on six 'cardiac' genes expressed in muscle (alpha-cardiac actin, CARP, CASQ2, troponin T2 cardiac [TNNT2], CUGBP2, and connexin 43). Comparison to a 27 time point murine muscle regeneration series and mdx muscle profiles showed that CARP and Cx43 were macrophage-associated, and TNNT2 activated-myoblast-associated. Upregulation of cardiac actin and CUGBP2 was not associated with muscle regeneration profiles, suggesting a more specific dysregulation induced by dystrophin deficiency. We found two Y-linked genes expressed solely in male muscle (RPS4Y, DDX3Y), and two autosomal genes expressed much more highly in female muscle (GRO2, ZNF91) (all comparisons P<0.01). Finally, we present the first web-accessible expression profiling database for all data, including image files (.dat), processed image files (.cel), and complete comparison files which are publicly available through a novel queriable web site, that permits query-by-gene across all profiles (http://microarray.cnmcresearch.org/pga). These data enumerate the full range of molecular changes associated downstream of dystrophin deficiency, and provide a web-accessible platform to study the specificity of transcriptional pathway alterations in muscle disease.

Animals↗

Risk factors and characteristics of ocular complications, and efficacy of autologous serum tears after haematopoietic progenitor cell transplantation.

The objective of the study was to evaluate the frequency and clinical characteristics of ocular complications and their risk factors, as well as autologous serum tears (AST) for the treatment of dry eye in these patients. Data from the files of 124 patients who had undergone allogeneic haematopoietic progenitor cell transplantation (HPCT) were evaluated. In addition, 33 HPCT patients were examined and their data were compared with controls. Analysis of tears and AST was performed. Dry eye manifestation occurred in 32% of patients and was positively correlated with age over 27 years (P = 0.05), peripheral blood progenitor cell transplant (P = 0.002), chronic graft-versus-host disease (P = 0.0027), and chronic or acute myeloid leukaemia (P = 0.001). Dry mouth and Schirmer test < 5 mm were predictive factors for dry eye in HPCT patients (P = 0.002 and odds ratio 3.9 and P = 0.007, odds ratio = 5.9, respectively). Microbiological analysis revealed that six of 11 AST samples were contaminated after 30 days of use. The present study supports the role of potential risk factors for ocular complications and key elements to detect alterations in the tear film from HPCT patients. In addition, AST contamination must be considered after longer periods of use.

Adolescent↗

Object-oriented data handler for sequence analysis software development.

We report an object-oriented data handler and supplementary tools for the development of molecular genetics application software for various sequence analyses. Our data handler has a flexible and expandable format that supports the most common data types for molecular genetic software. New data types can be constructed in an object-oriented manner from the basic units. The data handler includes an object library, a format-converting program and a viewer that can visualize simultaneously the data contained in several files to construct a general picture from separate data. This software has been implemented on an IBM PC-compatible personal computer.

Atrial Natriuretic Factor↗

A tool-kit for cDNA microarray and promoter analysis.

We describe two sets of programs for expediting routine tasks in analysis of cDNA microarray data and promoter sequences. The first set permits bad data points to be flagged with respect to a number of parameters and performs normalization in three different ways. It allows combining of result files into comprehensive data sets, evaluation of the quality of both technical and biological replicates and row and/or column standardization of data matrices. The second set supports mapping ESTs in the genome, identifying the corresponding genes and recovering their promoters, analyzing promoters for transcription factor binding sites, and visual representation of the results. The programs are designed primarily for Arabidopsis thaliana researchers, but can be adapted readily for other model systems. Availability and Supplementary information: http://www.personal.psu.edu/nhs109/Programs/

Algorithms↗

Long-term study of accommodative esotropia.

PURPOSE: Previous studies of accommodative esotropia have been hampered by bias-prone methods of data collection and analysis and by small sample size. The studies have conflicting conclusions, causing uncertain results. This study aims to determine long-term results of standard treatment of accommodative esotropia and identify predictors of outcome, while minimizing bias in data collection and analysis, using the largest possible sample size. METHODS: A research assistant collected data from all files of a large, long-established pediatric ophthalmology practice (M.M.P.). The assistant was given standardized collection forms that allowed inclusion of all patient data points over all visits. The assistant was masked as to study goals. She was instructed to include any patient with esotropia who had been prescribed glasses during treatment. Descriptive terms were converted to code numbers. A second, similarly masked research assistant entered data into a computerized database. Criteria for patient inclusion were designed to conform to earlier studies by I.H.L. and M.M.P. and were implemented by computer. RESULTS: The database totaled 1,307 patients (747,717 data points). Of these, 354 qualified for this analysis. A greater difference between near and distance esodeviation (AC/A relationship) correlated with a higher rate of deterioration of accommodative esotropia control (P<.0001). Deterioration also positively correlated with earlier age at onset, inferior oblique overaction, and amblyopia. CONCLUSIONS: This study agrees with our previous findings that a high AC/A relationship increases the likelihood of deterioration of accommodative esotropia, thus confirming the integrity of the database. This unique, unbiased dataset will be used for future analyses of esotropia.

Accommodation, Ocular↗

Effects of drug administration in pregnancy on children's school behaviour.

Files with prescription data were used to assess possible behavioural changes in children, whose mothers used benzodiazepines or neuroleptic drugs during the second half of their pregnancy. Prescriptions, bearing the identification number of women resident in one district of Prague, filed in pharmacies during 1974 and the first three months of 1975 represent the first part of the data. During 1984, children born in the appropriate earlier period were searched and linked with the earlier prescription data. A group of 68 children with possible exposure to neuroleptics and a group of 15 children possibly exposed to diazepam during the second half of their intrauterine development were found. Two groups of 55 and 7 children, respectively, born of mothers without exposure to these drugs, were chosen as controls. The teachers of classes attended by these children were addressed by a letter and asked to evaluate their behaviour at school. This was done by means of a form containing analogue scales evaluating different features of behaviour. Each child was compared with its control. The statistical evaluation with Student's t-test, regression analysis and analysis of variance did not reveal any significant difference between both groups and their controls.

Antipsychotic Agents↗

Efficient transmission of compressed data for remote volume visualization.

One of the goals of telemedicine is to enable remote visualization and browsing of medical volumes. There is a need to employ scalable compression schemes and efficient client-server models to obtain interactivity and an enhanced viewing experience. First, we present a scheme that uses JPEG2000 and JPIP (JPEG2000 Interactive Protocol) to transmit data in a multi-resolution and progressive fashion. The server exploits the spatial locality offered by the wavelet transform and packet indexing information to transmit, in so far as possible, compressed volume data relevant to the clients query. Once the client identifies its volume of interest (VOI), the volume is refined progressively within the VOI from an initial lossy to a final lossless representation. Contextual background information can also be made available having quality fading away from the VOI. Second, we present a prioritization that enables the client to progressively visualize scene content from a compressed file. In our specific example, the client is able to make requests to progressively receive data corresponding to any tissue type. The server is now capable of reordering the same compressed data file on the fly to serve data packets prioritized as per the client's request. Lastly, we describe the effect of compression parameters on compression ratio, decoding times and interactivity. We also present suggestions for optimizing JPEG2000 for remote volume visualization and volume browsing applications. The resulting system is ideally suited for client-server applications with the server maintaining the compressed volume data, to be browsed by a client with a low bandwidth constraint.

Algorithms↗

Software for tabular data protection.

In order for national statistical offices to maintain the trust of the public to collect data and publish statistics of importance to society and decision-making, it is imperative that respondents (persons or establishments) be guaranteed privacy and confidentiality in return for providing requested confidential data. Consequently, for most survey and census data, disclosure limitation techniques must be applied before the data are ready for public release. For microdata, examples of methods that can be used to identify respondents include directly extracting identifying information from microdata files or indirectly identifying respondents by matching a given file with an external file. For tabular data, respondents may be identified directly from small cell counts or respondent contributions to heavily concentrated cells of magnitude data may be closely approximated by the cell value. Indirect disclosure is possible in tables through manipulation of additive tabular relationships between cell values and totals, e.g. manipulating rows and column totals in a two-dimensional table. Two-dimensional statistical tables are a staple of official statistics. This paper describes a desktop software system that for the first time implements within a single framework four standard disclosure limitation techniques for protecting tabular data in two-dimensional tables: complementary cell suppression, minimum-distance controlled rounding, unbiased controlled rounding, and controlled rounding subject to subtotals constraints, and a fifth, new method: controlled tabular adjustment, and summarizes the five methods.

Computer Security↗

Distant recurrence in breast cancer. Survival expectations and first choice of chemotherapy regimen.

Controversial questions in recurrent breast cancer include the magnitude of the survival benefit of combination chemotherapy and the best choice of first line chemotherapy. Data from the files of the Danish Breast Cancer Cooperative Group (DBCG) show that with current systemic treatment median survival after distant recurrence is 19 months. Since historical data from the pre-chemotherapy era indicate a median survival of 12 months, the survival benefit of standard chemotherapy appears to be around 7 months in the average patient. The DBCG trial 80-R2 is the largest randomized trial of CAF (cyclophosphamide, doxorubicin, 5-fluorouracil) versus CMF (cyclophosphamide, methotrexate, 5-fluorouracil) in recurrent breast cancer. A review of this study and 6 other similar studies shows that CAF is clearly superior to CMF in terms of better tumor shrinkage, prolonged overall time to progression, and decreased need of secondary therapy. The adverse effects of the two treatments are largely comparable, but CAF causes severe alopecia and is more expensive than CMF. On balance, the existing evidence indicates that CAF rather than CMF should be chosen as first line chemotherapy in recurrent breast cancer.

Antineoplastic Combined Chemotherapy Protocols↗

[Quality of data acceptable for perinatal epidemiology surveillance: assessment of the health certificate at birth and the national obstetrics medical file. Study in three Seine-Maritime maternal wards].

Data from several sources could be used for perinatal epidemiology surveillance aimed at an assessment of regional programs such as those proposed by the Superior Committee for Public Health. A retrospective study of 561 births was conducted in three maternity wards in the French Seine Maritime department in order to evaluate the reliability of two data sources: the national obstetrics medical file and the health certificate at birth. The delivery room records were used as the gold standard. The sensitivity of the obstetrics file was better than that of the health certificate. With the obstetrics file, it was possible to identify almost all the vaginal route interventions, almost all the premature births and all the cesareans. With the health certificate, 39-58% of the vaginal route interventions, 61% of the premature births and 61-72% of the cesareans performed in the three wards studied were identified. The quality of data in the obstetrics file appears to be better than that in the health certificate but only concerns 40% of births in the geographical area studied. Inversely, the health certificate is theoretically delivered for all births (actually delivered for 93%). Integrating these two information systems could be an optimum solution.

Bias↗

usiGrabber: automating the curation of proteomics spectra data at scale, making large datasets ready for use in machine learning systems.

MOTIVATION: An unprecedented amount of mass spectrometry-based proteomics data is publicly available through repositories such as the PRoteomics IDEntifications Database (PRIDE), and the field is increasingly leveraging machine-learning approaches. However, the available data is not ready to be reused in a scalable way beyond the original acquisition purpose. Existing machine learning models commonly rely on a few manually curated datasets that require deep domain expertise and tedious technical work to construct. Importantly, these datasets have not been updated in recent years, so that newly published data remains inaccessible. We present usiGrabber, a scalable framework for assembling large proteomic datasets. usiGrabber is designed around portability and extensibility. It extracts spectra identification data from mzIdentML files, stores additional project-level metadata retrieved through the PRIDE API, indexes raw spectra using Universal Spectrum Identifiers (USIs), and offers download utilities to retrieve spectra data at scale. RESULTS: Within 49&#x2009;h, we parsed over 800 million peptide spectrum matches and corresponding USIs from over 1200 projects. As a proof of concept, we used usiGrabber to construct a phosphorylation-specific training dataset of nearly 11 million spectra in under 2 days and used it to retrain a binary phosphorylation classifier based on the AHLF model architecture. With a balanced accuracy of 0.78, our model achieves comparable performance to the original model on an independent test set, showing that automated data extraction is an alternative to manual curation of static datasets. AVAILABILITY AND IMPLEMENTATION: All code is available at https://github.com/usiGrabber/usiGrabber; the data are available at https://zenodo.org/records/18853258.

Machine Learning↗

[Dental consultations at a university hospital in Toulouse: the role for conservative treatments in oral health].

A retrospective study was carried out looking at the files of 481 patients who were seen for their first consultation at a university hospital dental school clinic. Their needs for conservative treatment were evaluated and the data from their files were analysed, including the following elements: demography, general health status, the expressed reason for the consultation, and the initially planned treatment. The descriptive analysis reveals a relatively young population group with an average age of 43.8 years, for whom 32% of them had existing or past medical conditions which necessitated that certain precautions be taken when prescribing conservative treatment. The most frequently given reasons for the consultation were the need for an annual dental check-up (41%) and pain (20%). 74% of the patients were in need of conservative treatments, and a priority for an initial treatment was cited for 48.4% of them.

Adult↗

[Epidemiologic study of the clinical features, diagnosis, and treatment of kidney trauma in Cantabria].

OBJECTIVE: This study reviews the clinical features, methods and treatment of renal trauma (RT). METHODS: The epidemiological study was conducted using data from 340 cases of RT that had been treated in the Department of Urology of the Marqués de Valdecilla University Hospital in Santander. All the information was filed in a data base specifically for the purpose. RESULTS: Statistical analyses of the data showed the incidence of the symptoms, diagnosis and treatment of RT. CONCLUSIONS: Radiological and ultrasound evaluation are still fundamental for the diagnosis of RT. Only 4% of these patients have died. Abdominal, thoracic and cranial injuries cause the most severe lesions. Concerning treatment, only 17% of the cases with RT required surgery.

Adolescent↗

Analysis of multiple-choice items.

This paper deals with scoring and analysis of multiple-choice exam questions. Data may be read by an optical mark reader scanner or from a file created by a system editor and stored in a format convenient to be analysed by a specially designed program 'MCQ'. In addition the data file created can be linked with the package SPSSx for pertinent statistical methods or SPSS Graphics for illustration of score and item distributions. The MCQ program comprises: (1) scoring; (2) item analysis: difficulty and discrimination indices; (3) testing the goodness of alternatives attached to each question. The method is applicable to multiple-choice questions with 4.5 greater than 5; true/false and combined exam questions. The method should prove useful for instructors to build up a balanced discriminant questions bank.

Computer Graphics↗

Longitudinal patterns of Medicare use by cause of death.

To study the use of health services before death for different causes, a 6-year file of Medicare use and cost data was linked to a file of death certificate information for persons dying at ages 65 years or over in 1979. Patterns of medical care use during the last years of life varied substantially by cause of death, reflecting the degree of chronicity of the disease that resulted in death and the nature of treatment. Persons dying of nephritis, chronic obstructive pulmonary disease, and diabetes mellitus incurred consistently high expenses for 6 years before death. Costs for cancer decedents were also high, especially in the last 2 years of life. Persons in their last 2 years of life have a considerable impact on Medicare expenses. An estimated 13 percent of annual Medicare expenses were attributable to persons who were within 2 years of death from heart disease and 10.7 percent to persons who were within 2 years of death from cancer.

Age Factors↗

An interactive data base system for assessing and managing outpatient experiences in residency training.

A computerized system is described for assessing and managing a residency program's outpatient experience for its residents. The utilization of MUMPS hierarchical files allows rapid interactive searches, making this system an effective tool for its intended uses. The rationale behind the design of the data files and searches and the usefulness of the data in residency training are discussed.

Ambulatory Care↗

Issues in cross-national comparisons of crash data.

While national accident record systems appear to provide an attractive method of comparing the effectiveness of traffic safety programs, differences in definitions, data collection methods and file structure may lead to misleading conclusions. Variations in crash definitions and methods for measuring alcohol involvement are described and illustrated with data from US crash files. Examples of international comparisons are provided.

Accidents, Traffic↗