Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Files”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Basics of flow cytometry.

In summary, a beginner requires fundamental knowledge about flow cytometric instrumentation in order to effectively use this technology. It is important to remember that flow cytometers are very complex instruments that are composed of four closely related systems. The fluidic system transports particles from a suspension through the cytometer for interrogation by an illumination system. The resulting light scattering and fluorescence is collected, filtered, and converted into electrical signals by the optical and electronics system. The data storage and computer control system saves acquired data and is also the user interface for controlling most instrument functions. These four systems provide a very unique and powerful analytical tool for researchers and clinicians. This is because they analyze the properties of individual particles, and thousands of particles can be analyzed in a matter of seconds. Thus, data for a flow cytometric sample are a collection of many measurements instead of a single bulk measurement. Basic knowledge of instrumentation is a tremendous aid to designing experiments that can be successfully analyzed using flow cytometry. For example, it is important to know the emission wavelength of the laser in the instrument that will be used for analysis. This wavelength is critical knowledge for selecting probes. It is also important to understand that a different range of wavelengths is detected for each fluorescent channel. This will aid selection of probes that are compatible with the flow cytometer. Understanding the complication that emission spectra overlap contributes to detection can be used to guide fluorochrome selections for multicolor analysis. All of these experiment design considerations that rely on knowledge of how flow cytometers work are a very practical and effective means of avoiding wasted time, energy, and costly reagents. Data analysis is a paramount issue in flow cytometry. Analysis includes interpreting as well as presenting data that has been stored in list-mode files. Data analysis is very graphically oriented. There are a number of types of graphic representation that are available to visually aid data analysis. Two standard types of displays are used. These data plots are one-parameter histograms and bivariate plots. A user must be familiar with these two fundamental types of display in order to effectively analyze data. Histograms are the most simple modes of data representation. Histograms allow visualization of a single acquired parameter. Mean fluorescence and distributional statistics can be obtained based on markers that the user can graphically set on the plot. Percentages of positively expressing particles relative to a control sample can also obtained in a similar manner. In addition, multiple histograms can be overlayed on one another to depict qualitative and quantitative differences in two or more samples. Two-parameter data plots are somewhat more complicated than histograms; however, they can yield more information. Two-parameter dot plots of FSC vs SSC allow visualization of both light-scattering parameters that are important for identifying populations of interest. Bivariate fluorescent plots allow discrimination of dual-labeled populations that might remain hidden if histograms were used to display fluorescent data. Two-parameter plots that combine one light-scattering parameter and a fluorescent parameter are useful for analyzing control samples to elucidate the origin of nonspecific binding. Data analysis is very graphically oriented. Experience and pattern recognition become important when using two-parameter data plots for qualitative as well as quantitative analysis. The technique of gating or drawing regions on dual parameter light-scatter plots allows one to exclude information and examine the population of interest by disallowing particles that might confound or interfere with analysis. This is one of the fundamental uses for gating. (ABSTRACT TRUNCATED)

Animals↗

A microcomputer data base system for maternal serum alpha fetoprotein screening programmed in BASIC.

A computer program was developed for the IBM PC computer to be used for maternal serum alpha fetoprotein screening (MSAFP). This program, written in BASIC, 1) maintains a data base for patients tested, 2) allows input and storage of MSAFP results, 3) calculates the gestational age dependent result in multiple of the median (MOM), 4) makes an individualized interpretation of the result based on maternal age, weight and diabetes, 5) makes an appropriate recommendation for subsequent action based on the result, and 6) prints a report containing the above information. In addition, the program will print a daily log of patients tested, provides a follow-up sheet to assist in tracking abnormal results to term, converts the data files into files which can be easily intergrated into more powerful data base programs, and generates a monthly statement for billing purposes. The program can be easily modified by someone with minimal training in the BASIC programming language.

Female↗

Longevity and efficiency associated with age structures of female pigs and herd management in commercial breeding herds.

Annual performance measurements, age structures of female pig inventories, and by-parity culling rates were abstracted from data files of 110 herds that participated in a data-share program in Japan. Parity at culling was used as a prime measurement of longevity, whereas pigs weaned x mated female(-1) x year(-1) (PWMFY) was a prime measurement of reproductive efficiency. High or low longevity herds were based on the greatest 50% of the herds or the remaining herds ranked by parity at culling, whereas high or low reproductive efficiency herds were grouped according to the greatest 50% of the herds or the remaining herds ranked by PWMFY. Measurements were analyzed as a 2 x 2 factorial arrangement, using the main effects of the 2 herd groups of longevity (high or low) and reproductive efficiency (high or low). Means of parity at culling and PWMFY were 4.6 (SD = 0.82) and 21.2 (SD = 3.02), respectively. The high longevity group had 1.27 greater parities at culling than the low longevity group (P < 0.05), but no differences between the high and low longevity groups were found in PWMFY (P = 0.21). No differences between the high and low efficiency groups were found in parity at culling (P = 0.50). No interactions between the longevity and efficiency groups were found on any longevity or efficiency measurement (P > 0.20). In herd management, the percentage of reserviced females and the percentage of multiple matings were associated with the longevity group and the efficiency group (P < 0.05). The high longevity group had lower culling rates in parity 0 to 6 than the low longevity group (P < 0.05), whereas no differences between the low and high efficiency groups were found in culling rates in parity 0 to 2 (P > 0.20). This study suggests that measures to achieve longevity and high reproductive efficiency in breeding herds do not conflict and that high reproductive efficiency and high longevity can be achieved.

Aging↗

The agricultural dispersal-valley drift spray drift modeling system compared with pesticide drift data.

The coupling of the valley drift (VALDRIFT) atmospheric dispersion/deposition model with the agricultural dispersal (AGDISP) aircraft wake model generates a modeling system for predicting the off-target drift of pesticides sprayed in a mountain valley. The approach uses the AGDISP near-field spray model to estimate the mass fraction of pesticide remaining airborne after initial application, then the VALDRIFT complex terrain model to estimate the drift of pesticide from the target area. The modeling system inputs include detailed spray information, a measure (or estimate) of winds in the valley, and the valley topographic characteristics; the results are pesticide concentrations throughout the valley atmosphere and pesticide deposition to the valley surface. The AGDISP and VALDRIFT models are operated independently, with the results from AGDISP being used as input to VALDRIFT through user-created data files. The modeling system was evaluated using pesticide drift data from spray trials conducted in the Mill Creek Canyon of Utah's Wasatch Mountains, USA, during the late spring of 1993. The predicted deposition compared within a factor of three of the observations (70% of the time) at all sampling locations extending several kilometers down-valley from the spray treatment block. The overall average ratio of predicted-to-observed deposition was 0.9.

Agriculture↗

An Assessment of Improvement in Reliability and Completeness of Mumbai Cancer Registry Data from 1964-1997.

The Mumbai Cancer Registry was established in 1964 with the aim of obtaining reliable morbidity and mortality data from precisely defined urban population. It was first and only such registry for merely two decades functioning in the country. Up to now more than 200,000 cancer cases are registered and with over 100,000 cancer deaths are recorded in data files. For studying improvements in the Mumbai Cancer Registry data, the data published in consecutive seven volumes (Vol.-II to Vol.-VIII) of "Cancer Incidence of Five Continents published by International Agency on Research on Cancer", Lyon, France have been used. For studying completeness of the data, the indicators 'Proportion of Deaths in Period'; 'Proportion of Death Certificates only' and stability of age incidence rates have been utilized. The indicators 'Proportion of cases registered on histological verification', 'The proportion of cases where age is not known', 'The flattening of age incidence curve' and 'Proportion of other and unspecified neoplasms can throw some light on the quality of data collected by the registry. There has been notable improvement in percentages of histological verification cases and substantial decrease in the proportion of death certificate alone cases in both the sexes over a period of time. Mortality Incidence ratio remained stable over a period of time in both the sexes. The proportion of cases where age is not known never exceeded 0.020% in either sex, for any site, for any period. The proportion of cases registered as other and unspecified sites, initially was around 8 to 9% then it has been dropped down to 5%. The crude incidence rates for all sites together are stable throughout the period of observation in both the sexes while age adjusted incidence rates show declining trend in both the sexes. There is no change in the pattern of age-specific incidence curves over a period of time in both the sexes. On examining various indices of reliability and completeness of Mumbai cancer registry data it can be concluded that, the data collected by this registry is quiet complete and reliable. While applying various checks for validity for a period from 1964-66 to 1993-97, it indicates that there is quiet improvement in almost all indices over a period of time in Mumbai cancer registry data.

Journal Article↗

[The data processing of neurophysiological examination and the diagnostic support system].

Neurophysiological examinations include EEG, EMG, Evoked Potential (EP) and ENG. In this paper, EEG and EP data processing, and the diagnostic support system using these findings were discussed. The FFT and AR model for analysis of frequency, and the pattern recognition for spike or spindle detection are the technical methods for application to EEG data processing of the first order. Amplitude or phase mapping techniques, as a second order data processing, were not only applied to the diagnosis of neurological disorders, but also used for detection of electrical equivalent dipole source localization which was inversely reconstructed in the cerebral cortex from the distribution of EEG potentials on the scalp. The averaging technique was used for detection of small evoked potentials (EP) such as SEP or ABR. The following items were included in the diagnostic support system applied to data processing. 1) Topographic analysis in the brain mapping of SEP to median nerve stimulation was applied to the identification of the central sulcus during a neurosurgical procedure to maintain the QOL of patient. 2) The usefulness of EEG automatic reporting system with Japanese sentences and brain topographic mapping was described. 3) Digital EEG in a data filing system using magneto optical disks was used for data analysis and diagnostic support system after EEG examination.

Algorithms↗

Maxsim, software for the analysis of multiple axonal arbors and their simulated activation.

In order to analyze the structural organization of complex axonal arbors reconstructed from histological serial sections, and to investigate the functional implications of their geometrical properties, we developed software providing the following facilities: (1) direct importation of data files generated by a commercially available 3-D light-microscopic reconstruction system, including routine procedures for identification and correction of data acquisition errors; (2) real-time 3-D rotations of the arbors in the stack of serial sections; (3) multiple interactive display modes; (4) possibility of modifying diameter and/or connectivity of different branches; (5) simulation of the invasion of the arbor by a single action potential initiated at any chosen point, and visualization of spatio-temporal profiles of activation; (6) extraction of quantitative data converted to standard file formats compatible with available mathematical software. All these tools can be applied to single or multiple axons, individually or simultaneously. The software, called Maxsim, is a highly flexible C-written program running on graphical workstations using the UNIX operating system and X-Window environment.

Animals↗

The hepatic transcriptome as a window on whole-body physiology and pathophysiology.

Transcriptomics can be a valuable aid to pathologists. The information derived from microarray studies may soon include the entire transcriptomes of most cell types, tissues and organs for the major species used for toxicology and human disease risk assessment. Gene expression changes observed in such studies relate to every aspect of normal physiology and pathophysiology. When interpreting such data, one is forced to look "far from the lamp post:' and in so doing, face one's ignorance of many areas of biology. The central role of the liver in toxicology, as well as in many aspects of whole-body physiology, makes the hepatic transcriptome an excellent place to start your studies. This article provides data that reveals the effects of fasting and circadian rhythm on the rat hepatic transcriptome, both of which need to be kept in mind when interpreting large-scale gene expression in the liver. Once you become comfortable with evaluating mRNA expression profiles and learn to correlate these data with your clinical and morphological observations, you may wonder why you did not start your studies of transcriptomics sooner. Additional study data can be viewed at the journal website at (www.toxpath.org). Two data files are provided in Excel format, which contain the control animal data from each of the studies referred to in the text,including normalized signal intensity data for each animal (n=5) in the 6-hour, 24-hour, and 5-day time points. These files are briefly described in the associated 'Readme' file, and the complete list of GenBank numbers and Affymetrix IDs are provided in a separate txt file. These files are available at http://taylorandfrancis.metapress.comlopenurl.asp?genre=journal&issn=0192-6233. Click on the issue link for 33(1), then select this article. A download option appears at the bottom of this abstract. In order to access the full article online, you must either have an individual subscription or a member subscription accessed through (www.toxpath.org).

Animals↗

Tools for loading MEDLINE into a local relational database.

BACKGROUND: Researchers who use MEDLINE for text mining, information extraction, or natural language processing may benefit from having a copy of MEDLINE that they can manage locally. The National Library of Medicine (NLM) distributes MEDLINE in eXtensible Markup Language (XML)-formatted text files, but it is difficult to query MEDLINE in that format. We have developed software tools to parse the MEDLINE data files and load their contents into a relational database. Although the task is conceptually straightforward, the size and scope of MEDLINE make the task nontrivial. Given the increasing importance of text analysis in biology and medicine, we believe a local installation of MEDLINE will provide helpful computing infrastructure for researchers. RESULTS: We developed three software packages that parse and load MEDLINE, and ran each package to install separate instances of the MEDLINE database. For each installation, we collected data on loading time and disk-space utilization to provide examples of the process in different settings. Settings differed in terms of commercial database-management system (IBM DB2 or Oracle 9i), processor (Intel or Sun), programming language of installation software (Java or Perl), and methods employed in different versions of the software. The loading times for the three installations were 76 hours, 196 hours, and 132 hours, and disk-space utilization was 46.3 GB, 37.7 GB, and 31.6 GB, respectively. Loading times varied due to a variety of differences among the systems. Loading time also depended on whether data were written to intermediate files or not, and on whether input files were processed in sequence or in parallel. Disk-space utilization depended on the number of MEDLINE files processed, amount of indexing, and whether abstracts were stored as character large objects or truncated. CONCLUSIONS: Relational database (RDBMS) technology supports indexing and querying of very large datasets, and can accommodate a locally stored version of MEDLINE. RDBMS systems support a wide range of queries and facilitate certain tasks that are not directly supported by the application programming interface to PubMed. Because there is variation in hardware, software, and network infrastructures across sites, we cannot predict the exact time required for a user to load MEDLINE, but our results suggest that performance of the software is reasonable. Our database schemas and conversion software are publicly available at http://biotext.berkeley.edu.

Database Management Systems↗

Can we use automated data to assess quality of hypertension care?

OBJECTIVE: To determine whether extractable blood pressure (BP) information available in a computerized patient record system (CPRS) could be used to assess quality of hypertension care independently of clinicians' notes. STUDY DESIGN: Retrospective cohort study of a random sample of hypertensive patients from 10 Department of Veterans Affairs (VA) sites across the country. METHODS: We abstracted BPs from electronic clinicians' notes for all medical visits of 981 hypertensive patients in 1999. We compared these with BP measurements available in a separate vitals signs file in the CPRS. We also evaluated whether assessments of performance varied by source by using patients' last documented BP reading. RESULTS: When the vital signs file and notes were combined, a BP measurement was taken for 71% of 6097 medical visits; 60% had a BP measurement only in the vital signs file. Combining sources, 43% of patients had a BP reading of less than 140/90 mm Hg; by site this varied (34%-51%). Vital signs file data alone yielded similar findings; site rankings by rates of BP control changed minimally. CONCLUSIONS: Current performance review programs collect clinical data from both clinicians' notes and automated sources as available. However, we found that notes contribute little information with respect to BP values beyond automated data alone. The VA's vital signs file is a prototypical automated data system that could make assessment of hypertension care more efficient in many settings.

Aged↗

User-definable bull's-eye database analysis.

Several quantitative bull's-eye database programs have been developed and employed successfully, but generally they restrict the user to limited types of quantitative analysis. We developed a type of bull's-eye analysis which facilitates user-defined processing, and then explored the effects of various types of processing on the comparisons of patient information with that of reference databases. Male and female bull's-eye database were generated from 32 normal patients using unweighted 2D prefiltering, ramp backprojection, unweighted 3D postfiltering, and peak value circumferential plotting (base method). The data from each patient were then reprocessed and compared to the databases by means of three different approaches: (1) using the base method, (2) using average as opposed to peak value profiles, and (3) using a resolution recovery prefilter instead of a smoothing prefilter. Significant differences in the number of apparently abnormal regions were found between the three methods. In other words, the type of single-photon emission tomography (SPET) processing affected the accuracy of comparisons between patient and database information. Because even sophisticated analysis can now be performed on personal computers, we conclude that, rather than a preprocessed data file, clinical "normal reference" information should consist of original SPET data (in a standard format, e.g., Interfile) from a series of documented normal patients. Each user could then generate reference bull's-eye database by applying his or her own clinical processing procedures to the data.

Humans↗

Data processing of the output from a Vickers M300 clinical chemistry analyser. Principles and implementation.

A suite of data processing programs is described, which takes the results' log from a Vickers Medical LTD M300 clinical chemistry analyser; corrects phasing errors; allows various types of recalibration (recalculation) of the data; and delivers the results to a general purpose laboratory data filing system (PHOENIX-ACHILLES). An important problem with the data is that results may be unphased with respect to their identification data for mechanical reasons. The principles underlying rephasing, and other requirements for handling the M300 data are described, together with the processes used by the operator. Deliberately, no ammendments were made to the supported software for the integral computer in the M300 analyser. The new software supplements the analyser by allowing the operator to summarise the corrections to the data which would have been made under manual conditions: the required corrections are then completed by the computer. About 70 working days were required to complete the program: this was much more than has been required to program data handling for several other analysers which do not have phasing problems.

Chemistry, Clinical↗

Real time data acquisition and analysis of cardiovascular experiments in dogs.

A hardware-software system is described which enables physiologists to utilize the DEC PDP-ll computer for data acquisition and analysis in the real-time experimental environment. The system allows for A/D conversion of up to 16 physiological parameters, as well as various calculations based upon these parameters such as heart rate, cardiac work, and derivatives and integrals of ventricular tension and pressure. By using a push-button box, the investigator can request a display of the parameters just acquired, a graphic display summarizing the results of the experiment up to the time of request, and also change various parameters such as gain factors, names of stages and erasures. At the end of the experiment, the computer prints a table summarizing the course of the experiment. A data file is written in a standardized format containing the essential data obtained during the experiment.

Animals↗

Ethical issues in sharing epidemiologic data.

Concurrent with the explosion in large data files and computers capable of handling both linkage of large data sets and analyzing multiple studies for meta-analysis, this decade has seen a rise in professional concern about the need for researchers to share their data. As scientific groups began to address this question, its importance and complexity became quickly apparent. In this paper recent developments on the ethics of data sharing in statistics, sociology, psychology, and other fields related to epidemiology are summarized, followed by a discussion on why data should be shared, what kinds of data should be shared, who among epidemiologists should be sharing data, when it is appropriate to share data, and how data sharing should be conducted.

Access to Information↗

Small protein biomarkers of culture in Bacillus spores detected using capillary liquid chromatography coupled with matrix assisted laser desorption/ionization mass spectrometry.

Capillary liquid chromatography (cLC) coupled with matrix-assisted laser desorption/ionization (MALDI) time-of-flight mass spectrometry (TOF-MS) was used to compare small proteins and peptides extracted from Bacillus subtilis spores grown on four different media. A single, efficient protein separation, compatible with MALDI-MS analysis, was employed to reduce competitive ionization between proteins, and thus interrogate more proteins than possible using direct MALDI-MS. The MALDI-MS data files for each fraction are assembled as two-dimensional data sets of retention time and mass information. This method of visualizing small protein data required careful attention to background correction as well as mass and retention time variability. The resulting data sets were used to create comparative displays of differences in protein profiles between different spore preparations. Protein differences were found between two different solid media in both phase bright and phase dark spore phenotype. The protein differences between two different liquid media were also examined. As an extension of this method, we have demonstrated that candidate protein biomarkers can be trypsin digested to provide identifying peptide fragment information following the cLC-MALDI experiment. We have demonstrated this method on two markers and utilized acid breakdown information to identify one additional marker for this organism. The resulting method can be used to identify discriminating proteins as potential biomarkers of growth media, which might ultimately be used for source attribution.

Amino Acid Sequence↗

NM-Win: a personal computer-based Microsoft Windows front-end to NONMEM IV.

A Microsoft Windows-based front-end, NM-Win, has been written to provide a more user-friendly environment to do nonlinear mixed effect modeling with the NONMEM program. NM-Win utilizes an object-oriented interface design which allows users to view and edit control, PRED, and/or data files using Windows Notepad. In addition, calls made to the Microsoft FORTRAN compiler and linker which generate the final NONMEN executable are performed simply by clicking the "Run NONMEN" button. During the executive step, interactions can be viewed in a window to check the progress of the run. Errors encountered while NONMEN or NM-TRAN is running are brought to a window for ease in debugging. Advanced options allow the user the flexibility of compiling user-written PRED files and creating linker response files. While the PC platform is not optimal for large data set or complex models, it does permit easier debugging and offers multitasking while Windows is running.

Algorithms↗

Delineating disability, labour force participation and employment restrictions among persons with psychosis.

OBJECTIVE: To delineate at a population level: activity restrictions, labour market participation, educational attainment, employment restrictions and employment characteristics of persons with psychosis compared with healthy non-disabled persons. METHOD: Confidentialized data files were provided by the Australian Bureau of Statistics. The data were collected in a national survey titled "Survey of Disability, Ageing and Carers, Australia 1998". Multi-stage sampling strategies obtained a probability sample of 42 664 individuals. Trained interviewers using ICD-10 computer-assisted interviews identified household residents with psychosis. RESULTS: Among householders with psychosis aged 15-64 years, 75.2% were non-participants in the labour market, 21.1% were employed and 3.7% were looking for work. Completing school years 10 and 11, and vocational training, appeared to offer an employment advantage. CONCLUSION: Persons with psychotic disorders have low rates of labour force participation and may benefit from greater participation in educational and vocational services. Implications for policy development are discussed.

Adolescent↗

Metabolomics spectral formatting, alignment and conversion tools (MSFACTs).

MOTIVATION: The amplified interest in metabolic profiling has generated the need for additional tools to assist in the rapid analysis of complex data sets. RESULTS: A new program; metabolomics spectral formatting, alignment and conversion tools, (MSFACTs) is described here for the automated import, reformatting, alignment, and export of large chromatographic data sets to allow more rapid visualization and interrogation of metabolomic data. MSFACTs incorporates two tools: one for the alignment of integrated chromatographic peak lists and another for extracting information from raw chromatographic ASCII formatted data files. MSFACTs is illustrated in the processing of GC/MS metabolomic data from different tissues of the model legume plant, Medicago truncatula. The results document that various tissues such as roots, stems, and leaves from the same plant can be easily differentiated based on metabolite profiles. Further, similar types of tissues within the same plant, such as the first to eleventh internodes of stems, could also be differentiated based on metabolite profiles. AVAILABILITY: Freely available upon request for academic and non-commercial use. Commercial use is available through licensing agreement http://www.noble.org/PlantBio/MS/MSFACTs/MSFACTs.html.

Cluster Analysis↗