Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Computing Methodologies”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 811 records · Page 45Linked to original sources

Exploiting large scale computing to construct high resolution linkage disequilibrium maps of the human genome.

UNLABELLED: Linkage disequilibrium (LD) maps increase power and precision in association mapping, define optimal marker spacing and identify recombination hot-spots and regions influenced by natural selection. Phase II of HapMap provides approximately 2.8-fold more single nucleotide polymorphisms (SNPs) than phase I for constructing higher resolution maps. LDMAP-cluster, is a parallel program for rapid map construction in a Linux environment used here to construct genome-wide LD maps with >8.2 million SNPs from the phase II data. AVAILABILITY: The LD maps, LDMAP-cluster and documentation are available from: http://www.som.soton.ac.uk/research/geneticsdiv/epidemiology/LDMAP. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Algorithms↗

Statistical genetics of an annual plant, Impatiens capensis. I. Genetic basis of quantitative variation.

Analysis of quantitative genetics in natural populations has been hindered by computational and methodological problems in statistical analysis. We developed and validated a jackknife procedure to test for existence of broad sense heritabilities and dominance or maternal effects influencing quantitative characters in Impatiens capensis. Early life cycle characters showed evidence of dominance and/or maternal effects, while later characters exhibited predominantly environmental variation. Monte Carlo simulations demonstrate that these jackknife tests of variance components are extremely robust to heterogeneous error variances. Statistical methods from human genetics provide evidence for either a major locus influencing germination date, or genes that affect phenotypic variability per se. We urge explicit consideration of statistical behavior of estimation and testing procedures for proper biological interpretation of statistical results.

Genetic Variation↗

Making sense of the metabolome using evolutionary computation: seeing the wood with the trees.

One should perhaps start off by asking the question, 'But what wood is it we want to see?' There are so many trees that make up the wood; within a post-genomics context, genes, transcripts, proteins, and metabolites are the more tangible ones. Rather than studying these components in isolation, a more holistic approach is to unravel the interactions between the myriad of subcellular components and this is vital to systems biology. Moreover, this will help define the phenotype of the organism under investigation. Metabolomics is complementary to transcriptomics and proteomics, and despite the immense metabolite diversity observed in plants, metabolomics has been embraced by the plant community and in particular for studying metabolic networks. Whilst post-genomic science is producing vast data torrents, it is well known that data do not equal knowledge and so the extraction of the most meaningful parts of these data is key to the generation of useful new knowledge. A metabolomics experiment is guaranteed to generate thousands of data points (e.g. samples multiplied by the levels of particular metabolites) of which only a handful might be needed to describe the problem adequately. Evolutionary computational-based methods such as genetic algorithms and genetic programming are ideal strategies for mining such high-dimensional data to generate useful relationships, rules, and predictions. This article describes these techniques and highlights their usefulness within metabolomics.

Algorithms↗

Demonstration of a word design strategy for DNA computing on surfaces.

A strategy for DNA computing on surfaces using linked sets of 'DNA words' that are short oligonucleotides (16mers) is proposed. The 16mer words have the format 5'-FFFFvvvvvvvvFFFF-3' in which 4-8 bits of data are stored in 8 variable ('v') base locations, and the remaining fixed ('F') base locations are used as a word label. Using a template and map strategy, a set of 108 8mers each of which possesses at least a 4 base mismatch with the complements to all the other members of the set (4bm complements) are identified for use as a variable base sequence set. In addition, sets of 4 and 12 word labels of the form ABCD....DCBA that are respectively 8bm and 6bm complements with each other are identified. The 16mers are chosen to have a G/C content of 50% in order to make the thermodynamic stability of the perfectly matched hybridized DNA duplexes similar; a simple pairwise additive method is used to estimate the perfect match and mismatch hybridization thermodynamics. A series of preliminary experiments are presented that use small arrays of 16mers attached to chemically modified gold surfaces and fluorescently labeled complements to study the hybridization adsorption and enzymatic manipulation of the oligonucleotides.

Base Sequence↗

Prediction of DNA single-strand conformation polymorphism: analysis by capillary electrophoresis and computerized DNA modeling.

We have analyzed previously three representative p53 single-point mutations by capillary-electrophoresis single-strand conformation polymorphism (CE-SSCP). In the current study, we compared our CE-SSCP results with the potential secondary structures predicted by an RNA/DNA-folding algorithm with DNA energy rules, used in conjunction with a computer analysis workbench called STRUCTURELAB. Each of these mutations produces measurable shifts in CE migration times relative to wild type. Using computerized folding analysis, each of the mutations was found to have a conformational difference relative to wild type, which accounts for the observed differences in CE migration. Additional properties exhibited in the CE electropherograms were also explained using the computerized analysis. These include the appearance of secondary peaks and the temperature dependence of the electrophoretic patterns. The results yield insight into the mechanism of SSCP and how the conditions of this measurement, especially temperature, may be optimized to improve the sensitivity of the SSCP method. The results may also impact other diagnostic methods, which would benefit by a better understanding of DNA single-strand conformation polymorphisms to optimize conditions for enzymatic cleavage and DNA hybridization reactions.

Base Sequence↗

Attributable risk ratio estimation from matched-pairs case-control data.

Explicit formulas are provided for estimating the attributable risk ratio among the exposed and the entire target population utilizing matched-pairs data. Large-sample standard errors and corresponding confidence intervals are provided. These estimates can be obtained from the cross-classification frequencies of matched pairs by disease and exposure status in the usual 2 X 2 table. The key to the development of these formulas lies in recognizing that attributable risk among the exposed is a direct function of the odds ratio, and population attributable risk is a direct function of the odds ratio and exposure prevalence among only the cases (assuming a rare disease). The formulas presented in this paper require only a calculator for computation. The methodology is illustrated with data from a matched-pairs case-control study of oral conjugated estrogens and endometrial cancer.

Contraceptives, Oral↗

Parallel FD-TD simulation of radiobase antennae.

The rigorous characterisation of the behaviour of a radiobase antenna for wireless communication systems is a hot topic both for antenna or communication system design and for radioprotection-hazard reasons. Such a characterisation deserves a numerical solution, and the use of a finite-difference time-domain (FD-TD) approach is an attractive candidate. Unfortunately, it has strong memory and CPU-time requirements. Numerical complexity can be successfully afforded by using parallel computing. The parallel implementation of the FD-TD code, individuating the theoretical lower bound for its parallel execution time are discussed and the findings achieved on the APE/Quadrics massively parallel systems are presented. Results obtained from the simulation of actual radiobase antennae, clearly demonstrate that massively parallel processing is a viable approach to solving electromagnetic problems, allowing the simulation of radiating devices which could not be modelled through conventional computing systems. The tests showed a sustained computational speed equal to 17% of the theoretical maximum.

Algorithms↗

Weighted-support vector machines for predicting membrane protein types based on pseudo-amino acid composition.

Membrane proteins are generally classified into the following five types: (1) type I membrane proteins, (2) type II membrane proteins, (3) multipass transmembrane proteins, (4) lipid chain-anchored membrane proteins and (5) GPI-anchored membrane proteins. Prediction of membrane protein types has become one of the growing hot topics in bioinformatics. Currently, we are facing two critical challenges in this area: first, how to take into account the extremely complicated sequence-order effects, and second, how to deal with the highly uneven sizes of the subsets in a training dataset. In this paper, stimulated by the concept of using the pseudo-amino acid composition to incorporate the sequence-order effects, the spectral analysis technique is introduced to represent the statistical sample of a protein. Based on such a framework, the weighted support vector machine (SVM) algorithm is applied. The new approach has remarkable power in dealing with the bias caused by the situation when one subset in the training dataset contains many more samples than the other. The new method is particularly useful when our focus is aimed at proteins belonging to small subsets. The results obtained by the self-consistency test, jackknife test and independent dataset test are encouraging, indicating that the current approach may serve as a powerful complementary tool to other existing methods for predicting the types of membrane proteins.

Algorithms↗

Efficient heterogeneous execution of Monte Carlo shielding calculations on a Beowulf cluster.

Recent work has been done in using a high-performance 'Beowulf' cluster computer system for the efficient distribution of Monte Carlo shielding calculations. This has enabled the rapid solution of complex shielding problems at low cost and with greater modularity and scalability than traditional platforms. The work has shown that a simple approach to distributing the workload is as efficient as using more traditional techniques such as PVM (Parallel Virtual Machine). In addition, when used in an operational setting this technique is fairer with the use of resources than traditional methods, in that it does not tie up a single computing resource but instead shares the capacity with other tasks. These developments in computing technology have enabled shielding problems to be solved that would have taken an unacceptably long time to run on traditional platforms. This paper discusses the BNFL Beowulf cluster and a number of tests that have recently been run to demonstrate the efficiency of the asynchronous technique in running the MCBEND program. The BNFL Beowulf currently consists of 84 standard PCs running RedHat Linux. Current performance of the machine has been estimated to be between 40 and 100 Gflop s(-1). When the whole system is employed on one problem up to four million particles can be tracked per second. There are plans to review its size in line with future business needs.

Computer Communication Networks↗

Neutron analysis of spent fuel storage installation using parallel computing and advance discrete ordinates and Monte Carlo techniques.

In the United States, the Nuclear Waste Policy Act of 1982 mandated centralised storage of spent nuclear fuel by 1988. However, the Yucca Mountain project is currently scheduled to start accepting spent nuclear fuel in 2010. Since many nuclear power plants were only designed for -10 y of spent fuel pool storage, > 35 plants have been forced into alternate means of spent fuel storage. In order to continue operation and make room in spent fuel pools, nuclear generators are turning towards independent spent fuel storage installations (ISFSIs). Typical vertical concrete ISFSIs are -6.1 m high and 3.3 m in diameter. The inherently large system, and the presence of thick concrete shields result in difficulties for both Monte Carlo (MC) and discrete ordinates (SN) calculations. MC calculations require significant variance reduction and multiple runs to obtain a detailed dose distribution. SN models need a large number of spatial meshes to accurately model the geometry and high quadrature orders to reduce ray effects, therefore, requiring significant amounts of computer memory and time. The use of various differencing schemes is needed to account for radial heterogeneity in material cross sections and densities. Two P3, S12, discrete ordinate, PENTRAN (parallel environment neutral-particle TRANsport) models were analysed and different MC models compared. A multigroup MCNP model was developed for direct comparison to the SN models. The biased A3MCNP (automated adjoint accelerated MCNP) and unbiased (MCNP) continuous energy MC models were developed to assess the adequacy of the CASK multigroup (22 neutron, 18 gamma) cross sections. The PENTRAN SN results are in close agreement (5%) with the multigroup MC results; however, they differ by -20-30% from the continuous-energy MC predictions. This large difference can be attributed to the expected difference between multigroup and continuous energy cross sections, and the fact that the CASK library is based on the old ENDF/B-II library. Both MC and SN calculations were run in parallel on a BEOWULF PC-cluster (eight processors). Timing results indicate that the SN calculation yielded a detailed dose distribution at over 318,426 points in -164 h. Unbiased continuous energy MC required 214 h to calculate dose rates with a 1% relative error in only 18 regions on the surface of the cask. The biased A3MCNP calculations yields dose rates with -0.8% relative error in only 2.5 h on one processor. This study demonstrates that a parallel code, such as the 3-D parallel SN transport code, PENTRAN can solve a complex large problem, such as the storage cask, accurately and efficiently. Moreover, this calculation was performed on a relatively inexpensive PC-cluster. Possible inadequacies of the CASK cross section library still need to be evaluated.

Computer-Aided Design↗

A cost-construction model to assess the total cost of an anesthesiology residency program.

BACKGROUND: Although the total costs of graduate medical education are difficult to quantify, this information may be of great importance for health policy and planning over the next decade. This study describes the total costs associated with the residency program at the University of Texas--Houston Department of Anesthesiology during the 1996-1997 academic year. METHODS: The authors used cost-construction methodology, which computes the cost of teaching from information on program description, resident enrollment, faculty and resident salaries and benefits, and overhead. Surveys of faculty and residents were conducted to determine the time spent in teaching activities; access to institutional and departmental financial records was obtained to quantify associated costs. The model was then developed and examined for a range of assumptions concerning resident productivity, replacement costs, and the cost allocation of activities jointly producing clinical care and education. RESULTS: The cost of resident training (cost of didactic teaching, direct clinical supervision, teaching-related preparation and administration, plus the support of the teaching program) was estimated at $75,070 per resident per year. This cost was less than the estimated replacement value of the teaching and clinical services provided by residents, $103,436 per resident per year. Sensitivity analysis, with different assumptions regarding resident replacement cost and reimbursement rates, varied the cost estimates but generally identified the anesthesiology residency program as a financial asset. CONCLUSIONS: In most scenarios, the value of the teaching and clinical services provided by residents exceeded the cost of the resources used in the educational program.

Anesthesiology↗

The effect of concurvity in generalized additive models linking mortality to ambient particulate matter.

In recent years, a number of studies have applied generalized additive models to time series data to estimate associations between exposure to air pollution and cardiorespiratory morbidity and mortality. If concurvity, the nonparametric analogue of multicollinearity, is present in the data, statistical software such as S-plus can seriously underestimate the variance of fitted model parameters, leading to significance tests with inflated type 1 error. This paper uses computer simulation and analyses of actual epidemiologic data to explore this underestimation of standard errors. We provide a method for assessing concurvity in data and an alternate class of models that is unaffected by concurvity. We argue that some degree of concurvity is likely to be present in all epidemiologic time series datasets and we explore through the use of meta-analysis the possible impact of concurvity on the existing body of work relating ambient levels of sulfate particles to mortality.

Air Pollutants↗

A resource facility for kinetic analysis: modeling using the SAAM computer programs.

Kinetic analysis and integrated system modeling have contributed significantly to understanding the physiology and pathophysiology of metabolic systems in humans and animals. Many experimental biologists are aware of the usefulness of these techniques and recognize that kinetic modeling requires special expertise. The Resource Facility for Kinetic Analysis (RFKA) provides this expertise through: (1) development and application of modeling technology for biomedical problems, and (2) development of computer-based kinetic modeling methodologies concentrating on the computer program Simulation, Analysis, and Modeling (SAAM) and its conversational version, CONversational SAAM (CONSAM). The RFKA offers consultation to the biomedical community in the use of modeling to analyze kinetic data and trains individuals in using this technology for biomedical research. Early versions of SAAM were widely applied in solving dosimetry problems; many users, however, are not familiar with recent improvements to the software. The purpose of this paper is to acquaint biomedical researchers in the dosimetry field with RFKA, which, together with the joint National Cancer Institute-National Heart, Lung and Blood Institute project, is overseeing SAAM development and applications. In addition, RFKA provides many service activities to the SAAM user community that are relevant to solving dosimetry problems.

Animals↗

Elderly injury: a profile of trauma experience in the Sunshine (Retirement) State.

OBJECTIVE: By using mandatory discharge data from a state agency, the records of 116,687 patients hospitalized for treatment of injury were evaluated to develop an epidemiologic and demographic profile of this population and to compare outcomes of patients treated in state-designated trauma centers (TC) with those treated in nontrauma centers (NTC). METHODS: Injury severity was calculated by using the International Classification Injury Severity Score methodology to compute individual diagnosis survival risk ratios from 698,187 reported diagnoses, and then by using these survival risk ratios to determine probability of survival for every patient. The population was then categorized by age, injury type, treatment facility designation, injury severity as indicated by probability of survival, and discharge disposition. Incidence of potentially preventable death was compared between TC and NTC, as was the effect on outcome of noninjury comorbidity. RESULTS: The average age of this population was 58 +/- 26 years with significant skew toward the elderly in NTC (mean age, 62 +/- 26 years). The most commonly encountered injuries likewise reflected the elderly nature of this population. Although 71.3% received care in NTC, the majority of severely injured were treated in TC. Potentially preventable mortality (>0.5) was significantly lower in TC. The effect of noninjury comorbidity on outcome was better managed by TC, both in terms of decreased mortality and in proportion of patients discharged home. CONCLUSION: These data demonstrate the unique characteristics of injury victims treated in the state of Florida and indicate that the developing trauma system is demonstrating productivity in terms of avoidance of preventable death, efficient management of noninjury comorbid problems, and more complete recovery as indicated by proportion of patients discharged to home.

Aged↗