Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Is it better to combine predictions?

We have compared the accuracy of the individual protein secondary structure prediction methods: PHD, DSC, NNSSP and Predator against the accuracy obtained by combing the predictions of the methods. A range of ways of combing predictions were tested: voting, biased voting, linear discrimination, neural networks and decision trees. The combined methods that involve 'learning' (the non-voting methods) were trained using a set of 496 non-homologous domains; this dataset was biased as some of the secondary structure prediction methods had used them for training. We used two independent test sets to compare predictions: the first consisted of 17 non-homologous domains from CASP3 (Third Community Wide Experiment on the Critical Assessment of Techniques for Protein Structure Prediction); the second set consisted of 405 domains that were selected in the same way as the training set, and were non-homologous to each other and the training set. On both test datasets the most accurate individual method was NNSSP, then PHD, DSC and the least accurate was Predator; however, it was not possible to conclusively show a significant difference between the individual methods. Comparing the accuracy of the single methods with that obtained by combing predictions it was found that it was better to use a combination of predictions. On both test datasets it was possible to obtain a approximately 3% improvement in accuracy by combing predictions. In most cases the combined methods were statistically significantly better (at P = 0.05 on the CASP3 test set, and P = 0.01 on the EBI test set). On the CASP3 test dataset there was no significant difference in accuracy between any of the combined method of prediction: on the EBI test dataset, linear discrimination and neural networks significantly outperformed voting techniques. We conclude that it is better to combine predictions.

Algorithms↗

Use of covariance analysis for the prediction of structural domain boundaries from multiple protein sequence alignments.

Current methods for identification of domains within protein sequences require either structural information or the identification of homologous domain sequences in different sequence contexts. Knowledge of structural domain boundaries is important for fold recognition experiments and structural determination by X-ray crystallography or nuclear magnetic resonance spectroscopy using the divide-and-conquer approach. Here, a new and conceptually simple method for the identification of structural domain boundaries in multiple protein sequence alignments is presented. Analysis of covariance at positions within the alignment is first used to predict 3D contacts. By the nature of the domain as an independent folding unit, inter-domain predicted contacts are fewer than intra-domain predicted contacts. By analysing all possible domain boundaries and constructing a smoothed profile of predicted contact density (PCD), true structural domain boundaries are predicted as local profile minima associated with low PCD. A training data set is constructed from 52 non-homologous two-domain protein sequences of known 3D structure and used to determine optimal parameters for the profile analysis. The alignments in the training data set contained 48 +/- 17 (mean +/- SD) sequences and lengths of 257 +/- 121 residues. Of the 47 alignments yielding predictions, 35% of true domain boundaries are predicted to within 15 amino acids by the local profile minimum with the lowest profile value. Including predictions from the second- and third-lowest local minima increases the correct domain boundary coverage to 60%, whereas the lowest five local minima cover 79% of correct domain boundaries. Through further profile analysis, criteria are presented which reliably identify subsets of more accurate predictions. Retrospective analysis of CASP3 targets shows predictions of sufficient accuracy to enable dramatically improved fold recognition results. Finally, a prediction is made for geminivirus AL1 protein which is in full agreement with biochemical data, yielding a plausible, novel threading result.

DNA-Binding Proteins↗

Is global outcome predictable in the rehabilitation of patients with musculoskeletal disorders? A pilot study.

Definition of prognostic factors for outcome quality is of increasing interest in rehabilitation medicine. The main question of this pilot study in 552 patients was whether global outcome could be predicted by a team-based chief physician specialized in physical medicine and rehabilitation (PMR), and whether other predictive factors would exist (ICIDH-2 levels, pain, working incapacity). Little data is available about the possibility of global prediction of prognosis in the rehabilitation of patients with musculoskeletal disorders. All 552 patients met each member of the rehabilitation team and key data from each patient was discussed at the rehabilitation conference within the first 2 days. On entry to the study, a chief physician specialized in PMR assessed the patient's key data, which was structured according to ICIDH-2 (ICF) and assessed quantitatively on a scale from zero to ten. Second, the PMR physician rated the expected global prognosis on the basis of ICIDH-2 and other key data, and in respect to the defined rehabilitation goals (see Table 2). At the same time, the patient and an assistant doctor (AD) assessed pain scores (VAS 0-10) and the actual working incapacity (%). These assessments were completed within the first 3 days and were repeated before discharge. Assessment of outcome was rated by both, separately, according to the above-mentioned scale. Different regression models were calculated, searching for significant differences between the numerous variables. In the regression models, the best predictor for outcome was the PMR physician assessment. Complete and good correspondence between prediction and outcome was obtained in 71.4% (42.1% and 29.3%, respectively) in the descriptive model. Quantitatively assessed ICIDH-2 levels, pain at entrance and working incapacity at entrance were not predictive factors for global outcome. The global outcome was rated as very good/good in 79.0% of cases by patient and in 75.1% cases by the AD, as moderate in 13.9% of cases by the patient and 18.4% of cases by the AD, and as poor/worsening in 7.1% of cases by the patient and in 6.5% of cases by the AD. Rating of outcome by the patient and the AD gave complete and good correspondence in 87.6% and no correspondence in only 2.6% of cases. Pain could be reduced highly significantly (P<0.001). There was a highly significant degree of correlation between quality of outcome and pain relief (outcome 'very good' and 'good', P<0.001; 'moderate', P=0.003; 'poor/worsening', not significant). Partial or complete reduction of working incapacity could be reached in 30% of the patients. This had no statistical influence on global outcome; neither did persistent working incapacity. Prediction of global outcome by a team-based PMR assessment seems to be a useful semiquantitative method with high predictive value. The method, including the critical point of validation, is currently being extensively discussed. Prediction is an integral process based on the high information grade of a multiprofessional rehabilitation team, the ICIDH-2 structures, the definition of rehabilitation goals, the knowledge and experience in bio-psycho-social medicine and the application of common sense. Rating of global outcome by the patient/AD is an integrative process as well. Pain relief is an important and very strong factor, with a high degree of influence on global outcome in musculoskeletal rehabilitation, probably by improving quality of life. Working incapacity is no reason for refusing patients rehabilitation and both improvement of working capacity and persistence of working incapacity, has no statistical influence on global outcome. Finally, the extent of the four ICIDH-2 levels, especially negative contextual factors, were not predictive, that is, they had no significant influence on global outcome in this study. In conclusion, prediction of global outcome by a team-based chief physician specialized in PMR is of high predictive value, practicable and useful for rehabilitation processes, quality assurance, insurance companies and health policies. To our knowledge, this is the first published study on this topic.

Adolescent↗

Mortality predictions in the intensive care unit: comparing physicians with scoring systems.

OBJECTIVE: Risk-prediction models offer potential advantages over physician predictions of outcomes in the intensive care unit (ICU). Our systematic review compared the accuracy of ICU physicians' and scoring system predictions of ICU or hospital mortality of critically ill adults. DATA SOURCE: MEDLINE (1966-2005), CINAHL (1982-2005), Ovid Healthstar (1975-2004), EMBASE (1980-2005), SciSearch (1980-2005), PsychLit (1985-2004), the Cochrane Library (Issue 1, 2005), PubMed "related articles," personal files, abstract proceedings, and reference lists. STUDY SELECTION: We considered all studies that compared physician predictions of ICU or hospital survival of critically ill adults to an objective scoring system, computer model, or prediction rule. We excluded studies if they focused exclusively on the development or economic evaluation of a scoring system, computer model, or prediction rule. DATA EXTRACTION AND ANALYSIS: We independently abstracted data and assessed study quality in duplicate. We determined summary receiver operating characteristic curves and areas under the summary receiver operating characteristic curves+/-se and summary diagnostic odds ratios. DATA SYNTHESIS: We included 12 observational studies of moderate methodological quality. The area under the summary receiver operating characteristic curves for seven studies was 0.85+/-0.03 for physician predictions compared with 0.63+/-0.06 for scoring system predictions (p=.002). Physicians' summary diagnostic odds ratios derived from the area under the summary receiver operating characteristic curves were significantly higher (12.43; 95% confidence interval 5.47, 27.11) than scoring systems' summary diagnostic odds ratios (2.25; 95% confidence interval 0.78, 6.52, p=.001). Combined results of all 12 studies indicated that physicians predict mortality more accurately than do scoring systems: ratio of diagnostic odds ratios (95% confidence interval) 1.92 (1.19, 3.08) (p=.007). CONCLUSIONS: Observational studies suggest that ICU physicians discriminate between survivors and nonsurvivors more accurately than do scoring systems in the first 24 hrs of ICU admission. The overall accuracy of both predictions of patient mortality was moderate, implying limited usefulness of outcome prediction in the first 24 hrs for clinical decision making.

Critical Illness↗

A neural-network based method for prediction of gamma-turns in proteins from multiple sequence alignment.

In the present study, an attempt has been made to develop a method for predicting gamma-turns in proteins. First, we have implemented the commonly used statistical and machine-learning techniques in the field of protein structure prediction, for the prediction of gamma-turns. All the methods have been trained and tested on a set of 320 nonhomologous protein chains by a fivefold cross-validation technique. It has been observed that the performance of all methods is very poor, having a Matthew's Correlation Coefficient (MCC) </= 0.06. Second, predicted secondary structure obtained from PSIPRED is used in gamma-turn prediction. It has been found that machine-learning methods outperform statistical methods and achieve an MCC of 0.11 when secondary structure information is used. The performance of gamma-turn prediction is further improved when multiple sequence alignment is used as the input instead of a single sequence. Based on this study, we have developed a method, GammaPred, for gamma-turn prediction (MCC = 0.17). The GammaPred is a neural-network-based method, which predicts gamma-turns in two steps. In the first step, a sequence-to-structure network is used to predict the gamma-turns from multiple alignment of protein sequence. In the second step, it uses a structure-to-structure network in which input consists of predicted gamma-turns obtained from the first step and predicted secondary structure obtained from PSIPRED.

Databases, Protein↗

Predicting the liver histology in chronic hepatitis C: how good is the clinician?

OBJECTIVE: Liver biopsy is believed to be necessary before antiviral treatment in hepatitis C. Studies have found symptoms and biochemistry poorly predictive of grade and stage. In practice, a combination of factors is used to anticipate histology. The aim of this study is to evaluate the ability of global clinical assessment to predict histology in hepatitis C. METHODS: Fifty-four consecutive patients referred to a university center for consideration of antiviral therapy were enrolled. Clinical and laboratory data were recorded as was a prediction of the inflammatory grade (0-3) and fibrotic stage (0-3), with fibrotic stage 3 referring to cirrhosis. Liver biopsies were read by a blinded pathologist. The predictive value of the clinical assessment and individual parameters was assessed. RESULTS: All predictions were < or = 1 point off the actual grade and stage. Thirty-six (66.7%) patients' grades and 41 (75.9%) patients' stages were exactly predicted. All four cirrhotic patients (sensitivity 100%, specificity 94%) and one case of hemochromatosis were correctly predicted. Spider nevi, organomegaly, white blood cell count < or = 4 x 10(9)/L, ALT > 120 U/L, bilirubin > 20 micromol/L, albumin < or = 35 g/L, and ferritin > 200 microg/L predicted grade > or =2. Stage > or =2 was associated with age > 40 yr, previous decompensation, spider nevi, organomegaly, white blood cell count < or = 4 x 10(9)/L, albumin < or = 35 g/L, platelets < or = 150 x 10(9)/L, and international normalized ratio > 1.2. Grade correlated with stage (Spearman coefficient = 0.54, p < 0.001). By multivariate analysis, ferritin plus spider nevi or hypoalbuminemia was independently predictive of inflammation. Spider nevi and thrombocytopenia, with either splenomegaly or hypoalbuminemia, were useful three-variable models for predicting fibrosis. The corresponding scoring systems produced useful likelihood ratios. CONCLUSIONS: Global clinical assessment mirroring clinical practice in a tertiary liver transplant center is moderately accurate in predicting grade and stage in hepatitis C. Liver biopsy is the current gold standard; however, the amount of new information gleaned is less than was perceived. The need for routine biopsy before antiviral treatment in hepatitis C should be reevaluated in a multicenter study.

Adult↗

Epstein-Barr virus genome may encode a protein showing significant amino acid and predicted secondary structure homology with glycoprotein B of herpes simplex virus 1.

We report significant sequence and predicted secondary structure homology between the herpes simplex virus 1 glycoprotein B (gB) and a protein predicted to be encoded by the BALF4 reading frame of Epstein-Barr virus (EBV). Homology was detectable at the DNA level and was highly significant at the protein level and when evolutionary substitution frequencies of amino acids in related proteins were taken into account. Hydropathic analyses predicted that the two proteins possess conserved N-terminal and C-terminal hydrophobic domains. The N-terminal hydrophobic domains share features in common with known cleavable membrane insertion signal sequences. The amino acid sequences of the C-terminal hydrophobic domains predict three adjacent membrane-spanning segments as had been previously predicted for gB. In an alignment of the two amino acid sequences, 247 of 903 gB residues had a matched pair in the BALF4 sequence, and 247 of 854 BALF4 residues were found to have a matched pair in the gB sequence. In addition, all 10 cysteine residues located outside the predicted signal sequence of both proteins were conserved, as were four predicted N-linked glycosylation sites. In all, 43% of the residues in the aligned sequences are predicted to possess equivalent secondary structures. gB is a virion envelope glycoprotein required for virus entry into cells. The domain of gB determining the rate of entry into cells has been mapped; the predicted structure of this domain in gB and the predicted EBV protein are almost identical. Similarly, the cytoplasmic domain of gB postulated to interact with submembrane proteins was also nearly identical in predicted structure to that of the EBV protein. These results suggest that EBV encodes a protein similar in structure and function to the herpes simplex virus 1 gB.

Amino Acid Sequence↗

CooPPS: a system for the cooperative prediction of protein structures.

Predicting the three-dimensional structure of proteins is a difficult task. In the last few years several approaches have been proposed for performing this task taking into account different protein chemical and physical properties. As a result, a growing number of protein structure prediction tools is becoming available, some of them specialized to work on either some aspects of the predictions or on some categories of proteins; however, they are still not sufficiently accurate and reliable for predicting all kinds of proteins. In this context, it is useful to jointly apply different prediction tools and combine their results in order to improve the quality of the predictions. However, several problems have to be solved in order to make this a viable possibility. In this paper a framework and a tool is proposed which allows: (i) definition of a common reference applicative domain for different prediction tools; (ii) characterization of prediction tools through evaluating some quality parameters; (iii) characterization of the performances of a team of predictors jointly applied over a prediction problem; (iv) the singling out of the best team for a prediction problem; and (v) the integration of predictor results in the team in order to obtain a unique prediction. A system implementing the various steps of the proposed framework (CooPPS) has been developed and several experiments for testing the effectiveness of the proposed approach have been carried out.

Algorithms↗

A valuable improvement of adult height prediction methods in short normal children.

OBJECTIVES: The potential benefit of growth hormone (GH) administration to increase adult height of normal children of short stature might be blurred by the accuracy and the precision of the prediction methods used to estimate final height before onset of therapy. The aim of the present study was to evaluate three prediction methods: Bayley-Pinneau (BP), Roche-Wainer-Thissen (RWT) and Tanner-Whitehouse Mark II (TW2) and to improve their accuracy and precision by exploring their correlation with various parameters obtained in peripubertal children with poor predicted adult height. STUDY DESIGN: Accuracy and precision of the prediction methods were evaluated retrospectively by comparing predicted adult heights estimated in 62 boys at 13.7 +/- 0.9 years and in 28 girls at 12.1 +/- 0.9 years of age, with their adult heights measured respectively at 20.7 +/- 2.6 years and 18.8 +/- 2.8 years. RESULTS: At the time of prediction, the height for chronological age was -2.07 +/- 0.68 standard deviation scores for boys and -2.15 +/- 0.6 years for girls. Measured adult heights were significantly lower than target heights (165.1 +/- 5.1 vs. 169.4 +/- 4.8 cm for boys; p < 0.001 and 153.1 +/- 3.9 vs. 156.3 +/- 5.0 cm for girls; p = 0.001). For boys, the BP method was the most accurate and also the most convenient with a predicted adult height of 164.7 +/- 5.0 cm and a small underestimation of 0.4 +/- 3.5 cm. For girls, the TW2 method was the most accurate with a predicted height of 152.4 +/- 3.7 cm with a little underestimation of 0.7 +/- 3.5 cm. There were no important differences between the precision of these methods. The use of a correction factor derived from the bone age delay at the time of prediction in boys and from the chronological age at the time of prediction in girls improved the accuracy of the predicted adult height. CONCLUSIONS: The use of a factor correcting the accuracy of the BP method in boys and of the TW2 method in girls should be valuable in assessing the potential benefit of GH therapy to increase adult height in short normal children.

Adolescent↗

Local protein structure prediction using discriminative models.

BACKGROUND: In recent years protein structure prediction methods using local structure information have shown promising improvements. The quality of new fold predictions has risen significantly and in fold recognition incorporation of local structure predictions led to improvements in the accuracy of results. We developed a local structure prediction method to be integrated into either fold recognition or new fold prediction methods. For each local sequence window of a protein sequence the method predicts probability estimates for the sequence to attain particular local structures from a set of predefined local structure candidates. The first step is to define a set of local structure representatives based on clustering recurrent local structures. In the second step a discriminative model is trained to predict the local structure representative given local sequence information. RESULTS: The step of clustering local structures yields an average RMSD quantization error of 1.19 A for 27 structural representatives (for a fragment length of 7 residues). In the prediction step the area under the ROC curve for detection of the 27 classes ranges from 0.68 to 0.88. CONCLUSION: The described method yields probability estimates for local protein structure candidates, giving signals for all kinds of local structure. These local structure predictions can be incorporated either into fold recognition algorithms to improve alignment quality and the overall prediction accuracy or into new fold prediction methods.

Algorithms↗

A two-stage approach for improved prediction of residue contact maps.

BACKGROUND: Protein topology representations such as residue contact maps are an important intermediate step towards ab initio prediction of protein structure. Although improvements have occurred over the last years, the problem of accurately predicting residue contact maps from primary sequences is still largely unsolved. Among the reasons for this are the unbalanced nature of the problem (with far fewer examples of contacts than non-contacts), the formidable challenge of capturing long-range interactions in the maps, the intrinsic difficulty of mapping one-dimensional input sequences into two-dimensional output maps. In order to alleviate these problems and achieve improved contact map predictions, in this paper we split the task into two stages: the prediction of a map's principal eigenvector (PE) from the primary sequence; the reconstruction of the contact map from the PE and primary sequence. Predicting the PE from the primary sequence consists in mapping a vector into a vector. This task is less complex than mapping vectors directly into two-dimensional matrices since the size of the problem is drastically reduced and so is the scale length of interactions that need to be learned. RESULTS: We develop architectures composed of ensembles of two-layered bidirectional recurrent neural networks to classify the components of the PE in 2, 3 and 4 classes from protein primary sequence, predicted secondary structure, and hydrophobicity interaction scales. Our predictor, tested on a non redundant set of 2171 proteins, achieves classification performances of up to 72.6%, 16% above a base-line statistical predictor. We design a system for the prediction of contact maps from the predicted PE. Our results show that predicting maps through the PE yields sizeable gains especially for long-range contacts which are particularly critical for accurate protein 3D reconstruction. The final predictor's accuracy on a non-redundant set of 327 targets is 35.4% and 19.8% for minimum contact separations of 12 and 24, respectively, when the top length/5 contacts are selected. On the 11 CASP6 Novel Fold targets we achieve similar accuracies (36.5% and 19.7%). This favourably compares with the best automated predictors at CASP6. CONCLUSION: Our final system for contact map prediction achieves state-of-the-art performances, and may provide valuable constraints for improved ab initio prediction of protein structures. A suite of predictors of structural features, including the PE, and PE-based contact maps, is available at http://distill.ucd.ie.

Algorithms↗

Empirical evaluation of prediction intervals for cancer incidence.

BACKGROUND: Prediction intervals can be calculated for predicting cancer incidence on the basis of a statistical model. These intervals include the uncertainty of the parameter estimates and variations in future rates but do not include the uncertainty of assumptions, such as continuation of current trends. In this study we evaluated whether prediction intervals are useful in practice. METHODS: Rates for the period 1993-97 were predicted from cancer incidence rates in the five Nordic countries for the period 1958-87. In a Poisson regression model, 95% prediction intervals were constructed for 200 combinations of 20 cancer types for males and females in the five countries. The coverage level was calculated as the proportion of the prediction intervals that covered the observed number of cases in 1993-97. RESULTS: Overall, 52% (104/200) of the prediction intervals covered the observed numbers. When the prediction intervals were divided into quartiles according to the number of cases in the last observed period, the coverage level was inversely proportional to the frequency (84%, 52%, 46% and 26%). The coverage level varied widely among the five countries, but the difference declined after adjustment for the number of cases in each country. CONCLUSION: The coverage level of prediction intervals strongly depended on the number of cases on which the predictions were based. As the sample size increased, uncertainty about the adequacy of the model dominated, and the coverage level fell far below 95%. Prediction intervals for cancer incidence must therefore be interpreted with caution.

Adolescent↗

Length of sick leave - why not ask the sick-listed? Sick-listed individuals predict their length of sick leave more accurately than professionals.

BACKGROUND: The knowledge of factors accurately predicting the long lasting sick leaves is sparse, but information on medical condition is believed to be necessary to identify persons at risk. Based on the current practice, with identifying sick-listed individuals at risk of long-lasting sick leaves, the objectives of this study were to inquire the diagnostic accuracy of length of sick leaves predicted in the Norwegian National Insurance Offices, and to compare their predictions with the self-predictions of the sick-listed. METHODS: Based on medical certificates, two National Insurance medical consultants and two National Insurance officers predicted, at day 14, the length of sick leave in 993 consecutive cases of sick leave, resulting from musculoskeletal or mental disorders, in this 1-year follow-up study. Two months later they reassessed 322 cases based on extended medical certificates. Self-predictions were obtained in 152 sick-listed subjects when their sick leave passed 14 days. Diagnostic accuracy of the predictions was analysed by ROC area, sensitivity, specificity, likelihood ratio, and positive predictive value was included in the analyses of predictive validity. RESULTS: The sick-listed identified sick leave lasting 12 weeks or longer with an ROC area of 80.9% (95% CI 73.7-86.8), while the corresponding estimates for medical consultants and officers had ROC areas of 55.6% (95% CI 45.6-65.6%) and 56.0% (95% CI 46.6-65.4%), respectively. The predictions of sick-listed males were significantly better than those of female subjects, and older subjects predicted somewhat better than younger subjects. Neither formal medical competence, nor additional medical information, noticeably improved the diagnostic accuracy based on medical certificates. CONCLUSION: This study demonstrates that the accuracy of a prognosis based on medical documentation in sickness absence forms, is lower than that of one based on direct communication with the sick-listed themselves.

Administrative Personnel↗

An adaptive prediction and detection algorithm for multistream syndromic surveillance.

BACKGROUND: Surveillance of Over-the-Counter pharmaceutical (OTC) sales as a potential early indicator of developing public health conditions, in particular in cases of interest to biosurvellance, has been suggested in the literature. This paper is a continuation of a previous study in which we formulated the problem of estimating clinical data from OTC sales in terms of optimal LMS linear and Finite Impulse Response (FIR) filters. In this paper we extend our results to predict clinical data multiple steps ahead using OTC sales as well as the clinical data itself. METHODS: The OTC data are grouped into a few categories and we predict the clinical data using a multichannel filter that encompasses all the past OTC categories as well as the past clinical data itself. The prediction is performed using FIR (Finite Impulse Response) filters and the recursive least squares method in order to adapt rapidly to nonstationary behaviour. In addition, we inject simulated events in both clinical and OTC data streams to evaluate the predictions by computing the Receiver Operating Characteristic curves of a threshold detector based on predicted outputs. RESULTS: We present all prediction results showing the effectiveness of the combined filtering operation. In addition, we compute and present the performance of a detector using the prediction output. CONCLUSION: Multichannel adaptive FIR least squares filtering provides a viable method of predicting public health conditions, as represented by clinical data, from OTC sales, and/or the clinical data. The potential value to a biosurveillance system cannot, however, be determined without studying this approach in the presence of transient events (nonstationary events of relatively short duration and fast rise times). Our simulated events superimposed on actual OTC and clinical data allow us to provide an upper bound on that potential value under some restricted conditions. Based on our ROC curves we argue that a biosurveillance system can provide early warning of an impending clinical event using ancillary data streams (such as OTC) with established correlations with the clinical data, and a prediction method that can react to nonstationary events sufficiently fast. Whether OTC (or other data streams yet to be identified) provide the best source of predicting clinical data is still an open question. We present a framework and an example to show how to measure the effectiveness of predictions, and compute an upper bound on this performance for the Recursive Least Squares method when the following two conditions are met: (1) an event of sufficient strength exists in both data streams, without distortion, and (2) it occurs in the OTC (or other ancillary streams) earlier than in the clinical data.

Algorithms↗

Evolutionary conservation analysis increases the colocalization of predicted exonic splicing enhancers in the BRCA1 gene with missense sequence changes and in-frame deletions, but not polymorphisms.

INTRODUCTION: Aberrant pre-mRNA splicing can be more detrimental to the function of a gene than changes in the length or nature of the encoded amino acid sequence. Although predicting the effects of changes in consensus 5' and 3' splice sites near intron:exon boundaries is relatively straightforward, predicting the possible effects of changes in exonic splicing enhancers (ESEs) remains a challenge. METHODS: As an initial step toward determining which ESEs predicted by the web-based tool ESEfinder in the breast cancer susceptibility gene BRCA1 are likely to be functional, we have determined their evolutionary conservation and compared their location with known BRCA1 sequence variants. RESULTS: Using the default settings of ESEfinder, we initially detected 669 potential ESEs in the coding region of the BRCA1 gene. Increasing the threshold score reduced the total number to 464, while taking into consideration the proximity to splice donor and acceptor sites reduced the number to 211. Approximately 11% of these ESEs (23/211) either are identical at the nucleotide level in human, primates, mouse, cow, dog and opossum Brca1 (conserved) or are detectable by ESEfinder in the same position in the Brca1 sequence (shared). The frequency of conserved and shared predicted ESEs between human and mouse is higher in BRCA1 exons (2.8 per 100 nucleotides) than in introns (0.6 per 100 nucleotides). Of conserved or shared putative ESEs, 61% (14/23) were predicted to be affected by sequence variants reported in the Breast Cancer Information Core database. Applying the filters described above increased the colocalization of predicted ESEs with missense changes, in-frame deletions and unclassified variants predicted to be deleterious to protein function, whereas they decreased the colocalization with known polymorphisms or unclassified variants predicted to be neutral. CONCLUSION: In this report we show that evolutionary conservation analysis may be used to improve the specificity of an ESE prediction tool. This is the first report on the prediction of the frequency and distribution of ESEs in the BRCA1 gene, and it is the first reported attempt to predict which ESEs are most likely to be functional and therefore which sequence variants in ESEs are most likely to be pathogenic.

Breast Neoplasms↗

Predicting abundance of desert riparian birds: validation and calibration of the Effective Area Model.

Reliable prediction of the effects of landscape change on species abundance is critical to land managers who must make frequent, rapid decisions with long-term consequences. However, due to inherent temporal and spatial variability in ecological systems, previous attempts to predict species abundance in novel locations and/or time frames have been largely unsuccessful. The Effective Area Model (EAM) uses change in habitat composition and geometry coupled with response of animals to habitat edges to predict change in species abundance at a landscape scale. Our research goals were to validate EAM abundance predictions in new locations and to develop a calibration framework that enables absolute abundance predictions in novel regions or time frames. For model validation, we compared the EAM to a null model excluding edge effects in terms of accurate prediction of species abundance. The EAM outperformed the null model for 83.3% of species (N=12) for which it was possible to discern a difference when considering 50 validation sites. Likewise, the EAM outperformed the null model when considering subsets of validation sites categorized on the basis of four variables (isolation, presence of water, region, and focal habitat). Additionally, we explored a framework for producing calibrated models to decrease prediction error given inherent temporal and spatial variability in abundance. We calibrated the EAM to new locations using linear regression between observed and predicted abundance with and without additional habitat covariates. We found that model adjustments for unexplained variability in time and space, as well as variability that can be explained by incorporating additional covariates, improved EAM predictions. Calibrated EAM abundance estimates with additional site-level variables explained a significant amount of variability (P < 0.05) in observed abundance for 17 of 20 species, with R2 values >25% for 12 species, >48% for six species, and >60% for four species when considering all predictive models. The calibration framework described in this paper can be used to predict absolute abundance in sites different from those in which data were collected if the target population of sites to which one would like to statistically infer is sampled in a probabilistic way.

Animals↗

Prediction of drug-drug interactions for AUCoral of high clearance drug from in vitro data: utilization of a microtiter plate assay and a dispersion model.

The purpose of this study was to propose a new method to predict in vivo drug-drug interactions (DDIs) for a high clearance drug from in vitro data. As the high clearance drug, NE-100 (N, N-dipropyl-2-[4-methoxy-3-(2-phenylethoxy)phenyl]ethylamine monohydrochloride) was used. First, approach based on I(u)/K(i) value was used for the prediction of DDIs between NE-100 and concomitant drugs. When the K(i) values (K(i-cal)) obtained from the microtiter plate (MTP) assay and the reported K(i) values (K(i-rep)) for these drugs were used to predict increases at levels of NE-100 AUC(oral) (AUC(oral) ratio), the AUC(oral) ratios from the I(u)/K(i-cal) correlated with those from the I(u)/K(i-rep). This result suggests that the K(i-cal) from the MTP assay can be used for prediction of DDIs instead of the K(i-rep) value. Second, a new approach combining the inhibition rate (R) calculated from the MTP assay and two physiological models was used to predict DDIs. When the AUC(oral) ratios of NE-100 by various drugs were predicted using the R value and the well-stirred model, the ratios were similar to those predicted using the I(u)/K(i). However, after co-administration of drugs such as quinidine, propafenone and thioridazine (potent inhibitors of CYP2D6), the NE-100 AUC(oral) ratios predicted from the dispersion model was much greater than those from well-stirred model. This result shows that application of the dispersion model to the prediction method using the R value might sensitively and precisely predict the increased levels of AUC(oral) by DDIs for high clearance drug, compared with the prediction method using I(u)/K(i) value.

Administration, Oral↗

Using ultrasound measurements to predict body composition of yearling bulls.

Carcass traits have been successfully used to determine body composition of steers. Body composition, in turn, has been used to predict energy content of ADG to compute feed requirements of individual animals fed in groups. This information is used in the Cornell value discovery system (CVDS) to predict DM required (DMR) for the observed animal performance. In this experiment, the prediction of individual DMR for the observed performance of group-fed yearling bulls was evaluated using energy content of gain, which was based on ultrasound measurements to estimate carcass traits and energy content of ADG. One hundred eighteen spring-born purebred and crossbred bulls (BW = 288 +/- 4.3 kg) were sorted visually into 3 marketing groups based on estimated days to reach USDA low Choice quality grade. The bulls were fed a common high-concentrate diet in 12 slatted-floor pens (9 to 10 head/pen). Ultrasound measurements including back-fat (uBF), rump fat, LM area (uLMA), and intramuscular fat were taken at approximately 1 yr of age. Carcass measurements including HCW, backfat over the 12th to 13th rib (BF), marbling score (MRB), and LM area (LMA) were collected for comparison with ultrasound data for predicting carcass composition. The 9th to 11th-rib section was removed and dissected into soft tissue and bone for determination of chemical composition, which was used to predict carcass fat and empty body fat (EBF). The predicted EBF averaged 23.7 +/- 4.0%. Multiple regression analysis indicated that carcass traits explained 72% of the variation in predicted EBF (EBF = 16.0583 + 5.6352 x BF + 0.01781 x HCW + 1.0486 x MRB - 0.1239 x LMA). Because carcass traits are not available on bulls intended for use as herd sires, another equation using predicted HCW (pHCW) and ultrasound measurements was developed (EBF = 39.9535 x uBF - 0.1384 x uLMA + 0.0867 x pHCW - 0.0897 x uBF x pHCW - 1.3690). This equation accounted for 62% of the variation in EBF. The use of an equation to predict EBF developed with steer composition data overpredicted the EBF predicted in these experiments (28.7 vs. 23.7%, respectively). In a validation study with 37 individually fed bulls, the use of the ultrasound-based equation in the CVDS to predict energy content of gain accounted for 60% of the variation in the observed efficiency of gain, with 1.5% bias, and identified 3 of the 4 most efficient bulls.

Adipose Tissue↗