Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

Attempting to predict hospital admission in acute asthma.

If accurate predictions of outcome could be made, the emergency care of patients with asthma would be expedited. To evaluate how well initial peak flow determinations predicted hospitalization when used alone and when combined with other clinical variables, we prospectively studied 200 visits for the care of acute asthma. Using discriminant analysis, we selected the variables that best predicted discharge or admission for the first 100 cases. The best predictive variables were initial peak flow, history of treatment in preceding 24 hours, age at onset of asthma, and number of previous hospitalizations for asthma. This combination correctly predicted admission or discharge for 82% of the 200 cases. Despite this overall accuracy, admission was not well predicted. In the first 100 cases, only six of the 18 admissions were correctly predicted, and in the second 100 cases, none of the 15 admissions were correctly predicted, Initial peak flow measurements, even when combined with other variables, cannot predict hospitalization well enough to be substituted for a therapeutic trial.

Acute Disease↗

Genomic and Developmental Models to Predict Cognitive and Adaptive Outcomes in Autistic Children.

IMPORTANCE: Although early signs of autism are often observed between 18 and 36 months of age, there is considerable uncertainty regarding future development. Clinicians lack predictive tools to identify those who will later be diagnosed with co-occurring intellectual disability (ID). OBJECTIVE: To predict ID in children diagnosed with autism. DESIGN, SETTING, AND PARTICIPANTS: This prognostic study involved the development and validation of models integrating genetic variants and developmental milestones to predict ID. Models were trained, cross-validated, and tested for generalizability across 3 autism cohorts: Simons Foundation Powering Autism Research (SPARK), Simons Simplex Collection, and MSSNG. Autistic participants were assessed older than 6 years of age for ID. Study data were analyzed from January 2023 to July 2024. EXPOSURES: Ages at attaining early developmental milestones, occurrence of language regression, polygenic scores for cognitive ability and autism, rare copy number variants, de novo loss-of-function and missense variants impacting constrained genes. MAIN OUTCOMES AND MEASURES: The out-of-sample performance of predictive models was assessed using the area under the receiver operating characteristic curve (AUROC), positive predictive values (PPVs), and negative predictive values (NPVs). RESULTS: A total of 5633 autistic participants (4574 male [81.2%]) were included in this analysis. On average, participants were diagnosed with autism at 4 (IQR, 3-7) years of age and assessed for ID at 11 (8-14) years of age, with 1159 participants (20.6%) being diagnosed with ID. The model integrating all predictors yielded an AUROC of 0.653 (95% CI, 0.625-0.681), and this predictive performance was cross-validated and generalized across cohorts. This modest performance reflected that only a subset of individuals carried large-effect variants, high polygenic scores, or presented delayed milestones. However, combinations of genetic variants that are typically not considered clinically relevant by diagnostic laboratories achieved PPVs of 55% and correctly identified 10% of individuals developing ID. The addition of polygenic scores to developmental milestones specifically improved NPVs rather than PPVs. Notably, the ability to stratify ID probabilities using genetic variants was up to 2-fold higher in individuals with delayed milestones compared with those with typical development. CONCLUSIONS AND RELEVANCE: Results of this prognostic study suggest that the growing number of neurodevelopmental condition-associated variants cannot, in most cases, be used alone for predicting ID. However, models combining different classes of variants with developmental milestones provide clinically relevant individual-level predictions that could be useful for targeting early interventions.

Humans↗

Processing and analysis of CASP3 protein structure predictions.

Livermore Prediction Center provides basic infrastructure for the CASP (Critical Assessment of Structure Prediction) experiments, including prediction processing and verification servers, a system of prediction evaluation tools, and interactive numerical and graphical displays. Here we outline the essentials of our approach, with discussion of the superposition procedures, definitions of basic measures, and descriptions of new methods developed to analyze predictions. Our primary focus is on the evaluation of three-dimensional models and secondary structure predictions. To put the results of the three prediction experiments held to date on the same footing, the latest CASP3 evaluation criteria were retrospectively applied to both CASP1 and CASP2 predictions. Finally, we give an overview of our website (http:/(/)PredictionCenter.llnl.gov), which makes the target structures, predictions, and the evaluation system accessible to the community.

Amino Acid Sequence↗

Structure classification-based assessment of CASP3 predictions for the fold recognition targets.

The sequences of at least 23 of the 43 CASP3 targets showed no significant similarity to the sequences of known structures. The experimental structures of all but three of these 23 targets revealed substantial similarities to known structures, with at least eleven of the target structures likely being distantly homologous to known structures. Nineteen of the 23 target structures were available at the time of the final CASP3 meeting in Asilomar in December 1998, whereas the experimental data on the protein folds of the remaining four targets were obtained afterwards. The predicted three-dimensional structures for each of the 23 targets were analyzed to select those predictions sharing with the experimental structures a similar overall fold and/or having correctly folded a substantial fraction of the target sequence. Initially, predicted models were numerically evaluated and the evaluation results aided the selection process. Each target structure was then classified to identify a minimal set of structural features characteristic to its protein fold and evolutionary superfamily. The predictions containing this set were assessed comparatively to find the best predictions for each target. The predictions of new folds were assessed separately. The total number of the selected 'correct' predictions and the quality of these predictions were used to compare the performance of different predictor teams and different prediction methods in the fold prediction/recognition category.

Bacterial Proteins↗

Antenatal prediction of fetal pH in growth restricted fetuses using computer analysis of the fetal heart rate.

We tested the accuracy of a mathematical model based on computer analysis of the fetal heart rate tracing in predicting umbilical artery pH at birth. In a previous report based on data on 38 growth-restricted fetuses, the second-order polynomial regression equation, umbilical artery pH = 7.28 + 0.002 (duration of episodes of low variation in minutes) + 0.00009 (duration of episodes of low variation in minutes), was retrospectively found to be the best model for the prediction of umbilical artery pH at birth. In the present study, this formula was prospectively tested in 29 growth restricted fetuses between 26 and 37 weeks of gestation from pregnancies with abnormal uterine and/or umbilical artery Doppler velocimetry. Computer analysis of the fetal heart rate tracing of 1 hour duration was performed within 1.5-6 hours of cesarean birth prior to the onset of labor. Umbilical artery cord blood was collected at birth with pH determined within 5 minutes of collection. Acidemia was defined as umbilical artery pH < 7.20, preacidemia pH 7.20-7.25 and nonacidemia pH > 7.25. Then, the data on all 67 growth-restricted fetuses were pooled to generate a new formula that was retrospectively assessed against the entire group. Values are reported as median (range). In the 29 prospectively evaluated cases, there was no statistical difference between the predicted and actual umbilical artery pH at birth [7.28 (7.1-7.29) vs. 7.28 (7.18-7.37), P = 0.57]. The median difference between the paired predicted and actual umbilical artery pH values was -0.001 (-0.10-0.08). The difference between the predicted and actual umbilical artery pH was zero and within +/- 0.04 in 17% (5/29) and 76% (22/29) of the cases, respectively. When the data on the 67 growth-restricted fetuses were pooled together the formula did not change. There was no difference between the predicted and actual umbilical artery pH at birth when the formula was applied to all 67 growth-restricted fetuses [7.28 (7.08-7.29) vs. 7.27 (6.97-7.37), P = 0.41]. The median difference between the paired predicted and actual pH values was -0.001 (-0.12-0.12). The difference between the predicted and actual umbilical artery pH was zero and within +/- 0.04 in 15% (10/67) and 74% (49/67) of the cases, respectively. The accuracy of the formula in correctly categorizing the umbilical artery pH at birth was: acidemia 67% (8/12), preacidemia 28% (8/29) and nonacidemia 80% (37/46), P < 0.0001. A mathematical formula using the computer analysis index of duration of episodes of low variation reliably predicted umbilical artery pH at birth. This type of noninvasive monitoring may allow for the antepartum estimation and continuous tracking of fetal pH.

Birth Weight↗

Bayesian probabilistic approach for predicting backbone structures in terms of protein blocks.

By using an unsupervised cluster analyzer, we have identified a local structural alphabet composed of 16 folding patterns of five consecutive C(alpha) ("protein blocks"). The dependence that exists between successive blocks is explicitly taken into account. A Bayesian approach based on the relation protein block-amino acid propensity is used for prediction and leads to a success rate close to 35%. Sharing sequence windows associated with certain blocks into "sequence families" improves the prediction accuracy by 6%. This prediction accuracy exceeds 75% when keeping the first four predicted protein blocks at each site of the protein. In addition, two different strategies are proposed: the first one defines the number of protein blocks in each site needed for respecting a user-fixed prediction accuracy, and alternatively, the second one defines the different protein sites to be predicted with a user-fixed number of blocks and a chosen accuracy. This last strategy applied to the ubiquitin conjugating enzyme (alpha/beta protein) shows that 91% of the sites may be predicted with a prediction accuracy larger than 77% considering only three blocks per site. The prediction strategies proposed improve our knowledge about sequence-structure dependence and should be very useful in ab initio protein modelling.

Artificial Intelligence↗

Clinical and sign prediction: the draw-a-person and female homosexuality.

The validity of predicting female homosexuality from empirical signs from the Draw-A-Person (DAP) was compared to the validity of psychologists' "blind" predictions from the same DAP protocols. Four specific DAP signs significantly predicted homosexual drawings from those of heterosexual controls; patterns derived from these signs were even better predictors. Two of four clinicians predicted sexual orientation greater than chance; one predicted as well as the optimal sign pattern. In view of the expected shrinkage involved in cross-validating the empirical signs and patterns, clinical prediction probably would equal or surpass the accuracy of statistical prediction. Individual differences in clinical prediction and the process of clinical prediction were discussed.

Adolescent↗

Plasma Proteomic Profiles Predict Individual Future Osteoarthritis Risk.

OBJECTIVE: Osteoarthritis (OA) is a widespread degenerative joint disease that causes a considerable socioeconomic burden. Despite progress in genetic and environmental insights, early diagnosis is still limited by the lack of evident symptoms during the initial phases and accurate biomarkers. This study aims to identify plasma proteins associated with future risk of OA and develop a predictive model. METHODS: We conducted a large-scale proteomic analysis of 45,307 participants from the UK Biobank, excluding those with baseline OA. Plasma samples were assayed using the Olink Explore Proximity Extension Assay targeting 1,463 unique proteins. Clinical variables and OA outcomes were extracted and linked to electronic health records. A predictive model was constructed using the LightGBM machine learning method, and SHapley Additive exPlanations (SHAP) were applied to evaluate the importance of variables. RESULTS: We identified a panel of proteins significantly associated with the risk of developing OA. Notably, after adjusting for multiple confounders, collagen type IX alpha 1 chain (COL9A1) and cartilage acidic protein 1 (CRTAC1) were the most significant predictors of incident OA, with hazard ratios of 1.54 (95% confidence interval [CI] 1.48-1.61) and 1.65 (95% CI 1.54-1.78), respectively. SHAP analysis allowed a profound interpretation of the contribution of each protein and clinical variable to the model, revealing the multifactorial nature of OA risk prediction. The temporal trajectories of plasma proteins indicated that the levels of COL9A1 and CRTAC1 began to deviate from normal for more than a decade before OA onset, suggesting their potential use in early detection strategies. The predictive model, developed using the LightGBM algorithm, integrated proteins with clinical covariates and demonstrated an area under the curve (AUC) of 0.729 for 5-year OA prediction, 0.721 for 10-year prediction, and 0.723 for all incident OA. The predictive accuracy of the model was further enhanced for hip and knee OA, achieving AUCs of 0.820 and 0.803 for 5-year predictions. CONCLUSION: Our study identified the role of plasma proteomics in predicting future OA risk, which could contribute to preemptive measures. The innovative model, which integrates proteomic biomarkers with clinical data, offers a potential tool for risk assessment, potentially optimizing OA management strategies and enhancing prevention efforts.

Humans↗

A simple noninvasive index for predicting long-term outcome of chronic hepatitis C after interferon-based therapy.

Changes in hepatic fibrosis after interferon-based therapy may be important in determining the long-term outcome of chronic hepatitis C (CHC). The use of liver biopsy for posttreatment assessment is not a viable option as a routine follow-up procedure. This study evaluated the predictive value of a simple noninvasive index, the aspartate aminotransferase (AST)-to-platelet ratio index assessed 6 months after end of treatment (APRI-M6). We evaluated APRI-M6, platelet-M6, AST-M6, and alpha-fetoprotein-M6 of 776 CHC patients with interferon-based therapy as well as the parameters at baseline of 562 untreated patients who were evaluated to predict the risk of hepatocellular carcinoma (HCC) and mortality, during a mean follow-up period of 4.75 (1.0-12.2) and 5.15 (1.0-16) years, respectively. Based on analysis of receiver operating characteristics (ROC) and using optimized cutoff point, the APRI-M6 and platelet-M6 had superior prediction models for long-term outcome with area under the curve of 0.870-0.875 and 0.824-0.847, respectively, and accuracy of 78%-81% and 76%-78%, respectively, for interferon-based-treated patients. The predictive values of all 4 parameters were poor in untreated patients. In subgroup analysis, the APRI-M6 provided a more consistent prediction ratio than platelet-M6 for sustained responders and cirrhosis-free subgroups; both parameters had similar prediction power for nonresponders and were unsatisfactory in patients with cirrhosis. According to Cox proportional hazards analysis, cirrhosis and APRI-M6 were the 2 most important factors for predicting HCC. In conclusion, APRI-M6 can accurately predict the long-term outcome of patients subjected to interferon-based treatment. Nevertheless, the data needs further validation, particularly since the predictive accuracy for patients with cirrhosis is low.

Antiviral Agents↗

Resting energy expenditure should be measured in patients with cirrhosis, not predicted.

Measurements of resting energy expenditure (REE) can be used to determine energy requirements. Prediction formulae can be used to estimate REE but have not been validated in cirrhotic patients. REE was measured, by indirect calorimetry, in 100 cirrhotic patients and 41 comparable healthy volunteers, and the results compared with estimates predicted using the Harris-Benedict, Schofield, Mifflin, Cunningham, and Owen formulae, and the disease-specific Müller formula. The mean (+/- 1 SD) measured REE in the healthy volunteers (1,590 +/- 306 kcal/24 h) was significantly greater than the mean Harris-Benedict, Mifflin, Cunningham, and Owen predictions but comparable with the mean Schofield prediction; individual predicted values varied widely from measured values (95% limits of agreement, -460 to +424 kcal). The mean measured REE in the cirrhotic patients was significantly greater than in the healthy volunteers (23.2 +/- 3. 8 cf 21.9 +/- 2.9 kcal/kg/24 h; P <.05). The mean measured REE in the cirrhotic patients (1,660 +/- 337 kcal/24 h) was significantly different from mean predicted values (Harris-Benedict, 1,532 +/- 252 kcal/24 h, P <.0001; Schofield, 1,575 +/- 254 kcal/24 h, P <.0005; Mifflin, 1,460 +/- 254 kcal/24 h, P <.0001; Cunningham, 1,713 +/- 252 kcal/24 h, P <.05; Owen, 1,521 +/- 281 kcal/24 h, P <.0001; Müller, 1,783 +/- 204 kcal/24 h, P <.0001); individual predicted values varied widely from measured values (95% limits of agreement, -632 to +573 kcal). Simple regression analysis showed that fat-free mass (FFM) was the strongest predictor of measured REE in the cirrhotic patients, accounting for 52% of the variation observed. However, a population-specific prediction equation, derived using stepwise regression analysis, which incorporated FFM, age, and Pugh's score, accounted for only 61% of the observed variation in measured REE. REE should, therefore, be measured in cirrhotic patients, not predicted.

Adult↗

The relative predictive performance of two theophylline pharmacokinetic dosing programs.

The predictive performance of 2 theophylline pharmacokinetic dosing programs (Abbott and Simkin) was evaluated using a group of 44 inpatients who had 2 serum concentrations (TSC) measured during hospitalization. Bias was assessed with the median prediction error (PE) and precision was assessed with the median absolute PE. The Abbott program was significantly less biased than the Simkin program in predicting the first TSC (PEs 0.1 and -1.3 micrograms/ml, respectively; p less than 0.05). No significant difference in bias was observed in predicting the second TSC, or in precision in predicting either the first or second TSC. Both programs exhibited small improvements in prediction precision when the first TSC was used to predict the second. Correlations of predicted versus measured TSC also improved with the second prediction. These programs may be useful in dosing theophylline; however, TSC monitoring and the application of sound clinical judgment are warranted.

Adult↗

The accuracy of predicting functional recovery in patients following a stroke, by physiotherapists and patients.

BACKGROUND AND PURPOSE: The potential for post-stroke recovery and the range of predictive variables has been studied extensively. Knowledge of these variables alongside other factors, such as performance in therapy and professional experience, enable ongoing predictions to be made by members of the rehabilitation team. Patients' own predictions for their recovery has yet to receive much attention in this area of research. The aim of this study was to compare the predictive accuracy of the physiotherapist and the stroke patient with regard to functional change during a period 6-12 weeks post-stroke. METHOD: The stroke sample (N = 29) came from two National Health Service Trusts as did the physiotherapists (N = 4). No comparisons were made between the hospitals and data was coded for anonymity. Estimations were made by both physiotherapists and patients regarding items on each of the three sections of The Rivermead Motor Assessment (RMA). Intra-class correlation coefficients (ICCs) were used to describe agreement of each set of predictions with the achieved RMA scores. The results reported here represent the main emphasis of the research; however, other areas were also screened (for example, change in cognition, language and quality of life) by use of basic standardized measures. Recovery was also compared to other known predictive variables, such as age, severity of stroke and urinary incontinence. RESULTS: At follow-up assessment it was found that both physiotherapists' and patients' predictions demonstrated high and significant agreement with the achieved RMA scores at 12 weeks (ICCs ranging from 0.727 to 0.968). Physiotherapists' predictions demonstrated marginally higher levels of agreement than patients' predictions. CONCLUSIONS: The degree of accuracy demonstrated by both physiotherapists and patients was considerable. The patient group was perhaps the more notable as no subject had had prior knowledge of a stroke. The implications in respect of lay persons' involvement in decision making and in the rehabilitative process, alongside the health professionals, are perhaps worthy of closer consideration.

Aged↗

Accuracy of prediction of walking for young stroke patients by use of the FIM.

BACKGROUND AND PURPOSE: Clinical prediction of walking outcome after a stroke is essential for effective discharge planning. However, its accuracy has hardly been explored. This study took place in a regional unit admitting patients with complex neurological disabilities for specialist inpatient rehabilitation. The aim was to compare predicted outcome (goal score) with achieved outcome (discharge score) on the seven-point locomotion subscale of the Functional Independence Measure (FIM), to evaluate its precision and identify factors influencing accuracy. METHOD: Admission, goal and discharge scores were analysed retrospectively for 141 subjects (90 M; 51 F) admitted consecutively to the Unit with median age 54 years (range 15-68 years) with median length of stay 13.6 weeks (range 3-35 weeks). RESULTS: Ninety subjects (64%) gained from two to six points; 50 subjects (35%) gained one point or showed no change. One patient deteriorated by two points. Excluding patients admitted with the highest score (FIM level 7), the overall level of agreement between predicted and discharge scores was moderate (weighted kappa 0.47). Prediction was accurate to +/- 1 point in 113 subjects (80%). Overprediction by > or = 2 points occurred in 16 subjects (11%) and underprediction by > or = 2 points in 12 subjects (9%). Analysis of the most-disabled cohort, admitted with FIM levels 1 or 2 scores, revealed a higher sensitivity for predicting 'independence' (FIM levels 5-7) (78%) than 'dependence' (FIM levels 1-4) (65%). Accuracy was not affected by age, gender or side of stroke. Inaccurate predictions were associated with lower admission FIM level scores (p = -0.26; p = 0.002) and a greater length of stay (p = 0.36; p < 0.001). Subjects with quadriplegia were more likely to have inaccurate outcome predictions made than those with hemiplegia (p = 0.025) and those with neglect were more likely to have inaccurate outcome predictions made than those without neglect (p = 0.017). CONCLUSION: Further investigation into clinical prediction and the variables which confound accuracy is needed for effective planning.

Adolescent↗

Improving protein secondary structure prediction with aligned homologous sequences.

Most recent protein secondary structure prediction methods use sequence alignments to improve the prediction quality. We investigate the relationship between the location of secondary structural elements, gaps, and variable residue positions in multiple sequence alignments. We further investigate how these relationships compare with those found in structurally aligned protein families. We show how such associations may be used to improve the quality of prediction of the secondary structure elements, using the Quadratic-Logistic method with profiles. Furthermore, we analyze the extent to which the number of homologous sequences influences the quality of prediction. The analysis of variable residue positions shows that surprisingly, helical regions exhibit greater variability than do coil regions, which are generally thought to be the most common secondary structure elements in loops. However, the correlation between variability and the presence of helices does not significantly improve prediction quality. Gaps are a distinct signal for coil regions. Increasing the coil propensity for those residues occurring in gap regions enhances the overall prediction quality. Prediction accuracy increases initially with the number of homologues, but changes negligibly as the number of homologues exceeds about 14. The alignment quality affects the prediction more than other factors, hence a careful selection and alignment of even a small number of homologues can lead to significant improvements in prediction accuracy.

Amino Acid Sequence↗

Prediction of alpha-turns in proteins using PSI-BLAST profiles and secondary structure information.

In this paper a systematic attempt has been made to develop a better method for predicting alpha-turns in proteins. Most of the commonly used approaches in the field of protein structure prediction have been tried in this study, which includes statistical approach "Sequence Coupled Model" and machine learning approaches; i) artificial neural network (ANN); ii) Weka (Waikato Environment for Knowledge Analysis) Classifiers and iii) Parallel Exemplar Based Learning (PEBLS). We have also used multiple sequence alignment obtained from PSIBLAST and secondary structure information predicted by PSIPRED. The training and testing of all methods has been performed on a data set of 193 non-homologous protein X-ray structures using five-fold cross-validation. It has been observed that ANN with multiple sequence alignment and predicted secondary structure information outperforms other methods. Based on our observations we have developed an ANN-based method for predicting alpha-turns in proteins. The main components of the method are two feed-forward back-propagation networks with a single hidden layer. The first sequence-structure network is trained with the multiple sequence alignment in the form of PSI-BLAST-generated position specific scoring matrices. The initial predictions obtained from the first network and PSIPRED predicted secondary structure are used as input to the second structure-structure network to refine the predictions obtained from the first net. The final network yields an overall prediction accuracy of 78.0% and MCC of 0.16. A web server AlphaPred (http://www.imtech.res.in/raghava/alphapred/) has been developed based on this approach.

Amino Acid Sequence↗

Prediction of interface residues in protein-protein complexes by a consensus neural network method: test against NMR data.

The number of structures of protein-protein complexes deposited to the Protein Data Bank is growing rapidly. These structures embed important information for predicting structures of new protein complexes. This motivated us to develop the PPISP method for predicting interface residues in protein-protein complexes. In PPISP, sequence profiles and solvent accessibility of spatially neighboring surface residues were used as input to a neural network. The network was trained on native interface residues collected from the Protein Data Bank. The prediction accuracy at the time was 70% with 47% coverage of native interface residues. Now we have extensively improved PPISP. The training set now consisted of 1156 nonhomologous protein chains. Test on a set of 100 nonhomologous protein chains showed that the prediction accuracy is now increased to 80% with 51% coverage. To solve the problem of over-prediction and under-prediction associated with individual neural network models, we developed a consensus method that combines predictions from multiple models with different levels of accuracy and coverage. Applied on a benchmark set of 68 proteins for protein-protein docking, the consensus approach outperformed the best individual models by 3-8 percentage points in accuracy. To demonstrate the predictive power of cons-PPISP, eight complex-forming proteins with interfaces characterized by NMR were tested. These proteins are nonhomologous to the training set and have a total of 144 interface residues identified by chemical shift perturbation. cons-PPISP predicted 174 interface residues with 69% accuracy and 47% coverage and promises to complement experimental techniques in characterizing protein-protein interfaces. .

Binding Sites↗

Prediction of protein secondary structure content using amino acid composition and evolutionary information.

Knowing protein structure and inferring its function from the structure are one of the main issues of computational structural biology, and often the first step is studying protein secondary structure. There have been many attempts to predict protein secondary structure contents. Previous attempts assumed that the content of protein secondary structure can be predicted successfully using the information on the amino acid composition of a protein. Recent methods achieved remarkable prediction accuracy by using the expanded composition information. The overall average error of the most successful method is 3.4%. Here, we demonstrate that even if we only use the simple amino acid composition information alone, it is possible to improve the prediction accuracy significantly if the evolutionary information is included. The idea is motivated by the observation that evolutionarily related proteins share the similar structure. After calculating the homolog-averaged amino acid composition of a protein, which can be easily obtained from the multiple sequence alignment by running PSI-BLAST, those 20 numbers are learned by a multiple linear regression, an artificial neural network and a support vector regression. The overall average error of method by a support vector regression is 3.3%. It is remarkable that we obtain the comparable accuracy without utilizing the expanded composition information such as pair-coupled amino acid composition. This work again demonstrates that the amino acid composition is a fundamental characteristic of a protein. It is anticipated that our novel idea can be applied to many areas of protein bioinformatics where the amino acid composition information is utilized, such as subcellular localization prediction, enzyme subclass prediction, domain boundary prediction, signal sequence prediction, and prediction of unfolded segment in a protein sequence, to name a few.

Amino Acid Sequence↗

A method for alpha-helical integral membrane protein fold prediction.

Integral membrane proteins (of the alpha-helical class) are of central importance in a wide variety of vital cellular functions. Despite considerable effort on methods to predict the location of the helices, little attention has been directed toward developing an automatic method to pack the helices together. In principle, the prediction of membrane proteins should be easier than the prediction of globular proteins: there is only one type of secondary structure and all helices pack with a common alignment across the membrane. This allows all possible structures to be represented on a simple lattice and exhaustively enumerated. Prediction success lies not in generating many possible folds but in recognizing which corresponds to the native. Our evaluation of each fold is based on how well the exposed surface predicted from a multiple sequence alignment fits its allocated position. Just as exposure to solvent in globular proteins can be predicted from sequence variation, so exposure to lipid can be recognized by variable-hydrophobic (variphobic) positions. Application to both bacteriorhodopsin and the eukaryotic rhodopsin/opsin families revealed that the angular size of the lipid-exposed faces must be predicted accurately to allow selection of the correct fold. With the inherent uncertainties in helix prediction and parameter choice, this accuracy could not be guaranteed but the correct fold was typically found in the top six candidates. Our method provides the first completely automatic method that can proceed from a scan of the protein sequence databanks to a predicted three-dimensional structure with no intervention required from the investigator. Within the limited domain of the seven helix bundle proteins, a good chance can be given of selecting the correct structure. However, the limited number of sequences available with a corresponding known structure makes further characterization of the method difficult.

Amino Acid Sequence↗