Search PubMedSearch

PubMed · 40366734

Missense variants pathogenicity annotation from homologous proteins.

Abstract

MOTIVATION: High-throughput DNA sequencing has revealed millions of single nucleotide variants (SNVs) in the human genome, with a small fraction linked to disease. The effect of missense variants, which alter the protein sequence, is particularly challenging to interpret due to the scarcity of clinical annotations and experimental information. While using conservation and structural information, current prediction tools still struggle to predict variant pathogenicity. In this study, we explored the pathogenicity of homologous missense variants-variants in equivalent positions across homologous proteins-focusing on proteins involved in autosomal dominant diseases. RESULTS: Our analysis of 2976 pathogenic and 17 555 non-pathogenic homologous variants demonstrated that pathogenicity can be extrapolated with 95% accuracy within a family, or up to 98% for closer homologs. Remarkably, the evaluation of 27 commonly used mutation predictor methods revealed that they were not fully capturing this biological feature. To facilitate the exploration of homologous variants, we created HomolVar, a web server that computationally predicts the pathogenesis of missense variants using annotations from homologous variants, freely available at https://rarevariants.org/HomolVar. Overall, these findings and the accompanying tool offer a robust method for predicting the pathogenicity of unannotated variants, enhancing genotype-phenotype correlations, and contributing to diagnosing rare genetic disorders. AVAILABILITY AND IMPLEMENTATION: HomolVar is freely available at https://rarevariants.org/HomolVar.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Gabriel Ruiz-Alías, Sergi Soldevila, Xavier Altafaj, Arnau Cordomí, Mireia Olivella. 2025-05-06. Missense variants pathogenicity annotation from homologous proteins.. https://doi.org/10.1093/bioinformatics%2Fbtaf305

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Reinforcement learning-based dynamic ensemble for missense variant effect prediction and tiered prioritization of VUS.

BACKGROUND: Accurate classification of missense variants remains a challenging task despite major advances in genomics. Numerous computational models have been developed to assist in variant classification, but often require repeated integration and benchmarking efforts. Ensemble methods have been proposed to overcome the limitations of single predictors, but mostly rely on fixed, predefined weights that constrain their ability to capture interactions among predictive signals. METHODS: We present GenixRL, a dynamic ensemble framework that reformulates model fusion as a reinforcement learning optimization problem. GenixRL uses a Q-learning agent to learn a policy that dynamically weights the probabilistic outputs of complementary predictors, including BayesDel (addAF and noAF), ClinPred, and MetaRNN. Replacing static weighting with policy learning allows GenixRL to adaptively identify optimal weightings and substantially improve classification accuracy. RESULTS: In benchmark evaluation against 25 state-of-the-art predictors, GenixRL achieved an AUROC of 0.9644 on an independent ClinVar dataset. On saturation genome editing assays for BRCA1 and BRCA2, GenixRL achieved the best performance and ranked highest on 14 of 17 clinically significant genes in a zero-shot evaluation. Applied to uncertain and conflicting ClinVar variants, GenixRL enabled tiered, evidence-based prioritization of hundreds of thousands of variants as likely pathogenic or pathogenic with high confidence, supported by orthogonal population evidence from gnomAD. CONCLUSION: GenixRL advances pathogenicity prediction for missense variants and provides an adaptive ensemble that sorts variants of uncertain significance into tiered candidates for expert curation and functional validation.

Mutation, Missense

Deep learning-based assessment of missense variants in the COG4 gene presented with bilateral congenital cataract.

OBJECTIVE: We compared the protein structure and pathogenicity of clinically relevant variants of the COG4 gene with AlphaFold2 (AF2), Alpha Missense (AM), and ThermoMPNN for the first time. METHODS AND ANALYSIS: The sequences of clinically relevant Cog4 missense variants (one novel identified p.Y714F and three pre-existing p.G512R, p.R729W and p.L769R from Uniprot Q9H9E3) were imported into AF2 for protein structural prediction, and the pathogenicity was estimated using AM and ThermoMPNN. Different pathogenicity metrics were aggregated with principal component analysis (PCA) and further analysed at three levels (amino acid position, substitution and post-translation) based on all possible Cog4 missense variants (n=14 915). RESULTS: Localised protein structural impact including change of conformation and amino acid polarity, breakage of hydrogen bond and salt-bridge, and formation of alpha-helix were identified among clinically relevant Cog4 variants. The global structural comparison with multidimensional scaling demonstrated variants with similar protein structures (AF2) tended to exhibit similar clinical and biological phenotypes. The Cog4 p.Y714F variant exhibited greater protein structural similarity to mutated Cog4 found in Saul‒Wilson syndrome (p.G512R) and shared similar clinical phenotype (congenital cataract and psychomotor retardation). PCA of included pathogenic metrics demonstrated p.Y714F occurred at a critical position in Cog4 amino acid sequence with disrupted post-translational phosphorylation. CONCLUSION: Deep learning algorithms, including AF2, AM and ThermoMPNN, can be useful for evaluating variant of uncertain significance (VUS) by structural and pathogenicity prediction. Despite classified as VUS (American College of Medical Genetics and Genomics criteria: PM1, PP4), the pathogenicity in this Cog4 variant cannot be ruled out and warrants further investigation.

Mutation, Missense