Search PubMedSearch

SEARCH · Search PubMed

Results for “ClinVar”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

29 records · Page 2Linked to original sources

The ClinGen Severe Combined Immunodeficiency Disease Variant Curation Expert Panel: Specifications for classification of variants in ADA, DCLRE1C, IL2RG, IL7R, JAK3, RAG1, and RAG2.

PURPOSE: This collaborative study, led by the Clinical Genome Resource Severe Combined Immunodeficiency Disease Variant Curation Expert Panel (ClinGen SCID-VCEP), implemented and adapted the American College of Medical Genetics and Genomics/Association for Molecular Pathology (ACMG/AMP) guidelines for interpreting germline variants in genes with established relationships to SCID. The effort focused on the 7 most common SCID-related genes identified by SCID newborn screening in North America: ADA, DCLRE1C, IL2RG, IL7R, JAK3, RAG1, and RAG2. METHODS: The SCID-VCEP conducted a rigorous review of variants that involved database analyses, literature review, and expert feedback to derive gene-specific modifications to the ACMG/AMP guidelines. These specifications were validated using a pilot set of 90 variants. RESULTS: Of these 90 variants, 25 were classified as pathogenic, 21 as likely pathogenic, 14 as variants of uncertain significance, 18 as likely benign, and 12 as benign. Seventeen variants with conflicting classifications in ClinVar were successfully resolved. The criteria included modifications to 20 of the 28 original ACMG/AMP criteria specific to SCID-related genes. CONCLUSION: The SCID-specific variant curation guidelines developed by the SCID-VCEP will enhance the precision of SCID genetic diagnosis and provide a robust framework for interpreting variants in SCID-related genes, contributing to appropriate treatment of SCID.

Humans

Exploring penetrance of clinically relevant variants in over 800,000 humans from the Genome Aggregation Database.

Incomplete penetrance, or absence of disease phenotype in an individual with a disease-associated variant, is a major challenge in variant interpretation. Studying individuals with apparent incomplete penetrance can shed light on underlying drivers of altered phenotype penetrance. Here, we investigate clinically relevant variants from ClinVar in 807,162 individuals from the Genome Aggregation Database (gnomAD), demonstrating improved representation in gnomAD version 4. We then conduct a comprehensive case-by-case assessment of 734 predicted loss of function variants in 77 genes associated with severe, early-onset, highly penetrant haploinsufficient disease. Here, we identify explanations for the presumed lack of disease manifestation in 701 of 734 variants (95%). Individuals with unexplained lack of disease manifestation in this set of disorders are rare, underscoring the need and power of deep case-by-case assessment presented here to minimize false assignments of disease risk, particularly in unaffected individuals with higher rates of secondary properties that result in rescue.

Humans

GUANinE v1.1 reveals complementarity of supervised and genomic language models.

There has been much debate about the benefits of supervised versus unsupervised learning on genomes. Determining which is better in what contexts requires developing comprehensive benchmarks spanning functional and evolutionary tasks. Importantly, such benchmarks need large sample sizes to enable well-powered ranking of models. Having developed and applied such a benchmark here (GUANinE v1.1), we conclusively demonstrate each paradigm offers key advantages and outperforms on certain tasks. In accordance with training, supervised sequence-to-function models exhibit strong performance when annotating functional states characterized by chromatin accessibility or histone marks, while self-supervised language models outperform on evolutionary conservation. Our hundreds of new evaluations in this v1.1 expansion provide evidence for a tradeoff between input context size and model parameter count for a fixed compute budget, which we depict with new metrics such as kiloparameters/base pair. We also construct two new large-scale variant interpretation tasks in v1.1: cadd-snv measuring deleteriousness, and clinvar-snv measuring clinical pathogenicity. We find that conservation scores, and by extension, genomic language models, predict deleteriousness well, but successfully translating deleteriousness predictions to pathogenicity remains challenging. GUANinE v1.1 newly evaluates dozens of pretrained genomic models, and we conclude that moderate-context hybrid or post-trained language models may define the next era of machine learning in genomics.

Genomics

Post-translational modification of proteins in the human testis development pathway.

BACKGROUND: The foetal testes produce the androgens necessary to masculinise the developing embryo and support the maturation of germ cells, that will eventually develop into sperm, thus ensuring future reproductive capacity. The testes develop from the bi-potential gonads in a highly orchestrated process resulting in the differentiation of a complex tissue with multiple cellular lineages. While recent transcriptomic and chromatin-based analyses of human foetal testes have provided an unprecedented level of insight into signalling pathways activated during this process, proteomic studies of the human foetal gonads remain limited. Proteins are active molecules and post-translational modification (PTM) of proteins influences protein activity, stability and localisation. Studies have shown that PTMs regulate critical proteins in testis development, and their disruptions are implicated in congenital disorders including differences of sex development (DSD), in which sex development is atypical. Despite this, the role and regulation of protein PTM during human testis development remains poorly understood due to limited access to human foetal gonadal tissue, a paucity of large-scale proteomics studies, and a lack of robust of human gonad in vitro models. OBJECTIVE AND RATIONALE: This review aims to provide a comprehensive analysis of validated PTMs affecting proteins critical for testicular development. We discuss PTMs with evidence for a role in normal testis development, and highlight those disrupted in DSD. We review emerging techniques, including proteomic technologies and organ modelling systems that may advance our understanding of PTMs in foetal testis development. We discuss challenges that have restricted the application of these technologies and how overcoming these will significantly improve our understanding of testis development and disease, diagnostics and patient outcomes. SEARCH METHODS: We searched PubMed and the University of Melbourne library for peer-reviewed English-language studies using keywords such as phosphorylation, SUMOylation, acetylation, ubiquitination alongside each protein of interest. PTM sites in proteins involved in testis development were identified using the PhosphoSitePlus database focusing those confirmed in in vitro or animal model studies. ClinVar and the Human Gene Mutation Database were used to identify patient variants that may disrupt PTM sites. OUTCOMES: Our review finds that proteins required for human foetal testis development are subject to extensive PTM. Several PTM sites and PTM-mediated pathways [e.g. MAPK (mitogen-activated protein kinase) pathway] are disrupted in patients with DSD or related conditions. While recent advances in proteomics technologies hold considerable promise, their application to human foetal gonads has been constrained by technical, ethical, and logistical challenges. Encouragingly, emerging high-sensitivity and low-input technologies, alongside stem cell-based approaches, offer viable pathways to overcoming these barriers. WIDER IMPLICATIONS: The relationship between gene regulation, protein expression, and cellular outcome is inherently non-linear, shaped by additional regulatory layers-most notably PTMs. The contribution of PTMs to human testis development in both typical and atypical contexts is a major knowledge gap. Addressing this gap has broad clinical and biological relevance: it may help improve genetic diagnosis or shed light on how proteins or pathways critical for testis development respond to environmental signals-an increasingly pressing question as declining global fertility rates bring testicular function under greater scrutiny. REGISTRATION NUMBER: N/A.

Humans

Integrating genetic predictors into subsequent breast cancer risk prediction in survivors of childhood cancer.

PURPOSE: Female survivors of childhood cancer are at high risk for developing breast cancer. The contributions of most general population primary breast cancer genetic predictors to this risk have not been explored. METHODS: Analyses included females who survived &#x2265;5 years after their childhood cancer diagnosis with available array (N&#x2009;=&#x2009;2096, subsequent breast cancer [SBC]=218) or whole-genome sequencing (WGS; N&#x2009;=&#x2009;3292, SBC=101) data from the Childhood Cancer Survivor Study and St. Jude Lifetime Cohort. We computed 99 externally-validated primary breast cancer polygenic risk scores (PRS). Using deep-coverage WGS, ClinVar-annotated pathogenic/likely pathogenic (P/LP) variants in breast cancer susceptibility genes were identified. Cox proportional hazards models assessed associations with SBC risk, adjusting for treatments and genetic ancestry. RESULTS: Among 5388 female survivors (genetic ancestry, European: N&#x2009;=&#x2009;4,752; African: N&#x2009;=&#x2009;444; East Asian: N&#x2009;=&#x2009;192), 319 developed SBC. Most (90.9%) PRSs were nominally associated with SBC risk (P&#x2009;<&#x2009;0.05), but effect sizes varied substantially. PRSs with superior discriminatory ability had greater genome-wide coverage (e.g., 6.4 million-variant PRS, HR per SD&#x2009;=&#x2009;1.71, 95% CI&#x2009;=&#x2009;1.43 to 2.05; P&#x2009;=&#x2009;4.2x10-9) and 7.7-fold higher odds (P&#x2009;=&#x2009;7.0x10-4) of including variants in multiple DNA damage repair pathways compared with PRSs with weaker risk associations. Among survivors with WGS, 1.6% carried P/LP variants in clinical testing panel genes, which was associated with a 7.4-fold greater risk (95% CI&#x2009;=&#x2009;3.16 to 17.19). Including genetic factors improved SBC risk prediction by age 40 (P&#x2009;<&#x2009;0.001) compared to treatment exposures alone. CONCLUSIONS: Externally-validated primary breast cancer genetic susceptibility predictors are relevant for SBC risk prediction and should be prioritized for risk stratification in survivors.

Journal Article

The Data Distillery: A Graph Framework for Semantic Integration and Querying of Biomedical Data.

The Data Distillery Knowledge Graph (DDKG) is a framework for semantic integration and querying of biomedical data across domains. Built for the NIH Common Fund Data Ecosystem, it supports translational research by linking clinical and experimental datasets in a unified graph model. Clinical standards such as ICD-10, SNOMED, and DrugBank are integrated through UMLS, while genomics and basic science data are structured using ontologies and standards such as HPO, GENCODE, Ensembl, STRING, and ClinVar. The DDKG uses a property graph architecture based on the UBKG infrastructure and supports ontology-based ingestion, identifier normalization, and graph-native querying. The system is modular and can be extended with new datasets or schema modules. We demonstrate its utility for informatics queries across eight use cases, including regulatory variant analysis, tissue-specific expression, biomarker discovery, and cross-species variant prioritization. The DDKG is accessible via a public interface, a programmatic API, and downloadable builds for local use.

Journal Article

BRCA1 Exon 11 Mutations in Breast Cancer: A Study From Pakistan.

Breast cancer ranks among the top causes of cancer-related deaths in women around the globe, with genetic mutations in the BRCA1 gene being a frequent cause of breast or ovarian cancer. This study investigates hotspot mutations in exon 11 of the BRCA1 gene among Pakistani women diagnosed with breast cancer. Thirty clinically diagnosed breast cancer patients, all women, were enrolled in the current study, and high-quality DNA was extracted from peripheral blood samples. Two of the twenty-five successfully sequenced samples had a homozygous missense variant (c.2312T&#x2009;>&#x2009;C: p.Leu771Ser) detected by Sanger sequencing after PCR amplification. Upon investigation in the ClinVar database, the identified variant showed conflicting interpretations of pathogenicity. Demographic data highlighted an early disease onset, showing that 56% of patients were under 50&#x2009;years of age. The need for genetic screening was further supported by the fact that 24% of the patients had a positive family history of cancer. Our study emphasizes the necessity of screening BRCA1 gene mutations to better understand the pathogenic potential of the identified variants in the Pakistani population.

Humans

Human Variation-Informed Prioritization of MPHOSPH6 in Lung Adenocarcinoma: A Source-Aware Multiomics Evidence Framework.

Moving from an association signal to a clinically credible biomarker requires several links that are often conflated: verified variant identity, aligned allelic effects, reproducible gene-level association, relevant cellular expression, and a plausible functional consequence. We developed a source-aware multiomics framework to assess MPHOSPH6 in lung adenocarcinoma (LUAD) while keeping those evidence classes separate. Six prespecified rsIDs were recovered from the harmonized TRICL LUAD dataset, of which five reached p < 5 &#xd7; 10 - 8. Only rs112333466 and rs76474922 were available with alignable alleles in FinnGen R10, and both showed concordant directions. Fixed-effect estimates were OR = 1.592 for rs112333466-T (95% CI, 1.401-1.809; p = 9.91 &#xd7; 10 - 13) and OR = 0.819 for rs76474922-C (95% CI, 0.773-0.867; p = 1.03 &#xd7; 10 - 11). In a prespecified two-variant GTEx v8 lung model, genetically predicted MPHOSPH6 expression was positively associated with LUAD in TRICL (Z = 3.341, p = 8.35 &#xd7; 10 - 4) and FinnGen (Z = 2.697, p = 0.0070). This gene-level result did not establish colocalization or connect MPHOSPH6 to the six susceptibility rsIDs. Patient-level analysis of 89,241 immune cells from six paired tumor and normal-adjacent lung samples found no significant difference in MPHOSPH6 pseudobulk abundance (exact paired Wilcoxon p = 0.3125). None of 688 lung-lineage pharmacogenomic tests remained significant after false-discovery-rate correction. Ten recorded MPHOSPH6 missense alleles, including five ClinVar variants of uncertain significance, were curated; structural analysis identified I58 at an experimental RNA-exosome interface and defined a focused perturbation series. MPHOSPH6 is therefore supported as a human-variation-informed candidate for functional evaluation, not as a validated LUAD biomarker, pathogenic gene, drug-response predictor, or therapeutic target.

Humans

Frequency of variants in Mendelian Alzheimer's disease genes within the Alzheimer's Disease Sequencing Project.

BackgroundPrior studies examined variants within presenilin-2 (PSEN2), presenilin-1 (PSEN1), and amyloid precursor protein (APP) genes. However, previously-reported clinically-relevant variants and other predicted damaging missense (DM) variants have not been characterized in a newer release of the Alzheimer's Disease Sequencing Project (ADSP).ObjectiveTo characterize previously-reported clinically-relevant variants and DM variants in PSEN2, PSEN1, APP within the participants from the ADSP.MethodsWe identified rare variants (MAF&#x2009;<&#x2009;1%) in PSEN2, PSEN1, and APP in 14,641 individuals with whole genome sequencing and 16,849 individuals with whole exome sequencing available (Ntotal&#x2009;=&#x2009;31,490). We additionally curated variants from ClinVar, OMIM, and Alzforum and report carriers of variants in clinical databases as well as predicted DM variants in these genes.ResultsWe detected 31 previously-reported clinically-relevant variants with alternate alleles observed within the ADSP: 4 variants in PSEN2, 25 in PSEN1, and 2 in APP. The overall variant carrier rate for the 31 clinically-relevant variants in the ADSP was 0.3%. We observed that 79.5% of the variant carriers were cases compared to 3.9% were controls. In those with AD, the mean age of onset of AD among carriers of these clinically-relevant variants was 19.6&#x2009;&#xb1;&#x2009;1.4 years earlier compared with noncarriers (p&#x2009;=&#x2009;7.8&#x2009;&#xd7;&#x2009;10-57). Additionally, we identified 197 rare variants (MAF&#x2009;<&#x2009;1%) within ADSP participants not reported in known clinical databases.ConclusionsA small proportion of individuals in the ADSP are carriers of a previously-reported clinically-relevant variant allele for AD and these participants have significantly earlier age of AD onset compared to noncarriers.

Humans

Elucidating the roles of SOD3 correlated genes and reactive oxygen species in rare human diseases using a bioinformatic-ontology approach.

Superoxide Dismutase 3 (SOD3) scavenges extracellular superoxide giving a hydrogen peroxide metabolite. Both Reactive Oxygen Species diffuse through aquaporins causing oxidative stress and biomolecular damage. SOD3 is differentially expressed in cancer and this research utilises Gene Expression Omnibus data series GSE2109 with 2,158 cancer samples. Genome-wide expression correlation analysis was conducted with SOD3 as the seed gene. Categorical SOD3 Pearson Correlation gene lists incrementing in correlation strength by 0.01 from &#x3c1;&#x2265;|0.34| to &#x3c1;&#x2265;|0.41| were extracted from the data. Positively and negatively SOD3 correlated genes were separated for each list and checked for significance against disease overlapping genes in the ClinVar and Orphanet databases via Enrichr. Disease causal genes were added to the relevant gene list and checked against Gene Ontology, Phenotype Ontology, and Elsevier Pathways via Enrichr before the significant ontologies containing causal and non-overlapping genes were reviewed with a literature search for possible disease and oxidative stress associations. 12 significant individually discriminated disorders were identified: Autosomal Dominant Cutis Laxa (p = 6.05x10-7), Renal Tubular Dysgenesis of Genetic Origin (p = 6.05x10-7), Lethal Arteriopathy Syndrome due to Fibulin-4 Deficiency (p = 6.54x10-9), EMILIN-1-related Connective Tissue Disease (p = 6.54x10-9), Holt-Oram Syndrome (p = 7.72x10-10), Multisystemic Smooth Muscle Dysfunction Syndrome (p = 9.95x10-15), Distal Hereditary Motor Neuropathy type 2 (p = 4.48x10-7), Congenital Glaucoma (p = 5.24x210-9), Megacystis-Microcolon-Intestinal Hypoperistalsis Syndrome (p = 3.77x10-16), Classical-like Ehlers-Danlos Syndrome type 1 (p = 3.77x10-16), Retinoblastoma (p = 1.9x10-8), and Lynch Syndrome (p = 5.04x10-9). 35 novel (21 unique) genes across 12 disorders were identified: ADNP, AOC3, CDC42EP2, CHTOP, CNN1, DES, FOXF1, FXR1, HLTF, KCNMB1, MTF2, MYH11, PLN, PNPLA2, REST, SGCA, SORBS1, SYNPO2, TAGLN, WAPL, and ZMYM4. These genes are proffered as potential biomarkers or therapeutic targets for the corresponding rare diseases discussed.

Humans

Unique signatures of highly constrained genes across publicly available genomic databases.

PURPOSE: Publicly available genomic databases are critical in understanding human genetic variation. They also provide unique insights into patterns of genetic constraints and their relationship with human disease. METHODS: We utilized one of the largest publicly available databases, Genome Aggregate Database, to determine genes that are highly constrained for only loss-of-function, only missense, and both loss-of-function/missense variants. We identified their unique signatures and explored their causal relationship with human diseases. Those genes were also evaluated for chromosomal location, tissue-level expression, Gene Ontology analysis, and gene family categorization using multiple publicly available databases. RESULTS: We identified unique patterns of inheritance, protein size, and enrichment in distinct molecular pathways for those constrained genes associated with human disease. In addition, we identified genes that are currently not known to cause human disease, which may be excellent gene discovery candidates. CONCLUSION: We elucidate biological pathways of highly constrained genes that expand our understanding of critical cellular proteins. The findings can also advance research in rare diseases.

Humans