Search PubMedSearch

Biomedical subjects

Mary L Marazita

Publications and source records attributed to Mary L Marazita.

5 recordsLinked to original sources

Trio-based GWAS reveals loci associated with different forms of isolated cleft lip.

Orofacial clefts (OFCs) are the most common craniofacial birth defect and comprise a diverse group of traits with complex and heterogeneous etiologies. Genetic studies of OFCs typically approach this diversity by stratifying cases into broad diagnostic classes, including cleft lip (CL), cleft palate (CP), and cleft lip with palate (CLP). Although this strategy has yielded important insights into OFC risk, it ignores the phenotypic heterogeneity within each subtype. CL exhibits marked phenotypic variability, involving differences in alveolar involvement, laterality, and sidedness that may reflect distinct etiologies. Given this phenotypic diversity within CL, we assembled a multi-ancestry cohort of 837 nonsyndromic CL case-parent trios with whole-genome sequencing and detailed phenotyping. We performed genome-wide association scans (GWAS) via transmission disequilibrium tests for CL overall and for 14 CL subtypes defined by involvement of the alveolus (with and without), laterality (uni- and bilateral), and sidedness (left and right). We identified four genome-wide significant loci. Two loci, IRF6 and 8q24.21, were both detected in the overall CL GWAS. PLCB1/PLCB4 and MAFB were detected in GWASs of alveolar cleft involvement and CL left sidedness, respectively. These subtype-specific associations were followed by case-only comparisons that reflect the presence or absence of alveolus cleft or left-sided bias of CL to confirm the specificity of the association signal to the particular subtype. Our results provide evidence of within-class CL subtype-specific genetic links for loci previously discussed in the context of primary OFC classes and demonstrate the value of granular OFC subtype characterization to capture trait-specific associations.

Alveolus Cleft

Comprehensive analysis of de novo variants across 2,497 orofacial cleft trios reveals novel genetic drivers of disease.

BACKGROUND: Orofacial clefts (OFCs) and other palate abnormalities (PAs) are among the most common birth defects worldwide and are characterized by the abnormal formation of the lip and/or palate. Genetic studies have traditionally classified OFC cases as either syndromic, involving OFCs alongside other congenital anomalies, or nonsyndromic, which represent the majority of cases and occur in isolation. Emerging genomic evidence indicates that genes traditionally associated with syndromic forms of OFC can also harbor variants contributing to isolated cases, challenging the notion of a strict dichotomy between these categories and supporting their integration for gene discovery. METHODS: In this study, we applied multiple analytic approaches to characterize the genetic architecture of OFC and PAs by integrating genomic data from 2,497 trios with probands diagnosed with an OFC (n=2,080) or PA (n=417). We compared these findings across OFC subtypes and syndromic status with those from 5,515 control trios to identify enriched biological pathways and mechanisms and to prioritize candidate genes using variant burden testing. RESULTS: We observed a significant enrichment of de novo protein-truncating and damaging missense variants in cases compared to controls (OR = 2.17, p = 1.21×10-32), with particularly strong signals in biologically relevant gene sets involving OFC-associated, constrained, Mendelian disorder, and mouse candidate genes. Variant burden testing identified 39 OFC risk genes at FDR ≤ 0.05, which we then integrated with 593 established OFC genes to interrogate the functional underpinnings of OFC via network analysis. This analysis revealed 309 high-order interactor genes not previously associated with OFC. Notably, this OFC network clustered into ten distinct biological pathways, with nucleosome-associated genes showing significant enrichment among cases in our cohort (OR = 14.8, p = 8.1×10-4). In a final integrative step, we combined evidence across all analyses to nominate 231 candidate genes, 32 of which contained at least two deleterious de novo variants in our cohort. CONCLUSIONS: These findings underscore the value of integrating diverse OFC and PA subtypes, syndromic status, and variant classes to elucidate the genetic architecture of these disorders, highlighting both phenotypic expansion of known disease genes and the emergence of novel gene-phenotype associations.

De Novo Variant Enrichment

Fundamentals of FAIR biomedical data analyses in the cloud using custom pipelines.

As the biomedical data ecosystem increasingly embraces the findable, accessible, interoperable, and reusable (FAIR) data principles to publish multimodal datasets to the cloud, opportunities for cloud-based research continue to expand. Besides the potential for accelerated and diverse biomedical discovery that comes from a harmonized data ecosystem, the cloud also presents a shift away from the standard practice of duplicating data to computational clusters or local computers for analysis. However, despite these benefits, researcher migration to the cloud has lagged, in part due to insufficient educational resources to train biomedical scientists on cloud infrastructure. There exists a conceptual lack especially around the crafting of custom analytic pipelines that require software not pre-installed by cloud analysis platforms. We here present three fundamental concepts necessary for custom pipeline creation in the cloud. These overarching concepts are workflow and cloud provider agnostic, extending the utility of this education to serve as a foundation for any computational analysis running any dataset in any biomedical cloud platform. We illustrate these concepts using one of our own custom analyses, a study using the case-parent trio design to detect sex-specific genetic effects on orofacial cleft (OFC) risk, which we crafted in the biomedical cloud analysis platform CAVATICA.

Cloud Computing

Clustering individuals using INMTD: a novel versatile multi-view embedding framework integrating omics and imaging data.

MOTIVATION: Combining omics and images can lead to a more comprehensive clustering of individuals than classic single-view approaches. Among the various approaches for multi-view clustering, nonnegative matrix tri-factorization (NMTF) and nonnegative Tucker decomposition (NTD) are advantageous in learning low-rank embeddings with promising interpretability. Besides, there is a need to handle unwanted drivers of clusterings (i.e. confounders). RESULTS: In this work, we introduce a novel multi-view clustering method based on NMTF and NTD, named INMTD, which integrates omics and 3D imaging data to derive unconfounded subgroups of individuals. According to the adjusted Rand index, INMTD outperformed other clustering methods on a synthetic dataset with known clusters. In the application to real-life facial-genomic data, INMTD generated biologically relevant embeddings for individuals, genetics, and facial morphology. By removing confounded embedding vectors, we derived an unconfounded clustering with better internal and external quality; the genetic and facial annotations of each derived subgroup highlighted distinctive characteristics. In conclusion, INMTD can effectively integrate omics data and 3D images for unconfounded clustering with biologically meaningful interpretation. AVAILABILITY AND IMPLEMENTATION: INMTD is freely available at https://github.com/ZuqiLi/INMTD.

Cluster Analysis

Optimized phenotyping of complex morphological traits: enhancing discovery of common and rare genetic variants.

Genotype-phenotype (G-P) analyses for complex morphological traits typically utilize simple, predetermined anatomical measures or features derived via unsupervised dimension reduction techniques (e.g. principal component analysis (PCA) or eigen-shapes). Despite the popularity of these approaches, they do not necessarily reveal axes of phenotypic variation that are genetically relevant. Therefore, we introduce a framework to optimize phenotyping for G-P analyses, such as genome-wide association studies (GWAS) of common variants or rare variant association studies (RVAS) of rare variants. Our strategy is two-fold: (i) we construct a multidimensional feature space spanning a wide range of phenotypic variation, and (ii) within this feature space, we use an optimization algorithm to search for directions or feature combinations that are genetically enriched. To test our approach, we examine human facial shape in the context of GWAS and RVAS. In GWAS, we optimize for phenotypes exhibiting high heritability, estimated from either family data or genomic relatedness measured in unrelated individuals. In RVAS, we optimize for the skewness of phenotype distributions, aiming to detect commingled distributions that suggest single or few genomic loci with major effects. We compare our approach with eigen-shapes as baseline in GWAS involving 8246 individuals of European ancestry and in gene-based tests of rare variants with a subset of 1906 individuals. After applying linkage disequilibrium score regression to our GWAS results, heritability-enriched phenotypes yielded the highest SNP heritability, followed by eigen-shapes, while commingling-based traits displayed the lowest SNP heritability. Heritability-enriched phenotypes also exhibited higher discovery rates, identifying the same number of independent genomic loci as eigen-shapes with a smaller effective number of traits. For RVAS, commingling-based traits resulted in more genes passing the exome-wide significance threshold than eigen-shapes, while heritability-enriched phenotypes lead to only a few associations. Overall, our results demonstrate that optimized phenotyping allows for the extraction of genetically relevant traits that can specifically enhance discovery efforts of common and rare variants, as evidenced by their increased power in facial GWAS and RVAS.

Humans