Search PubMedSearch

Biomedical subjects

Soraya Bardien

Publications and source records attributed to Soraya Bardien.

3 recordsLinked to original sources

CNV-Finder: Streamlining Copy Number Variation Discovery.

Copy Number Variations (CNVs) play pivotal roles in the etiology of complex diseases and are variable across diverse populations. Understanding the association between CNVs and disease susceptibility is significant in disease genetics research and often requires analysis of large sample sizes. One of the most cost-effective and scalable methods for detecting CNVs is based on normalized signal intensity values, such as Log R Ratio (LRR) and B Allele Frequency (BAF), from Illumina genotyping arrays. In this study, we present CNV-Finder, a novel pipeline integrating deep learning techniques on array data, specifically a Long Short-Term Memory (LSTM) network, to expedite the large-scale identification of CNVs within predefined genomic regions. This facilitates efficient prioritization of samples for time-consuming or costly subsequent analyses such as Multiplex Ligation-dependent Probe Amplification (MLPA), short-read, and long-read whole genome sequencing. We incorporate four genes to establish our methods-Parkin (PRKN), Leucine Rich Repeat And Ig Domain Containing 2 (LINGO2), Microtubule Associated Protein Tau (MAPT), and alpha-Synuclein (SNCA)-which may be relevant to neurological diseases such as Alzheimer's disease (AD), Parkinson's disease (PD), Progressive Supranuclear Palsy (PSP), or related disorders such as essential tremor (ET). By training our models on expert-annotated samples and validating them across diverse cohorts, including those from the Global Parkinson's Genetics Program (GP2) and additional dementia-specific databases, we demonstrate the efficacy of CNV-Finder in accurately detecting deletions and duplications. Our pipeline outputs app-compatible files for visualization within CNV-Finder's interactive web application. This interface enables researchers to review predictions and filter displayed samples by model prediction values, LRR range, and variant count in order to explore or confirm results. Our pipeline integrates this human feedback to enhance model performance and reduce false positive rates. Through a series of comprehensive analyses and validations using visual inspection, MLPA, short-read, and long-read sequencing data, we demonstrate the robustness and adaptability of CNV-Finder in identifying CNVs with regions of varied size, probe density, and noise. Our findings highlight the significance of contextual understanding and human expertise in enhancing the precision of CNV identification, particularly in complex genomic regions like 17q21.31. The CNV-Finder pipeline is a scalable, publicly available resource for the scientific community, available on GitHub (https://github.com/GP2code/CNV-Finder; DOI 10.5281/zenodo.14182563). CNV-Finder not only expedites accurate candidate identification but also significantly reduces the manual workload for researchers, enabling future targeted validation and downstream analyses in regions or phenotypes of interest.

Copy Number Variation (CNV)

Genome-wide association study of copy number variations in Parkinson's disease.

OBJECTIVE: To investigate the impact of copy number variations (CNVs) on Parkinson's disease (PD) pathogenesis using genome-wide data and explore their role in sporadic PD. METHODS: We analyzed CNV data from 11,035 PD patients (including 2,731 early-onset PD (EOPD)) and 8,901 controls from the COURAGE-PD consortium using a sliding window CNV-GWAS and genome-wide burden analysis. The independent dataset from the Global Parkinson Genetics Program (GP2) consisted of 23,089 cases and 18,824 controls were used to validate our initial findings. RESULTS: The exploratory dataset identifies multiple CNV regions associated with PD risk. The nominated CNV loci were not confirmed in an independent dataset, except that only a deletion in the PRKN gene, a well-established EOPD locus, remained genome-wide significant and robustly supported. CNV burden analysis showed a higher prevalence of CNVs in PD-related genes in patients compared to controls (OR=1.56 [1.18-2.09], p=0.0013), with PRKN showing the highest burden (OR=1.47 [1.10-1.98], p=0.026). Patients with CNVs in PRKN had an earlier disease onset. Burden analysis with controls and EOPD patients showed similar results. INTERPRETATION: The largest CNV-based GWAS on PD highlights both the promise and pitfalls of array-based CNV detection in PD and underscores the relevance of whole-genome sequencing approaches in resolving the role of CNV in PD. The array-based findings are prone towards false positive findings that might arise either from platform limitations and/or cohort biases. Future studies require improved genotyping resolution and rigorous cross-cohort validation to reliably assess CNV contributions to PD risk.

Journal Article

eVOC: a controlled vocabulary for unifying gene expression data.

Expression data contribute significantly to the biological value of the sequenced human genome, providing extensive information about gene structure and the pattern of gene expression. ESTs, together with SAGE libraries and microarray experiment information, provide a broad and rich view of the transcriptome. However, it is difficult to perform large-scale expression mining of the data generated by these diverse experimental approaches. Not only is the data stored in disparate locations, but there is frequent ambiguity in the meaning of terms used to describe the source of the material used in the experiment. Untangling semantic differences between the data provided by different resources is therefore largely reliant on the domain knowledge of a human expert. We present here eVOC, a system which associates labelled target cDNAs for microarray experiments, or cDNA libraries and their associated transcripts with controlled terms in a set of hierarchical vocabularies. eVOC consists of four orthogonal controlled vocabularies suitable for describing the domains of human gene expression data including Anatomical System, Cell Type, Pathology and Developmental Stage. We have curated and annotated 7016 cDNA libraries represented in dbEST, as well as 104 SAGE libraries,with expression information,and provide this as an integrated, public resource that allows the linking of transcripts and libraries with expression terms. Both the vocabularies and the vocabulary-annotated libraries can be retrieved from http://www.sanbi.ac.za/evoc/. Several groups are involved in developing this resource with the aim of unifying transcript expression information.

Animals