Search PubMedSearch

SEARCH · Search PubMed

Results for “copy number complexity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

CNV-Finder: Streamlining Copy Number Variation Discovery.

Copy Number Variations (CNVs) play pivotal roles in the etiology of complex diseases and are variable across diverse populations. Understanding the association between CNVs and disease susceptibility is significant in disease genetics research and often requires analysis of large sample sizes. One of the most cost-effective and scalable methods for detecting CNVs is based on normalized signal intensity values, such as Log R Ratio (LRR) and B Allele Frequency (BAF), from Illumina genotyping arrays. In this study, we present CNV-Finder, a novel pipeline integrating deep learning techniques on array data, specifically a Long Short-Term Memory (LSTM) network, to expedite the large-scale identification of CNVs within predefined genomic regions. This facilitates efficient prioritization of samples for time-consuming or costly subsequent analyses such as Multiplex Ligation-dependent Probe Amplification (MLPA), short-read, and long-read whole genome sequencing. We incorporate four genes to establish our methods-Parkin (PRKN), Leucine Rich Repeat And Ig Domain Containing 2 (LINGO2), Microtubule Associated Protein Tau (MAPT), and alpha-Synuclein (SNCA)-which may be relevant to neurological diseases such as Alzheimer's disease (AD), Parkinson's disease (PD), Progressive Supranuclear Palsy (PSP), or related disorders such as essential tremor (ET). By training our models on expert-annotated samples and validating them across diverse cohorts, including those from the Global Parkinson's Genetics Program (GP2) and additional dementia-specific databases, we demonstrate the efficacy of CNV-Finder in accurately detecting deletions and duplications. Our pipeline outputs app-compatible files for visualization within CNV-Finder's interactive web application. This interface enables researchers to review predictions and filter displayed samples by model prediction values, LRR range, and variant count in order to explore or confirm results. Our pipeline integrates this human feedback to enhance model performance and reduce false positive rates. Through a series of comprehensive analyses and validations using visual inspection, MLPA, short-read, and long-read sequencing data, we demonstrate the robustness and adaptability of CNV-Finder in identifying CNVs with regions of varied size, probe density, and noise. Our findings highlight the significance of contextual understanding and human expertise in enhancing the precision of CNV identification, particularly in complex genomic regions like 17q21.31. The CNV-Finder pipeline is a scalable, publicly available resource for the scientific community, available on GitHub (https://github.com/GP2code/CNV-Finder; DOI 10.5281/zenodo.14182563). CNV-Finder not only expedites accurate candidate identification but also significantly reduces the manual workload for researchers, enabling future targeted validation and downstream analyses in regions or phenotypes of interest.

Copy Number Variation (CNV)

PARTAGE: Parallel analysis of replication timing and gene expression.

The human genome is partitioned into functional compartments that replicate at specific times during the S-phase. This temporal program, referred to as replication timing (RT), is co-regulated with the 3D genome organization, is cell type-specific, and changes during development in coordination with gene expression. Moreover, RT alterations are linked to abnormal gene expression, genome instability, and structural variation in multiple diseases, including cancer. However, mechanistic links between RT, large-scale 3D genome architecture, and transcriptional regulation remain poorly understood. A major limitation is that current approaches require the separate profiling of RT and transcriptomes from independent batches of samples, obscuring the complex co-regulation between the epigenome and transcriptome. Here, we developed PARTAGE, a multiomics approach that enables joint profiling of copy number variation (CNV), RT, and gene expression from the same sample, providing a more accurate integrative view of the complex relationships between RT and gene regulation.

Journal Article

Comprehensive molecular profiling of multiple myeloma identifies refined copy number and expression subtypes.

Multiple myeloma is a treatable, but currently incurable, hematological malignancy of plasma cells characterized by diverse and complex tumor genetics for which precision medicine approaches to treatment are lacking. The Multiple Myeloma Research Foundation's Relating Clinical Outcomes in Multiple Myeloma to Personal Assessment of Genetic Profile study ( NCT01454297 ) is a longitudinal, observational clinical study of newly diagnosed patients with multiple myeloma (n = 1,143) where tumor samples are characterized using whole-genome sequencing, whole-exome sequencing and RNA sequencing at diagnosis and progression, and clinical data are collected every 3 months. Analyses of the baseline cohort identified genes that are the target of recurrent gain-of-function and loss-of-function events. Consensus clustering identified 8 and 12 unique copy number and expression subtypes of myeloma, respectively, identifying high-risk genetic subtypes and elucidating many of the molecular underpinnings of these unique biological groups. Analysis of serial samples showed that 25.5% of patients transition to a high-risk expression subtype at progression. We observed robust expression of immunotherapy targets in this subtype, suggesting a potential therapeutic option.

Humans

Optical genome mapping improves clinical interpretation of constitutional copy-number gains and reduces their VUS burden.

PURPOSE: Genomic structure of copy-number gains is critical for their clinical interpretation but cannot be determined by chromosomal microarray (CMA) analysis, which does not provide information about chromosomal location and orientation of multiplied regions. We thus hypothesized that in CMA testing gains have higher probability than losses to be classified as variants of uncertain significance (VUS) and that structural information from optical genome mapping (OGM) may improve their interpretation. METHODS: Using a χ2 test, we assessed the association between classification of copy-number variants as VUS and their type (gains vs losses) in a cohort of 4073 CMA cases. Thirty-three VUS gains involving disease-associated genes were characterized by OGM to evaluate if OGM data enable their more conclusive clinical interpretation. RESULTS: The proportion of variants reported as VUS compared with likely pathogenic/pathogenic was significantly higher for gains than losses, confirming their increased VUS burden. OGM successfully determined genomic structure for all 33 copy-number gains, showing that 26 of 33 were tandem duplications and 7 of 33 were complex rearrangements. Structural information facilitated clinical interpretation in majority of the cases; it supported benign nature for 27 of 33 gains and was inconclusive or supported pathogenic role for 6 of 33. An estimated 20% of reported VUS gains would not have been reportable if we had OGM data. CONCLUSION: We illustrate a specific advantage of OGM compared with CMA: in addition to detecting both copy-number variants and balanced rearrangements, OGM improves clinical interpretation of copy-number gains by providing structural information and is thus expected to significantly decrease their VUS burden.

Humans

Long-read based detection of large copy number variants with potential functional significance using the ContextSV structural variant caller.

Long-read sequencing enables improved detection of structural variants (SVs) in the human genome due to its substantially increased read lengths. However, currently widely used long-read SV callers primarily rely on alignment-based evidence, limiting their ability to detect large and complex SVs and potentially missing disease-relevant events. To address these limitations, we developed ContextSV, a framework that integrates alignment evidence with copy number predictions derived from sequencing coverage and single-nucleotide variant allele frequencies to improve SV detection, particularly for large copy number variants (CNVs). We additionally developed ContextScore, a machine learning-based classification model to assign SV confidence scores based on genomic context features and integrated it within ContextSV. Through benchmarking analyses on both simulated and real datasets, we demonstrate that ContextSV improves detection of large CNVs and inversions that may be missed by existing long-read SV callers. We further illustrate its utility by identifying and experimentally validating multiple large SVs in the KOLF2.1J reference stem cell line that were not detected by other methods. Collectively, our results demonstrate that ContextSV serves as a valuable complement to existing long-read SV detection approaches by improving sensitivity for large and clinically relevant SVs.

Humans

Using cancer profiles to identify synthetic lethal therapeutic targets and predictive biomarkers in cancer gene dependency data.

MOTIVATION: Large scale loss-of-function screens utilising CRISPR or siRNA can provide profound insights into the importance of individual genes for the survival of a cancer cell and can drive the identification of therapeutic targets and biomarkers, and the development of targeted drugs. However, the analysis of these data and the substantial bodies of metadata that relate to them, is technically challenging and typically requires substantial expertise in data science and computer coding. RESULTS: To facilitate the analysis of cancer gene dependency data by cancer biologists and clinical scientists, we have developed DepMine-a computational toolkit providing a powerful system for framing complex queries relating cancer gene dependency to the underlying genetic changes that occur in cancer cells. DepMine identifies synthetic lethal relationships between putative target genes and complex 'cancer profiles' built from user-specified combinations of mutations, copy-number variation, and expression levels, and can refine these to optimal biomarker definitions for target dependency. AVAILABILITY: The Python implementation of DepMine and associated data files can be obtained at https://github.com/UOSbioinformaticslab/depmine and is free to academics and Not-For-Profit organisations. The DepMine release referenced in this paper is archived as DOI: 10.5281/zenodo.19570601.

Humans

Analysis of chromosomes, nucleic acids, and polypeptides in hamster cells transformed by herpes simplex virus type 2.

Syrian hamster embryo fibroblasts were oncogenically transformed by UV-inactivated Herpes simplex type 2. Eighteen clones were isolated shortly after transformation occurred. Two clones and their tumor derivatives were studied using several techniques. The karyotype analysis revealed different chromosome patterns in the two clones and a tendency toward hypodiploidy in the tumor derivatives. All of these cell lines were shown by molecular hybridization to contain 40% of the HSV-genome in several copies. The viral DNA sequence complexity was retained in the tumor derivatives, but a decrease in the copy number was observed. Viral RNA's were detected by in situ hybridization in all the lines that were tested. Viral antigens could be observed in these transformed cells by immunofluorescence. Finally, polypeptide analysis showed three differences between normal and transformed cells.

Animals

Evolutionary Diversification and Functions of the Candidate Male Killing Gene wmk.

Symbiont-mediated male killing (MK) is a mechanism that selectively eliminates male offspring, often by disrupting sex-specific developmental processes. In Drosophila melanogaster, the WO-mediated killing gene wmk from Wolbachia prophage WO transgenically reproduces the MK phenotype, yet how the gene evolves and functions across diverse Wolbachia has not been systematically investigated. We analyzed 32 Wolbachia genomes available in the NCBI database to study wmk homologs across different arthropod hosts, reproductive parasitism functions, and Wolbachia supergroups. First, we report at least five distinct wmk phylogenetic clusters (Types I to V), often organized in multigenic dyads or triads. Second, among MK Wolbachia, there is a significantly higher number of wmk genes and diversity in Lepidoptera strains than in Drosophila strains, which exclusively harbor wmk Types I and III. Third, there are three patterns of wmk sequence and genomic organizational changes in Drosophila MK strains that associate with different evolutionary trajectories underpinning the MK phenotype. Fourth, single and combinatory transgenic expression of Types I and III in D. melanogaster uncovers male-biased lethality associated with Type I; however, dual expression of the Types together elicits a major reduction in offspring number. Fifth, wmk genes have low expression level across D. melanogaster developmental stages relative to the cifA and cifB genes, which could explain why cytoplasmic incompatibility is expressed in this system. These findings establish a complex and phylogenetically informed genetic basis of wmk-induced lethality, highlighting the role of gene copy number and expression, wmk Types, and host background in shaping the phenotype.

Animals

Optical genome mapping enhanced by refined variant interpretation in pediatric acute lymphoblastic leukemia.

Reliable detection of structural variants (SVs) and copy number variations (CNVs) is crucial in the contemporary diagnostics of pediatric B-cell acute lymphoblastic leukemia (B-ALL). However, limitations of commonly used conventional and molecular cytogenetic methods may hinder the accurate genetic characterization of patients. Optical genome mapping (OGM) offers a reliable alternative by enabling high-resolution, genome-wide detection of CNVs and SVs. Chromosomal aberrations were screened using OGM in 51 children with B-ALL. The results were compared with those of karyotyping, fluorescence in situ hybridization (FISH), digital multiplex ligation-dependent probe amplification (digitalMLPA), and targeted RNA sequencing (RNA-seq). OGM data showed high congruency with karyotyping and FISH findings, detecting clinically relevant variants beyond G-banding results and unraveling a complex KMT2A fusion undetected by FISH. Gene fusions involved in complex ETV6::RUNX1 translocations, but not detected by RNA-seq, were confirmed using FISH. Normalization of OGM copy number values with DNA-index-improved concordance with FISH-derived copy numbers in near-tri/tetraploid cases. In the peripheral regions of OGM variants (fringe-zones), a novel evaluation strategy called 'FriZone' was applied, which significantly improved the concordance between OGM and digitalMLPA. In addition, a co-segregation analysis revealed strong associations between ETV6::RUNX1 fusion and deletions of ETV6, RAG2, and NR3C2. OGM uncovered complex rearrangements undetected by widely used methods in 15% of cases, improving genetic classification and risk stratification in 10% of the patients. The FriZone analysis and normalization by DNA-index provide a refined, more accurate approach to OGM variant interpretation, facilitating the efficient application of OGM in clinical diagnostics. © 2026 The Author(s). The Journal of Pathology published by John Wiley & Sons Ltd on behalf of The Pathological Society of Great Britain and Ireland.

Humans

MarkerMatch: a proximity-based probe-matching algorithm for joint analysis of copy-number variants from different genotyping arrays.

MOTIVATION: Copy-number variants (CNVs) are a form of genetic structural variation with increasing importance in complex human disorders. Both DNA sequencing and microarray data can be used to detect CNVs, which can be used in genetic association tests. Unlike genotypes, CNV detection in microarrays requires the use of observed intensity signals at each probe, which limits the imputability for analyses that span multiple array types. Thus far, a consensus set of probes (those present on all arrays) has been used to circumvent the problem of differing array-specific sensitivities. This has led to excessive reduction in overall sensitivity since arrays can have an undesirably low probe overlap. To overcome this limitation, we developed MarkerMatch, a proximity-based algorithm that matches probes across different genotyping microarrays to maximize the number of probes considered in the CNV calling algorithm, thereby increasing the resolution and sensitivity while preserving precision. RESULTS: By analyzing CNV calls from 4906 individuals genotyped across three different arrays, we show that the MarkerMatch approach improves sensitivity by increasing the density of probes available for CNV calling while maintaining precision or improving it relative to the current practice (e.g. use of consensus probes only). We further demonstrate that MarkerMatch matches the CNV detection from current practice in terms of F1 score and PPV for larger CNVs. We also optimize MarkerMatch parameters, DMAX and Method, and find an optimal DMAX setting at 10 kb, with no clear optimal candidate based on Method, indicating that parameters for this metric should be determined on a use case basis. AVAILABILITY: The R package for MarkerMatch is available at: https://github.com/FranjoIM/MarkerMatch. The code used for analysis and implementation is available at: https://doi.org/10.5281/zenodo.18460979. The live notebook is available at https://fivankovic.notion.site/2026-markermatch.

DNA Copy Number Variations

Spatial Total RNA Sequencing of Formalin-Fixed Paraffin-Embedded Tissue by spRandom-seq.

The molecular pathogenesis of infectious diseases and cancer is orchestrated by nanoscale of host and microbial RNA transcripts within the tissue microenvironment. Nevertheless, spatially resolving the comprehensive transcriptional landscape within complex clinical tissues, like formalin-fixed paraffin-embedded (FFPE) specimens, still poses a formidable challenge. Here, we present spRandom-seq, a random primer-based spatial total RNA sequencing technology designed to spatially resolve complete transcriptomes from host, bacteria, and even nanoscale viruses in FFPE tissues. Capitalizing on the random primer design, our technology not only facilitated the discovery of specific lncRNAs and alternative splicing events in mouse brain and olfactory bulb, but also delineated pronounced spatial heterogeneity in clinical FFPE sections-across distinct tumor regions in breast cancer and microbial infection sites in Klebsiella pneumoniae-infected tissues. Importantly, integrated analysis of host and viral RNAs in FFPE samples from hepatitis B virus (HBV)‑positive hepatocellular carcinoma (HCC) demonstrated that complement and coagulation pathways were specifically activated across expansive HBV‑infected tumor areas, which also exhibited an increased burden of copy number variations (CNVs). Owing to its compatibility with existing spatial transcriptomics platforms and minimal operational complexity, spRandom-seq represents a practical and scalable approach for clinical pathology applications and infection diagnostics.

Paraffin Embedding

MarkerMatch: A Proximity-Based Probe-Matching Algorithm for Joint Analysis of Copy-Number Variants from Different Genotyping Arrays.

MOTIVATION: Copy-number variants (CNVs) are a form of genetic structural variation with increasing importance in complex human disorders. Both DNA sequencing and microarray data can be used to call CNVs, which can be used in association tests, such as association between CNV number and disease status. Unlike genotypes, CNV detection in microarrays requires the use of observed intensity signals at each probe, which limits the imputability for analyses that span multiple array types. Thus far, a consensus set of probes (the intersection encompassing the probes that occur in common on all arrays) has been used to circumvent the problem of differing array-specific sensitivities. This has, however, led to excessive reduction in overall sensitivity of CNV calls as arrays can have an undesirably low overlap of probe sets. To overcome this limitation, we developed MarkerMatch, a proximity-based algorithm that matches probes across different genotyping microarrays to maximize the number of probes considered in the CNV calling algorithm, thereby increasing the resolution and sensitivity while preserving precision. RESULTS: By analyzing CNV calls from 4,906 individuals genotyped across three different arrays (Global Screening Array, Omni2.5 array, and Omni Express Exome array), we show that the MarkerMatch approach improves sensitivity by increasing the density of probes available for CNV calling while maintaining precision or improving it relative to the current practice (e.g., use of consensus probes only). We further demonstrate that MarkerMatch exceeds the output from current practice in terms of F1 score, Fowlkes-Mallows index, and Jaccard index. We also optimize MarkerMatch parameters, D MAX and Method, and find an optimal D MAX setting at 10kb, with no clear optimal candidate based on Method, indicating that parameters for this metric should be determined on a use case basis.

Journal Article

Pleomorphic Liposarcoma: Comprehensive Genomic Analysis of 39 Cases With Comparison to Other Genomically Complex Sarcomas.

Pleomorphic liposarcoma (PLPS) is an aggressive high-grade sarcoma that often shows diverse morphological features and can mimic high-grade undifferentiated pleomorphic sarcoma (UPS)/spindle cell sarcoma or myxofibrosarcoma (MFS), especially when pleomorphic lipoblasts are sparse. The molecular profile of PLPS is distinct from well differentiated/dedifferentiated liposarcoma and myxoid liposarcoma. In this study, we investigate 39 cases of PLPS by comprehensive genomic profiling, occurring in 32 patients with available molecular data. Cases were reviewed and morphologic parameters-lipoblastic component, UPS-like, and MFS-like areas were estimated. The genomic findings were collected and compared to UPS and MFS groups studied using the same platform. The cohort included 15 females and 17 males, with a median age of 56.5 (range, 34-78). The lower extremity (n = 17) was the most common site involved, followed by upper extremity (n = 5) and pelvis (n = 5). UPS-like and MFS-like patterns were the most common morphologic variants, ranging from 15% to 95% and 20% to 90%, respectively. TP53 (87%) and RB1 (51%) mutations and copy number alterations were the most common alterations seen, followed by ATRX (36%). Compared to UPS and MFS, TP53 and RB1 gene alterations were significantly more common in PLPS. Conversely, CDKN2A/B deletions were infrequent in PLPS. Survival analysis showed that MYC amplification was associated with significantly shorter overall survival in PLPS. Among histologic variants, CYSLTR2 alterations were found to be highest in cases with predominantly pleomorphic lipoblasts; additionally, strong correlations were found between gene alteration frequencies of MFS and MFS-like PLPS, and between UPS and UPS-like PLPS. RB1 allele-specific copy number analysis showed loss of heterozygosity in 82% of cases. Our cohort of PLPS showed a complex molecular landscape with distinct genetic alterations, histologic correlations, and clinical outcomes, highlighting its unique position among genomically complex sarcomas and providing insights that may inform future diagnostic and therapeutic approaches.

Humans

The MTORC1 signaling pathway related gene POLR3G serves as a potential prognostic biomarker in Hepatocellular Carcinoma.

This study aims to investigate the prognostic significance and potential biological functions of the MTORC1 signaling pathway-associated gene POLR3G in Hepatocellular carcinoma (HCC). A prognostic risk model for HCC was developed by integrating HCC-related datasets and associated clinical data obtained from The Cancer Genome Atlas (TCGA) database. The GSVA website was employed to analyze the model genes across pan-cancer datasets, focusing on copy number variations (CNV), single nucleotide variations (SNV), methylation differences, drug sensitivity and immune cell infiltration profiles. Subsequently, we examined the expression levels and prognostic significance of POLR3G in HCC. Utilizing Spearman correlation analysis, we identified genes associated with POLR3G. Furthermore, Gene Set Enrichment Analysis (GSEA) was employed to elucidate the potential signaling pathways in which POLR3G may be involved. The relationship between POLR3G expression and immune cell abundance in HCC samples was assessed using the ssGSEA algorithm. Finally, the impact of POLR3G on HCC cell proliferation was validated through CCK-8 and EDU cell proliferation assays. Through univariate Cox regression analysis and LASSO regression analysis, we established a prognostic risk model for HCC comprising 13 genes. The analysis revealed that individuals categorized in the low-risk group had a markedly improved overall survival probability relative to those in the high-risk group. POLR3G exhibited a markedly elevated expression in HCC tissues when compared to adjacent normal tissues. The expression of POLR3G was correlated with tumor grade, and elevated POLR3G expression was associated with poor prognosis in HCC patients. Furthermore, the expression level of POLR3G was found to be correlated with the level of immune cell infiltration. Knockdown of POLR3G significantly inhibited the proliferative capacity of hepatocellular carcinoma cells. The findings suggest that POLR3G may serve as a potential biomarker influencing the prognosis of hepatocellular carcinoma patients by modulating the tumor immune microenvironment.

Humans

Integrative Genomic Profiling of Newly Diagnosed Prostate Cancers Progressing on Surveillance.

OBJECTIVE: To identify molecular features associated with earlier progression to definitive therapy amongst patients with localized prostate cancer (PCa) managed on active surveillance (AS). METHODS: We performed a retrospective pilot study of 7 patients with low- to intermediate-risk PCa undergoing serial multiparametric MRI (mpMRI)-targeted biopsies of the same lesion while on AS, who all proceeded to definitive therapy. Time-to-treatment (TTT) was defined as years from first biopsy on AS to definitive therapy. Laser-capture microdissection was used to separate tumor epithelium, benign glands, high-grade prostatic intraepithelial neoplasia, and stroma in each biopsy specimen. DNA from the tumor and matched benign tissue underwent whole-exome sequencing, and RNA from all compartments underwent whole-transcriptome sequencing. Somatic mutations and copy-number alterations were compared across serial biopsies and used to reconstruct phylogenies and quantify clonal complexity. RESULTS: Tumors exhibited substantial intratumoral heterogeneity, and in 3 of 6 paired cases, serial mpMRI-targeted biopsies showed discordant somatic profiles consistent with sampling distinct major clones over time. By contrast, no single gene-level alteration, and few large-scale chromosomal events, were associated with TTT. High clonal complexity, defined as ≥3 subclones, was associated with significantly shorter TTT than low complexity (median 1.9 vs 7.2 years; P = .0082). Exploratory pathway analyses of individual tissue components suggested TTT-associated differences in inflammatory signaling and stromal-epithelial cross-talk. CONCLUSION: In this small, hypothesis-generating cohort, clonal complexity was more closely associated with earlier definitive therapy than individual genomic alterations. Larger prospective studies are needed to validate whether multiomic measures of clonal architecture can improve AS risk stratification.

Humans

The SMN locus in the T2T era: Structure, gene conversion, and clinical implications.

Long-read sequencing, paralog-aware variant calling, and telomere-to-telomere (T2T) human genome assemblies now enable the resolution of copy-, haplotype-, and nucleotide-level complexities in segmentally duplicated loci, which were previously inaccessible with short-read sequencing. In this review, we highlight how current technologies and analysis methods reveal extensive diversity in copy number (CN), structure, and gene conversion within the spinal muscular atrophy-associated survival motor neuron (SMN) locus. We summarize how understanding population-level structural variation could be translated into clinical practice, where a nucleotide-level view of the SMN locus may refine prognostic accuracy beyond SMN2 CN and explain variable treatment responses. Finally, we discuss how the approaches and methodologies required to study the SMN locus may be applied elsewhere, providing a scaffold to characterize other complex human genetic regions.

Humans

Effect of estrogen on gene expression in the chick oviduct. Effect of estrogen on the sequence and population complexity of chick oviduct poly(A)-containing RNA.

Total cellular RNA preparations were isolated from chicken oviducts at three different development stages: (a) immature chicks which were chronically stimulated with estrogen; (b) estrogen-stimulated chicks which were then withdrawn from hormone for 12 days; and (c) laying hens. Total cellular RNA containing 3'-poly(A) sequences (poly(A)-RNA) were than isolated from these preparations using oligo(dT)-cellulose chromatography. The number average nucleotide length of the poly(A)-RNA preparations in each case was approximately 2000 nucleotides. The number average nucleotide length of the poly(A) residues at the 3'-terminal end of each RNA preparation was approximately 70 adenylate residues. Complementary DNA (cDNA) copies to each preparation of poly(A)-RNA were synthesized using avian myeloblastosis virus RNA-directed DNA polymerase. The cDNApoly(A) preparations were then utilized in DNA excess hybridization experiments to analyze the complexity of the DNA sequences from which these RNAs were transcribed. Approximately 22% of each of the total cellular poly(A)-RNAs were transcribed from repeated DNA sequences (average repeat frequency of 35 copies/genome) while the remaining majority were transcribed from single copy or unique sequence DNA. It was possible to estimate the number of different poly(A)-RNA sequences per cell by analyzing the kinetics of hybridization of these cDNApoly(A) preparations to total cellular poly(A)-RNA extracts under conditions of RNA excess. The results revealed that 41% of the poly(A)-RNA from laying hen oviduct consisted of, on the average, three different sequences/cell, each of which was present in approximately 25,000 copies/cell. The remainder of the poly(A)-RNA in this tissue consisted of approximately 25,000 different sequences/cell, which were present largely in only two or three copies/cell. A somewhat similar sequence complexity was found for oviduct cells prepared from estrogen-stimulated chicks. We estimated that there were approximately 20,000 different poly(A)-RNA sequences/cell, each represented in only one to two copies/cell. However, there were five sequences which were present, on the average, in a concentration of 5600 copies/cell. The poly(A)-RNAs from hormone-wtihdrawn tissue, on the other hand, had a lower sequence complexity. There were only approximately 10,000 different poly(A)-RNA sequences/cell, each present in about three copies/cell. Furthermore, the few sequences present in a great abundance in hen and hormone-stimulated tissues were apparently absent in oviduct tissue from hormone-wtihdrawn chicks, suggesting that the intracellular concentrations of these high frequency RNA sequences are dependent on estrogen.

Animals

Genomic Profiling of Anophthalmia/Microphthalmia-Associated CNVs Reveals Complex Genotype-Phenotype Correlations and Incomplete Penetrance.

BACKGROUND: Anophthalmia/microphthalmia (A/M) is a severe congenital ocular malformation characterized by the complete absence or small size of the eye bulb. Interpreting copy number variations (CNVs) in A/M is challenged by variable genotype-phenotype correlations and reduced penetrance. This study investigated the genetic etiology of A/M-associated CNVs. METHODS: Genomic profiling was performed on four unrelated families presenting with ocular anomalies or harboring A/M-susceptible CNVs. Variants were evaluated by integrating American College of Medical Genetics and Genomics (ACMG) guidelines with clinical phenotypes and familial segregation. RESULTS: An inherited 8.13 Mb deletion (8p23.3p23.1) in Patient 1 was excluded due to genotype-phenotype mismatch. Patients 2 and 3 harbored de novo pathogenic deletions involving OTX2 (14q22.3) and SOX2 (3q26.33), causing typical A/M. Case 4 revealed a 14q22.2q23.1 deletion encompassing OTX2 in a fetus and mother without ocular anomalies, consistent with the incomplete penetrance of OTX2-related microphthalmia. Thus, CNV-induced haploinsufficiency causes A/M with high phenotypic variability. CONCLUSION: Accurate CNV interpretation requires robust genotype-phenotype correlation and careful assessment of incomplete penetrance to prevent diagnostic pitfalls and improve genetic counseling.

Female