Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “structural variants”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

The antimicrobial effect of a structural variant of subtilin against outgrowing Bacillus cereus T spores and vegetative cells occurs by different mechanisms.

Subtilin is a ribosomally synthesized antimicrobial peptide that contains several unusual amino acids as a result of posttranslational modifications. Site-directed mutagenesis was employed to construct a structural variant of subtilin in which the unusual dehydroalanine (Dha) residue at position 5 was changed to alanine. Proton nuclear magnetic resonance spectroscopy, amino acid composition, and N-terminal sequence analysis established that the mutation did not disrupt posttranslational processing of the precursor peptide. This mutant subtilin was devoid of antimicrobial activity as assessed by its lack of inhibitory effects on outgrowth of Bacillus cereus T spores. However, this same mutant subtilin was fully active with respect to its ability to induce lysis of vegetative B. cereus T cells. Because an intact Dha-5 residue is required in the one instance but not in the other, it was concluded that the molecular mechanism by which subtilin inhibits (without lysis) spore outgrowth is not the same as the mechanism by which it inhibits (with lysis) vegetative cells.

Amino Acid Sequence↗

Structural variant discovery and diagnostic impact in rare diseases from short-read and long-read sequencing.

Rare diseases collectively affect 1 in 10 individuals, yet current genetic testing fails to identify a causal variant for most cases. At present, cytogenetic methods and/or sequencing approaches such as exome (ES) or short-read genome sequencing (srGS) represent the state-of-the-art for comprehensive clinical discovery of sequence and structural variants (SVs), including copy number variants, balanced SVs, complex SVs, and tandem repeats (TRs). Recently, long-read genome sequencing (lrGS), coupled with multiomics data, has presented great promise to resolve variation in genomic regions recalcitrant to characterization by srGS such as highly repetitive simple repeat sequences and segmental duplications. However, there are few guidelines to enable clinical interpretation of genetic variation in these highly repetitive genomic regions, and the enthusiasm of the field in adopting lrGS has made it difficult to assess the true added diagnostic yield of this technology due to widely variable and inconsistently applied analytic pipelines and variable degrees of pre-screening by ES or srGS. Here, we investigated the contribution of SVs to rare diseases using srGS as a front-line strategy when paired with highly sensitive SV discovery and evaluate the added diagnostic yield of incorporating lrGS for a subset of cases. Our srGS analysis encompassed 1,462 families (3,450 individuals) recruited through the Broad Institute Center for Mendelian Genetics and the Genomics Research to Elucidate the Genetics of Rare Diseases (GREGoR) programs. Diagnostic SVs were identified in 5.4% of cases (79/1,462), of which 80% were uniquely detectable by srGS compared to standard cytogenetic techniques. For 96 families (including 10 families with a heterozygous variant observed in a known recessive gene of clinical relevance), we performed lrGS with methylation profiling, as well as long-read transcriptomic analyses in a subset of 20 trios. Analyses with lrGS yielded over 25,000 SVs per genome, 63% of which were not captured by srGS, along with an additional ~200 rare SNV/indels per genome not previously captured and 12 differentially methylated regions per genome. Among these, we identified only one diagnostic variant not interpreted by srGS, an apparently mosaic de novo SNV in CASK that was absent in the srGS callset due to allelic imbalance. No new diagnoses were supported by long-read transcriptomics or episignatures. In this well characterized rare disease cohort, the added diagnostic yield was thus 1.04% (1/96 families). Following a systematic literature review of prior lrGS studies, we find that most reported diagnoses were detectable by srGS and that our added diagnostic yield is consistent with those prior studies. These studies emphasize the significant impact of comprehensive SV discovery in rare disease cases and further demonstrate the power for increased discovery of novel genomic variation and episignatures from lrGS. Nonetheless, they also serve to temper expectations of dramatic diagnostic advances in rare disease patients until there is more extensive annotation of the functional and clinical impact of all coding and noncoding variation uniquely accessible to lrGS with extensive reference databases spanning highly repetitive genomic sequencing that could be enabled by this transformative technology.

Journal Article↗

Dissecting genetic architecture and improving machine learning‑based genomic prediction of flowering time in Osmanthus fragrans by integrating structural variants.

Sweet osmanthus (Osmanthus fragrans), a traditional ornamental plant in China, exhibits substantial variation in autumn flowering time, which significantly affects landscape application and cultivation efficiency. Here, we performed a genome-wide association study on 127 resequenced accessions classified into early, intermediate, and late flowering types, using a set of 2,325,410 single-nucleotide polymorphisms (SNPs) and 246,824 structural variants (SVs). By integrating SNP/insertion and deletion (Indel) and SV data with weighted gene co-expression network analysis, machine learning, and genomic prediction, we dissected the genetic architecture of flowering time. We identified 24 associated SNP/Indels and six SVs, mapping to 30 candidate genes, including known flowering regulators FLK, LOS1, Y14, MIF2, and GID1B. These genes showed tissue-specific expression, with some responding to low temperature. The two hub genes, GUX1 and LYG027904, were located within modules of the co-expression network associated with low-temperature treatment. Haplotype analysis revealed a specific three-SNP haplotype associated with late flowering and linked to LOS1, and epistatic interactions among combined genotypes contributed to phenotypic variation. Notably, integrating SVs with SNP/Indels improved genomic prediction accuracy; the gradient boosting decision tree model outperformed other machine learning algorithms, achieving a mean accuracy of 0.859 and an AUC > 0.8 (where AUC is area under receiver operating characteristic curve) for all flowering types. These findings provide insights into the genetic mechanisms underlying flowering time variation in O. fragrans, offer candidate genes and haplotypes for molecular breeding, and highlight the value of integrating SVs with machine learning for genomic prediction in woody ornamentals.

Machine Learning↗

Two structural variants of Nek2 kinase, termed Nek2A and Nek2B, are differentially expressed in Xenopus tissues and development.

Nek2 kinase, a NIMA-related kinase, has been suggested to play both meiotic and mitotic roles in mammals, but its function(s) during development is poorly understood. We have isolated here cDNAs encoding a Xenopus homolog of mammalian Nek2 and have shown that Xenopus Nek2 has two structural variants, termed Nek2A and Nek2B. Nek2A, most likely a C-terminally spliced form, corresponds to the previously described human and mouse Nek2, while Nek2B is most probably a novel, C-terminally unspliced form of Nek2. As a consequence of this (probable) alternative splicing, Nek2B lacks the C-terminal 70-amino-acid sequence of Nek2A, which contains a PEST sequence (or a motif for rapid degradation). Western blot analysis reveals that Nek2A is expressed predominantly in the testis (presumably in spermatocytes) and very weakly in the stomach and, during development, only after the neurula stage. By contrast, Nek2B is expressed mainly in the ovary and in both primary and secondary oocytes and early embryos up to the neurula stage. These results suggest that Nek2A and Nek2B may play both meiotic and mitotic roles, but in a spatially and temporally complementary manner during Xenopus development, and that Nek2B, rather than Nek2A (or the conventional form of Nek2), may play an important role in early development. We discuss the possibility that a counterpart of Xenopus Nek2B might also exist and function in early mammalian development.

Alternative Splicing↗

A Staphylococcus aureus lipoteichoic acid (LTA) derived structural variant with two diacylglycerol residues.

Based on 1,2-O-isopropylidene-sn-glycerol five chiral building blocks containing differently modified glycerol residues were required for the synthesis of the target molecule 2. One of these building blocks is diacylglyceryl beta-gentiobioside carrying a phosphite residue at 6b-O position. Ligation of these five building blocks led to the desired glycerol phosphate backbone to which d-alanyl residues were attached, thus generating after O-deprotection the target molecule 2, a bisamphiphilic structural variant of Staphylococcus aureus LTA. This compound displayed higher potency in terms of cytokine release by human blood leukocytes than the monoamphiphilic variant LTA.

Carbohydrate Conformation↗

Expression of murine leukemia virus envelope glycoprotein gp69/71 on mouse thymocytes. Evidence for two structural variants distinguished by presence vs. absence of GIX antigen.

Thymocytes of several mouse strains were tested for expression of the gp69/71 envelope component of murine leukemia virus by surface iodination, followed by immunoprecipitation and sodium dodecyl sulfate (SDS)-polyacrylamide gel electrophoresis. Theses strains included two congenic lines differing from their partner stocks with respect to expression of GIX antigen demonstrable in the cytoxicity assay. We conclude that:(a) two structural variants of gp69/71 can be expressed on mouse thymocytes, (b) these are distinguishable by a small difference in mobility in SDS gels, (c) one carries GIX antigen and the other not, (d) they are coded, or their expression is regulated, by different chromosomal loci that are not closely linked, and (e) both can be expressed together on the thymocytes of inbred mice. In the intact thymocyte plasma membrane, the sites of group-specific antigen shared by the two gp69/71 variants, unlike the GIX type specificity carried by only one of them, are probably inaccessible to antibody.

Alleles↗

Lack of genetically determined structural variants of the human serotonin-1E (5-HT1E) receptor protein points to its evolutionary conservation.

Using single strand conformational analysis, we screened the complete coding sequence of the serotonin-1E (5-HT1E) receptor gene for the presence of DNA sequence variation in a sample of 157 unrelated individuals. We detected only a silent C-->T transition at the third position of codon 177. The lack of significant mutations leading to structural variants of the human 5-HT1E receptor protein points to a high evolutionary conservation of this receptor protein.

Base Sequence↗

[Interrelationship between structural variants of the apolipoprotein B and ischemic heart disease and plasma lipid levels].

Xba I and EcoR I polymorphism of the apolipoprotein B (APOB) gene was studied by PCR. A significant increase in the frequency of allele X+ and haplotype H+E+ was demonstrated in patients with coronarographically documented coronary heart disease (CHD) over that of the general population. Association of allele E- with increased levels of serum triglycerides was found. The results provide evidence about the contribution of structural variants of the APOB gene to determining CHD.

Adult↗

Comparative kinetic analysis of structural variants of the hairpin ribozyme reveals further potential to optimize its catalytic performance.

The hairpin ribozyme derived from the minus strand of the satellite RNA associated with the tobacco ringspot virus is one of the small catalytic RNAs that has been shown to catalyze trans-cleavage reactions. There is much interest in designing hairpin ribozymes with improved catalytic activity for the development of new therapeutic agents. Extensive mutagenesis studies as well as in vitro selection experiments have been performed to define the structure and optimize its catalytic activity. This communication describes a comparative kinetic analysis of four structural variants, introduced, either alone, or in combination, into the hairpin ribozyme. We have shown that extension of the helix 2 from 4 to 6 bp resulted in a significant decrease in K(M). Furthermore, the combination of this extension with the simultaneous stabilization of helix 4, led to a more than two-fold increase in the catalytic efficiency. This variant showed a 15-fold reduction in the K(M) value in respect to the wild-type ribozyme. This could be of great interest for the in vivo application of this catalytic motif. The 9-bp enlargement of helix 4 implied about a three-fold improvement in the catalytic activity. Similarly, the U39C substitution brought up the efficiency of the ribozyme slightly. However, introduction of nucleotides at the hinge region between A and B domains reduced the catalytic activity. This reduction was gradually increased with the number of nucleotides. Results obtained with variants carrying more than one modification always agreed with the ones obtained from each single variant.

Base Sequence↗

A novel structural variant of the human beta 4 integrin cDNA.

The ability of the alpha 6 beta 4 integrin to function as a laminin receptor appears to be cell-type dependent. We reported that this integrin functions as a laminin receptor on clone A cells, a colon carcinoma cell line (Lee et al., J. Cell Biol., 117:671-678), but this integrin may not function as a laminin receptor on all cell types in which it is expressed. One potential mode of alpha 6 beta 4 regulation resides in the beta 4 cytoplasmic domain because structural variants of this domain exist. We isolated beta 4 clones from a clone A cDNA library and identified a 21 bp (7aa), in-frame deletion not previously reported. This 7aa variant is located within a region that exhibits a relatively high degree of homology (42%) with the 70aa insert previously reported by Tamura et al. (J. Cell Biol., 111:1593-1604). One major difference between these two regions is that the region we have highlighted does not contain the four potential serine/threonine phosphorylation sites that are present in the 210 bp (70aa) insert. PCR analysis revealed that the 7aa variant is also expressed in RNA obtained from normal colon and placenta.

Amino Acid Sequence↗

Profiling of early gene expression induced by erythropoietin receptor structural variants.

The development of erythroid progenitor cells is triggered via the expression of the erythropoietin receptor (EPOR) and its activation by erythropoietin. The function of the resulting receptor complex depends critically on the presence of activated JAK2, and the complex contains a large number of signaling molecules recruited to eight phosphorylated tyrosine residues. Studies using mutant receptor forms have demonstrated that truncated receptors lacking all tyrosines are able to support red blood cell development with low efficiency, whereas add-back mutants containing either Tyr343 or Tyr479 reconstitute EPOR signaling and erythropoiesis in vivo. To study the contribution of tyrosines to receptor function, we analyzed the activation of essential signaling pathways and early gene induction promoted by different receptor structural variants using human epidermal growth factor receptor/murine EPOR hybrids. In our experiments, receptors lacking all tyrosine residues or the JAK2-binding site did not induce mitogenic and anti-apoptotic signaling, whereas add-back mutant receptors containing single tyrosine residues (Try343 and Tyr479) supported the activation of these functions efficiently. Profiling of early gene expression using cDNA array hybridization revealed that (i) the high redundancy in the activation of signaling pathways is continued at the level of transcription; (ii) the expression of many genes targeted by the wild-type receptor is not supported by add-back mutants; and (iii) a small set of genes are exclusively induced by add-back receptors. We report the identification of several early genes that have not been implicated in the EPOR-dependent response so far.

Animals↗

The arrangement of ribosomal RNA genes in Schistosoma mansoni. Identification of polymorphic structural variants.

The two large ribsomal RNA subunits of Schistosoma mansoni are encoded within a 10 000-base sequence, which is tandemly repeated in the schistosome genome. Restriction endonuclease digestion with Bam HI cuts the rRNA gene into three fragments, which have been clones separately in pBR322 and used to constract a physical map of the gene. The sequence encoding the smaller rRNA subunit is about 2000 bases in length and is situated on the 5' sie of the sequence encoding the larger subunit, which is about 4000 bases. Approximately 4000 bases of the rRNA gene are spacer and do not code for mature rRNA. There are approximately 100 copies of the rRNA gene per haploid genome of which about 10% exhibit length heterogeneity, as judged by the hybridization of the rDNA plasmids to restriction endonuclease digests of genomic DNA. These structural variants appear to contain additional DNA sequences at more than one site within the gene and are polymorphic within the species S. mansoni.

Base Sequence↗

Frequency and distribution of structural variants of hemoglobin and thalassemic states in Western Japan.

Hemolysates from 100,000 people who visited the Kyushu University Hospital and affiliated hospitals during the past 15 years were screened for hemoglobinopathies using electrophoresis on thin-layer starch gel; those exhibiting an abnormality were characterized further on clinical, biochemical, and genetic grounds. Of about 97,000 adult and 3,140 cord blood samples, 29 contained electrophoretically detectable abnormalities in the heterozygous condition. Another 17 samples had quantitative changes in the levels of the minor hemoglobin components. Of the thalassemic conditions, 12 involved beta-thalassemia, 3 alpha-thalassemia, 1 delta beta-thalassemia, and 1 delta-thalassemia. Among 45 carriers of beta-thalassemia from 12 families, 5 were noted to have thalassemia intermedia since they exhibited much more severe hemolytic syndromes than those with typical beta-thalassemia minor. The frequency with which we could detect a structural variant of Hb A in the adults by electrophoresis was one in 3,800 samples. About one in 8,000 carried a beta-thalassemia gene.

Adult↗

Strains of Actinomyces naeslundii and Actinomyces viscosus exhibit structurally variant fimbrial subunit proteins and bind to different peptide motifs in salivary proteins.

Oral strains of Actinomyces spp. express type 1 fimbriae, which are composed of major FimP subunits, and bind preferentially to salivary acidic proline-rich proteins (APRPs) or to statherin. We have mapped genetic differences in the fimP subunit genes and the peptide recognition motifs within the host proteins associated with these differential binding specificities. The fimP genes were amplified by PCR from Actinomyces viscosus ATCC 19246, with preferential binding to statherin, and from Actinomyces naeslundii LY7, P-1-K, and B-1-K, with preferential binding to APRPs. The fimP gene from the statherin-binding strain 19246 is novel and has about 80% nucleotide and amino acid sequence identity to the highly conserved fimP genes of the APRP-binding strains (about 98 to 99% sequence identity). The novel FimP protein contains an amino-terminal signal peptide, randomly distributed single-amino-acid substitutions, and structurally different segments and ends with a cell wall-anchoring and a membrane-spanning region. When agarose beads with CNBr-linked host determinant-specific decapeptides were used, A. viscosus 19246 bound to the Thr42Phe43 terminus of statherin and A. naeslundii LY7 bound to the Pro149Gln150 termini of APRPs. Furthermore, while the APRP-binding A. naeslundii strains originate from the human mouth, A. viscosus strains isolated from the oral cavity of rat and hamster hosts showed preferential binding to statherin and contained the novel fimP gene. Thus, A. viscosus and A. naeslundii display structurally variant fimP genes whose protein products are likely to interact with different peptide motifs and to determine animal host tropism.

Actinomyces↗

Long-read Sequences Mapped to a Complete Reference Genome Uncover Uncaptured Structural Variants across the Beta-globin Cluster in Africans with Sickle Cell Disease.

African genomes are marked by extensive complexity in the number and distribution of variants, yet remain under-represented in genetic databases and the human reference genome. This gap in representation limits the broad application of genomic medicine. Sickle cell disease (SCD) - one of the most common monogenic diseases - has its highest prevalence in Africa, and variation in disease severity has consistently been linked to the beta-globin locus, including levels of fetal hemoglobin (HbF). Modulation of HbF is central to current SCD gene therapies; however, the inherent complexity and variation at the locus in African genomes presents a challenge to translating these advances to Africa. Here, we align long-read single molecule sequences (LRS) targeted to the beta-globin region to the hg38 and T2T-CHM13v2 genome references in 40 individuals with SCD, predominantly recruited from three African countries. We demonstrate that the expanded T2T-CHM13v2 reference sequence at this locus reduces Structural Variant (SV) calls by 70% and uncovers uncaptured single nucleotide variants (SNVs). Across the cluster we report 343 SVs and 196 SNVs that have not been previously reported, including in LRS data from the All of Us project. By including African populations from ethnolinguistic groups that have not been previously surveyed we improve variant resolution and bolster evidence for observed variation. Finally, we identify a common ∼4kb insertion locus overlapping the HBB promoter among individuals with high HbF. These results demonstrate the utility of combining a comprehensive reference genome with LRS in African populations to uncover genomic variation at disease-associated loci.

SNV↗

[Haplotypes of the beta-globulin locus in Czechs and Slovaks with beta-thalassemia and structurally variant hemoglobins].

In 29 Czech and Slovak families with the most frequent and newly identified beta-thalassaemic alleles and with some structural haemoglobin variants (Hb E, Hb Haná, Hb Santa Ana) haplotypes of the beta-globin locus of alleles with these mutations were identified. In most instances haplotypes I and V were involved which were found in 57% of the patients. The bond of the most common beta-thalassaemic mutation: IVS-I-1, IVS-I-110, CD 39 (C-T), IVS-II-745, IVS-I-6 with alleles with the same haplotypes as in the mediterranean region suggests a mediterranean origin of these mutations. In Hb Santa Ana a hitherto not described haplotype was identified (-(+)-(-)-(+3), indicating a de novo origin of the mutation. Also in newly identified beta-thalassaemic mutations in CD 7/8 (+G), in CD 38/39 (-C) and in HbE and Hb Haná de novo development is probable.

Czech Republic↗

Human hypoxanthine-guanine phosphoribosyltransferase. Demonstration of structural variants in lymphoblastoid cells derived from patients with a deficiency of the enzyme.

We have explored the possibility of using cultured lymphoblasts from patients with a deficiency of hypoxanthine-guanine phosphoribosyltransferase (HPRT) as a source of cells for the isolation and characterization of mutant forms of the enzyme. HPRT from lymphoblasts derived from six male patients of five unrelated HPRT-deficient families was highly purified and characterized with regard to: (a) level of immunoreactive protein, (b) absolute specific activity, (c) isoelectric point, (d) migration during nondenaturing polyacrylamide gel electrophoresis, and (e) apparent subunit molecular weight. There experiments were performed on small quantities of lymphoblasts using several micromethods involving protein blot analysis of crude extracts as well as isolation and characterization of enzyme labeled in culture with radioactive amino acids. The lymphoblast enzymes from four of the patients exhibited structural and functional abnormalities that were similar to the recently described abnormalities found with the highly purified erythrocyte enzymes from these same patients. In addition, a previously undescribed HPRT variant was isolated and characterized from lymphoblasts derived from two male siblings. This unique variant has been called HPRT Ann Arbor. We conclude that lymphoblastoid cell lines can be used as a source of cells for the detection, isolation, and characterization of structural variants of human HPRT.

Adolescent↗

Genome assembly comparison identifies structural variants in the human genome.

Numerous types of DNA variation exist, ranging from SNPs to larger structural alterations such as copy number variants (CNVs) and inversions. Alignment of DNA sequence from different sources has been used to identify SNPs and intermediate-sized variants (ISVs). However, only a small proportion of total heterogeneity is characterized, and little is known of the characteristics of most smaller-sized (<50 kb) variants. Here we show that genome assembly comparison is a robust approach for identification of all classes of genetic variation. Through comparison of two human assemblies (Celera's R27c compilation and the Build 35 reference sequence), we identified megabases of sequence (in the form of 13,534 putative non-SNP events) that were absent, inverted or polymorphic in one assembly. Database comparison and laboratory experimentation further demonstrated overlap or validation for 240 variable regions and confirmed >1.5 million SNPs. Some differences were simple insertions and deletions, but in regions containing CNVs, segmental duplication and repetitive DNA, they were more complex. Our results uncover substantial undescribed variation in humans, highlighting the need for comprehensive annotation strategies to fully interpret genome scanning and personalized sequencing projects.

Base Sequence↗