Search PubMedSearch

SEARCH · Search PubMed

Results for “coding variants”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Idiopathic neonatal arterial ischaemic stroke: a trio-based whole-exome sequencing study.

OBJECTIVE: To assess the contribution of rare coding genetic variants to idiopathic neonatal arterial ischaemic stroke (NAIS). DESIGN: Observational genetic study using trio-based whole-exome sequencing (WES). SETTING: Multicentre study. PATIENTS: 23 newborns diagnosed with idiopathic NAIS and their biological parents. INTERVENTIONS: WES-trio with a customised workflow for filtering and interpreting variants in de novo autosomal dominant and recessive inheritance models. MAIN OUTCOME MEASURES: Identification of pathogenic (P) or likely pathogenic (LP) variants potentially associated with NAIS. RESULTS: We identified 28 unique rare de novo variants in 28 genes across 23 newborns with NAIS. Under the autosomal recessive model, no candidate genes were identified. No common P/LP variant across the 23 newborns was detected. In-silico predictors and comprehensive knowledge-driven analysis highlighted PIK3CD (p.Gln431Arg) as a candidate gene in one patient with perforant stroke. However, no more cases were identified with PIK3CD variants, and functional studies are warranted to assess its pathogenicity impact. CONCLUSIONS: Trio-based WES did not identify a monogenic cause for idiopathic NAIS. Coding variants therefore appear unlikely to explain the underlying genetic base of the disease. Furthermore, PIK3CD (p.Gln431Arg) may contribute to perforant stroke, although it requires further association evidence. As the potential role of non-coding or structural variants in NAIS remains possible, genome-wide long-read sequencing approaches may provide further insights into the genetic architecture of this condition.

Humans

Integrative Long-Read Multi-Omics of a Patient With GPI Deficiency: A Molecular Case Study of a Candidate Dual-Effect GPI Variant.

The molecular determinants of phenotypic severity in red cell enzymopathies are often obscured by the disconnect between coding sequence variants and their regulatory landscapes. Here we present a single-patient molecular case study that uses an integrative multi-omic approach-combining short-read WGS, PacBio HiFi long-read sequencing, native CpG methylation profiling, and Iso-Seq full-length transcriptomics-to characterize a severe, transfusion-dependent hemolytic anaemia. We identified a compound heterozygous state in the glucose-6-phosphate isomerase (GPI) gene, with no wild-type allele present. One allele (Haplotype 1) carried a missense variant (p.His191Arg); the other (Haplotype 2) carried a distinct missense variant, c.1414C>T (p.Arg472Cys), previously reported as biochemically unstable. Long-read phasing placed the two variants in trans. Allele-resolved transcript counts showed a directionally consistent but statistically non-significant trend toward higher expression of Haplotype 2 across two Iso-Seq replicates. Notably, the c.1414C>T transition abolishes a local CpG dinucleotide; in a small number of haplotype-2 reads spanning this position, the corresponding cytosine on the wild-type/Haplotype-1 background was methylated. We did not measure GPI protein abundance, enzymatic activity, or stability in this patient, and we do not establish that methylation at this site regulates GPI transcription. On the basis of these correlative observations in a single patient, we propose-as a hypothesis for future testing-that a coding variant might simultaneously perturb protein stability and disrupt a local epigenetic mark, and we outline the experiments required to test whether such a dual effect contributes to disease. This case illustrates the value of integrative long-read multi-omics for generating mechanistic hypotheses about variants of uncertain significance, while underscoring that causal claims require dedicated functional validation.

Humans

Molecular cloning of a dog thyrotropin (TSH) receptor variant.

A clone coding for a variant form of thyrotropin receptor was isolated from a dog thyroid cDNA library. It was characterized by a 75 bp deletion in the coding region, additionally to minor modifications in the 3' untranslated region. The corresponding 25 amino acids deletion is located in the long NH2 terminal extracellular domain which is characteristic of the glycoprotein hormone receptors. This region of the protein is composed of imperfect repeats and the deletion corresponds exactly to one of the repeat units. This suggests that the repeats correspond to individual exons in the thyrotropin receptor chromosomal gene. It is not known whether the deletion of the repeat and the concomitant suppression of one of the N-glycosylation sites of the molecule do alter the receptor function.

Amino Acid Sequence

Complete coding sequences of cDNAs of four variants of rabbit skeletal muscle troponin T.

Four variants of troponin T (TnT) cDNAs have been isolated and sequenced. These cDNAs have been derived from rabbit skeletal muscle, the most widely studied source of troponin, of a 11-day-old animal. One variant (TnT-1) contains the complete coding sequence, while in three variants the coding sequences are truncated at the 5' termini. The previously published amino acid sequence differs from the present cDNA-derived sequences at three locations. At least two, possibly all, of them are probably accounted for by errors in peptide sequencing. The present results are consistent with the two types of alternative splicing of TnT genes, both being first reported on the rat gene. (1) Highly variable sequences in the amino-terminal region are accounted for by the alternative splicing of exons 4-8 in an interchangeable but not mutually exclusive manner. (2) In the carboxyl-terminal region, the alternative splicing of two exons 17 (beta-type) or 16 (alpha-type) in mutually exclusive manner is consistent with the difference between all the four cDNAs, which express exon 17, and the previously published peptide sequence (derived from the adult muscle) in which exon 16 is present. This variation also corresponds to the finding in chicken skeletal muscle that the choice of exon 16 or 17 may be dependent on developmental stages. Finally, a sequence is observed corresponding to an extra exon or exons between exons 5 and 6. This sequence is shorter than that of the chicken skeletal muscle gene and is not detected in the rat skeletal muscle gene.

Amino Acid Sequence

Analysis of structure and conservation for supporting functional evaluation of PMS2 missense variants.

Germline defects in mismatch repair (MMR) genes are known to significantly increase the risk of developing certain types of cancers, notably colorectal and endometrial cancers. These conditions are characterized under Lynch syndrome. Accurate diagnosis of this predisposition, along with meaningful predictive testing for family members, necessitates the identification of pathogenic variants. However, classifying small coding genetic variants identified in cancer patients is very challenging, specifically in the case of PMS2 variants, since PMS2 pathogenic variants display a lower penetrance and less severe phenotype and therefore a lower tumor burden in affected families. We have assembled clinical data on four PMS2 missense variants of uncertain significance (VUS) identified in 23 patients (p.(Asp286Gly), p.(Asn335Ser), p.(Ile679Thr) and p.(Arg799Trp)). For these variants, functional testing was performed (RNA splicing, protein stability and catalytic activity). Since many protein ortholog sequences and accurate predictive models from AlphaFold2 are available, we also included a systematic analysis of residue conservation and structural role (ConStruct assessment). Overall, our findings indicate that p.(Asp286Gly) and p.(Arg799Trp) behave similarly to wild-type PMS2 and are thus probably neutral. In contrast, p.(Asn335Ser) and p.(Ile679Thr) conferred defects in protein expression or MMR activity. These could be explained by the relevant roles of these amino acids in MLH1-PMS2-N-terminal dimerization (p.Asn335) and C-terminal dimerization (p.Ile679). Our data thus suggest that p.(Asp286Gly) and p.(Arg799Trp) are benign, while the tumor risk in the other two variants remains to be established. Taken together, we suggest roadmaps for the individualized evaluation of difficult uncertain variants by comprising information from all available sources.

Humans

Rhinovirus infection of airway epithelial cells uncovers the non-ciliated subset as a likely driver of genetic susceptibility to childhood-onset asthma.

Asthma is a complex disease caused by genetic and environmental factors. Epidemiological studies have shown that in children, wheezing during rhinovirus infection (a cause of the common cold) is associated with asthma development during childhood. This has led scientists to hypothesize there could be a causal relationship between rhinovirus infection and asthma or that RV-induced wheezing identifies individuals at increased risk for asthma development. However, not all children who wheeze when they have a cold develop asthma. Genome-wide association studies (GWAS) have identified hundreds of genetic variants contributing to asthma susceptibility, with the vast majority of likely causal variants being non-coding. Integrative analyses with transcriptomic and epigenomic datasets have indicated that T cells drive asthma risk, which has been supported by mouse studies. However, the datasets ascertained in these integrative analyses lack airway epithelial cells. Furthermore, large-scale transcriptomic T cell studies have not identified the regulatory effects of most non-coding risk variants in asthma GWAS, indicating there could be additional cell types harboring these "missing regulatory effects". Given that airway epithelial cells are the first line of defense against rhinovirus, we hypothesized they could be mediators of genetic susceptibility to asthma. Here we integrate GWAS data with transcriptomic datasets of airway epithelial cells subject to stimuli that could induce activation states relevant to asthma. We demonstrate that epithelial cultures infected with rhinovirus significantly upregulate childhood-onset asthma-associated genes. We show that this upregulation occurs specifically in non-ciliated epithelial cells. This enrichment for genes in asthma risk loci, or 'asthma heritability enrichment' is also significant for epithelial genes upregulated with influenza infection, but not with SARS-CoV-2 infection or cytokine activation. Additionally, cells from patients with asthma showed a stronger heritability enrichment compared to cells from healthy individuals. Overall, our results suggest that rhinovirus infection is an environmental factor that interacts with genetic risk factors through non-ciliated airway epithelial cells to drive childhood-onset asthma.

Preprint

Functional phenotyping of genomic variants using joint multiomic single-cell DNA-RNA sequencing.

Genetic variants (both coding and noncoding) can impact gene function and expression, driving disease mechanisms such as cancer progression. The systematic study of endogenous genetic variants is hindered by inefficient precision editing tools, combined with technical limitations in confidently linking genotypes to gene expression at single-cell resolution. We developed single-cell DNA-RNA sequencing (SDR-seq) to simultaneously profile up to 480 genomic DNA loci and genes in thousands of single cells, enabling accurate determination of coding and noncoding variant zygosity alongside associated gene expression changes. Using SDR-seq, we associate coding and noncoding variants with distinct gene expression in human induced pluripotent stem cells. Furthermore, we demonstrate that in primary B cell lymphoma samples, cells with a higher mutational burden exhibit elevated B cell receptor signaling and tumorigenic gene expression. SDR-seq provides a powerful platform to dissect regulatory mechanisms encoded by genetic variants, advancing our understanding of gene expression regulation and its implications for disease.

Humans

Whole genome sequence analysis of low-density lipoprotein cholesterol across 246 K individuals.

BACKGROUND: Rare genetic variation provided by whole genome sequence datasets has been relatively less explored for its contributions to human traits. Meta-analysis of sequencing data offers advantages by integrating larger sample sizes from diverse cohorts, thereby increasing the likelihood of discovering novel insights into complex traits. Furthermore, emerging methods in genome-wide rare variant association testing further improve power and interpretability. RESULTS: Here, we conduct the largest meta-analysis of whole genome sequencing for low-density lipoprotein cholesterol (LDL-C), a therapeutic target for coronary artery disease, analyzing data from 246 K participants and integrating 1.23B variants from the UK Biobank and the Trans-Omics for Precision Medicine (TOPMed) program. We identify numerous rare coding and non-coding gene associations related to LDL-C, with replication across 86 K participants in All of Us. Our findings are based on single-variant analyses, rare coding and non-coding variant aggregation tests, and sliding window approaches. Through this comprehensive analysis, we identify 704 novel single-variant associations, 25 novel rare coding variant aggregates, 28 novel rare non-coding variant aggregates, and one novel sliding window aggregate. CONCLUSIONS: This study provides a meta-analysis framework for large-scale whole genome sequence association analyses from diverse population groups, yielding novel rare non-coding variant associations.

Humans

Extracting and calibrating evidence of variant pathogenicity from population biobank data.

Genomic medicine requires a robust evidence base of variant phenotypic impacts, which remains incomplete even in extensively studied genes with monogenic disease associations. Here, we evaluated the broad potential of using population cohort data to identify evidence that can be used in variant assessment. Across 41 genes related to 18 clinically actionable monogenic phenotypes, we calculated variant-level odds ratios of disease enrichment using data from 469,803 UK Biobank participants. We found significant differences in odds ratio values between ClinVar-labeled pathogenic and benign variants in 11 phenotypes, spanning both common and rare disorders. To facilitate clinical translation, we calibrated the strength of evidence provided by variant-level odds ratios to align with American College of Medical Genetics and Genomics and the Association for Molecular Pathology (ACMG/AMP) interpretation guidelines (PS4 criterion) and found that odds ratios may reach "moderate," "strong," or "very strong" evidence, varying by phenotype and gene. Overall, we found that 2.6% (N = 12,350) of participants harbor a rare variant of uncertain significance (VUS) with at least moderate evidence of pathogenicity-an indication of potentially unrecognized disease risk. Finally, by incorporating computational and functional data alongside population-based odds ratios, we identified variants that met the criteria for clinical reclassification. Notably, using this approach, we identified that 12.4% of rare VUSs in LDLR seen in participants meet diagnostic criteria to be classified as likely pathogenic, demonstrating its potential to scale the reclassification of VUSs.

Humans

Blood mitochondrial heteroplasmic variants and cognitive performance in late midlife: REGARDS study.

BACKGROUND: Studies linking mitochondrial DNA (mtDNA) variants to cognition yielded inconsistent findings, and the underlying mechanisms remain unclear. We investigated whether mtDNA heteroplasmic variants were associated with cognitive outcomes, including the Montreal Cognitive Assessment (MoCA), in 197 late midlife adults from the Reasons for Geographic and Racial Differences in Stroke (REGARDS) cohort with complete data. METHODS: MtDNA was sequenced from blood using targeted deep sequencing. Adjusted linear and mixed-effects models examined the associations by functional regions, genes, total variant burden, nonsynonymous variants, and control regions. RESULTS: Heteroplasmic variants in the control region (β = -0.44, 95% CI: -0.83, -0.05, p = 0.027) and transfer RNA (tRNA) genes (β = -1.34, 95% CI: -2.58, -0.11, p = 0.034) were associated with MoCA baseline scores. Individual variants in cytochrome c oxidase subunit 1 (CO1) (β = -1.51, 95% CI: -2.54, -0.47, p = 0.005), NADH dehydrogenase subunit 1 (ND1) (β = -2.63, 95% CI: -4.56, -0.70, p = 0.008), and Displacement Loop (D-LOOP2) (β = -2.25, 95% CI: -4.20, -0.30, p = 0.025) was associated with reduced baseline MoCA scores. The ND6 (β = −1.23, 95% CI: −2.09, − 0.37, p = 0.006), ND4 (β = −1.11, 95% CI: −2.02, − 0.20, p = 0.018), ATP Synthase Membrane Subunit 8 (ATP8; β = −1.38, 95% CI: −2.63, − 0.13, p = 0.031), and D-LOOP1 (β = −0.61, 95% CI: −1.20, − 0.01, p = 0.045) genes suggested a potential association with executive function. Longitudinal Animal Fluency Test (AFT) scores were inversely associated with heteroplasmic variants in coding regions (β = -0.10, 95% CI: -0.19, -0.006, p = 0.049), the total number of variants (β = -0.06, 95% CI: -0.11, -0.003, p = 0.037) and total nonsynonymous variants (β = -0.11, 95% CI: -0.21, -0.01, p = 0.040). Variants in the control region were associated with the greatest decline in verbal fluency (β = −0.20, 95% CI: −0.39 to − 0.002, p = 0.049). No associations were observed between mitochondrial variants and verbal memory performance or the MoCA composite scores. CONCLUSIONS: Our study indicates that mitochondrial variants measured in blood may provide insight into cognitive function during midlife. However, additional studies are needed to validate these associations and to address potential power limitations in our study.

Humans

Domain mapping of disease mutations reveals pathogenic SORL1 variants in Alzheimer's disease.

BACKGROUND: Protein truncating variants (PTVs) in SORL1 are observed almost exclusively in Alzheimer’s Disease (AD) cases, but the effect of rare SORL1 missense variants is unclear. METHODS: To identify high-priority missense variants (HPVs), we applied ‘domain mapping of disease mutations’ for the 637 unique coding SORL1 variants detected in 18,959 AD-cases and 21,893 non-demented controls. RESULTS: In this sample, PTVs and HPVs associated with respectively a 35- and 10-fold increased risk of early onset AD and 17- and 6-fold increased risk of overall AD. The median age at onset (AAO) of PTV- and HPV-carriers was 62 and 64 years, and APOE-genotype contributed to AAO-variability. The median AAO of PTV- and HPV-carriers is ~8–10 years earlier than wild-type SORL1 carriers, matched for APOE-genotype. Specific HPVs are highly penetrant and lead to earlier AAOs than PTVs, suggesting possible dominant negative effects. CONCLUSION: Our results justify a debate on whether HPV carriers should be considered for clinical counseling.

Humans

Characterization of a cDNA clone coding for a sea urchin histone H2A variant related to the H2A.F/Z histone protein in vertebrates.

A cDNA clone coding for a sea urchin histone H2A variant has been isolated. The coding region of the clone has been sequenced and the sequence found to be closely related to the H2A.F sequence in chickens. The nucleotide sequence of the sea urchin H2A.F/Z is 74% conserved when compared to chicken H2A.F and 51% conserved compared to sea urchin H2A early and 60% compared to sea urchin H2A late. The nucleotide-derived amino acid comparisons show that H2A.F/Z is 97% homologous with H2A.F in chickens and 57% and 56% homologous when compared to sea urchin H2A early and late respectively. There are between 3-6 copies of the H2A.F/Z sequence in the S. purpuratus genome. The H2A.F/Z gene sequence codes for the previously identified H2A.Z protein. All embryonic stages and adult tissues tested contain mRNA for H2A.F/Z. The mRNA appears in the poly A+ RNA fraction after chromatography over oligo dT cellulose.

Age Factors

Amplification by the polymerase chain reaction of a specific target sequence in the gene coding for Escherichia coli verotoxin (VTe variant).

Synthetic oligonucleotide primers were used in a polymerase chain reaction (PCR) protocol to target a specific sequence in the gene coding for the A subunit of Escherichia coli verotoxin (VTe-variant, VTev). This PCR protocol permits the VTe-variant target sequence to be distinguished from closely related sequences in the same coding regions for type 1, type 2, and type 2 variant E. coli verotoxins. This procedure will be a valuable adjunct to other DNA amplification techniques currently being used for molecular epidemiological studies of verotoxigenic E. coli.

Bacterial Toxins

Opacity genes in Neisseria gonorrhoeae: control of phase and antigenic variation.

The chromosome of N. gonorrhoeae contains several complete expression genes coding for variant opacity proteins. DNA sequence analysis of two opacity genes derived from the same locus (opaE1) of two isogenic gonococcal variants reveals common and variable regions in these genes. Genomic blotting experiments using synthetic probes suggest gene conversion as a principle for the assembly of variant sequence information in opacity genes. The 5' region of opacity genes is composed of identical pentameric pyrimidine units (CTCTT) encoding the hydrophobic portion of the opacity leader peptide. This coding repeat is variable in a given locus with respect of the number of pentameric units. While all expression loci in a single cell are constitutively transcribed, the production of opacity proteins is determined by the coding repeat sequence on the translational level.

Amino Acid Sequence

Nucleotide sequence and gene organization of sea urchin mitochondrial DNA.

The 15,650 base-pair mitochondrial genome of the sea urchin Strongylocentrotus purpuratus has been cloned and sequenced. It exhibits a novel organization that suggests the primacy of post-transcriptional gene regulation. The same 13 polypeptides, two rRNAs and 22 tRNAs are encoded as in other animal mitochondrial DNAs, but are organized with extreme economy; non-coding information between genes is almost completely absent, some stop codons are generated post-transcriptionally and tRNA sequences are interspersed between only a minority of other structural genes. The genome uses a variant genetic code, in which AAA specifies asparagine, ATA isoleucine, TGA tryptophan and AGN serine, and has an unusual pattern of codon bias. The order of genes shows several differences from that of vertebrates. The genes for the large (16 S) ribosomal RNA and for NADH dehydrogenase subunit 4L (ND4L) are in different positions, located respectively between those encoding ND2 and cytochrome oxidase subunit I (COI) and between COI and COII. This organization is conserved amongst at least four regular echinoids diverging by some 225 million years. Most tRNA genes are also in different positions. The only long unassigned sequence in the genome (121 base-pairs) is located within a cluster of 15 tRNA genes. It contains elements resembling some of those found in the displacement (D) loop of vertebrate mtDNAs, notably polypurine/polypyrimidine tracts that may play a role in regulating transcription and the initiation of replication. The separation of the ribosomal RNA genes from each other and from the putative control region imposes special demands on the transcription of the genome.

Animals

Drosophila melanogaster tRNAVal3b genes and their allogenes.

Drosophila tRNAVal3b genes have been analyzed with respect to their nucleotide sequence and in vitro transcription efficiency. Plasmid pDt78R contains a single tRNA gene derived from the major tRNAVal3b gene cluster at chromosome band 84D. Its sequence corresponds to that of the tRNAVal3b. Two other plasmids, pDt41R and pDt48, each contain a tRNAVal3b-like gene from the minor tRNAVal3b gene cluster at chromosome bands 90BC. They contain the expected CAC anticodon, but their sequence differs from the tRNA at four positions. In homologous cell-free extracts, the tRNAVal3b variant genes in pDt41R and pDt48 are transcribed an order of magnitude more efficiently than the tRNAVal3b gene in pDt78R. However, the variant genes do not appear to contribute significantly to the in vivo tRNA pool [Larsen et al.: Mol. Gen. Genet. 185 (1982) 390-396]. We propose the term allogenes to describe families of related DNA sequences that may code for variant forms of a standard tRNA isoaccepting species.

Chromosome Mapping

Structural analysis of a variant clone of Snyder-Theilen feline sarcoma virus.

A variant clone of Snyder-Theilen feline sarcoma virus (ST-FeSV) encoding a polyprotein with a molecular weight of approximately 104 kDa (P104) was compared to the P85 encoding prototype clone of ST-FeSV. Analysis of chimeric genes constructed with the viral oncogenes of the two clones indicated that the variant clone coded for a larger polyprotein than the prototype clone because of genetic differences in its 3' portion. Comparative DNA sequence analysis revealed that one nucleotide just upstream of the termination condon TGA in the prototype proviral DNA was deleted from the variant clone resulting in a 468-bp larger open reading frame. Furthermore, it appeared that the U3 regions of the long terminal repeats (LTRs) of the variant clone contained an insertion of 71 bp as compared to the LTRs of the prototype clone. In addition, both clones differed also from each other with respect to genetic sequences deleted from their env gene regions.

Base Sequence