Search PubMedSearch

SEARCH · Search PubMed

Results for “coding variants”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Nucleotide sequence and gene organization of sea urchin mitochondrial DNA.

The 15,650 base-pair mitochondrial genome of the sea urchin Strongylocentrotus purpuratus has been cloned and sequenced. It exhibits a novel organization that suggests the primacy of post-transcriptional gene regulation. The same 13 polypeptides, two rRNAs and 22 tRNAs are encoded as in other animal mitochondrial DNAs, but are organized with extreme economy; non-coding information between genes is almost completely absent, some stop codons are generated post-transcriptionally and tRNA sequences are interspersed between only a minority of other structural genes. The genome uses a variant genetic code, in which AAA specifies asparagine, ATA isoleucine, TGA tryptophan and AGN serine, and has an unusual pattern of codon bias. The order of genes shows several differences from that of vertebrates. The genes for the large (16 S) ribosomal RNA and for NADH dehydrogenase subunit 4L (ND4L) are in different positions, located respectively between those encoding ND2 and cytochrome oxidase subunit I (COI) and between COI and COII. This organization is conserved amongst at least four regular echinoids diverging by some 225 million years. Most tRNA genes are also in different positions. The only long unassigned sequence in the genome (121 base-pairs) is located within a cluster of 15 tRNA genes. It contains elements resembling some of those found in the displacement (D) loop of vertebrate mtDNAs, notably polypurine/polypyrimidine tracts that may play a role in regulating transcription and the initiation of replication. The separation of the ribosomal RNA genes from each other and from the putative control region imposes special demands on the transcription of the genome.

Animals

Drosophila melanogaster tRNAVal3b genes and their allogenes.

Drosophila tRNAVal3b genes have been analyzed with respect to their nucleotide sequence and in vitro transcription efficiency. Plasmid pDt78R contains a single tRNA gene derived from the major tRNAVal3b gene cluster at chromosome band 84D. Its sequence corresponds to that of the tRNAVal3b. Two other plasmids, pDt41R and pDt48, each contain a tRNAVal3b-like gene from the minor tRNAVal3b gene cluster at chromosome bands 90BC. They contain the expected CAC anticodon, but their sequence differs from the tRNA at four positions. In homologous cell-free extracts, the tRNAVal3b variant genes in pDt41R and pDt48 are transcribed an order of magnitude more efficiently than the tRNAVal3b gene in pDt78R. However, the variant genes do not appear to contribute significantly to the in vivo tRNA pool [Larsen et al.: Mol. Gen. Genet. 185 (1982) 390-396]. We propose the term allogenes to describe families of related DNA sequences that may code for variant forms of a standard tRNA isoaccepting species.

Chromosome Mapping

Structural analysis of a variant clone of Snyder-Theilen feline sarcoma virus.

A variant clone of Snyder-Theilen feline sarcoma virus (ST-FeSV) encoding a polyprotein with a molecular weight of approximately 104 kDa (P104) was compared to the P85 encoding prototype clone of ST-FeSV. Analysis of chimeric genes constructed with the viral oncogenes of the two clones indicated that the variant clone coded for a larger polyprotein than the prototype clone because of genetic differences in its 3' portion. Comparative DNA sequence analysis revealed that one nucleotide just upstream of the termination condon TGA in the prototype proviral DNA was deleted from the variant clone resulting in a 468-bp larger open reading frame. Furthermore, it appeared that the U3 regions of the long terminal repeats (LTRs) of the variant clone contained an insertion of 71 bp as compared to the LTRs of the prototype clone. In addition, both clones differed also from each other with respect to genetic sequences deleted from their env gene regions.

Base Sequence

Role of gag sequence in the biochemical properties and transforming activity of the avian sarcoma virus UR2-encoded gag-ros fusion protein.

The transforming protein P68gag-ros of avian sarcoma virus UR2 is a transmembrane tyrosine protein kinase molecule with the gag portion protruding extracellularly. To investigate the role of the gag moiety in the biochemical properties and biological functions of the P68gag-ros fusion protein, retroviruses containing the ros coding sequence of UR2 were constructed and analyzed. The gag-free ros protein was expressed from one of the mutant retroviruses at a level 10 to 50% of that of the wild-type UR2. However, the gag-free ros-containing viruses were not able to either transform chicken embryo fibroblasts or induce tumors in chickens. The specific tyrosine protein kinase activity of gag-free ros protein is about 10- to 20-fold reduced as judged by in vitro autophosphorylation. The gag-free ros protein is still capable of associating with membrane fractions including the plasma membrane, indicating that sequences essential for recognition and binding membranes must be located within ros. Upon passages of the gag-free mutants, transforming and tumorigenic variants occasionally emerged. The variants were found to have regained the gag sequence fused to the 5' end of the ros, apparently via recombination with the helper virus or through intramolecular recombination between ros and upstream gag sequences in the same virus construct. All three variants analyzed code for gag-ros fusion protein larger than 68 kDa. The gag-ros recombination junction of one of the transforming variants was sequenced and found to consist of a p19-p10-p27-ros fusion sequence. We conclude that the gag sequence is essential for the transforming activity of P68gag-ros but is not important for its membrane association.

Amino Acid Sequence

Sequence optimization targeting mRNA stability enhances monoclonal antibody titers in CHO cells.

This study presents a DNA sequence optimization approach that integrates mRNA stability as a tunable design parameter to enhance monoclonal antibody expression in Chinese hamster ovary (CHO) cells. A comprehensive combinatorial library of synonymous coding-sequence variants of an IgG1 light chain was integrated as single copies at a defined genomic locus in CHO cells with identical regulatory elements. Steady-state mRNA abundance, quantified by deep sequencing of gDNA and mRNA, served as a proxy for mRNA stability. These data were used to train a machine learning model that predicts mRNA abundance from coding sequence using embeddings from a pre-trained nucleotide transformer. This abundance predictor, together with established translational metrics, was incorporated into a genetic algorithm for multi-objective codon optimization. As proof-of-concept, we optimized sequences encoding Trastuzumab to either maximize or minimize the abundance criterion and obtained benchmark sequences from two commercial providers. Using targeted integration, we generated CHO cell lines and measured protein titer and cell-specific productivity. Sequences optimized for high abundance significantly increased intracellular mRNA levels (+41%), protein titer (+59%), and cell-specific productivity (+85%) relative to low-abundance designs, while viable cell densities remained comparable. Compared to commercial benchmarks, high-abundance sequences achieved significantly higher titer (+70%) and cell-specific productivity (+98%). These findings establish mRNA stability as a practical and complementary design parameter for codon optimization in monoclonal antibody production, with potential applicability to other proteins and expression systems.

CHO

Characterization of F107 fimbriae of Escherichia coli 107/86, which causes edema disease in pigs, and nucleotide sequence of the F107 major fimbrial subunit gene, fedA.

F107 fimbriae were isolated and purified from edema disease strain 107/86 of Escherichia coli. Plasmid pIH120 was constructed, which contains the gene cluster that codes for adhesive F107 fimbriae. The major fimbrial subunit gene, fedA, was sequenced. An open reading frame that codes for a protein with 170 amino acids, including a 21-amino-acid signal peptide, was found. The protein without the signal sequence has a calculated molecular mass of 15,099 Da. Construction of a nonsense mutation in the open reading frame of fedA abolished both fimbrial expression and the capacity to adhere to isolated porcine intestinal villi. In a screening of 28 reference edema disease strains and isolates from clinically ill piglets, fedA was detected in 24 cases (85.7%). In 20 (83.3%) of these 24 strains, fedA was found in association with Shiga-like toxin II variant genes, coding for the toxin that is characteristic for edema disease strains of E. coli. The fimbrial subunit gene was not detected in enterotoxigenic E. coli strains. Because of the capacity of E. coli HB101(pIH120) transformants to adhere to isolated porcine intestinal villi, the high prevalence of fedA in edema disease strains, and the high correlation with the Shiga-like toxin II variant toxin-encoding genes, we suggest that F107 fimbriae are an important virulence factor in edema disease strains of E. coli.

Amino Acid Sequence

The restriction of codon ambiguity on the basis of known variants.

The genetic code may be used to formulate the nucleotide sequence of a messenger RNA from the known amino acid sequence of a protein. Unfortunately, the degeneracy of the code means that there will be ambiguity in the nucleotide assignments in a third or more of the positions. A simple procedure is given that utilizes the information of known genetic variants to reduce that ambiguity. Problems associated with silent polymorphism are treated. The human alpha and beta hemoglobins are used to exemplify the technique. A total of 68 nucleotides in the two sequences are thereby made less ambiguous. One reduction leads to a nucleotide inconsistent with the result of the recently published beta hemoglobin sequence.

Base Sequence

Missense variants in human forkhead transcription factors reveal determinants of forkhead DNA bispecificity.

Recognition of specific DNA sequences by transcription factors (TFs) is a key step in transcriptional control of gene expression. While most forkhead (FH) TFs bind either an FKH (RYAAAYA) or an FHL (GACGC) recognition motif, some FHs can bind both motifs. Mechanisms that control whether an FH is monospecific vs. bispecific have remained unknown. Screening a library of 12 reference FH proteins, 61 naturally occurring missense variants including clinical variants, and 22 designed mutant FHs for DNA-binding activity using universal ("all 10-mer") protein-binding microarrays revealed non-DNA-contacting residues that control mono- vs. bispecificity. Variation in non-DNA-contacting amino acid residues of TFs is associated with human traits and may play a role in the evolution of TF DNA-binding activities and gene regulatory networks.

Humans

Human hepatic lipase mutations and polymorphisms.

Human hepatic lipase (HL) is a 477 residue glycoprotein that hydrolyzes triglycerides from plasma lipoproteins. Familial HL deficiency is a rare recessive disorder that is characterized by premature atherosclerosis and abnormal circulating lipoproteins. While studying the HL gene from the world's index family with HL deficiency, we identified four coding sequence variants of HL, one in each of exons 4, 5, 6, and 8. In this report we present the genetic basis for two new HL gene variants, one in each of exons 3 and 5. All six HL DNA variants are single base pair changes. Two variants (at codons 133 and 202) are diallelic DNA polymorphisms that are silent at the amino acid level. One variant (V73M) is an allele that defines an uncommon HL isoprotein. One variant (N193S) has two alleles of approximately equal frequency in the population that specify two common HL isoproteins. Two variants (S267F and T383M) are rare mutations found to date only in HL deficient subjects and their relatives. Of the six HL variants described to date, only S267F and T383M are associated with hyperlipidemia.

Alleles

Nucleotide sequence and expression of two cDNA coding for two histone H2B variants of maize.

The complete amino acid sequences of two variants of histone H2B of maize were deduced from the cDNAs isolated from a maize cDNA library. The two encoded proteins are 150 (H2B(1)) and 149 (H2B(2)) amino acids long and shows the classical organization of H2B histones. The hydrophobic C-terminal region is highly conserved as compared to that of the animal counterparts with only 21 changes (13 conservative) among the 90 residues. Between the N-terminal part and the C-terminal region we note the presence of a basic cluster (9 residues) characteristic of histones H2B. The N-terminal third is extended as compared to the animal consensus H2B and has the same size as the H2B histone of wheat. Up to 9 acidic residues and a five time repeated pentapeptide PA/KXE/KK are present in this region. Southern-blot hybridization showed that the H2B histones are encoded by a multigenic family like the other core histones (H3 and H4) of plants. The general expression pattern of these genes was not significantly different from that of the H3 and H4 genes neither in germinating seeds nor in different tissues of adult maize.

Amino Acid Sequence

Sequence and gene organization of the chicken mitochondrial genome. A novel gene order in higher vertebrates.

The 16,775 base-pair mitochondrial genome of the white Leghorn chicken has been cloned and sequenced. The avian genome encodes the same set of genes (13 proteins, 2 rRNAs and 22 tRNAs) as do other vertebrate mitochondrial DNAs and is organized in a very similar economical fashion. There are very few intergenic nucleotides and several instances of overlaps between protein or tRNA genes. The protein genes are highly similar to their mammalian and amphibian counterparts and are translated according to the same variant genetic code. Despite these highly conserved features, the chicken mitochondrial genome displays two distinctive characteristics. First, it exhibits a novel gene order, the contiguous tRNA(Glu) and ND6 genes are located immediately adjacent to the displacement loop region of the molecule, just ahead of the contiguous tRNA(Pro), tRNA(Thr) and cytochrome b genes, which border the displacement loop region in other vertebrate mitochondrial genomes. This unusual gene order is conserved among the galliform birds. Second, a light-strand replication origin, equivalent to the conserved sequence found between the tRNA(Cys) and tRNA(Asn) genes in all vertebrate mitochondrial genomes sequenced thus far, is absent in the chicken genome. These observations indicate that galliform mitochondrial genomes departed from their mammalian and amphibian counterparts during the course of evolution of vertebrate species. These unexpected characteristics represent useful markers for investigating phylogenetic relationships at a higher taxonomic level.

Amino Acid Sequence

Generation of a neutralization-resistant variant of HIV-1 is due to selection for a point mutation in the envelope gene.

Transmission and growth of HIV-1 produced from the biologically active clone HTLV-III/HXB2D in the constant presence of a neutralizing antiserum yielded a viral population specifically resistant to neutralization by the same antiserum. Molecular clones MX-1 and -2, containing the entire envelope gene, were obtained from cultures of the resistant variant. The coding regions for the large envelope protein and most of the transmembrane envelope protein of two such clones were substituted for the homologous segment of HXB2D. Infectious viruses from these constructs were also specifically resistant to neutralization by the selecting antiserum. The exchanged fragment contained only one base change, resulting in an Ala----Thr replacement at position 582. When this substitution was introduced into HXB2D it conferred the resistant phenotype. Thus, small differences may be selected for in vivo by the host immune response and result in relatively large differences in susceptibility of the virus to such a response.

Antibodies, Viral

Rhinovirus infection of airway epithelial cells uncovers the non-ciliated subset as a likely driver of genetic risk to childhood-onset asthma.

Asthma is a complex disease caused by genetic and environmental factors. Studies show that wheezing during rhinovirus infection correlates with childhood asthma development. Over 150 non-coding risk variants for asthma have been identified, many affecting gene regulation in T cells, but the effects of most risk variants remain unknown. We hypothesized that airway epithelial cells could also mediate genetic susceptibility to asthma given they are the first line of defense against respiratory viruses and allergens. We integrated genetic data with transcriptomics of airway epithelial cells subject to different stimuli. We demonstrate that rhinovirus infection significantly upregulates childhood-onset asthma-associated genes, particularly in non-ciliated cells. This enrichment is also observed with influenza infection but not with severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) or cytokine activation. Overall, our results suggest that rhinovirus infection is an environmental factor that interacts with genetic risk factors through non-ciliated airway epithelial cells to drive childhood-onset asthma.

Humans

3D epigenome of glial cell types in developing human cortex.

The human cortex is complex and heterogeneous, undergoing extensive expansion during development1,2. Our prior study of neurogenesis, including radial glia (RG), intermediate progenitor cells, excitatory neurons and interneurons demonstrated that chromatin looping underlies transcriptional regulation for lineage-specific genes, shedding light on how non-coding genetic variants contribute to neuropsychiatric disorders by means of cell-type-specific gene regulation3. RG have a crucial role in generating cellular diversity through both neurogenesis and gliogenesis and can be further classified into ventricular RG (vRG) and outer RG (oRG)4,5. Given their significance in cortical development, we conducted a comprehensive three-dimensional (3D) epigenomic analysis of four main glial populations, including vRG, oRG, oligodendrocyte precursor cells and microglia, from the mid-gestational human neocortex. By integrating gene expression, chromatin accessibility, DNA methylation and 3D chromatin interactions, we identified cell-type-specific candidate cis-regulatory elements (cCREs) and validated their regulatory function using transgenic mouse embryos. Using machine learning, we prioritized 112 schizophrenia risk variants within glia cCREs and further confirmed the predicted vRG enhancer disruption by the rs4449074 risk allele in vivo. Finally, oRG cCREs are enriched for human accelerated regions compared with other cCREs and a subset of human accelerated regions show activity differences from their chimpanzee orthologues that interact with genes involved in neuronal development. Our findings advance the understanding of human-specific gene regulation during corticogenesis.

Journal Article

Telomere conversion in trypanosomes.

Activation of the gene coding for variant surface glycoprotein (VSG) 118 in Trypanosoma brucei proceeds via a duplicative transposition to a telomeric expression site. The resulting active expression-linked extra copy (ELC) is usually flanked by DNA that lacks sites for most restriction enzymes and that is thought to interfere with the cloning of the ELC as recombinant DNA in Escherichia coli. We have circumvented this problem by cloning an aberrant 118 ELC gene, flanked at the 3'-side by at least 1 kb DNA, that contains restriction enzyme sites. Our analysis shows that this DNA and the 3'-end of the 118 ELC gene are derived from another VSG gene (1.1006) that is permanently located at a telomeric position. We propose that the 3'-end of the 1.1006 gene and (all of) its 3' flanking sequence moved to the expression site by a telomere conversion. Such a telomere conversion can also account for the appearance of an extra copy of the 1.1006 gene detected in a sub-population of our trypanosome strain.

Amino Acid Sequence

Optimizing Control Definitions in Opioid Use Disorder Genetic Research Using Electronic Health Records.

Amidst the opioid crisis, understanding the genetic basis of opioid use disorder (OUD) is crucial for identifying biological mechanisms and intervention points. However, genome-wide association studies (GWASs) have been hampered by inadequate sample sizes and often the use of control populations not assessed for prior opioid exposure. Because opioid exposure is a prerequisite for the development of OUD, consideration of exposure history in controls is important. Electronic health record data (EHR) paired with genomic information allow a broader sampling of patients with OUD and exposed controls. We leveraged data across two healthcare systems to evaluate the impact of using controls not screened for opioid exposure ('generic') versus minimally opioid-exposed control ('exposed'). First, at the phenotypic level, we conducted phenome-wide association studies (PheWAS) to compare the medical comorbidity profiles of OUD cases when using generic versus exposed controls. While PheWAS results for OUD-related comorbidities were more pronounced when using the generic group, 83% of the disease associations were overlapping and of similar effect sizes. Second, at the genetic level, we conducted GWAS (cases vs. generic; cases vs. exposed) and assessed differences in genetic correlations and degrees of phenotypic misclassification. Genetic results were concordant across control groups based on heritability (generic: 0.16 ± 0.07 vs. 0.10 ± 0.07), associations with the coding OPRM1 variant rs1799971 (pgeneric = 8.83E-03 vs. pexposed = 1.83E-02) and genetic correlations with prior OUD GWAS (rg-generic = 0.83 ± 0.26 vs. rg-exposed = 0.78 ± 0.27). Although GWASs were limited by sample size (Ngeneric = 6269, Nexposed = 6365), compared to an independent OUD GWAS (N = 425 944), the dilution value for the two GWAS was not different from 1, suggesting no major impact of phenotypic misclassification. This study represents the first effort to enhance OUD genetic research through optimization of control definitions using EHR data. Generic controls ascertained within the US health systems, where exposure to prescription opioids is high, offer a practical alternative for genetic studies of OUD.

Humans

Characterization of pyrimidine deoxyribonucleoside kinase (thymidine kinase) and thymidylate kinase as a multifunctional enzyme in cells transformed by herpes simplex virus type 1 and in cells infected with mutant strains of herpes simplex virus.

Pyrimidine deoxyribonucleoside kinase (thymidine kinase [TK]) was purified from two herpes simplex virus type 1 (HVS-1)-transformed TK-deficient mouse (LMTK-) cell lines and from LMTK- cells infected with HSV-1 mutant viruses coding for variant TK enzymes. These preparations exhibited normal or variant virus-induced thymidylate kinase activities correlating with their relative TK activities. Neither virus-induced activity was detected in LMTK- cells infected with an HSV-1 TK-deficient mutant. These results suggest that HSV-1 thymidylate kinase activity and TK activity are mediated by the same protein.

Animals

Genetic Analysis of Genomic and Methylomic Variation and Identification of Multi-Trait Mutants in Rice Carried on Chang'e-5.

Global food security is facing challenges from population growth to diminishing arable land. Space mutation breeding holds promise for overcoming the variation limitations in conventional breeding; however, the mutagenic effects of the deep-space environment on rice and the transgenerational inheritance patterns of induced variations remain unclear. In this study, rice seeds carried by the Chang'e-5 spacecraft were used as materials. Whole-genome sequencing and whole-genome bisulfite sequencing were performed on the first (SP1) and second generations (SP2) of space-mutagenized plants after their return to Earth. The results showed that the number of genomic variants in the SP2 generation increased significantly compared with SP1, and SNPs, homozygous sites, and variants in coding regions were more heritable. The genome-wide methylation level was elevated in the SP2 generation, and among differentially methylated cytosines, those in the CG context exhibited the highest heritability. Furthermore, large-scale screening for nitrogen efficiency, tolerance to PEG-induced stress, and germination-stage cold resistant mutants was conducted in the SP2 generation, and phenotypic validation was performed in the third generation (SP3). By integrating multi-omics analyses of representative mutants to mine candidate genes, a number of heritable elite mutants were obtained, and seven candidate genes for key traits were identified. This study systematically elucidates the transgenerational inheritance patterns of deep-space-induced variation in rice. The multi-trait mutants obtained provide valuable germplasm resources for gene cloning and breeding applications in rice.

DNA methylation