Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Variant identification”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

Identification of a novel splice variant: human SGT1B (SUGT1B).

We identified a novel splice variant of human SGT1 (SUGT1), a suppressor of the G2 allele of SKP1, by analysis of 8 human EST clones whose open reading frame encoded 365 amino acids. We termed this variant SGT1B (SUGT1B) and the original SGT1A (SUGT1A). The putative SGT1B and SGT1A proteins are 91% identical, and both contain a tetratricopeptide repeat (TPR) domain, two variable regions, a CS domain, and a SGS domain. The NCBI human genome database showed that SGT1B and SGT1A are located on chromosome band 13q14.13. SGT1B contains an additional 33 amino acids encoded by a region between exons 5 and 6 of SGT1A and lacks Ser110 of SGT1A. Immunoblotting using antibodies to N-terminal (amino acids 1-157) and C-terminal (amino acids 182-333) regions of SGT1A detected 2 bands whose sizes corresponded to those predicted for SGT1A and SGT1B. Overexpression experiments confirmed this finding. Additional immunoblot analysis demonstrated that both are highly expressed in human brain, liver, lung, and testis.

Alternative Splicing↗

Identification of two alternate splice variants of a novel serine protease expressed in steroidogenic tissues.

During the search for the serine protease that cleaves pro-gamma-melatropin to stimulate adrenal growth, we identified another novel protease, which we called Adrenal mitochondrial protease (AmP). In situ hybridisation detected AmP transcripts in steroidogenic tissues such as the brain, testis, in ovarian follicles as well as in the adrenal cortex. Full length cloning identified two splice variants differing by a 222 nucleotide insertion in the 5' end of the short variant. The shorter variant codes for a 371 amino acid protein of 40.7 kDa and computer analysis predicts it to be targeted to the cytosol while the longer 445 amino acid protein of 48.4 kDa is mitochondrial. Cellular targeting was confirmed by tagging with GFP. The short variant was clearly cytosolic however, the cells expressing AmP-Long had large vacuoles, possibly as a result of distended (apoptotic?) mitochondria. Due to the mitochondrial localisation of the long variant of the protease and its expression in steroidogenic tissues, it may be expected to be involved in the steroidogenic pathway, possibly by cleaving steroidogenic acute regulatory protein (StAR). We investigated this by co-transfecting AmP-Long with StAR and F2 plasmid into COS-1 cells and measuring the effect on pregnenolone production. It was found that AmP-Long has no effect on steroidogenesis nor cleaves StAR as was shown by western blot analysis using StAR antibody.

Adrenal Glands↗

Identification and characterization of novel variant major surface glycoprotein gene families in rat Pneumocystis carinii.

The major surface glycoprotein (MSG) is an abundant, immunodominant protein on the surface of the opportunistic pathogen Pneumocystis carinii. The current study identified two novel variant MSG (vMSG) gene families in rat P. carinii that are closely related to but distinct from MSG. These gene families encode proteins of approximately 90 kDa (v1MSG) and approximately 115 kDa (v2MSG). Compared with MSG, v1MSG is characterized by a deletion near the carboxyl terminus. The predicted v1MSG and v2MSG proteins are highly homologous to MSG at the carboxyl, but not the amino, terminus. Like MSG, they are cysteine-rich. Approximately 10% of the apparent molecular weight is due to N-linked glycosylation. Southern blotting studies demonstrated that, like MSG, v1MSG and v2MSG are the products of multicopy gene families. However, unlike MSG, each vMSG gene encodes a signal peptide, suggesting that the regulation of vMSG is different from that of MSG.

5' Untranslated Regions↗

Identification of a novel splice variant of C3G which shows tissue-specific expression.

C3G is a guanine nucleotide-releasing protein that binds to the Src homology 3 (SH3) domain of the adapter protein Crk. In this study, we isolated cDNAs coding for rat C3G. Northern blot analysis of RNA from various rat tissues and cell lines showed a major transcript of about 7 kb which was present at the highest level in testis. A comparison of the amino acid sequence (derived from the cDNA sequence) of rat C3G with the human form showed 87.3% sequence identity. The principal difference was the presence of an additional 51 amino acids in the rat C3G sequence after the fifth PXXP motif. This difference may be attributable to alternative splicing of the primary transcript. This interpretation was supported by reverse transcription-polymerase chain reaction (RT-PCR) assays, which resulted in two products differing by 153 bp. The RT-PCR analysis of RNA from various rat tissues showed that the relative expression levels of the two splice forms were variable. The form of C3G with the insertion of 51 amino acids (named C3G-2) was present in rat testis at a high level and, to a lesser extent, in brain, but it was not seen (or was present at a very low level) in other rat tissues and certain rat and mouse cell lines. This expression pattern of the C3G-2 form was confirmed by Northern blotting using the insert region as a probe. The C3G-1 form, without the insertion of 51 amino acids, was present in almost all rat tissues except testis and in cell lines of rat, mouse, or human origin. Thus, in rat cells, we have identified a novel splice variant of C3G. The expression pattern indicates that the form of C3G described here is likely to serve a tissue- or cell-specific physiological function.

Amino Acid Sequence↗

Identification of a new thyroglobulin variant: a guanine-to-adenine transition resulting in the substitution of arginine 2510 by glutamine.

We analyzed thyroglobulin (Tg) reverse transcription polymerase chain reaction (RT-PCR) products from three congenital goiters and three normal thyroid tissues by Taq I digestion. Tg coding sequences were amplified from position 57 to 8448 in 12 amplification fragments. A Taq I restriction fragment length polymorphism was detected in the most 3' RT-PCR product (nt 7584 through 8448). Data from the sequence showed a G-->A transition (nt 7627) causing the disappearance of the Taq I site in position 7625. It produced the substitution of arginine for a glutamine at position 2510. Afterwards, we established that the glutamine allele is present in normal unrelated individuals, with an allelic frequency of 62%. This Tg variant is thus widely represented in the human population. The available sequence information from rat and bovine Tg showed the presence, in both, of glutamine at position 2510.

Adenine↗

DeepGeSeq: deep learning library for genomic sequence modeling and analysis.

MOTIVATION: Deep learning methods have demonstrated significant potential in genomics, enabling broad applications such as sequence activity prediction, regulatory rule identification, and variant effect quantification. However, their widespread adoption is often hindered by the steep computational learning curve required for model construction, training, and downstream biological interpretation. Here, we introduce DeepGeSeq, a user-friendly Deep-learning library tailored for Genomic Sequence modeling and analysis. RESULTS: By integrating state-of-the-art architectural modules, DeepGeSeq streamlines the entire deep learning workflow, requiring minimal user input via a simple configuration file and an intuitive agentic skill. We comprehensively validate the efficacy of DeepGeSeq through diverse case studies, encompassing pipeline verification using synthetic datasets, the reproduction and application of established models, and model fine-tuning coupled with biological interpretation on user-defined data. Furthermore, we demonstrate DeepGeSeq's versatility in domain-specific applications, including single-cell ATAC-seq modeling for cell-type clustering, and MPRA data modeling coupled with in silico saturation mutagenesis to dissect cis-regulatory elements. Ultimately, DeepGeSeq bridges the gap between computational complexity and biological discovery, providing an accessible resource that facilitates the development and broad application of deep learning methods in genomics research. AVAILABILITY AND IMPLEMENTATION: https://github.com/JiaqiLi1024/DeepGeSeq.

Deep Learning↗

Genetic studies of brown adipocyte induction.

We seek to discover an effective method for utilizing thermogenesis to reduce the caloric load in obese individuals. Experimental evidence indicates that nonshivering thermogenesis is the most effective cellular and biochemical mechanism known for reducing excessive adiposity. In this presentation, we describe our experiments aimed at understanding how nonshivering thermogenesis can be induced. In addition, these experiments have led to a genetic approach for the identification of variant genes that coordinate the expression of pathways of gene transcription that are associated with brown adipocyte induction.

Adipocytes↗

Identification of human apolipoprotein E variant gene: apolipoprotein E7 (Glu244,245----Lys244,245).

Apolipoprotein E (apoE) is one of the protein moieties of the human serum lipoproteins. Three major isoforms of apoE (apoE2, apoE3, and apoE4) and minor variant isoforms (apoE1, apoE5, and apoE7) have been detected by isoelectric focusing. In this study we have cloned the apoE7 gene from a patient with the apoE3/E7 phenotype associated with hypertriglyceridemia and diabetes mellitus. DNA sequencing revealed that the apoE7 gene has two base substitutions (G----A) changing Glu244,245----Lys244,245, compared with the apoE3 gene. The replacement of the two amino acids is consistent with the result of isoelectric focusing of the apoE7 isoprotein, which shifts to four positively charged units compared with the apoE3 isoprotein.

Amino Acid Sequence↗

Allelic variation and light-responsive regulation of FaMYB10-2 underlie tissue-specific anthocyanin accumulation in strawberry.

Anthocyanins critically determine fruit color, nutrition, and stress resilience in cultivated strawberry (Fragaria × ananassa), directly influencing consumer preference. Despite complex genetic and environmental regulation of their biosynthesis, the basis for tissue-specific pigmentation, notably the widespread occurrence of red skin and pale flesh, remains poorly understood. We integrated genomic, transcriptomic, and functional analyses across 200 cultivars to dissect receptacle pigmentation regulation. Approaches included FaMYB10-2 allele mining, promoter structural variant (SV) identification, expression profiling, regulatory interaction assays, and characterization of upstream light-responsive factors. FaMYB10-2 was identified as the key R2R3-MYB regulator of fruit anthocyanin biosynthesis. Alleles FaMYB10-2.2 and FaMYB10-2.3 encode truncated proteins retaining bHLH-binding capacity but lacking activation domains, functioning as dominant-negative repressors. A promoter SV 986 bp upstream of FaMYB10-2 was associated with reduced pale fruit due to cis-regulatory divergence. The SV (Alt) allele is prevalent in Asian cultivars, while the Ref allele is enriched in Western germplasm. Crucially, a light-responsive FaHYH-FaWRKY71 cascade activates FaMYB10-2 and structural genes haplotype-dependently, compensating for weak MYB activity in the skin. Our findings reveal a multilayered regulatory system integrating allelic variation, cis-regulatory divergence, and environmental signals, advancing anthocyanin understanding and providing engineering targets for polyploid crop color improvement.

Fragaria↗

The missense genetic polymorphisms of human CYP2A13: functional significance in carcinogen activation and identification of a null allelic variant.

Cytochrome P450 2A13 (CYP2A13), an enzyme predominantly expressed in human respiratory tissues, is highly efficient for the metabolic activation of two suspected human lung carcinogens 4-(methylnitrosamino)-1-(3-pyridyl)-1-butanone (NNK) and aflatoxin B1 (AFB1). Functional genetic polymorphisms of CYP2A13 may therefore be an important factor in human susceptibility to related lung cancers. Among the reported CYP2A13 polymorphisms with missense variations, only CYP2A13*2 variant (containing either a single or double variation of R25Q and R257C) was studied for its NNK-metabolizing activity. The present study demonstrated that there was no remarkable difference in AFB1- and NNK-induced toxicity between the Flp-In Chinese Hamster Ovary (CHO) cells stably expressing wild-type CYP2A13 and the cells expressing the individual polymorphic variants R25Q, D158E, R257C, R25Q/R257C, V323L, F453Y, and R494C. In contrast, cells transfected with R101Q variant complementary DNA (cDNA), same as the vector control cells, showed no significant death even at highest concentrations of AFB1 (10microM) and NNK (200microM). This result correlated with the lack of CYP2A13 protein in the R101Q-CHO cells, although the genomic integration of transfected R101Q cDNA and the expression of R101Q messenger RNA were clearly demonstrated in these stable transfectants. Consistent with the possibility that the variation might reduce the protein stability, R101Q variant protein expressed in insect cells showed a loss of P450 peak and coumarin 7-hydroxylase activity as well as an increased susceptibility to limited protein digestion. Thus, the R101Q polymorphic change results in a null allelic variant of CYP2A13. Our results should be useful in designing and interpreting molecular epidemiological studies related to CYP2A13 genetic polymorphisms.

Aflatoxin B1↗

Update on genetics of inflammatory bowel disease.

Complex genetic disorders such as inflammatory bowel disease (IBD) result from the interplay between multiple genetic and environmental risk factors. The recent identification of variants of the CARD15/NOD2 protein as contributing to Crohn disease represents a major advance in defining disease pathogenesis. CARD15/NOD2 is expressed in monocytes and is capable of activating nuclear factor kappa B (NF-kappaB). Crohn disease-associated mutations in CARD15/NOD2 predominate in its C-terminus leucine-rich repeat domain, which is required for bacterial lipopolysaccharide-dependent induction of NF-kappaB activity. The relative risk of developing Crohn disease is estimated to be in the range of 2 to 3 in people carrying one mutation and 20 to 40 in people carrying two mutations in CARD15/NOD2. Homozygote and compound heterozygote carriers of CARD15/NOD2 mutations are characterized by an earlier age of onset, less involvement of the left colon, and positive association with stricturing disease. However, even carriers of two CARD15/NOD2 mutations have limited disease penetrance (ie, only a minority will develop the disease), suggesting that additional interacting genes and environmental triggers are required for disease expression. Several additional genetic regions have been implicated through genetic linkage and association studies.

Journal Article↗

Genotyping human polymorphic arylamine N-acetyltransferase: identification of new slow allotypic variants.

Arylamine N-acetyltransferase catalyses the N-acetylation of primary arylamine and hydrazine drugs and chemicals. N-acetylation is subject to a polymorphism and humans can be categorized as either fast or slow acetylators according to their ability to N-acetylate polymorphic substrates in vivo. Previously, slow acetylation has been linked to four distinct polymorphic N-acetyltransferase (pnat) alleles each of which contains one or more point mutations within the coding region of the pnat gene. One new rare slow variant of pnat has been identified by cloning and sequencing the pnat DNA from an individual whose NAT phenotype was determined by in vivo acetylation of the polymorphic substrate sulphamethazine. This allele, designated S1c, differs from the wild type fast allele at nucleotide positions 341 and 803. A second new rare slow allotypic variant, designated S3, has been identified by resistance of the pnat specific DNA to digestion with the restriction enzymes Fok I and Bam HI. A method of genotyping individuals for the arylamine N-acetyltransferase (NAT) polymorphism is presented which correctly predicts the phenotype of greater than 95% (21 of 22) of individuals as measured by the extent of acetylation of sulphamethazine in urine. This refined genotyping method was applied to a clinical population of 48 Caucasians with classical or definite rheumatoid arthritis each receiving daily between 150 and 500 mg of the anti-rheumatic drug, D-penicillamine. There is no difference in the N-acetyltransferase phenotype of the individuals who developed proteinuria and the control group with no adverse effects.

Adult↗

A novel method for rapid genotypic identification of alpha 1-antitrypsin variants.

There is worldwide growing awareness of alpha 1-antitrypsin deficiency (AATD), a major hereditary disorder in Caucasians. The gold standard for laboratory diagnosis of AATD is thin-layer isoelectrofocusing (IEF), which is labor intensive and should be performed in reference laboratories. The aim of this study was to find an easy, fast, and cheap method for detecting alpha1-antitrypsin S and Z variants, the most frequent variants associated with AATD. The novel method herein described is based on SexAI/Hpy99I RFLP. We studied samples from 90 subjects enrolled in the Italian National Registry for AATD, previously typed by isoelectrofocusing. We found a complete agreement among our results, IEF, and genotypes obtained by standard methods. We concluded that this novel method combines efficiency, ease, swiftness, and low cost.

DNA Primers↗

Identification and functional characterization of variants in human concentrative nucleoside transporter 3, hCNT3 (SLC28A3), arising from single nucleotide polymorphisms in coding regions of the hCNT3 gene.

INTRODUCTION: Human concentrative nucleoside transporter 3, hCNT3 (SLC28A3), which mediates transport of purine and pyrimidine nucleosides and a variety of antiviral and anticancer nucleoside drugs, was investigated to determine if there are single nucleotide polymorphisms in the coding regions of the hCNT3 gene. METHODS AND RESULTS: Ninety-six DNA samples from Caucasians (Coriell Panel) were sequenced and sixteen variants in exons and flanking intronic regions were identified, of which five were coding variants; three of these were non-synonymous (S5N, L131F, Y513F) and were further investigated for functional alterations of the resulting recombinant proteins in Saccharomyces cerevisiae and Xenopus laevis oocytes. In yeast, immunostaining and fluorescence quantitation of the reference (wild-type) and variant CNT3 proteins showed similar levels of expression. Kinetic studies were undertaken in yeast with a high through-put semi-automated assay process; reference hCNT3 exhibited Km values of 1.7+/-0.3, 3.6+/-1.3, 2.2+/-0.7, and 2.1+/-0.6 muM and Vmax values of 1402+/-286, 1310+/-113, 1020+/-44, and 1740+/-114 pmol/mg/min, respectively, for uridine, cytidine, adenosine and inosine. Similar Km and Vmax values were obtained for the three variant proteins assayed in yeast under identical conditions. All of the characterized hCNT3 variants produced in oocytes retained sodium and proton dependence of uridine transport based on measurements of radioisotope flux and two-electrode voltage-clamp studies. CONCLUSION: These results suggested a high degree of conservation of function for hCNT3 in the Caucasian population.

Adenosine↗

Epstein--Barr virus gene polymorphisms in Chinese Hodgkin's disease cases and healthy donors: identification of three distinct virus variants.

Epstein--Barr virus (EBV) is associated with several malignancies. Specific EBV gene variants, e.g. the BamHI f configuration, a C-terminal region 30 bp deletion in the latent membrane protein-1 (LMP1) gene (del-LMP) and the loss of an XhoI site in LMP1 (XhoI-loss), are found in Chinese cases of nasopharyngeal carcinoma (NPC), suggesting that EBV sequence variation may be involved in oncogenesis. In order to understand better the epidemiology of these EBV variants, they were studied in virus isolates from EBV-positive Chinese cases of Hodgkin's disease (HD; n=71) and donor throat washings from healthy CHINESE: Sequencing was performed of 15 representative EBV isolates, including the first analysis of the LMP1 promoter in Asian wild-type EBV isolates. The following observations were made. (i) Three EBV LMP1 variants were identified, designated Chinese groups (CG) 1--3. In both EBV-associated HD and in healthy Chinese, CG1-like viruses showing del-LMP1 and XhoI-loss were predominant. (ii) CG1viruses were distinct from European and African variants, suggesting that this profile is useful for epidemiological studies. (iii) Specific patterns of mutations were present in the LMP1 promoter in both CG1 and CG2. (iv) The BamHI f variant was not found in Chinese HD, in contrast to Chinese NPC and European HD. This study confirms that EBV isolates in Chinese HD and other tumours differ from those reported in Western cases. However, this reflects the predominant virus strain present in the healthy Chinese population, suggesting that these are geographically restricted polymorphisms rather than tumour-specific strains.

Asian People↗

Typing of intimin (eae) genes from enteropathogenic Escherichia coli (EPEC) isolated from children with diarrhoea in Montevideo, Uruguay: identification of two novel intimin variants (muB and xiR/beta2B).

A total of 71 enteropathogenic Escherichia coli (EPEC) strains isolated from children with diarrhoea in Montevideo, Uruguay, were characterized in this study. PCR showed that 57 isolates carried eae and bfp genes (typical EPEC strains), and 14 possessed only the eae gene (atypical EPEC strains). These EPEC strains belonged to 21 O : H serotypes, including eight novel serotypes not previously reported among human EPEC in other studies. However, 72% belonged to only four serotypes: O55:H- (six strains), O111:H2 (13 strains), O111:H- (14 strains) and O119:H6 (18 strains). Nine intimin types, namely, alpha1 (two O142 strains), beta1 (29 strains, including 13 O111:H2 and 14 O111:H-), gamma1 (three O55:H- strains), theta (five strains, including three strains with H40 antigen), kappa (two strains), epsilon1 (one strain), lambda (one strain), muB (six strains of serotypes O55:H51 and O55:H-) and xiR/beta2B (22 strains, including 18 O119:H6) were detected among the 71 EPEC strains. The authors have identified two novel intimin genes (muB and xiR/beta2B) in typical EPEC strains of serotypes O55:H51/H- and O119:H6/H-. The complete nucleotide sequences of the novel muB and xiR/beta2 variant genes were determined. PFGE typing after XbaI DNA digestion was performed on 44 representative EPEC strains. Genomic DNA fingerprinting revealed 44 distinct restriction patterns and the strains were clustered in 12 groups. Only 15 strains clustered in six groups of closely related (similarity>85%) PFGE patterns, suggesting the prevailing clonal diversity among EPEC strains isolated from children with diarrhoea in Montevideo.

Adhesins, Bacterial↗

ViewGene: a graphical tool for polymorphism visualization and characterization.

The human genome project is producing an enormous amount of sequence data, based on which single base changes between individuals can be identified. Unfortunately, computer tools that were adequate for sequence assembly are less than ideal for the characterization of polymorphism data [single nucleotide (snp) or insertion/deletion (indel)] and other sequence features, and their relationship to each other. We have developed viewGene as a flexible tool that takes input from a number of sequence formats and analysis programs (Genbank, FASTA, RepeatMasker, Cross match, BLAST, user-defined data) to construct a sequence reference scaffold that can be viewed through a simple graphical interface. polymorphisms generated from many sources can be added to this scaffold through the same sequence formats, with a variety of options to control what is displayed. Large amounts of polymorphism data can be organized so that patterns and haplotypes can be readily discerned. In our laboratory, viewGene has been used to view annotated genbank records, find nonrepetitive sequence fragments for polymorphism detection, and visualize similarity search results. Manipulation, cross-referencing, and haplotype viewing of snp data are essential for quality assessment and identification of variants associated with genetic disease, and viewGene provides all three of these important functions.

Amino Acid Sequence↗