Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic Structural Variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Evolutionary changes in the fungal carbamoyl-phosphate synthetase small subunit gene and its associated upstream open reading frame.

The Neurospora crassa arg-2 and the Saccharomyces cerevisiae ortholog CPA1 encode the arginine-specific carbamoyl-phosphate synthetase (CPS-A) small subunit. Arginine decreases synthesis of this subunit through the action of a 5' upstream open reading frame in the mRNA that encodes a cis-regulatory element, the arginine attenuator peptide (AAP), which stalls ribosomes in response to arginine. We performed a comparative analysis of the genomic structure and predicted peptide sequence of the AAP and CPS-A small subunit across many fungi. Differences at the genomic level included variation in intron number and position within the AAP and CPS-A coding regions and differences in known regulatory motifs. Although differences exist in AAP sequence, there were three absolutely conserved amino acid residues in the predicted peptide, including an aspartic acid crucial for arginine-dependent regulation of arg-2 and CPA1. A diverged Basidiomycete AAP was shown to retain function as an Arg-specific negative regulator of translation.

Amino Acid Motifs↗

Histone H3 variants and their potential role in indexing mammalian genomes: the "H3 barcode hypothesis".

In the history of science, provocative but, at times, controversial ideas have been put forward to explain basic problems that confront and intrigue the scientific community. These hypotheses, although often not correct in every detail, lead to increased discussion that ultimately guides experimental tests of the principal concepts and produce valuable insights into long-standing questions. Here, we present a hypothesis, the "H3 barcode hypothesis." Hopefully, our ideas will evoke critical discussion and new experimental approaches that bear on general topics, such as nuclear architecture, epigenetic memory, and cell-fate choice. Our hypothesis rests on the central concept that mammalian histone H3 variants (H3.1, H3.2, and H3.3), although remarkably similar in amino acid sequence, exhibit distinct posttranslational "signatures" that create different chromosomal domains or territories, which, in turn, influence epigenetic states during cellular differentiation and development. Although we restrict our comments to H3 variants in mammals, we expect that the more general concepts presented here will apply to other histone variant families in organisms that employ them.

Amino Acid Sequence↗

An invertible element of DNA controls phase variation of type 1 fimbriae of Escherichia coli.

The expression of type 1 fimbriae (pili) of Escherichia coli is turned on and off at the transcriptional level at a high frequency (10(-3) per cell per generation) in a process termed phase variation. Using Southern blot and DNA sequence analysis, we have detected a genomic rearrangement in the switch region immediately upstream of the fimbrial structural gene. This rearrangement involves an invertible 314-base-pair segment of DNA whose alternating orientation apparently results in the on-and-off activation of a promoter that determines the state of fimbrial expression.

Base Sequence↗

ATRX encodes a novel member of the SNF2 family of proteins: mutations point to a common mechanism underlying the ATR-X syndrome.

It was shown recently that mutations of the ATRX gene give rise to a severe, X-linked form of syndromal mental retardation associated with alpha thalassaemia (ATR-X syndrome). In this study, we have characterised the full-length cDNA and predicted structure of the ATRX protein. Comparative analysis shows that it is an entirely new member of the SNF2 subgroup of a superfamily of proteins with similar ATPase and helicase domains. ATRX probably acts as a regulator of gene expression. Definition of its genomic structure enabled us to identify four novel splicing defects by screening 52 affected individuals. Correlation between these and previously identified mutations with variations in the ATR-X phenotype provides insights into the pathophysiology of this disease and the normal role of the ATRX protein in vivo.

Amino Acid Sequence↗

Inversion within the haloalkaliphilic virus phi Ch1 DNA results in differential expression of structural proteins.

The sequence of phi Ch1 contains an open reading frame (int1) in the central part of its genome that belongs to the lambda integrase family of site-specific recombinases. Sequence similarities to known integrases include the highly conserved tetrad R-H-R-Y. The flanking sequences of int1 contain several direct repeats of 30 bp in length (IR-L and IR-R), which are orientated in an inverted direction. Here, we show that a recombination active region exists in the genome of phi Ch1: the number of those repeats, non-homologous regions within the repeat clusters IR-L and IR-R and the orientation of the int1 gene vary in a given virus population. Within this study, we identified circular intermediates, composed of the int1 gene and the inwards orientated repeat regions IR-L and IR-R, which could be involved in the recombination process itself. IR-L and IR-R are embedded within ORF34 and ORF36 respectively. As a consequence of the inversion within this region of phi Ch1, the C-terminal parts of the proteins encoded by ORF34 and 36 are exchanged. Both proteins, expressed in Escherichia coli, interact with specific antisera against whole virus particles, indicating that they could be parts of phi Ch1 virions. Expression of the protein(s) in Natrialba magadii could be detected 98 h after inoculation, which is similar to other structural proteins of phi Ch1. Taken together, the data show that the genome of phi Ch1 contains an invertible region that codes for a recombinase and structural proteins. Inversion of this segment results in a variation of these structural proteins.

Amino Acid Sequence↗

Association between patterns of nucleotide variation across the three fibrinogen genes and plasma fibrinogen levels: the Coronary Artery Risk Development in Young Adults (CARDIA) study.

BACKGROUND: Previous genotype-phenotype association studies of fibrinogen have been limited by incomplete knowledge of genomic sequence variation within and between major ethnic groups in FGB, FGA, and FGG. METHODS: We characterized the linkage disequilibrium patterns and haplotype structure across the human fibrinogen gene locus in European- and African-American populations. We analyzed the association between common polymorphisms in the fibrinogen genes and circulating levels of both 'functional' fibrinogen (measured by the Clauss clotting rate method) and total fibrinogen (measured by immunonephelometry) in a large, multi-center, bi-racial cohort of young US adults. RESULTS: A common haplotype tagged by the A minor allele of the well-studied FGB-455 G/A promoter polymorphism (FGB 1437) was confirmed to be strongly associated with increased plasma fibrinogen levels. Two non-coding variants specific to African-American chromosomes, FGA 3845 A and FGG 5729 G, were each associated with lower plasma fibrinogen levels. In European-Americans, a common haplotype tagged by FGA Thr312Ala and several other variant alleles across the fibrinogen gene locus was strongly associated with decreased fibrinogen levels as measured by functional assay, but not by immunoassay. Overall, common polymorphisms within the three fibrinogen genes explain < 2% of the variability in plasma fibrinogen concentration. CONCLUSIONS: In young adults, fibrinogen multi-locus genotypes are associated with plasma fibrinogen levels. The specific single nucleotide polymorphism and haplotype patterns for these associations differ according to population and also according to phenotypic assay. It is likely that a substantial proportion of the heritable component of plasma fibrinogen concentration is due to genetic variation outside the three fibrinogen genes.

Adolescent↗

Molecular biology and clinical implication of hepatitis C virus.

Hepatitis C virus (HCV) was first described in 1989 as the putative viral agent of non-A non-B hepatitis. It is a member of the Flaviviridae family and has been recognized as the major causative agent of chronic liver disease, including chronic active hepatitis, cirrhosis and hepatocellular carcinoma. HCV is a positive RNA virus with a genome containing approximately 9500 nucleotides. It has an open reading frame that encodes a large polyprotein of about 3000 amino acids and is characterized by extensive genetic diversity. HCV has been classified into at least 6 major genotypes with many subtypes and circulates within an infected individual as a number of closely related but distinct variants known as quasispecies. This article reviews aspects of the molecular biology of HCV and their clinical implication.

3' Untranslated Regions↗

A three-generation approach in biodemography is based on the developmental profiles and the epigenetics of female gametes.

We suggest that there are three premises underlying the need for biodemographic analyses of three-generations: 1.) To describe the structure of the genome, we need to use (apart from mutations) other kinds of heritable changes such as those mediated by facultative elements (variations) and epigenetic alterations. 2.) There are many reasons to analyze individual development and its deviations, such as the biodemographic perspective of fertilization - but also including all long-term intra-generational events of oogenesis and meiosis (beginning with the embryogenesis of the individual's mother - or during the grandmother's pregnancy). 3.) We need to explore the reality that every fertilized egg links - physically and genetically - three successive generations. We focus on genetic and epigenetic events, which start during egg cell lineage determination in F(n-2) gestation and which influence the developmental profile of F(n) generation cohorts. The three-generation approach in epidemiology and biodemography is important so that we might increase our understanding of the effects of environmental forces, such as viral epidemics, and of catastrophes, such as the Chernobyl accident. It is also important for evaluating the processes of senescence and the determinants of human disease.

Animals↗

Alternative splicing generates variants in important functional domains of human slow skeletal troponin T.

We provide the first nucleotide sequence information for the slow isoform of troponin T (TnT). Sequence and hybridization analyses revealed that a single slow TnT gene present in the human genome gives rise to at least two different slow TnT variants by alternative splicing. The observed variations in slow TnT splicing generated major structural differences between the two corresponding slow TnT proteins in a domain that is likely to be involved in critical interactions with troponin C, troponin I, and tropomyosin in the thin filament. Corresponding variations have not been found for fast or for cardiac TnT. The comparison of splicing patterns for fast, cardiac, and slow TnT reveals that the splicing pattern for each isoform is unique. These features raise important questions of why and how all the individual members of the closely related TnT gene family developed such complex but different schemes of alternative splicing to create sets of variant proteins. This unusual familial trait is not known in any other muscle or nonmuscle multigene family.

Amino Acid Sequence↗

Effect of breeding structure on population genetic parameters in Drosophila.

The breeding structure of populations has been neglected in studies of Drosophila, even though Wright and Dobzhansky's pioneering work on the genetics of natural populations was an attempt to tackle what they regarded as an essential factor in evolution. We compared the breeding structure of sympatric populations of D. melanogaster and D. simulans, two sibling species that are widely used in evolutionary studies. We recorded changes in population density and microsatellite variation patterns for 3 years in a temperate environment of southwestern France. Results were distinctively different in the two species. Maximum population levels in summer and in autumn were similar and fluctuated greatly over years, each species being in turn the most abundant. However, genetic data showed that D. melanogaster made up a continuous breeding population in time and space of practically infinite effective size. D. simulans was fragmented into isolates with a local effective size of between 50 and 350 individuals. A consequence of this was that, while a local sample provided a reliable estimate of regional genetic variability in D. melanogaster, a sample from the same area provided an underestimate of this parameter in D. simulans. In practical terms, this means that variations in breeding structure should be accounted for in sampling schemes and in designing evolutionary genetic models. More generally, this suggests the existence of differential reactions to local environments that might contribute to several genomic differences observed between these species.

Animals↗

Sequences homologous to variant antigen mRNA spliced leader in Trypanosomatidae which do not undergo antigenic variation.

Trypanosomes which parasitize mammals have evolved mechanisms to evade immune attack, such as the occupation of 'safe' intracellular sites (for example, Trypanosoma cruzi), or antigenic variation, exemplified by the salivarian trypanosomes (for example, Trypanosoma brucei). Antigenic variation is mediated by sequential expression of single variant surface glycoprotein (VSG) genes, and often involves transposition of the active gene. Every VSG transcript examined shares the same 5' terminal 35-nucleotide leader sequence. In T. brucei, this leader is encoded within a 1.4-kilobase unit tandemly reiterated to form a large array. It is hypothesized that this array is distantly linked to the expressed VSG gene and functions as a multiple promoter of VSG gene transcription, restricting transcription to that gene which, through genomic rearrangement, is placed downstream from the array. Leader and structural gene sequences are presumably juxtaposed by RNA splicing. Here we show that several trypanosomatids, both those which undergo antigenic variation (Trypanosoma congolense and Trypanosoma vivax) and those which do not (T. cruzi and Leptomonas collosoma), contain reiterated sequences homologous to the T. brucei spliced leader (SL). These results suggest that the SL, although utilized in VSG gene expression, is an ancestral sequence also used in the expression of other trypanosomatid genes.

Animals↗

Deficiencies of human complement component C4A and C4B and heterozygosity in length variants of RP-C4-CYP21-TNX (RCCX) modules in caucasians. The load of RCCX genetic diversity on major histocompatibility complex-associated disease.

The complement component C4 genes located in the major histocompatibility complex (MHC) class III region exhibit an unusually complex pattern of variations in gene number, gene size, and nucleotide polymorphism. Duplication or deletion of a C4 gene always concurs with its neighboring genes serine/threonine nuclear protein kinase RP, steroid 21-hydroxylase (CYP21), and tenascin (TNX), which together form a genetic unit termed the RCCX module. A detailed molecular genetic analysis of C4A and C4B and RCCX modular arrangements was correlated with immunochemical studies of C4A and C4B protein polymorphism in 150 normal Caucasians. The results show that bimodular RCCX has a frequency of 69%, whereas monomodular and trimodular RCCX structures account for 17.0 and 14.0%, respectively. Three quarters of C4 genes harbor the endogenous retrovirus HERV-K(C4). Partial deficiencies of C4A and C4B, primarily due to gene deletions and homoexpression of C4A proteins, have a combined frequency of 31.6%. This is probably the most common variation of gene dosage and gene size in human genomes. The seven RCCX physical variants create a great repertoire of haplotypes and diploid combinations, and a heterozygosity frequency of 69.4%. This phenomenon promotes the exchange of genetic information among RCCX constituents that is important in homogenizing the structural and functional diversities of C4A and C4B proteins. However, such length variants may cause unequal, interchromosomal crossovers leading to MHC-associated diseases. An analyses of the RCCX structures in 22 salt-losing, congenital adrenal hyperplasia patients revealed a significant increase in the monomodular structure with a long C4 gene linked to the pseudogene CYP21A, and bimodular structures with two CYP21A, which are likely generated by recombinations between heterozygous RCCX length variants.

Adrenal Hyperplasia, Congenital↗

Multiple QTLs influence variation in paraoxonase 1 activity in Mexican Americans.

Paraoxonase 1 (PON1), a high-density-lipoprotein-associated enzyme known to protect against cellular damage from toxic agents, may also have antioxidant properties. Although the importance of the influence of the PON1 structural locus on chromosome 7q21-22 for variation in the concentration and activity of the enzyme is well-documented, the contribution of other loci is poorly understood. Based on the recent observations of at least one additional quantitative trait locus (QTL) for PON1 activity in pedigreed baboons, we conducted a whole-genome linkage screen for QTLs other than the PON1 structural locus that may influence PON1 activity in humans. We measured PON1 activity in frozen serum for 1,406 individuals in more than 40 extended pedigrees from the San Antonio Family Heart Study (SAFHS). We used a maximum-likelihood-based variance decomposition approach implemented in SOLAR to test for QTLs that may influence PON1 activity. In addition to a QTL for which we detected the strongest, significant evidence (LOD = 31.41) at or near the PON1 structural locus on chromosome 7q21-22, we also localized at least one additional significant QTL on chromosome 12 (LOD = 3.56). Furthermore, we detected suggestive evidence for two more PON-related QTLs on chromosomes 17 and 19. We have provided evidence that other genes, in addition to the well-known ones on chromosome 7, play a role in influencing normal variation in PON1 activity.

Adult↗

Genetic and haplotype diversity among wild-derived mouse inbred strains.

With the completion of the mouse genome sequence, it is possible to define the amount, type, and organization of the genetic variation in this species. Recent reports have provided an overview of the structure of genetic variation among classical laboratory mice. On the other hand, little is known about the structure of genetic variation among wild-derived strains with the exception of the presence of higher levels of diversity. We have estimated the sequence diversity due to substitutions and insertions/deletions among 20 inbred strains of Mus musculus, chosen to enable interpretation of the molecular variation within a clear evolutionary framework. Here, we show that the level of sequence diversity present among these strains is one to two orders of magnitude higher than the level of sequence diversity observed in the human population, and only a minor fraction of the sequence differences observed is found among classical laboratory strains. Our analyses also demonstrate that deletions are significantly more frequent than insertions. We estimate that 50% of the total variation identified in M. musculus may be recovered in intrasubspecific crosses. Alleles at variants positions can be classified into 164 strain distribution patterns, a number exceeding those reported and predicted in panels of classical inbred strains. The number of strains, the analysis of multiple loci scattered across the genome, and the mosaic nature of the genome in hybrid and classical strains contribute to the observed diversity of strain distribution patterns. However, phylogenetic analyses demonstrate that ancient polymorphisms that segregate across species and subspecies play a major role in the generation of strain distribution patterns.

Animals↗

Structure and evolution of Paramecium hemoglobin genes.

Hemoglobin (Hb) genes have been cloned from three different species of ciliated protists, P. multimicronucleatum, P. triaurelia and P. jenningsi. Southern blotting of the genomic DNAs using the P. caudatum Hb cDNA showed both intraspecies variation in different stocks of P. caudatum and interspecies variation within the genus Paramecium. The isolated Hb genes were composed of 118, 117 and 117 codons, and interrupted by a short intron with 27, 29 and 29 bp at the same position, in P. multimicronucleatum, P. triaurelia and P. jenningsi, respectively. This suggests that the one-intron and two-exon structure has been conserved in the Hb genes in this genus. The amino acid sequences of the Paramecium Hbs were more than 87% identical to one another and homologous to those from the other ciliated protists Tetrahymena thermophila and T. pyriformis, the green alga Chlamydomonas eugametos, and the cyanobacterium Nostoc commune Hbs, all of which consist of about 120 amino acid residues (120-aa group). In particular, the amino acid sequences of the P. triaurelia and P. jenningsi Hbs were the same, although there were 20 nucleotide differences between the coding regions in the two genes. A maximum likelihood inference as to the phylogenetic relationships among these genes suggests that the Paramecium Hbs genes have evolved more rapidly than the other genes in the 120-aa group, and that P. triaurelia and P. genningsi are sibling species and the P. aurelia complex became a small cell after it separated from P. jenningsi.

Amino Acid Sequence↗

Mitochondrial DNA structure and expression in specialized subtypes of mammalian striated muscle.

Mitochondrial DNA (mt DNA) in cells of vertebrate organisms can assume an unusual triplex DNA structure known as the displacement loop (D loop). This triplex DNA structure forms when a partially replicated heavy strand of mtDNA (7S mtDNA) remains annealed to the light strand, displacing the native heavy strand in this region. The D-loop region contains the promoters for both heavy- and light-strand transcription as well as the origin of heavy-strand replication. However, the distribution of triplex and duplex forms of mtDNA in relation to respiratory activity of mammalian tissues has not been systematically characterized, and the functional significance of the D-loop structure is unknown. In comparisons of specialized muscle subtypes within the same species and of the same muscle subtype in different species, the relative proportion of D-loop versus duplex forms of mtDNA in striated muscle tissues of several mammalian species demonstrated marked variation, ranging from 1% in glycolytic fast skeletal fibers of the rabbit to 65% in the mouse heart. There was a consistent and direct correlation between the ratio of triplex to duplex forms of mtDNA and the capacity of these tissues for oxidative metabolism. The proportion of D-loop forms likewise correlated directly with mtDNA copy number, mtRNA abundance, and the specific activity of the mtDNA (gamma) polymerase. The D-loop form of mtDNA does not appear to be transcribed at greater efficiency than the duplex form, since the ratio of mtDNA copy number to mtRNA was unrelated to the proportion of triplex mtDNA genomes. However, tissues with a preponderance of D-loop forms tended to express greater levels of cytochrome b mRNA relative to mitochondrial rRNA transcripts, suggesting that the triplex structure may be associated with variations in partial versus full-length transcription of the heavy strand.

Animals↗

PubMind: literature-based genetic variant extraction and functional annotation using large language models.

Biomedical literature contains extensive functional knowledge on genetic variants, but much remains inaccessible in unstructured text. Existing resources such as ClinVar and HGMD remain limited by coverage, submission bias, update frequency, and sparse annotation. We develop PubMind, an artificial intelligence (AI)&#xa0;framework that uses large language models (LLMs)&#xa0;to triage and extract variant-function-disease associations and supporting evidence from biomedical text. PubMind captures single-nucleotide, copy-number, structural, and gene-fusion variants, and normalizes records to genomic and transcriptomic coordinates. Benchmarking shows >90% accuracy for variant recognition and 99% precision for disease extraction. Applied to >41 million PubMed abstracts and >5 million full-text articles, PubMind generates PubMind-DB, a database of ~1.3 million unique variants with contextual annotations, accessible via web interface and API. Only ~10% of PubMind variants overlap with ClinVar, and >80% of them&#xa0;show concordant pathogenicity labels. PubMind transforms unstructured biomedical text into structured genomic knowledge, advancing variant interpretation for precision medicine.

Large Language Models↗