Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic Structural Variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,639 records · Page 91Linked to original sources

Genetic structure and epidemiology of Ascaris populations: patterns of host affiliation in Guatemala.

In Guatemalan villages people commonly rear pigs, and both hosts may be infected with Ascaris. This study was designed to ask whether both humans and pigs are potential hosts in a single parasite transmission cycle in such villages, or alternatively, if there are two separate transmission cycles, one involving pigs and one involving human hosts. Parasites were collected from both host species from locations in the north and south of Guatemala. Allelic variation in the nuclear genome of Ascaris was measured using enzyme electrophoresis, while mitochondrial DNA (mtDNA) sequence variation was quantified using restriction mapping. Low levels of enzyme polymorphism were found in Ascaris, but allele frequencies at two loci, mannose phosphate isomerase and esterase, suggest that there is little gene exchange between parasite populations from humans and pigs. MtDNA haplotypes fall into two distinct clusters which differ in sequence by 3-4%; the two clusters broadly correspond to worms collected from humans and those collected from pigs. However, some parasites collected from humans have mtDNA characteristic of the 'pig Ascaris' haplotype cluster, while some parasites collected from pigs have mtDNA characteristic of the 'human Ascaris' haplotype cluster. These shared haplotypes are unlikely to represent contemporary cross-infection events. Patterns of phylogenetic similarity and geographical distribution of these haplotypes suggest, instead, that they are the result of two historical introgressions of mtDNA between the two host-associated Ascaris populations. The results clearly demonstrate that Ascaris from humans and pigs are involved in separate transmission cycles in Guatemala.

Alleles↗

Assigning genomic sequences to CATH.

We report the latest release (version 1.6) of the CATH protein domains database (http://www.biochem.ucl. ac.uk/bsm/cath ). This is a hierarchical classification of 18 577 domains into evolutionary families and structural groupings. We have identified 1028 homo-logous superfamilies in which the proteins have both structural, and sequence or functional similarity. These can be further clustered into 672 fold groups and 35 distinct architectures. Recent developments of the database include the generation of 3D templates for recognising structural relatives in each fold group, which has led to significant improvements in the speed and accuracy of updating the database and also means that less manual validation is required. We also report the establishment of the CATH-PFDB (Protein Family Database), which associates 1D sequences with the 3D homologous superfamilies. Sequences showing identifiable homology to entries in CATH have been extracted from GenBank using PSI-BLAST. A CATH-PSIBLAST server has been established, which allows you to scan a new sequence against the database. The CATH Dictionary of Homologous Superfamilies (DHS), which contains validated multiple structural alignments annotated with consensus functional information for evolutionary protein superfamilies, has been updated to include annotations associated with sequence relatives identified in GenBank. The DHS is a powerful tool for considering the variation of functional properties within a given CATH superfamily and in deciding what functional properties may be reliably inherited by a newly identified relative.

Amino Acid Sequence↗

Single nucleotide polymorphisms in TNFSF15 confer susceptibility to Crohn's disease.

The inflammatory bowel diseases (IBDs), Crohn's disease (CD) and ulcerative colitis, are chronic inflammatory disorders of the digestive tract. The pathogenesis of IBD is complicated, and it is widely accepted that immunologic, environmental and genetic components contribute to its etiology. To identify genetic susceptibility factors in CD, we performed a genome-wide association study in Japanese patients and controls using nearly 80,000 gene-based single nucleotide polymorphism (SNP) markers and investigated the haplotype structure of the candidate locus in Japanese and European patients. We identified highly significant associations (P = 1.71 x 10(-14) with odds ratio of 2.17) of SNPs and haplotypes within the TNFSF15 (the gene encoding tumor necrosis factor superfamily, member 15) genes in Japanese CD patients. The association was confirmed in the study of two European IBD cohorts. Interestingly, a core TNFSF15 haplotype showing association with increased risk to the disease was common in the two ethnic groups. Our results suggest that the genetic variations in the TNFSF15 gene contribute to the susceptibility to IBD in the Japanese and European populations.

Adult↗

Structure of the elastin gene.

The isolation and characterization of cDNAs encompassing the full length of chicken, cow, rat and human elastin mRNA have led to the elucidation of the primary structure of the respective tropoelastins. Large segments of the sequence are conserved but there are also considerable variations which range in extent from relatively small alterations, such as conservative amino acid substitutions, to variation in the length of hydrophobic segments and largescale deletions and insertions. In general, smaller differences are found among mammalian tropoelastins and greater ones between chicken and mammalian tropoelastins. Although only a single elastin gene is found per haploid genome, the primary transcript is subject to considerable alternative splicing, resulting in multiple tropoelastin isoforms. Functionally distinct hydrophobic and cross-link domains of the protein are encoded in separate exons which alternate in the gene. The introns of the human gene are rich in Alu repetitive sequences, which may be the site of recombinational events, and there are also several dinucleotide repeats, which may exhibit polymorphism and, therefore, be effective genetic markers. The 5' flanking region is G+C rich and contains potential binding sites for numerous modulating factors, but no TATA box or functional CAAT box. The basic promoter is contained within a 136 bp segment and transcription is initiated at multiple sites. These findings suggest that the regulation of elastin gene expression is complex and takes place at several levels.

Alternative Splicing↗

Restriction map of Chinese hamster mitochondrial DNA containing replication coordinates: comparison with Syrian hamster mitochondrial genome.

A precise physical map, containing the structurally and operationally defined D-loop origin, terminal region, and direction of heavy-strand replication, has been constructed for mitochondrial DNA (mtDNA) from ovary (CHO-KI) and lung cells of Chinese hamster (Cricetulus griseus, 2 N = 22), and compared with our previously established genome coordinates for mtDNA from Syrian hamster (Mesocricetus auratus, 2 N = 44). All four HpaI sites in Cricetulus are conserved in Mesocricetus (8 sites). Extensive variation exists for hexanucleotides cleaved by EcoRI, HindIII, PstI, KpnI and BamHI. Sequence divergence between Chinese and Syrian hamster mtDNAs, as reflected from analysis of the mapped recognition sites for these six endonucleases, is estimated as 5-9% base substitutions. mtDNAs from both hamster and several other mammalian species contain a commonly conserved HpaI site in the region of light strand initiation.

Animals↗

Estimating genome conservation between crop and model legume species.

Legumes are simultaneously one of the largest families of crop plants and a cornerstone in the biological nitrogen cycle. We combined molecular and phylogenetic analyses to evaluate genome conservation both within and between the two major clades of crop legumes. Genetic mapping of orthologous genes identifies broad conservation of genome macrostructure, especially within the galegoid legumes, while also highlighting inferred chromosomal rearrangements that may underlie the variation in chromosome number between these species. As a complement to comparative genetic mapping, we compared sequenced regions of the model legume Medicago truncatula with those of the diploid Lotus japonicus and the polyploid Glycine max. High conservation was observed between the genomes of M. truncatula and L. japonicus, whereas lower levels of conservation were evident between M. truncatula and G. max. In all cases, conserved genome microstructure was punctuated by significant structural divergence, including frequent insertion/deletion of individual genes or groups of genes and lineage-specific expansion/contraction of gene families. These results suggest that comparative mapping may have considerable utility for basic and applied research in the legumes, although its predictive value is likely to be tempered by phylogenetic distance and genome duplication.

Crops, Agricultural↗

Nucleoid structure and partition in Methanococcus jannaschii: an archaeon with multiple copies of the chromosome.

We measured different cellular parameters in the methanogenic archaeon Methanococcus jannaschii. In exponential growth phase, the cells contained multiple chromosomes and displayed a broad variation in size and DNA content. In most cells, the nucleoids were organized into a thread-like network, although less complex structures also were observed. During entry into stationary phase, chromosome replication continued to termination while no new rounds were initiated: the cells ended up with one to five chromosomes per cell with no apparent preference for any given DNA content. Most cells in stationary phase contained more than one genome equivalent. Asymmetric divisions were detected in stationary phase, and the nucleoids were found to be significantly more compact than in exponential phase.

Cell Cycle↗

Human mutations in glucose 6-phosphate dehydrogenase reflect evolutionary history.

Glucose 6-phosphate dehydrogenase (G6PD) is a cytosolic enzyme encoded by a housekeeping X-linked gene whose main function is to produce NADPH, a key electron donor in the defense against oxidizing agents and in reductive biosynthetic reactions. Inherited G6PD deficiency is associated with either episodic hemolytic anemia (triggered by fava beans or other agents) or life-long hemolytic anemia. We show here that an evolutionary analysis is a key to understanding the biology of a housekeeping gene. From the alignment of the amino acid (aa) sequence of 52 glucose 6-phosphate dehydrogenase (G6PD) species from 42 different organisms, we found a striking correlation between the aa replacements that cause G6PD deficiency in humans and the sequence conservation of G6PD: two-thirds of such replacements are in highly and moderately conserved (50-99%) aa; relatively few are in fully conserved aa (where they might be lethal) or in poorly conserved aa, where presumably they simply would not cause G6PD deficiency. This is consistent with the notion that all human mutants have residual enzyme activity and that null mutations are lethal at some stage of development. Comparing the distribution of mutations in a human housekeeping gene with evolutionary conservation is a useful tool for pinpointing amino acid residues important for the stability or the function of the corresponding protein. In view of the current explosive increase in full genome sequencing projects, this tool will become rapidly available for numerous other genes.

Amino Acid Sequence↗

G4SNVHunter: An R/Bioconductor Package for Evaluating SNV-Induced Disruption of G-Quadruplex Structures Leveraging the G4Hunter Algorithm.

G-quadruplexes (G4s) are nucleic acid secondary structures with important regulatory functions. Single-nucleotide variants (SNVs), one of the most common forms of genetic variation, can potentially impact the formation of G4 structures if they occur within G4 regions. However, there is currently a lack of software tools specifically designed to assess such effects. Here, we present an R/Bioconductor package named G4SNVHunter, which enables rapid detection of variants that may disrupt G4 structures. This tool, based on the core principles of the G4Hunter algorithm, can provide precise quantitative assessment of the propensity for G4 formation within genomic sequences. Specialized experimental methods can then be designed based on the results provided by G4SNVHunter to further verify the specific functions of the affected G4 structures, facilitating deeper insights into the biological impacts of genetic variants from the perspective of G4 structures. To showcase the functionality of the G4SNVHunter package, we analyzed the Neandertal and Denisovan archaic introgressed variants detected by the Sprime software, and identified approximately 5,800 variants located within G4 regions, among which around 230 may impair G4 structure formation propensity. The source code for the G4SNVHunter package has been publicly released under the MIT license at https://github.com/rongxinzh/G4SNVHunter and https://bioconductor.org/packages/devel/bioc/html/G4SNVHunter.html.

G-Quadruplexes↗

[The genetic map of HLA].

The major histocompatibility complex (MHC) gene products are known to play a fundamental role in foreign antigen presentation to the cellular immune system. The human major histocompatibility complex (HLA) spans about 4 million base pairs (4 Mbp) of DNA at chromosome position 6p21.3 and is one of the most intensively studied regions of the human genome, containing over 70 known genes. The HLA can be divided up into three regions: the class I sequences at the telomeric end of the complex; the class II loci at the centromeric end; and, between these, the class III genes including those for the complement components. The analysis of HLA has focused primarily on their structure and function. However, the quantitative differences in gene expression are also important, since quantitative variations of cell surface HLA influences immunoregulation and the expression of immunologically mediated diseases. Recently, biochemical and genetic studies have revealed the existence of tissue-specific cis-acting regulatory gene elements and trans-acting protein factors, capable of binding to the cis-acting regions, that govern HLA gene expression.

Chromosome Mapping↗

Strain diversity and conserved genome elements in Strawberry mild yellow edge virus.

The complete nucleotide sequence of Strawberry mild yellow edge virus isolate D74, and the sequences of a 878 nt region of the coat protein and flanking regions of twenty three isolates of SMYEV were obtained and analysed. The full sequence of the aphid transmissible strain D74 was deduced and found to have an 86% sequence identity to the non-transmissible Agrobacterium infectious MY18 strain. In contrast to isolate MY18 the 5' terminal nucleotides (GAAAAC) of D74 are typical of those from other potexviruses. However, both MY18 and D74 have a non-AUG initiation codon for ORF2 encoding the triple gene block protein 1 (TGB1), and an overlapping TGB3 and coat protein (CP), features unique in the Potexvirus genus. Other conserved features of the genome, including stem-loop structures in the untranslated regions, and motifs common to related viruses are described. The previously postulated ORF6 of MY18 is absent in twenty of the isolates sequenced, including D74. Phylogenetic analysis places all isolates in one of three distinct groups/strains named here I (type-D74), II (type-9Redland), and III (type-MY18), with the majority of isolates, including all European isolates tested, belonging to strain I (type-D74).

Amino Acid Motifs↗

IS231-MIC231 elements from Bacillus cereus sensu lato are modular.

Summary IS231A was originally discovered in Bacillus thuringiensis as a typical 1.6 kb insertion sequence (IS) displaying 20 bp inverted repeats (IR) flanking a transposase gene. A first major variation of this canonical organization was found in MIC231A1. This mobile insertion cassette (MIC), delineated by IS231A-related extremities, contained an active d-stereospecific endopeptidase (adp) gene instead of a transposase. Interestingly, it was shown that MIC231A1 can be mobilized in trans by the IS231A transposase. In this paper, we show that this family of IS231-MIC231 elements can be extended to a broad range of related entities displaying higher levels of structural complexity. Several IS231A-like elements contained, upstream of their transposase gene, passenger genes coding for putative antibiotic resistances or regulatory factors. Furthermore, the diversity of the MIC231 elements ranged from empty cassettes to structures carrying up to three passenger genes. Among these, MIC231V carried, in addition to an adp gene, an active fosfomycin resistance determinant. In vivo transposition assays showed that MIC231V is also trans-activated by the IS231A transposase. These results lend further support to the potential contribution of these modular mobile elements to the genome plasticity of the Bacillus cereus/B. thuringiensis group.

Bacillus cereus↗

Distribution of tandem repeat polymorphism within minisatellite MS621 (D5S110).

The minisatellite MS621 (D5S110) is a highly polymorphic tandem repeat locus which maps to the distal region of human chromosome 5p. Major repeat unit variants at D5S110 differ not by base substitutions but by differences in the repetition of an 11 bp sequence motif found within each repeat. The two major classes of repeat unit thus contain two ('dimeric' = D-type) or three ('trimeric' = T-type) copies of this short motif. Mapping the distribution of these D- and T-type repeat units within alleles has allowed the analysis of the structural basis of allelic variation and of one de novo mutation. In contrast to previous studies of some other highly polymorphic minisatellites, this analysis provided no clear evidence for polarized variability at D5S110.

Animals↗

Comparative genomics of the keratin-associated protein (KAP) gene clusters in human, chimpanzee, and baboon.

We have previously identified a cluster of 16 genes that encode hair-specific proteins, called keratin-associated proteins (KAPs), located on human Chromosome (Chr) 21q22.3. Here, we have identified similar KAP gene clusters in two primates, chimpanzee and baboon. DNA sequence comparison revealed the common cluster structure consisting of 16 KAP genes for these three primates, but a significant difference was found in the baboon. Baboon possesses a new KAP gene not found in human and chimpanzee, whereas one KAP gene ( KRTAP18.12) that exists in human and chimpanzee was lost in baboon, making no change in the total number of KAP genes. Interestingly, the sequence for coding regions are highly variable among species owing to insertions and deletions, resulting in variation of gene size. On the contrary, the sequences for the 5' upstream region are highly conserved among species. These findings suggest that the ancestral KAP gene cluster was composed of 17 genes before the divergence of Old World monkeys (baboon) to the anthropoid (human and chimpanzee).

Amino Acid Sequence↗

Challenging the dogma: the hidden layer of non-protein-coding RNAs in complex organisms.

The central dogma of biology holds that genetic information normally flows from DNA to RNA to protein. As a consequence it has been generally assumed that genes generally code for proteins, and that proteins fulfil not only most structural and catalytic but also most regulatory functions, in all cells, from microbes to mammals. However, the latter may not be the case in complex organisms. A number of startling observations about the extent of non-protein-coding RNA (ncRNA) transcription in the higher eukaryotes and the range of genetic and epigenetic phenomena that are RNA-directed suggests that the traditional view of the structure of genetic regulatory systems in animals and plants may be incorrect. ncRNA dominates the genomic output of the higher organisms and has been shown to control chromosome architecture, mRNA turnover and the developmental timing of protein expression, and may also regulate transcription and alternative splicing. This paper re-examines the available evidence and suggests a new framework for considering and understanding the genomic programming of biological complexity, autopoietic development and phenotypic variation.

Alternative Splicing↗

A novel type of EWS-CHOP fusion gene in myxoid liposarcoma.

The cytogenetic hallmark of myxoid type and round cell type liposarcoma consists of reciprocal translocation of t(12;16)(q13;p11) and t(12;22)(q13;q12), which results in fusion of TLS/FUS and CHOP, and EWS and CHOP, respectively. Nine structural variations of the TLS/FUS-CHOP chimeric transcript have been reported, however, only two types of EWS-CHOP have been described. We describe here a case of myxoid liposarcoma containing a novel EWS-CHOP chimeric transcript and identified the breakpoint occurring in intron 13 of EWS. Reverse transcription-polymerase chain reaction and direct sequence showed that exon 13 of EWS was in-frame fused to exon 2 of CHOP. Genomic analysis revealed that the breaks were located in intron 13 of EWS and intron 1 of CHOP.

Adult↗

A survey of genetic diversity and reproductive biology of Puya raimondii (Bromeliaceae), the endangered queen of the Andes.

Puya raimondii Harms is an outstanding giant rosette bromeliad found solely around 4000 m above sea level in the Andes. It flowers at the end of an 80 - 100-year or even longer life cycle and yields an enormous (4 - 6 m tall) spike composed of from 15,000 to 20,000 flowers. It is endemic and currently endangered, with populations distributed from Peru to the north of Bolivia. A genomic DNA marker-based analysis of the genetic structure of eight populations representative of the whole distribution of P. raimondii in Peru is reported in this paper. As few as 14 genotypes out of 160 plants were detected. Only 5 and 18 of the 217 AFLP marker loci screened were polymorphic within and among these populations, respectively. Four populations were completely monomorphic, each of the others displayed only one to three polymorphic loci. Less than 4 % of the total genomic variation was within populations and genetic similarity among populations was as high as 98.3 %. Results for seven cpSSR marker loci were in agreement with the existence of a single progenitor. Flow cytometry of seed nuclear DNA content and RAPD marker segregation analysis of progeny plantlets demonstrated that the extremely uniform genome of P. raimondii populations is not compatible with agamospermy (apomixis), but consistent with an inbreeding reproductive strategy. There is an urgent need for a protection programme to save not only this precious, isolated species, but also the unique ecosystem depending on it.

Bromeliaceae↗

Unifying multimodal single-cell data with a mixture-of-experts β-variational autoencoder framework.

Multimodal single-cell assays profile complementary layers of cell state, but integration is complicated by modality mismatch, sparsity, and uneven cohort coverage. Here, we present Unified Variational Inference (UniVI), a scalable mixture-of-experts β-variational autoencoder that learns a shared latent space while preserving modality-specific structure. UniVI couples modality-specific encoders/decoders with a shared latent prior and a symmetric cross-modal alignment objective, enabling consistent integration of paired measurements without curated feature-link graphs or preannotated reference atlases; optional supervised heads can be added when labels are available. Across paired RNA-protein (CITE-seq) and RNA-chromatin (10x Genomics Multiome, SHARE-seq) data spanning human PBMCs and mouse back skin-a nonhematopoietic tissue with continuous differentiation hierarchies-UniVI produces coherent embeddings, improves label transfer, and enables cross-modal reconstruction and denoising. Extending to trimodal measurements, UniVI maintains robust three-way alignment among RNA, chromatin accessibility, and surface proteins (TEA-seq), and accommodates DNA methylation in a paired scNMT-seq mouse gastrulation proof-of-concept under beta-binomial likelihoods. Performance degrades gracefully under severe cell type imbalance and in the presence of modality-exclusive populations. In an acute myeloid leukemia mosaic design, a paired RNA-protein bridge anchors independent RNA-only and protein+genotype cohorts, revealing genotype-associated neighborhoods that sharpen with mutation-aware fine-tuning. UniVI thus provides a flexible, interpretable framework for multimodal integration across paired, trimodal, and mosaic study designs and supports practical reference-to-query projection in partially observed studies.

Journal Article↗