Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic Structural Variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,657 records · Page 92Linked to original sources

Unifying multimodal single-cell data with a mixture-of-experts β-variational autoencoder framework.

Multimodal single-cell assays profile complementary layers of cell state, but integration is complicated by modality mismatch, sparsity, and uneven cohort coverage. Here, we present Unified Variational Inference (UniVI), a scalable mixture-of-experts β-variational autoencoder that learns a shared latent space while preserving modality-specific structure. UniVI couples modality-specific encoders/decoders with a shared latent prior and a symmetric cross-modal alignment objective, enabling consistent integration of paired measurements without curated feature-link graphs or preannotated reference atlases; optional supervised heads can be added when labels are available. Across paired RNA-protein (CITE-seq) and RNA-chromatin (10x Genomics Multiome, SHARE-seq) data spanning human PBMCs and mouse back skin-a nonhematopoietic tissue with continuous differentiation hierarchies-UniVI produces coherent embeddings, improves label transfer, and enables cross-modal reconstruction and denoising. Extending to trimodal measurements, UniVI maintains robust three-way alignment among RNA, chromatin accessibility, and surface proteins (TEA-seq), and accommodates DNA methylation in a paired scNMT-seq mouse gastrulation proof-of-concept under beta-binomial likelihoods. Performance degrades gracefully under severe cell type imbalance and in the presence of modality-exclusive populations. In an acute myeloid leukemia mosaic design, a paired RNA-protein bridge anchors independent RNA-only and protein+genotype cohorts, revealing genotype-associated neighborhoods that sharpen with mutation-aware fine-tuning. UniVI thus provides a flexible, interpretable framework for multimodal integration across paired, trimodal, and mosaic study designs and supports practical reference-to-query projection in partially observed studies.

Journal Article↗

TCR gene polymorphisms and autoimmune disease.

Autoimmunity may result from abnormal regulation within the immune system. As the T cell is the principal regulator of the immune system and its normal function depends on immune recognition or self/non-self discrimination, abnormalities of the idiotypic T-cell receptor (TCR) may be one cause of autoimmune disease. The TCR is a clonally distributed, cell-surface heterodimer which binds peptide antigen when complexed with HLA molecules. In order to recognize the variety of antigens it may possibly encounter, the TCR, by necessity, is a diverse structure. As with immunoglobulin, it is the variable domain of the TCR which interacts with antigen and exhibits the greatest amount of amino acid variability. The underlying genetic basis for this structural diversity is similar to that described for immunoglobulin, with TCR diversity relying on the somatic recombination, in a randomly imprecise manner, of smaller gene segments to form a functional gene. There are a large number of gene segments to choose from (particularly the TCRAV, TCRAJ and TCRBV gene segments) and some of these also exhibit allelic variation. Finally, polymorphisms in non-coding regions of TCR genes, leading to biased recombination or expression, are also beginning to be recognized. All these factors contribute to the polymorphic nature of the TCR, in terms of both structure and repertoire formation. It follows that inherited abnormalities in either coding or regulatory regions of TCR genes may predispose to aberrant T-cell function and autoimmune disease. This review will outline the genomic organization of the TCR genes, the genetic mechanisms responsible for the generation of diversity, and the results of investigations into the association between germline polymorphisms and autoimmune disease.

Alleles↗

LTR retrotransposons and flowering plant genome size: emergence of the increase/decrease model.

Long Terminal Repeat (LTR) retrotransposons are ubiquitous components of plant genomes. Because of their copy-and-paste mode of transposition, these elements tend to increase their copy number while they are active. In addition, it is now well established that the differences in genome size observed in the plant kingdom are accompanied by variations in LTR retrotransposon content, suggesting that LTR retrotransposons might be important players in the evolution of plant genome size, along with polyploidy. The recent availability of large genomic sequences for many crop species has made it possible to examine in detail how LTR retrotransposons actually drive genomic changes in plants. In the present paper, we provide a review of the recent publications that have contributed to the knowledge of plant LTR retrotransposons, as structural components of the genomes, as well as from an evolutionary genomic perspective. These studies have shown that plant genomes undergo genome size increases through bursts of retrotransposition, while there is a counteracting process that tends to eliminate the transposed copies from the genomes. This process involves recombination mechanisms that occur either between the LTRs of the elements, leading to the formation of solo-LTRs, or between direct repeats anywhere in the sequence of the element, leading to internal deletions. All these studies have led to the emergence of a new model for plant genome evolution that takes into account both genome size increases (through retrotransposition) and decreases (through solo-LTR and deletion formation). In the conclusion, we discuss this new model and present the future prospects in the study of plant genome evolution in relation to the activity of transposable elements.

Flowers↗

[Complete nucleotide sequence analysis of Bombyx mori Densonucleosis virus type 3 VD2 (China isolate)].

The genome of Bombyx mori densonucleosis virus type 3 (China isolate) contains two kinds different single-stranded linear DNA molecules (VD1, VD2). The VD2 was purified and cloned into the pUC119 vector, and the complete nucleotide sequence of VD2 was determined. Sequence analysis showed that the VD2 genome sequence consisted of 6022 nts, including inverted terminal repeats (ITRs) of 524 nts. In the viral genome, two major open reading frames (ORF1 and ORF2) in the plus strand and one minor ORF (ORF3) in the complementary strand were identified. Computer analysis suggested the plus stand ORF1 and the minus strand ORF3 most likely encode the major nonstructural protein, while the plus stand ORF2 most likely encodes the major structural protein. Comparing the complete genome sequence of BmDNV-3 VD2 with that of BmDNV-2 VD2 (Yamanashi isolate) demonstrated that they shared 97.7% genome sequence in the VD2 region, with substitutions of 132 nucleotides, deletions of 11 nucleotides and insertions of 2 nucleotides in the VD2 region of BmDNV-3. The results suggested that BmDNV-3 is closely related to BmDNV-2, but with some differences, giving a better understanding about the variation of the viruses and providing clues to their evolution.

Animals↗

Mg and Mc: mutations within the amino-terminal region of glycophorin A.

M and N are the two common ("normal") alleles at the MN locus of the MNSs blood group system. The antigens M and N that they determine are located within the amino-terminal region of glycophorin A. In the serologically active and glycosylated (*) fragment of glycophorin AN the sequence is Leu-Ser*-Thr*-Thr*-Glu-, and in that of glycophorin AM it is Ser-Ser*-Thr*-Thr*-Gly-. Mg and Mc are very rare ("variant") alleles of M and N; as to the corresponding antigens, Mg is serologically quite distinct from M and N, while Mc is a compound of both. Erythrocytes of genotypes MgN, MgM, MgMg, and McM, which were the object of the present study, contain normal amounts of glycophorin A in their membrane. In glycophorin AMg the amino-terminal sequence is related to that of glycophorin AN by substitution of asparagine for threonine in position 4, and it is nonglycosylated: Leu-Ser-Thr-Asn-Glu-. The corresponding structure of glycophorin AMc is Ser-Ser*-Thr*-Thr*-Glu-; it is thus closely related to that of glycophorin AN and AM, by substitution of the amino acids in positions 1 or 5, respectively. All of these substitutions can be explained by single base changes. The distinctions in chemical structure not only confirm the location of M and N in this region of glycophorin A, because they are the only differences observed, but also indicate, because they are correlated with the distinctions in antigenic specificity, that M and N are structural genes coding for amino acid sequences. The finding that Mc contains structural features of both M and N suggests that these two forms of glycophorin A have evolved from a common ancestral gene by single base substitutions at sites in the genome coding for amino acids in positions 1 and 5 of the sequence. Carbohydrate structures, however, are also necessary for full expression of antigens M and N. Glycosylation during biosynthesis of residues within the polypeptide appears to depend on a particular protein structure.

Alleles↗

Physical map of SeMNPV baculovirus DNA: an AcMNPV genomic variant.

A physical map of the plaque-purified SeMNPV-25 baculovirus DNA was constructed with HindIII, EcoRI, PstI, SstII, BstEII, BamHI, KpnI, and SmaI by multiple-enzyme digestion and DNA-DNA hybridization. The orientation of the physical map was to the AcMNPV polyhedrin gene, and a total of 104 restriction sites were ordered. The genome size was 131.89 kilobase pairs (87.05 X 10(6) Da). The physical maps of SeMNPV-25 DNA and AcMNPV-E2 DNA were compared, and they were similar except in four regions. Since the differences between the physical maps of SeMNPV-25 and AcMNPV-E2 were reconciled, the BstEII and SstII fragments for AcMNPV-E2 were also ordered. In addition, reiterated sequences in the SeMNPV-25 genome were identified and located on the physical map.

Base Sequence↗

Enzyme polymorphism and genetic population structure in Escherichia coli and Shigella.

Electrophoretically demonstrable variation in 12 enzymes was studied in more than 1 600 isolates of Escherichia coli from human and animal sources and in 123 strains of the four species of Shigella. All 12 enzymes were polymorphic; and the number of allozymes (mobility variants), which were equated with alleles, averaged 9.3 per locus in E. coli. For Shigella species, the mean number of alleles was 2.9 per locus. Some 77% of the allozymes recorded in Shigella were shared with E. coli. A total of 302 unique genotypic combinations of alleles over the 12 loci (electrophoretic types, ETs) was distinguished, of which 279 represented E. coli and 23 were Shigella. Among electrophoretic types, mean allelic diversity per locus was 0.52 for E. coli and 0.29 for Shigella. It was estimated that there are, on the average, about 0.3 detectable codon differences per locus between pairs of strains of E. coli and Shigella, which is roughly equivalent to 1.2 amino acid differences per enzyme. Evidence that the enzyme loci studied are a random sample of the genome is provided by a significant positive correlation between estimates of genetic divergence between pairs of strains obtained by DNA reassociation tests and estimates of genetic distance between the same strains based on electrophoresis. A principal components analysis of allozyme profiles revealed that the 302 ETs fall into three overlapping clusters, reflecting strong non-random associations of alleles, largely at four loci. Each of the four ETs of E. coli that have been most frequently recovered from natural populations has an allozyme profile that is very similar to, or identical with, the hypothetical modal ET of one of the groups. ETs of Shigella fall into two of the groups. No biological significance can at present bbe attributed to the genetic structure revealed by Multilocus electrophoretic techniques. The electrophoretic data are fully compatible with other molecular and more conventional evidence of a close affinity between E. coli and Shigella, and they raise questions regarding the present assignments of certain strains to species. In support of evidence from DNA reassociation tests and serotyping, the present study suggests that S. sonnei is homogeneous in chromosomal genotype.

Adult↗

Genome-wide identification of NBS genes in japonica rice reveals significant expansion of divergent non-TIR NBS-LRR genes.

A complete set of candidate disease resistance ( R) genes encoding nucleotide-binding sites (NBSs) was identified in the genome sequence of japonica rice ( Oryza sativaL. var. Nipponbare). These putative R genes were characterized with respect to structural diversity, phylogenetic relationships and chromosomal distribution, and compared with those in Arabidopsis thaliana. We found 535 NBS-coding sequences, including 480 non-TIR (Toll/IL-1 receptor) NBS-LRR (Leucine Rich Repeat) genes. TIR NBS-LRR genes, which are common in A. thaliana, have not been identified in the rice genome. The number of non-TIR NBS-LRR genes in rice is 8.7 times higher than that in A. thaliana, and they account for about 1% of all of predicted ORFs in the rice genome. Some 76% of the NBS genes were located in 44 gene clusters or in 57 tandem arrays, and 16 apparent gene duplications were detected in these regions. Phylogenetic analyses based both NBS and N-terminal regions classified the genes into about 200 groups, but no deep clades were detected, in contrast to the two distinct clusters found in A. thaliana. The structural and genetic diversity that exists among NBS-LRR proteins in rice is remarkable, and suggests that diversifying selection has played an important role in the evolution of R genes in this agronomically important species. (Supplemental material is available online at http://gattaca.nju.edu.cn.)

Amino Acid Sequence↗

Phenotypic variation in molecular mimicry between Helicobacter pylori lipopolysaccharides and human gastric epithelial cell surface glycoforms. Acid-induced phase variation in Lewis(x) and Lewis(y) expression by H. Pylori lipopolysaccharides.

Helicobacter pylori is an important gastroduodenal pathogen of humans whose survival in the gastric environment below pH 4 is dependent on bacterial production of urease, whereas above pH 4 urease-independent mechanisms are involved in survival, but that remain to be elucidated fully. Previous structural investigations on the lipopolysaccharides (LPSs) of H. pylori have shown that the majority of these surface glycolipids express partially fucosylated, glucosylated, or galactosylated N-acetyllactosamine (LacNAc) O-polysaccharide chains containing Lewis(x) (Le(x)) and/or Lewis(y) (Le(y)), although some strains also express type 1 determinants, Lewis(a), Lewis(b), and H-1 antigen. In this study, we investigated acid-induced changes in the structure and composition of LPS and cellular lipids of the genome-sequenced strain, H. pylori 26695. When grown in liquid medium at pH 7, the O-chain consisted of a type 2 LacNAc polysaccharide, which was glycosylated with alpha-1-fucose at O-3 of the majority of N-acetylglucosamine residues forming Le(x) units, including chain termination by a Le(x) unit. However, growth in liquid medium at pH 5 resulted in production of a more complex O-chain whose backbone of type 2 LacNAc units was partially glycosylated with alpha L-fucose, thus forming Le(x), whereas the majority of the nonfucosylated N-acetylglucosamine residues were substituted at O-6 by alpha-D-galactose residues, and the chain was terminated by a Le(y) unit. In contrast, detailed chemical analysis of the core and lipid A components of LPS and analysis of cellular lipids did not show significant differences between H. pylori 26695 grown at pH 5 and 7. Although putative molecular mechanisms affecting Le(x) and Le(y) expression have been investigated previously, this is the first report identifying an environmental trigger inducing phase variation of Le(x) and Le(y) in H. pylori that can aid adaptation of the bacterium to its ecological niche.

Carbohydrate Conformation↗

Ha-ras rare alleles in breast cancer susceptibility.

Over the last several years, evidence has accumulated to support the idea that rare Ha-ras polymorphisms are associated with inherited susceptibility to certain human cancers. A recent epidemiologic study conducted at our institution found a significant association specifically with breast cancer, although the mechanism underlying this relationship remains unclear. We have proposed that rare Ha-ras alleles are markers of a genomic instability that predisposes to breast cancer. To address this hypothesis, we are investigating the relationship between the presence of rare alleles and another form of instability, gene amplification, and are developing new methodologies both to improve VNTR allele length detection and to characterize the internal repeat sequence variations of the various alleles. These studies should enable us to more clearly define the role of this region in cancer development by delineating VNTR structure and function and the mechanisms of rare allele generation. Ultimately, we hope to identify VNTR characteristics that will permit more accurate cancer risk assessment.

Alleles↗

Structure and genomic organization of centromeric repeats in Arabidopsis species.

Centromeric repetitive sequences were isolated from Arabidopsis halleri ssp. gemmifera and A. lyrata ssp. kawasakiana. Two novel repeat families isolated from A. gemmifera were designated pAge1 and pAge2. These repeats are 180 bp in length and are organized in a head-to-tail manner. They are similar to the pAL1 repeats of A. thaliana and the pAa units of A. arenosa. Both A. gemmifera and A. kawasakiana possess the pAa, pAge1 and pAge2 repeat families. Sequence comparisons of different centromeric repeats revealed that these families share a highly conserved region of approximately 50 bp. Within each of the four repeat families, two or three regions showed low levels of sequence variation. The average difference in nucleotide sequence was approximately 10% within families and 30% between families, which resulted in clear distinctions between families upon phylogenetic analysis. FISH analysis revealed that the localization patterns for the pAa, pAge1 and pAge2 families were chromosome specific in A. gemmifera and A. kawasakiana. In one pair of chromosomes in A. gemmifera, and three pairs of chromosomes in A. kawasakiana, two repeat families were present. The presence of three families of centromeric repeats in A. gemmifera and A. kawasakiana indicates that the first step toward homogenization of centromeric repeats occurred at the chromosome level.

Arabidopsis↗

Molecular cloning and functional expression of a human thyrotropin-releasing hormone (TRH) receptor gene.

In this study, we isolated genomic DNA fragments coding for the human thyrotropin-releasing hormone (TRH) receptor. Analysis of the nucleotide sequence revealed that the human TRH receptor gene had an exon-intron structure comprising at least two exons. A polypeptide encoded by the gene consisted of 398 amino acid residues with putative seven transmembrane domains. It showed high homology as a whole amino acid sequence with the rat and mouse TRH receptors except for considerable variation in the C-terminal region. Chromosomal mapping study indicated that the human TRH receptor gene was assigned to chromosome 8. Chinese hamster ovary (CHO) cells transfected with a DNA fragment containing the coding regions of the human TRH receptor bound with [3H]TRH. This binding was inhibited by adding unlabeled TRH in a dose-dependent fashion. Scatchard analysis indicated that the transfected CHO cells expressed a single class of high affinity binding sites at a dissociation constant (Kd) of approximately 1 nM. These results demonstrated that the isolated gene encoded a specific TRH receptor with high affinity.

Amino Acid Sequence↗

Modes and clustering for time-warped gene expression profile data.

MOTIVATION: The study of the dynamics of regulatory processes has led to increased interest for the analysis of temporal gene expression level data. To address the dynamics of regulation, expression data are collected repeatedly over time. It is difficult to statistically represent the resulting high-dimensional data. When regulatory processes determine gene expression, time-warping is likely to be present, i.e. the sample of gene expression trajectories reflects variation not only in terms of the expression amplitudes, but also in terms of the temporal structure of gene expression. RESULTS: A non-parametric time-synchronized iterative mean updating technique is proposed to find an overall representation that corresponds to a mode of a sample of expression profiles, viewed as a random sample in function space. The proposed algorithm explores the application of previous work of Hall and Heckman to genome-wide expression data and provides an extension that includes random time-warping with the aim to synchronize timescales across genes. The proposed algorithm is universally applicable for the construction of modes for functional data with time-warping. We demonstrate the construction of mode functions for a sample of Drosophila gene expression data. The algorithm can be applied to define clusters among the observed trajectories of gene expression, without any kind of prior non-time-warped clustering, as illustrated in the numerical example.

Adaptation, Physiological↗

A highly divergent gene cluster in honey bees encodes a novel silk family.

The pupal cocoon of the domesticated silk moth Bombyx mori is the best known and most extensively studied insect silk. It is not widely known that Apis mellifera larvae also produce silk. We have used a combination of genomic and proteomic techniques to identify four honey bee fiber genes (AmelFibroin1-4) and two silk-associated genes (AmelSA1 and 2). The four fiber genes are small, comprise a single exon each, and are clustered on a short genomic region where the open reading frames are GC-rich amid low GC intergenic regions. The genes encode similar proteins that are highly helical and predicted to form unusually tight coiled coils. Despite the similarity in size, structure, and composition of the encoded proteins, the genes have low primary sequence identity. We propose that the four fiber genes have arisen from gene duplication events but have subsequently diverged significantly. The silk-associated genes encode proteins likely to act as a glue (AmelSA1) and involved in silk processing (AmelSA2). Although the silks of honey bees and silkmoths both originate in larval labial glands, the silk proteins are completely different in their primary, secondary, and tertiary structures as well as the genomic arrangement of the genes encoding them. This implies independent evolutionary origins for these functionally related proteins.

Amino Acid Sequence↗

Polymorphic proteins encoded within BZLF1 of defective and standard Epstein-Barr viruses disrupt latency.

These experiments identify an Epstein-Barr virus-encoded gene product, called ZEBRA (BamHI fragment Z Epstein-Barr replication activator) protein, which activates a switch between the latent and replicative life cycle of the virus. Our previous work had shown that the 2.7-kilobase-pair WZhet piece of rearranged Epstein-Barr virus DNA from a defective virus activated replication when introduced into cells with a latent genome, but it was not clear whether a protein product was required for the phenomenon. We now use deletional, site-directed, and chimeric mutagenesis, together with gene transfer, to show that a 43-kilodalton protein, encoded in the BZLF1 open reading frame of het DNA, is responsible for this process. The rearrangement in defective DNA does not contribute to the structural gene for the protein. Similar proteins with variable electrophoretic mobility (37 to 39 kilodaltons) were encoded by BamHI Z fragments from standard, nondefective Epstein-Barr virus genomes. Plasmids expressing the ZEBRA proteins from B95-8 and HR-1 viruses were less efficient at activating replication in D98/HR-1 cells than those which contained the ZEBRA gene from the defective virus. It is not yet known whether these functional differences are due to variations in expression of the plasmids or to intrinsic differences in the activity of these polymorphic polypeptides.

Cell Line↗

Advances in the detection of ploidy differences in cancer by in situ hybridization.

Three main techniques allow the detection of changes in the cellular genomic content. The karyotyping procedure on metaphase spreads can give specific information on chromosome number and structural chromosome changes, but analyses are restricted to a limited number of chromosome spreads. Furthermore, cell culturing of (in particular solid) cancer specimens can result in selection of a minor tumour cell population with a high proliferative capacity. On the other hand, flow cytometry allows the analyses of large numbers of cells, but does not detect small variations in the DNA content or structural changes. The fluorescent in situ hybridization (FISH) procedure combines the advantages of the two former procedures, in that relatively large numbers of cells can be analysed easily and specific chromosomal changes can be detected.

DNA Probes↗

[The effect of rye chromosomes on callus induction and regeneration in callus cultures of immature embryos of wheat-rye substitution lines, Triticum aestivum L. cultivar Saratovskaia 29/Secale cereale L. cultivar Onokhoiskaia].

The effect of individual rye chromosomes on the induction of callus and the character of its regenerating capacity was studied with cultured immature embryos of wheat-rye (Triticum aestivum L. cv. Saratovskaya 29-Secale cereale L. cv. Onokhoiskaya) substitution lines. The genotypic diversity of the substitution lines proved to significantly affect variation of parameters characterizing the major types of callus cultures, that is, frequencies of embryogenic calli, which are capable of shoot regeneration, and of morphogenic calli, which produce root structures. Functioning in the genotypic background of common wheat cultivar Saratovskaya, chromosomes 2R and 3R of rye cultivar Onokhoiskaya stimulated significantly the induction of embryogenic callus highly capable of shoot regeneration. Rye chromosome 2R present in place of chromosome 2D in the common wheat genome suppressed the induction of callus producing root structures. Rye chromosomes 1R and 6R suppressed the induction of embryogenic callus capable of shoot regeneration.

Chromosomes, Plant↗

Human population genetic structure and diversity inferred from polymorphic L1(LINE-1) and Alu insertions.

BACKGROUND/AIMS: The L1 retrotransposable element family is the most successful self-replicating genomic parasite of the human genome. L1 elements drive replication of Alu elements, and both have had far-reaching impacts on the human genome. We use L1 and Alu insertion polymorphisms to analyze human population structure. METHODS: We genotyped 75 recent, polymorphic L1 insertions in 317 individuals from 21 populations in sub-Saharan Africa, East Asia, Europe and the Indian subcontinent. This is the first sample of L1 loci large enough to support detailed population genetic inference. We analyzed these data in parallel with a set of 100 polymorphic Alu insertion loci previously genotyped in the same individuals. RESULTS AND CONCLUSION: The data sets yield congruent results that support the recent African origin model of human ancestry. A genetic clustering algorithm detects clusters of individuals corresponding to continental regions. The number of loci sampled is critical: with fewer than 50 typical loci, structure cannot be reliably discerned in these populations. The inclusion of geographically intermediate populations (from India) reduces the distinctness of clustering. Our results indicate that human genetic variation is neither perfectly correlated with geographic distance (purely clinal) nor independent of distance (purely clustered), but a combination of both: stepped clinal.

Alu Elements↗