Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic Structural Variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43Linked to original sources

[Epigenetics of the sperm cell].

In addition to genetic information, the spermatozoon carries another type of information, named epigenetic, which is not associated with variations of the DNA sequence. In somatic cells, it is now generally admitted that epigenetic information is not only regulated by DNA methylation but also involves modifications of the genome structure, or epigenome. During male germ cell maturation, the epigenome is globally re-organized, since most histones, which are associated to DNA in somatic cells, are removed and replaced by sperm specific nuclear proteins, the protamines, responsible for the tight compaction of the sperm DNA. However, a small proportion of histones, and probably other proteins, are retained within the sperm nucleus, and the structure of the sperm genome is actually heterogeneous. This heterogeneity of the sperm epigenome could support an epigenetic information, transmitted to the embryo, which could be crucial for its development. Although it is nowadays possible to appreciate the global structure of the sperm genome, the precise constitution of the sperm epigenome remains unknown. In particular, very recent data suggest that specific regions of the genome could be associated with particular proteins and define specific structures. This structural partitioning of the sperm genome could convey important epigenetic information, crucial for the embryo development.

DNA↗

Chromosomal fragility, structural rearrangements and mobile element activity may reflect dynamic epigenetic mechanisms of importance in neurobehavioural genetics.

Advances in human genome analyses have not yet allowed identification of specific genetic mechanisms underlying the expression of human neurobehavioural disorders. There is an increasing awareness that several genes may contribute to behavioural phenotypes and these genes appear to interact in as yet undetermined ways. It has been suggested that the problem needs elucidation from an epigenetic, gene expression perspective. Cytogenetic instability manifesting as chromosomal fragile sites, translocations, duplications, deletions and inversions, when co-occurring with neurobehavioural disorders, may offer a doorway to the investigation of such chromatin level, regulatory region, epigenetic processes. Due to earlier indications of non-specificity of chromosomal aberrations, poor phenotype:genotype correlations and a shift to analysing candidate coding regions on high resolution map level, the only utility of chromosomal breakpoints came to be seen as harbouring possible candidate genes of interest when segregating together with particular neurobehavioural disorders. More recent findings of the expression of highly specific subsets of fragile sites in association with Tourette and Rett syndromes need to be extended to other neurobehavioural disorders to ascertain whether observed patterns can be considered representative of 'chromatin endophenotypes' correlating with discrete sets of neurobehavioural symptoms. Environmental/epigenetic factors could affect the chromatin characteristics of the genome arising through DNA strand breakage, mobile element activity and retroinsertion, establishing new architectural features of regulatory control networks very rapidly in comparison to coding region evolution rates. Microarray-based techniques for the genome-wide mapping of in vivo protein-DNA interactions offer increasingly comprehensive views of genetic and epigenetic regulatory networks. It may be informative to include functionally significant chromatin structural variation analyses when considering candidate genes for neurobehavioural disorders.

Cell Cycle↗

Genomic structure, promoter analysis and expression of the porcine (Sus scrofa) TLR4 gene.

Toll-like receptor 4 (TLR4) is essential for initiating the innate response to lipopolysaccharide (LPS) from Gram-negative bacteria by acting as a signal-transducing receptor. As the pig industry faces a unique array of related pathogens, it is anticipated that the genotype of swine TLR4 could be of crucial importance in future strategies aimed at improving genetic resistance to infectious diseases. In order to help in investigating TLR4 as a candidate disease-resistance gene in pigs, we established its genomic structure and produced sufficient flanking intronic sequences to enable simple PCR amplification of the coding portions of the gene. Expression in different porcine tissues was studied and showed splicing variations in mRNA sequences. The cDNA sequence for poTLR4 contains an open reading frame of 2526bp that codes for 841 aa, 98 and 568bp in the 5'- and 3'-UTRs, respectively. Overall, the general organization of porcine, human, murine, and avian TLR4 genes is quite similar: three exons with the third one very long. A high level of conservation of the size and the sequence, especially for the two last exons and particularly in the sequence corresponding to the LRRs and TIR domain, is observed between species. The important antimicrobial properties of these proteins may account for a conservative selection pressure on these TLR4 coding sequences. Several putative binding sites described in the human and murine promoter of TLR4 genes have been identified in the 5'-flanking region of poTLR4. Conversely, this region lacks a TATA box, consensus initiator sequences, or GC-rich regions. The basic sequence data gathered will allow the establishment of an inventory of naturally occurring variation in porcine TLR4, so that alleles can be tested for disease association studies.

Animals↗

Phenylketonuria screening registry as a resource for population genetic studies.

BACKGROUND: Neonatal screening for metabolic diseases, involving samples stored on filter paper (Guthrie spots), provides a potential resource for genetic epidemiological studies. OBJECTIVE: To develop a method to make these dried blood spots available for large scale genetic epidemiology. METHODS: DNA from untraceable Guthrie spots was extracted using a saponin and chelex-100 based method and preamplified by improved primer preamplification. Analyses were done on 38 samples each of fresh, 10, and 25 year old Guthrie spots and the success rate determined for PCR amplification for five amplicon lengths. RESULTS: The method was applicable even on 25 year old samples. The success rate was 100% for 100 bp amplicons and 80% for 396 bp amplicons. Ninety four Guthrie samples were genotyped, including carriers of two different PKU mutations; all carriers were found (six R158Q, four R252W), with no false positives. Finally, 2132 anonymous samples from the Swedish PKU registry were extracted and preamplified and the allele frequencies of APOepsilon4, PPARgamma Pro12Ala, and the CCR5 32 bp deletion determined. Local variations in allele frequencies suggested subpopulation structuring. There was a significant difference (p<0.01) in regional allele frequencies for the CCR5 32 bp deletion in the Swedish population. CONCLUSION: Whole genome amplification makes it feasible to conduct large genetic epidemiological studies using PKU screening registries.

Alleles↗

A pan-cancer multi-omic SuperLearner for regulated cell death survival topologies.

INTRODUCTION: Regulated cell death (RCD) pathways influence tumor progression and immune modulation. We previously constructed a signature database mapping 25 RCD forms across seven multi-omic layers and 33 tumor types (CancerRCDShiny). Despite their ability to identify risk populations, translating these signatures into personalized clinical workflows requires a shift from cohort stratification to individualized risk mapping by modeling patient risk (survival topologies) to capture the non-linear dynamics of RCD signatures. METHODS: We engineered a pan-cancer multi-omic SuperLearner pipeline across 33 cancer types. Phase I performed zero-leakage harmonization and groupwise imputation to prevent cross-cohort amalgamation. Phase II deployed Elastic Net-regularized Cox regression as a CANARY diagnostic to map proportional hazards failures. Strata with a 35% missingness barrier entered Phase III, deploying a Quadripartite ensemble: Random Survival Forests, XGBoost, Survival-Boruta, and Multi-Task Logistic Regression, fused within an Elastic Net Multi-View Meta-Learner (MVL), with post-hoc TreeSHAP and LIME interpretability. RESULTS: The CANARY diagnostic demonstrated the structural invalidity of pan-cancer geometric proportional hazards. Across 96 admissible strata, Phase III executed algorithmic displacement: continuous multi-omic topologies suppressed static genomic mutations and copy number variations (85.7% vs. 0.0% apex retention). The MVL stabilized predictions against extreme variance; LIME surrogate validations (R 2&#x202f;<&#x202f;0.10) confirmed the systematic failure of linear interpretative proxies. N-dimensional TreeSHAP interaction mapping exposed synergistic and antagonistic rescue trajectories defining individualized Survival Topologies, which were invisible to additive models. The architecture was deployed as CancerRCDPredictor, a digital molecular tumor board with integrated LLM capabilities. The MVL SuperLearner achieved a median C-index of 0.749 (IQR: 0.722-0.836) across 96 modelable strata, with 95% bootstrap confidence intervals confirming precision (median width: 0.052) and permutation significance in 93.8% of strata (p&#x202f;<&#x202f;0.001). External CPTAC validation across ten cancer types demonstrated significant cross-cohort generalizability in clear cell renal carcinoma (KIRC; C-index 0.675, p&#x202f;=&#x202f;0.017) and modest performance across the remaining adequately powered cancers (median 0.582), underscoring the need for larger multi-institutional validation cohorts. CONCLUSION: This pan-cancer multi-omic SuperLearner bypasses linear topological failures, advancing beyond generalized stratification to establish a deterministically mapped architecture for predicting RCD-related survival topologies. Through the CancerRCDPredictor interface, multi-omic insights translate into individualized survival topology exploration, providing a foundation for future precision oncology validation.

SuperLearner↗

Molecular genetics of mosquito resistance to malaria parasites.

Malaria parasites are transmitted by the bite of an infected mosquito, but even efficient vector species possess multiple mechanisms that together destroy most of the parasites present in an infection. Variation between individual mosquitoes has allowed genetic analysis and mapping of loci controlling several resistance traits, and the underlying mechanisms of mosquito response to infection are being described using genomic tools such as transcriptional and proteomic analysis. Malaria infection imposes fitness costs on the vector, but various forms of resistance inflict their own costs, likely leading to an evolutionary tradeoff between infection and resistance. Plasmodium development can be successfully completed onlyin compatible mosquito-parasite species combinations, and resistance also appears to have parasite specificity. Studies of Drosophila, where genetic variation in immunocompetence is pervasive in wild populations, offer a comparative context for understanding coevolution of the mosquito-malaria relationship. More broadly, plants also possess systems of pathogen resistance with features that are structurally conserved in animal innate immunity, including insects, and genomic datasets now permit useful comparisons of resistance models even between such diverse organisms.

Animals↗

A structural motif in the variant surface glycoproteins of Trypanosoma brucei.

The variable domain of the trypanosome variant surface glycoprotein (VSG) ILTat 1.24 has been shown by X-ray crystallography to resemble closely the structures of VSG MITat 1.2, despite their low sequence similarity. Specific structural features of these VSGs, including substitution of carbohydrate for an alpha-helix, can be found in other VSG sequences. Thus antigenic variation in trypanosomes is accomplished by sequence variation, not gross structural alteration; the extensive sequence differences among VSGs may be required for another reason, such as the avoidance of recognition by helper T cells. Additionally, VSG sequences are found to define families, within a VSG superfamily, which have evolved in the trypanosome genome.

Amino Acid Sequence↗

Haplotype parsing: methods for extracting information from human genetic variations.

While the shared consensus genetic sequence of our species contains a great deal of information about our common biology, there is also much to be learned from the subtle genetic variations across our species. These variations are believed to be generally of little or no direct functional significance and predominantly reflect the chance accumulation of small genetic changes since our emergence as a species. Therefore, they carry little useful information when observed in a single individual. When tallied across a whole population though, these chance mutations can teach us a great deal about our evolutionary history and the patterns of inheritance in particular individuals. In particular, frequently observed patterns of single nucleotide polymorphisms (SNPs) in a population can identify segments of chromosome that have been passed down largely intact through long stretches of our evolution. Finding these frequently conserved chromosomal segments, or haplotypes, and developing methods to identify haplotype patterns in particular individuals, will in turn help us to identify those particular segments that carry genetic factors influencing risk for many common human diseases. To make the best use of this data, we will need to develop new models for the encoding of information in genome variations--the "language of genetic variation"--and new algorithms for fitting datasets to those models. This article surveys past work by the author and colleagues on this problem, utilising computational methods for locating frequent patterns in haploid sequence data, and "parsing" sequences so as to optimally explain them given the knowledge of the general population structure. The author's recent work in this area has been compiled into a set of computational tools available at http://www-2.cs.cmu.edu/~russells/software/hapmotif.html.

Algorithms↗

Variable human minisatellite-like regions in the Mycobacterium tuberculosis genome.

Mycobacterial interspersed repetitive units (MIRUs) are 40-100 bp DNA elements often found as tandem repeats and dispersed in intergenic regions of the Mycobacterium tuberculosis complex genomes. The M. tuberculosis H37Rv chromosome contains 41 MIRU loci. After polymerase chain reaction (PCR) and sequence analyses of these loci in 31 M. tuberculosis complex strains, 12 of them were found to display variations in tandem repeat copy numbers and, in most cases, sequence variations between repeat units as well. These features are reminiscent of those of certain human variable minisatellites. Of the 12 variable loci, only one was found to vary among genealogically distant BCG substrains, suggesting that these interspersed bacterial minisatellite-like structures evolve slowly in mycobacterial populations.

Base Sequence↗

Conservation of the mosaic structure of the four internal transcribed spacers and localisation of the rrn operons on the Streptococcus pneumoniae genome.

The detection of heterogeneity of the 16S-23S ribosomal intergenic transcribed spacer (ITS) region has become rather common over the past years for identification and typing purposes of bacteria. The ITS not only varies in sequence and length, but also in number of alleles per genome and in their position on the chromosome together with the ribosomal clusters. The ITS characterisation has allowed discrimination of several species within a genus and variation in ITS sequences between the multiple rrn operons present within a genome may be as high or greater than between strains of the same species or subspecies. It is important to understand the variability of ITS sequences in a given genome to gain insights into bacterial physiology and taxonomy. The present study describes the possibility to type Streptococcus pneumoniae by PCR-ribotyping of the spacer region, the determination of the molecular structure of the ITS, and the determination of the number and localisation of rrn operons in this microorganism. Our results show that the genome of S. pneumoniae contains four ribosomal operons, showing the same genomic organisation among strains, each containing a single ITS allele of 270 bp. The ITS sequence presents a mosaic organisation of blocks highly conserved intra- and inter-species within the genus Streptococcus, giving no possibility for variations to arise.

Amino Acid Sequence↗

Structure of a polymorphic repeat at the CACNA1C schizophrenia locus.

Genetic variation within intron 3 of the CACNA1C calcium channel gene is associated with schizophrenia and other neuropsychiatric disorders, but analysis of the causal variants and their effect is complicated by a nearby variable-number tandem repeat (VNTR). Here, we explored the structure and population variability of the CACNA1C intron 3 VNTR using 155 long-read genome assemblies from 78 diverse individuals. Based on sequence differences among repeat units, we clustered individual sequences into 7 VNTR structural alleles called Types. Three Types were related through large duplications, but the other Types diverged much earlier such that only 12 repeat units at the 5' end of the VNTR were shared across most Types. The most diverged Types were rare and present only in individuals with African ancestry, but a multiallelic structural polymorphism was present across populations at different frequencies, consistent with expansion of the VNTR preceding the emergence of early hominins. We demonstrated that this polymorphism was in complete linkage disequilibrium with fine-mapped schizophrenia variants from genome-wide association studies (GWAS), and that this risk haplotype was associated with decreased CACNA1C gene expression in the brain. Our work suggests that sequence variation within a human-specific VNTR affects gene expression, and provides a detailed characterization of new alleles at a flagship neuropsychiatric locus.

Variable-number tandem repeat↗

Comparative genome analysis of Campylobacter jejuni using whole genome DNA microarrays.

Whole genome DNA microarrays were constructed and used to investigate genomic diversity in 18 Campylobacter jejuni strains from diverse sources. New algorithms were developed that dynamically determine the boundary between the conserved and variable genes. Seven hypervariable plasticity regions (PR) were identified in the genome (PR1 to PR7) containing 136 genes (50%) of the variable gene pool. When comparisons were made with the sequenced strain NCTC11168, the number of absent or divergent genes ranged from 2.6% (40 genes) to 10.2% (163) and in total 16.3% (269) of the genes were variable. PR1 contains genes important in the utilisation of alternative electron acceptors for respiration and may confer a selective advantage to strains in restricted oxygen environments. PR2, 3 and 7 contain many outer membrane and periplasmic proteins and hypothetical proteins of unknown function that might be linked to phenotypic variation and adaptation to different ecological niches. PR4, 5 and 6 contain genes involved in the production and modification of antigenic surface structures.

Algorithms↗

Application of molecular biology to mental illness. Analysis of genomic DNA and brain mRNA.

Techniques in molecular biology and genetics have made it possible to systematically study gene effects in human disease. The number of gene clusters specifically encoding human brain structure and function is probably about 1,600 or half of all clusters. Evolutionary effects such as linkage disequilibrium and conservation of exons (DNA encoding structural proteins) as well as the fact that there are a tractable number of gene clusters involved, tend to make it quite likely that DNA pathology or DNA variation (polymorphism) predisposing to mental illness can be detected. Genes involved in mental illness can be detected either by studying DNA obtained from blood samples (genomic DNA) directly or by the analysis of mRNA and proteins from suitable cell or tissue preparations. The study of gene expression in the human brain is still in its infancy, nevertheless there are some hints that non-poly-adenylated mRNAs may be important in brain development and certain transcribed sequences may have a specific role in gene expression of the brain. The advantage of studying genomic DNA by the use of linkage and association analysis in multiply affected families is that it will, in the end, almost certainly yield a positive result for a disease with a substantial genetic input. Analysis of gene products from tissues such as brain could in theory detect specific disease genes but the approach will also identify genes secondarily affected by the disease process. Differentiation of genes that are primarily causing mental illness from those that are secondarily affected can be carried out by using such candidate genes as linkage markers in multiply affected families.

Animals↗

First nationwide full-genome characterisation of human-derived Andes virus in Chile: a retrospective genomic epidemiology study.

BACKGROUND: Andes virus (ANDV) is the only hantavirus known to transmit between humans and causes hantavirus cardiopulmonary syndrome in Chile and Argentina. In Chile, ANDV genomic diversity remains incompletely characterised. This study aimed to characterise the genetic diversity, geographical structure, and molecular signatures of ANDV using human clinical samples collected over a 13-year period (2011-24). METHODS: We conducted a retrospective genomic epidemiology study of ANDV infections in Chile. Clinical samples from patients with confirmed ANDV, collected between March 9, 2011, and June 27, 2024, were analysed and sequenced. Clinical and epidemiological data were obtained from diagnostic laboratories and surveillance programmes. Consensus sequences for the S, M, and L segments were generated, and genetic clustering and divergence were assessed using phylogenetic inference and variant calling. FINDINGS: We analysed clinical samples from 58 infected individuals and identified two major genomic variants of ANDV with distinct geographical distributions, defined by regionally structured patterns of nucleotide and amino acid substitutions across the S, M, and L segments: ANDV Chi-North (central Chile) and ANDV-South (southern Chile). No consistent clustering by clinical severity was observed, and no recurrent non-synonymous substitutions were uniquely associated with severe disease. Substitutions previously associated with person-to-person transmission in outbreaks in Argentina were not consistently observed in Chilean sequences, including in four person-to-person transmission cases. Although some substitutions described in ANDV-like viruses were present in the Chi-North lineage, this lineage remained phylogenetically distinct and geographically restricted to central Chile. INTERPRETATION: To our knowledge, this study provides the first nationwide genomic characterisation of human-derived ANDV in Chile. The identification of geographically structured variants indicates that ANDV diversity in Chile is driven by regional diversification rather than clinical outcome. The absence of consistent amino acid signatures associated with disease severity or person-to-person transmission suggests that these phenotypes are unlikely to be explained by viral genetic variation alone. These findings refine current understanding of ANDV evolution and highlight the need for continued integrated genomic surveillance in endemic regions. FUNDING: Agencia Nacional de Investigaci&#xf3;n y Desarrollo de Chile and National Institutes of Health.

Humans↗

Molecular characterization of the porcine deleted in malignant brain tumors 1 gene (DMBT1).

The human gene deleted in malignant brain tumors 1 (DMBT1) is considered to play a role in tumorigenesis and pathogen defense. It encodes a protein with multiple scavenger receptor cysteine-rich (SRCR) domains, which are involved in recognition and binding of a broad spectrum of bacterial pathogens. The SRCR domains are encoded by highly homologous repetitive exons, whose number in humans may vary from 8 to 13 due to genetic polymorphism. Here, we characterized the porcine DMBT1 gene on the mRNA and genomic level. We assembled a 4.5 kb porcine DMBT1 cDNA sequence from RT-PCR amplified seminal vesicle RNA. The porcine DMBT1 cDNA contains an open reading frame of 4050 nt. The transcript gives rise to a putative polypeptide of 1349 amino acids with a calculated mass of 147.9 kDa. Compared to human DMBT1, it contains only four N-terminal SRCR domains. Northern blotting revealed transcripts of approximately 4.7 kb in size in the tissues analyzed. Analysis of ESTs suggested the existence of secreted and transmembrane variants. The porcine DMBT1 gene spans about 54 kb on chromosome 14q28-q29. In contrast to the characterized cDNA, the genomic BAC clone only contained 3 exons coding for N-terminal SRCR domains. In different mammalian DMBT1 orthologs large interspecific differences in the number of SRCR exons and utilization of the transmembrane exon exist. Our data suggest that the porcine DMBT1 gene may share with the human DMBT1 gene additional intraspecific variations in the number of SRCR-coding exons.

Amino Acid Sequence↗

The genome sequence of the food-borne pathogen Campylobacter jejuni reveals hypervariable sequences.

Campylobacter jejuni, from the delta-epsilon group of proteobacteria, is a microaerophilic, Gram-negative, flagellate, spiral bacterium-properties it shares with the related gastric pathogen Helicobacter pylori. It is the leading cause of bacterial food-borne diarrhoeal disease throughout the world. In addition, infection with C. jejuni is the most frequent antecedent to a form of neuromuscular paralysis known as Guillain-Barré syndrome. Here we report the genome sequence of C. jejuni NCTC11168. C. jejuni has a circular chromosome of 1,641,481 base pairs (30.6% G+C) which is predicted to encode 1,654 proteins and 54 stable RNA species. The genome is unusual in that there are virtually no insertion sequences or phage-associated sequences and very few repeat sequences. One of the most striking findings in the genome was the presence of hypervariable sequences. These short homopolymeric runs of nucleotides were commonly found in genes encoding the biosynthesis or modification of surface structures, or in closely linked genes of unknown function. The apparently high rate of variation of these homopolymeric tracts may be important in the survival strategy of C. jejuni.

Amino Acid Sequence↗

Guide to the draft human genome.

There are a number of ways to investigate the structure, function and evolution of the human genome. These include examining the morphology of normal and abnormal chromosomes, constructing maps of genomic landmarks, following the genetic transmission of phenotypes and DNA sequence variations, and characterizing thousands of individual genes. To this list we can now add the elucidation of the genomic DNA sequence, albeit at 'working draft' accuracy. The current challenge is to weave together these disparate types of data to produce the information infrastructure needed to support the next generation of biomedical research. Here we provide an overview of the different sources of information about the human genome and how modern information technology, in particular the internet, allows us to link them together.

Amino Acid Sequence↗