Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Structural genome variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

A transcriptomic analysis of the phylum Nematoda.

The phylum Nematoda occupies a huge range of ecological niches, from free-living microbivores to human parasites. We analyzed the genomic biology of the phylum using 265,494 expressed-sequence tag sequences, corresponding to 93,645 putative genes, from 30 species, including 28 parasites. From 35% to 70% of each species' genes had significant similarity to proteins from the model nematode Caenorhabditis elegans. More than half of the putative genes were unique to the phylum, and 23% were unique to the species from which they were derived. We have not yet come close to exhausting the genomic diversity of the phylum. We identified more than 2,600 different known protein domains, some of which had differential abundances between major taxonomic groups of nematodes. We also defined 4,228 nematode-specific protein families from nematode-restricted genes: this class of genes probably underpins species- and higher-level taxonomic disparity. Nematode-specific families are particularly interesting as drug and vaccine targets.

Animals↗

Identification of genomic organisation, sequence variants and analysis of the role of the human dishevelled 1 gene in late onset Alzheimer's disease.

Alzheimer's disease (AD) is a disorder characterised by a progressive deterioration in memory and other cognitive functions. Neurofibrillary tangles (NFT) are a major pathological hallmark of AD, these are aggregations of paired helical filaments (PHF) comprised of the hyperphosphorylated microtubule associated protein tau. Several kinases, such as glycogen synthase kinase 3 beta (GSK3beta) and c-Jun N-terminal kinase (JNK), phosphorylate tau at sites that are phosphorylated in PHF. Dishevelled 1 (DVL1) is thought to act as a positive regulator of the wnt signalling pathway, and inhibits GSK3beta activity preventing beta-catenin degradation and thus allowing wnt target gene expression. JNK activation is also regulated by DVL1, however it is unclear if this is via the wnt signalling pathway. These observations suggest a central role for DVL1 in tau phosphorylation and AD and led us to investigate DVL1 as a candidate gene for this disorder. We determined the genomic structure of the DVL1 gene by sequencing and data mining and searched for sequence variations in the coding sequences and flanking introns. The DVL1 gene spans a region of approximately 13.8 kb (not including the 5' untranslated region) and is encoded by 15 exons. Analysis of over 4.3 kb of sequence, including 98% of exonic sequences and introns 2, 3, 6, 7, 9, 10, 11 and 12, revealed there to be six rare (< or =6%) sequence variations. None of these had any association with late onset AD. This would suggest that polymorphic variations in the coding sequences of DVL1 are not important in AD. However further analysis of regulatory regions may lead to the identification of other sequence variations which may be implicated in AD.

Adaptor Proteins, Signal Transducing↗

Simple stochastic birth and death models of genome evolution: was there enough time for us to evolve?

MOTIVATION: The distributions of many genome-associated quantities, including the membership of paralogous gene families can be approximated with power laws. We are interested in developing mathematical models of genome evolution that adequately account for the shape of these distributions and describe the evolutionary dynamics of their formation. RESULTS: We show that simple stochastic models of genome evolution lead to power-law asymptotics of protein domain family size distribution. These models, called Birth, Death and Innovation Models (BDIM), represent a special class of balanced birth-and-death processes, in which domain duplication and deletion rates are asymptotically equal up to the second order. The simplest, linear BDIM shows an excellent fit to the observed distributions of domain family size in diverse prokaryotic and eukaryotic genomes. However, the stochastic version of the linear BDIM explored here predicts that the actual size of large paralogous families is reached on an unrealistically long timescale. We show that introduction of non-linearity, which might be interpreted as interaction of a particular order between individual family members, allows the model to achieve genome evolution rates that are much better compatible with the current estimates of the rates of individual duplication/loss events.

Computer Simulation↗

Genetic predisposition to low bone mass is paralleled by an enhanced sensitivity to signals anabolic to the skeleton.

The structure of the adult skeleton is determined, in large part, by its genome. Whether genetic variations may influence the effectiveness of interventions to combat skeletal diseases remains unknown. The differential response of trabecular bone to an anabolic (low-level mechanical vibration) and a catabolic (disuse) mechanical stimulus were evaluated in three strains of adult mice. In low bone-mineral-density C57BL/6J mice, the low-level mechanical signal caused significantly larger bone formation rates (BFR) in the proximal tibia, but the removal of functional weight bearing did not significantly alter BFR. In mid-density BALB/cByJ mice, mechanical stimulation also increased BFR, whereas disuse significantly decreased BFR. In contrast, neither anabolic nor catabolic mechanical signals influenced any index of bone formation in high-density C3H/HeJ mice. Together, data from this study indicate that the sensitivity of trabecular tissue to both anabolic and catabolic stimuli is influenced by the genome. Extrapolated to humans, these results may explain in part why prophylaxes for low bone mass are not universally effective, yet also indicate that there may be a genotypic indication of people who are at reduced risk of suffering from bone loss.

Adaptation, Physiological↗

The repertoire of protein kinases encoded in the draft version of the human genome: atypical variations and uncommon domain combinations.

BACKGROUND: Phosphorylation by protein kinases is central to cellular signal transduction. Abnormal functioning of kinases has been implicated in developmental disorders and malignancies. Their activity is regulated by second messengers and by the binding of associated domains, which are also influential in translocating the catalytic component to their substrate sites, in mediating interaction with other proteins and carrying out their biological roles. RESULT: Using sensitive profile-search methods and manual analysis, the human genome has been surveyed for protein kinases. A set of 448 sequences, which show significant similarity to protein kinases and contain the critical residues essential for kinase function, have been selected for an analysis of domain combinations after classifying the kinase domains into subfamilies. The unusual domain combinations in particular kinases suggest their involvement in ubiquitination pathways and alternative modes of regulation for mitogen-activated protein kinase kinases (MAPKKs) and cyclin-dependent kinase (CDK)-like kinases. Previously unexplored kinases have been implicated in osteoblast differentiation and embryonic development on the basis of homology with kinases of known functions from other organisms. Kinases potentially unique to vertebrates are involved in highly evolved processes such as apoptosis, protein translation and tyrosine kinase signaling. In addition to coevolution with the kinase domain, duplication and recruitment of non-catalytic domains is apparent in signaling domains such as the PH, DAG-PE, SH2 and SH3 domains. CONCLUSIONS: Expansion of the functional repertoire and possible existence of alternative modes of regulation of certain kinases is suggested by their uncommon domain combinations. Experimental verification of the predicted implications of these kinases could enhance our understanding of their biological roles.

Amino Acid Motifs↗

Human resistin gene: molecular scanning and evaluation of association with insulin sensitivity and type 2 diabetes in Caucasians.

Insulin resistance is strongly associated with obesity, but even among obese subjects insulin sensitivity varies widely. Recently, a new adipocyte hormone, resistin, was identified, shown to reduce insulin-mediated glucose uptake, and shown to be increased in obese mice. We used the chromosome 19 draft sequence to determine the genomic structure of human resistin and to screen the exons, introns, and flanking sequences for variation. We screened 44 subjects with type 2 diabetes and 20 nondiabetic family members who were at the extremes of insulin sensitivity. We identified eight noncoding single nucleotide polymorphisms (SNPs) and one GAT microsatellite repeat. Three SNPs, which were in incomplete linkage disequilibrium with each other and had allelic frequencies exceeding 5%, were selected for further study. No SNP was associated with type 2 diabetes, but the SNP in the promoter region was a significant determinant of insulin sensitivity index (P = 0.04) among nondiabetic family members who had undergone iv glucose tolerance tests. The three common SNPs showed statistical significance as determinants of insulin sensitivity index (P < 0.01) in interaction with body mass index. Noncoding SNPs in the resistin gene may influence insulin sensitivity in interaction with obesity, but this finding will need to be confirmed in other populations.

Base Sequence↗

Mitochondrial DNA homeostasis: A novel therapeutic target for neurodegenerative diseases.

The mitochondrial genomic homeostasis is essential for the function of the oxidative phosphorylation system and cellular homeostasis. Mitochondrial DNA is particularly susceptible to aging-related oxidative stress due to the lack of a histone coat. Disturbances in mitochondrial DNA may contribute to functional decline during the aging process and in neurodegenerative diseases, leading to further impairment of mitochondrial DNA and initiating a vicious cycle. To date, it remains unclear how disturbed mitochondrial DNA is involved in the etiology of pathological aging and neurodegenerative diseases. The purpose of this review is to clarify the crucial roles of mitochondrial DNA homeostasis in the pathogenesis of neurodegenerative diseases. Mitochondrial DNA is distributed within nucleoids and is then transcribed into polycistronic mitochondrial DNA molecules within the mitochondrial granule region. Within the ultrastructure of the mitochondrial nucleoid and granule, a group of essential mitochondrial proteins involved in DNA replication, DNA transcription, RNA translation, RNA surveillance, and RNA degradation plays a crucial role in maintaining mitochondrial structure, genome integrity, and mitochondrial DNA processing. The uniparentally inherited mitochondrial DNA undergoes heritable polyploid variations, which include homoplasmy and heteroplasmy. Accumulating mitochondrial DNA alterations, such as deletions, point mutations, and methylations, occur during the pathogenic processes of neurodegenerative diseases. The increased mitochondrial DNA alterations can be propagated by the rise of deleterious heteroplasmy in neurodegenerative diseases, ultimately resulting in impairment to the oxidative phosphorylation system, biogenesis defects, and cellular metabolic dysfunction. Therefore, developing appropriate gene editing tools to rectify aberrant alterations in mitochondrial DNA and targeting the key proteins involved in maintaining mitochondrial DNA homeostasis can be considered promising therapeutic strategies for neurodegenerative diseases. Although therapeutic strategies targeting mitochondrial DNA in diseases show great potential, challenges related to efficacy and safety require a better understanding of the mechanisms underlying mitochondrial DNA alterations in aging and neurodegenerative diseases.

Alzheimer&#x2019;s disease↗

LdCompare: rapid computation of single- and multiple-marker r2 and genetic coverage.

UNLABELLED: The scale of genetic-variation datasets has increased enormously and the linkage equilibrium (LD) structure of these polymorphisms, particularly in whole-genome association studies, is of great interest. The significant computational complexity of calculating single- and multiple-marker correlations at a genome-wide scale remains challenging. We have developed a program that efficiently characterizes whole-genome LD structure on large number of SNPs in terms of single- and multiple-marker correlations. AVAILABILITY: LdCompare is licensed under the GNU General Public License (GPL). Source code, documentation, testing datasets and precompiled executables are available for download at: http://www.affymetrix.com/support/developer/tools/devnettools.affx

Algorithms↗

Human insulin genome sequence map, biochemical structure of insulin for recombinant DNA insulin.

Insulin is a essential molecule for type I diabetes that is marketed by very few companies. It is the first molecule, which was made by recombinant technology; but the commercialization process is very difficult. Knowledge about biochemical structure of insulin and human insulin genome sequence map is pivotal to large scale manufacturing of recombinant DNA Insulin. This paper reviews human insulin genome sequence map, the amino acid sequence of porcine insulin, crystal structure of porcine insulin, insulin monomer, aggregation surfaces of insulin, conformational variation in the insulin monomer, insulin X-ray structures for recombinant DNA technology in the synthesis of human insulin in Escherichia coli.

Animals↗

Genomic Insights Into Local Adaptation Across Heterogeneous Understory Habitats and Climate Change Vulnerability.

Understanding adaptive evolution and survival risks in understory herbs is crucial for the effective conservation of biodiversity. How environmental gradients shape species local adaptation patterns is not well understood, nor is how populations of understory herbs respond to a changing climate. In this study, we conducted population genomic analyses of Adenocaulon himalaicum (Asteraceae) with a pan-East Asian distribution, representing a good model for dominant understory herbs to elucidate adaptation mechanisms in heterogeneous forest ecosystems. Based on 34,398 putatively neutral single nucleotide polymorphisms (SNPs) across 27 populations, we identified three genetic lineages accompanied by high levels of genetic differentiation between populations. Our isolation by environment results (IBE) indicated a significant effect of environmental gradients on genomic variation of A. himalaicum (r&#x2009;=&#x2009;0.18, p&#x2009;=&#x2009;0.03). To decompose the relative contributions of climate, geography and population structure in explaining genetic variance, our partial RDA found that the prominent contribution of environmental effects (climatic and soil variables) explained 29% and 36% of the neutral and adaptive genetic variation, respectively. Using two genotype-environment association (GEA) methods, we identified 13 SNPs as candidates for core climate-related adaptation loci, with two of these loci further validated by qRT-PCR experiments. Projections of spatiotemporal genomic vulnerability under different future climate scenarios revealed that populations in the southeastern edge of the Himalayas, near the Sichuan Basin, the southernmost region of Northeast China and the northern Korean Peninsula, as well as northern Japan, were identified as the most vulnerable and should be prioritised for conservation. Therefore, our current study provides the genomic foundations for conservation and management strategies to elucidate how these understory herbs cope with future climate changes.

Climate Change↗

ASPIC: a web resource for alternative splicing prediction and transcript isoforms characterization.

Alternative splicing (AS) is now emerging as a major mechanism contributing to the expansion of the transcriptome and proteome complexity of multicellular organisms. The fact that a single gene locus may give rise to multiple mRNAs and protein isoforms, showing both major and subtle structural variations, is an exceptionally versatile tool in the optimization of the coding capacity of the eukaryotic genome. The huge and continuously increasing number of genome and transcript sequences provides an essential information source for the computational detection of genes AS pattern. However, much of this information is not optimally or comprehensively used in gene annotation by current genome annotation pipelines. We present here a web resource implementing the ASPIC algorithm which we developed previously for the investigation of AS of user submitted genes, based on comparative analysis of available transcript and genome data from a variety of species. The ASPIC web resource provides graphical and tabular views of the splicing patterns of all full-length mRNA isoforms compatible with the detected splice sites of genes under investigation as well as relevant structural and functional annotation. The ASPIC web resource-available at http://www.caspur.it/ASPIC/--is dynamically interconnected with the Ensembl and Unigene databases and also implements an upload facility.

Algorithms↗

Human Ro60 (SSA2) genomic organization and sequence alterations, examined in cutaneous lupus erythematosus.

BACKGROUND: The Ro 60 kDa protein (Ro60 or SSA2) is the major component of the Ro ribonucleoprotein (Ro RNP) complex, to which an immune response is a specific feature of several autoimmune diseases. The genomic organization and any sequence variation within the DNA encoding Ro60 are unknown. OBJECTIVES: To characterize the Ro60 gene structure and to assess whether any sequence alterations might be associated with serum anti-Ro antibody in subacute cutaneous lupus erythematosus (SCLE), thus potentially providing new insight into disease pathogenesis. METHODS: The cDNA sequence for Ro60 was obtained from the NCBI database and used for a BLAST search for a clone containing the entire genomic sequence. The intron-exon borders were confirmed by designing intronic primer pairs to flank each exon, which were then used to amplify genomic DNA for automated sequencing from 36 caucasian patients with SCLE (anti-Ro positive) and 49 with discoid LE (DLE, anti-Ro negative), in addition to 36 healthy caucasian controls. RESULTS: Heteroduplex analysis of polymerase chain reaction (PCR) products from patients and controls spanning all Ro60 exons (1-8) revealed a common bandshift in the PCR products spanning exon 7. Sequencing of the corresponding PCR products demonstrated an A > G substitution at nucleotide position 1318-7, within the consensus acceptor splice site of exon 7 (GenBank XM001901). The allele frequencies were major allele A (0.71) and minor allele G (0.29) in 72 control chromosomes, with no significant differences found between SCLE patients, DLE patients and controls. CONCLUSIONS: The genomic organization of the DNA encoding the Ro60 protein is described, including a common polymorphism within the consensus acceptor splice site of exon 7. Our delineation of a strategy for the genomic amplification of Ro60 forms a basis for further examination of the pathological functions of the Ro RNP in autoimmune disease.

Antibodies, Antinuclear↗

Recurrent structural variation and recent turnover at the 17q21.31 locus in humans and great apes.

The 17q21.31 locus in humans harbors several complex structural haplotypes including a ~970kb inversion. Different inversion haplotypes have been associated with susceptibility to microdeletions causing Koolen-de Vries syndrome and variation in fecundity and recombination rates. Here, using 210 haplotype-resolved human genome assemblies and pangenome graph-based approaches we characterize 11 distinct structural haplotypes, several of which have not been previously described. Extending our analyses to a set of haplotype-resolved great-ape genomes, we characterize the structure of an independent inversion in chimpanzees which extends an additional 650kb, encompasses 5 additional genes, and is ~2 million years younger than the human inversion. We further determine that gorillas exhibit an independent duplication of the KANSL1 gene which may predispose them to Koolen-de Vries syndrome causing microdeletions. Using short read sequencing data we characterize 17q21.31 haplotype diversity worldwide in ~5174 individuals from 107 populations finding increased frequencies of KANSL1 duplication-containing haplotypes in both European and South Asian populations as well as 8 double recombination events between inverted and non-inverted haplotypes ranging in size from 20-180kb. Finally, using 626 ancient Eurasian human genomes we show the frequency of haplotypes containing KANSL1 duplications has increased ~6-fold over the past 12 thousand years in Europe. Together, our results highlight the dynamics, complexity, and recurrent, independent evolution of a medically relevant locus across humans and great apes.

Journal Article↗

A combined analysis of genomic and primary protein structure defines the phylogenetic relationship of new members if the T-box family.

T-box genes form an ancient family of putative transcriptional regulators characterized by a region of homology to the DNA-binding domain of the murine Brachyury (T) gene product. This T-box domain is conserved from Caenorhabditis elegans to human, and mutations in T-box genes have been associated with developmental defects in Drosophila, zebrafish, mice, and humans. Here we report the identification of three novel murine T-box genes and an investigation of their evolutionary relationship to previously known family members by studying the genomic structure of the T-box. All T-box genes from nematodes to humans possess a characteristic central intron that presumably was inherited from a common ancestral precursor. Two additional intron positions are also conserved with the exception of two nematode T-box genes. Subsequent intron insertions, potential deletions, and/or intron sliding formed a structural basis for the divergence into distinct subfamilies and a substrate for length variations of the T-box domain. In mice, the 11 T-box genes known to date can be grouped into seven subfamilies. Genes assigned to the same subfamily by genomic structure show related expression patterns. We propose a model for the phylogenetic relationships within the gene family that provides a rationale for classifying new T-box genes and facilitates interspecific comparisons.

Amino Acid Sequence↗

Comparative genomics and evolutionary dynamics of Saccharomyces cerevisiae Ty elements.

The availability of the complete genome sequence of Saccharomyces cerevisiae provides the unique opportunity to study an entire genomic complement of retrotransposons from an evolutionary perspective. There are five families of yeast retrotransposons, Ty1-Ty5. We have conducted a series of comparative sequence analyses within and among S. cerevisiae Ty families in an effort to document the evolutionary forces that have shaped element variation. Our results indicate that within families Ty elements vary little in terms of both size and sequence. Furthermore, intra-element 5'-3' long terminal repeat (LTR) sequence comparisons indicate that almost all Ty elements in the genome have recently transposed. For each family, solo LTR sequences generated by intra-element recombination far outnumber full length insertions. Taken together, these results suggest a rapid genomic turnover of S. cerevisiae Ty elements. The closely related Ty1 and Ty2 are the most numerous elements in the genome. Phylogenetic analysis of full length insertions reveals that reverse transcriptase mediated recombination between Ty1 and Ty2 elements has generated a number of hybrid Ty1/2 elements. These hybrid Ty1/2 elements have similar genomic structures with chimeric LTRs and chimeric TYB (pol) genes. Analysis of the levels of nonsynonymous (Ka) and synonymous (Ks) nucleotide variation indicates that Ty1 and Ty2 coding regions have been subject to strong negative (purifying) selection. Distribution of Ka and Ks on Ty1, Ty2 and Ty1/2 phylogenies reveals evidence of negative selection on both internal and external branches. This pattern of variation suggests that the majority of full length Ty1, Ty2 and Ty1/2 insertions represent active or recently active element lineages and is consistent with a high level of genomic turnover. The evolutionary dynamics of S. cerevisae Ty elements uncovered by our analyses are discussed with respect to selection among elements and the interaction between the elements and their host genome.

Amino Acid Sequence↗

Investigation of malaria susceptibility determinants in the IFNG/IL26/IL22 genomic region.

Interferon-gamma, encoded by IFNG, is a key immunological mediator that is believed to play both a protective and a pathological role in malaria. Here, we investigate the relationship between IFNG variation and susceptibility to malaria. We began by analysing West African and European haplotype structure and patterns of linkage disequilibrium across a 100 kb genomic region encompassing IFNG and its immediate neighbours IL22 and IL26. A large case-control study of severe malaria in a West Africa population identified several weak associations with individual single-nucleotide polymorphisms in the IFNG and IL22 genes, and defined two IL22 haplotypes that are, respectively, associated with resistance and susceptibility. These data provide a starting point for functional and genetic analysis of the IFNG genomic region in malaria and other infectious and inflammatory conditions affecting African populations.

Black People↗

Sequence variations in the public human genome data reflect a bottlenecked population history.

Single-nucleotide polymorphisms (SNPs) constitute the great majority of variations in the human genome, and as heritable variable landmarks they are useful markers for disease mapping and resolving population structure. Redundant coverage in overlaps of large-insert genomic clones, sequenced as part of the Human Genome Project, comprises a quarter of the genome, and it is representative in terms of base compositional and functional sequence features. We mined these regions to produce 500,000 high-confidence SNP candidates as a uniform resource for describing nucleotide diversity and its regional variation within the genome. Distributions of marker density observed at different overlap length scales under a model of recombination and population size change show that the history of the population represented by the public genome sequence is one of collapse followed by a recent phase of mild size recovery. The inferred times of collapse and recovery are Upper Paleolithic, in agreement with archaeological evidence of the initial modern human colonization of Europe.

Databases, Nucleic Acid↗

Mapping Cerebellar Morphology in 15q11.2 CNV Carriers Using Normative Modeling.

Copy number variations (CNVs) at the 15q11.2 locus of the human genome have been associated with altered brain structure and increased risk for neurodevelopmental and neuropsychiatric disorders. The cerebellum is increasingly seen as a crucial brain region for neurodevelopmental conditions, yet the effects of 15q11.2 CNVs on cerebellar morphology remain largely unclear. Importantly, 15q11.2 CNVs shows reduced or incomplete penetrance (meaning that not all CNV carriers are affected) and variable expressivity (meaning that symptoms may differ between individuals with the same genetic alteration). Thus, there is a need to not only assess group differences, but also to quantify anatomical variability at the individual level. Here, we address these issues using normative models of brain anatomy trained on large datasets (n > 52k, age range: 3-85) to assess both group and individual-level deviations in cerebellar anatomy in carriers of 15q11.2 deletions (n = 120, mean [SD] age= 64.95 [7.58]) and duplications (n = 149, mean [SD] age=64.31 [7.21]), compared to non-carriers (n = 19,028, mean [SD] age=64.31 [7.58]). Group-level case-control analyses revealed significantly smaller total and regional cerebellar volumes in both deletion and duplication carriers, though with small effect sizes. Individual-level deviation analyses, capturing pronounced alterations in specific individuals, revealed a heterogeneous pattern among carriers. Overall, our findings suggest that CNVs at the 15q11.2 locus exert modest and highly individualized effects on cerebellar morphology.

15q11.2↗