Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic Structural Variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Conserved RNA secondary structures in viral genomes: a survey.

SUMMARY: The genomes of RNA viruses often carry conserved RNA structures that perform vital functions during the life cycle of the virus. Such structures can be detected using a combination of structure prediction and co-variation analysis. Here we present results from pilot studies on a variety of viral families performed during bioinformatics computer lab courses in past years.

Algorithms↗

Minisatellite variant repeat (MVR) mapping: analysis of 'null' repeat units at D1S8.

Minisatellite variant repeat mapping by PCR (MVR-PCR) is a new approach to studying variation in human DNA which analyses interspersion patterns of variant repeats within minisatellite arrays. MVR-PCR has been applied to the hypervariable human minisatellite D1S8 which contains two major classes of variant 29bp repeat units designated a-type and t-type. The MVR-PCR assay uses a- or t-type specific primers, together with an amplimer at a fixed site in the DNA flanking the minisatellite, to reveal the interspersion patterns of variant repeats along an allele. Extreme levels of variation are seen both in the internal structures of individual alleles and in the digital code generated from the two superimposed alleles in total genomic DNA. However, occasional repeat units fail to amplify in MVR-PCR, signifying the existence of further repeat sequence variants termed 'null' or O-type repeats. Although not significant in individual identification, correct genotyping of null repeats is important when using MVR digital codes in parentage analysis. We have therefore characterised these null repeats and show that most null repeats share a common variant repeat sequence. We discuss the possible origins of null repeats and their application to paternity testing and the analysis of minisatellite evolution.

Alleles↗

Mass spectrometry and tandem mass spectrometry, alone or after liquid chromatography, for analysis of polymerase chain reaction products in the detection of genomic variation.

The availability of the sequences of entire bacterial and human genomes has opened up tremendous opportunities in biomedical research. The next stage in genomics will include utilizing this information to obtain a clearer understanding of molecular diversity among pathogens (helping improved identification and detection) and among normal and diseased people (e.g. aiding cancer diagnosis). To delineate such differences it may sometimes be necessary to sequence multiple representative genomes. However, often it may be adequate to delineate structural differences between genes among individuals. This may be readily achieved by high-throughput mass spectrometry analysis of polymerase chain reaction products.

Chromatography, High Pressure Liquid↗

H2AX: tailoring histone H2A for chromatin-dependent genomic integrity.

During the last decade, chromatin research has been focusing on the role of histone variability as a modulator of chromatin structure and function. Histone variability can be the result of either post-translational modifications or intrinsic variation at the primary structure level: histone variants. In this review, we center our attention on one of the most extensively characterized of such histone variants in recent years, histone H2AX. The molecular phylogeny of this variant seems to have run in parallel with that of the major canonical somatic H2A1 in eukaryotes. Functionally, H2AX appears to be mainly associated with maintaining the genome integrity by participating in the repair of the double-stranded DNA breaks exogenously introduced by environmental damage (ionizing radiation, chemicals) or in the process of homologous recombination during meiosis. At the structural level, these processes involve the phosphorylation of serine at the SQE motif, which is present at the very end of the C-terminal domain of H2AX, and possibly other PTMs, some of which have recently started to be defined. We discuss a model to account for how these H2AX PTMs in conjunction with chromatin remodeling complexes (such as INO80 and SWRI) can modify chromatin structure (remodeling) to support the DNA unraveling ultimately required for DNA repair.

Amino Acid Sequence↗

Contribution of sequence variation in Drosophila actins to their incorporation into actin-based structures in vivo.

Actin is a highly conserved protein important for many cellular functions including motility, contraction in muscles and intracellular transport. Many eukaryotic genomes encode multiple actin protein isoforms that differ from each other by only a few residues. We addressed whether the sequence differences between actin paralogues in one species affect their ability to integrate into the large variety of structures generated by filamentous actin. We thus ectopically expressed all six Drosophila actins as fusion proteins with green fluorescent protein (GFP) in a variety of embryonic, larval and adult fly tissues. We found that each actin was able to integrate into most actin structures analysed. For example, in contrast to studies in mammalian cells, the two Drosophila cytoplasmic actins were incorporated into muscle sarcomeres. However, there were differences in the efficiency with which each actin was incorporated into specific actin structures. The most striking difference was observed within the Z-lines of the sarcomeres: one actin was specifically excluded and we mapped this feature to one or both of two residues within the C-terminal half of the protein. Thus, in Drosophila, the primary sequence of different actins does affect their ability to incorporate into actin structures, and so specific GFPactins may be used to label certain actin structures particularly well.

Actins↗

An integrated human immunoglobulin germline resource linking allele diversity to expressed repertoire structure.

Human immunoglobulin (IG) loci are highly polymorphic, yet existing germline resources remain noisy and incomplete, limiting our ability to link inherited variation to antibody repertoires. Here, we integrate high-fidelity long-read genomic sequencing with matched adaptive immune receptor repertoire sequencing (AIRR-seq) to construct HUSA, a population-scale, evidence-resolved germline resource. Using a conservative allele inference framework, HUSA expands current references more than three-fold, identifying over 1300 alleles while preserving allele-level evidence provenance across genomic and repertoire data. By linking genotype and expressed repertoires within individuals, we show that coding-region similarity predicts the structure of adjacent recombination signal sequences and leader regions, revealing that IG alleles are organized as linked cis-regulatory units associated with differences in recombination context and allele usage. These results define key germline constraints shaping repertoire formation and establish a robust, genotype-aware foundation for the analysis of immune receptor repertoires.

Journal Article↗

Genomic-based revelation of genetic structure and adaptive characterization of Schizopygopsis malacanthus in the Jinsha River and Yalong River.

BACKGROUND: As a highly specialized class of schizothoracine fishes, Schizopygopsis malacanthus has attracted much attention due to its widespread distribution. To investigate the impact of the Qinghai‒Tibet movement on S. malacanthus, we analyzed the genetic evolutionary history of this species. RESULTS: These results showed that there was a high level of genetic differentiation between Jinsha River (JSR) populations and Yalong River (YLR) populations. The genetic diversity of intra-YLR populations was higher than that of the intra-JSR populations. There was gene exchange of the Suwalong population to the Huoqu and Ganzi populations. Furthermore, both of the JSR and YLR populations exhibited a gradual increase in the genetic differentiation index from low to high altitudes, and the effective population of high-elevation populations has gradually expanded. In high-altitude populations, the selected genes were enriched in DNA repair, light transduction, and energy metabolism, reflecting the genetic basis for their migration to higher altitudes. CONCLUSIONS: S. malacanthus populations had the higher genetic differentiation and genetic diversity in the JSR and its main tributary YLR. Therefore, we should preserve high-elevation natural river sections as much as possible and reserve habitats for their migration and diffusion.

Animals↗

Sequence analysis of 22 kDa-like alpha-coixin genes and their comparison with homologous zein and kafirin genes reveals highly conserved protein structure and regulatory elements.

Several genomic and cDNA clones encoding the 22 kDa-like alpha-coixin, the alpha-prolamin of Coix seeds, were isolated and sequenced. Three contiguous 22 kDa-like alpha-coixin genes designated alpha-3A, alpha-3B and alpha-3C were found in the 15 kb alpha-3 genomic clone. The alpha-3A and alpha-3C genes presented in-frame stop codons at position +652. The two genes with truncated ORFs are flanking the alpha-3B gene, suggesting that the three alpha-coixin genes may have arisen by tandem duplication and that the stop codon was introduced before the duplication. Comparison of the deduced amino acid sequences of alpha-coixin clones with the published sequences of 22 kDa alpha-zein and 22 kDa-like alpha-kafirin revealed a highly conserved protein structure. The protein consists of an N-terminus, containing the signal peptide, followed by ten highly conserved tandem repeats of 15-20 amino acids flanked by polyglutamines, and a short C-terminus. The difference between the 22 kDa-like alpha-prolamins and the 19 kDa alpha-zein lies in the fact that the 19 kDa protein is exactly one repeat motif shorter than the 22 kDa proteins. Several putative regulatory sequences common to the zein and kafirin genes were identified within both the 5' and 3' flanking regions of alpha-3B. Nucleotide sequences that match the consensus TATA, CATC and the ca. -300 prolamin box are present at conserved positions in alpha-3B relative to zein and kafirin genes. Two putative Opaque-2 boxes are present in alpha-3B that occupies approximately the same positions as those identified for the 22 kDa alpha-zein and alpha-kafirin genes. Southern hybridization, using a fragment of a maize Opaque-2 cDNA clone as a probe, confirmed the presence of Opaque-2 homologous sequences in the Coix and sorghum genomes. The overall results suggest that the structural and regulatory genes involved in the expression of the 22 kDa-like alpha-prolamin genes of Coix, sorghum and maize, originated from a common ancestor, and that variations were introduced in the structural and regulatory sequences after species separation.

Amino Acid Sequence↗

Identification of a nitroimidazo-oxazine-specific protein involved in PA-824 resistance in Mycobacterium tuberculosis.

PA-824 is a promising new compound for the treatment of tuberculosis that is currently undergoing human trials. Like its progenitors metronidazole and CGI-17341, PA-824 is a prodrug of the nitroimidazole class, requiring bioreductive activation of an aromatic nitro group to exert an antitubercular effect. We have confirmed that resistance to PA-824 (a nitroimidazo-oxazine) and CGI-17341 (a nitroimidazo-oxazole) is most commonly mediated by loss of a specific glucose-6-phosphate dehydrogenase (FGD1) or its deazaflavin cofactor F420, which together provide electrons for the reductive activation of this class of molecules. Although FGD1 and F420 are necessary for sensitivity to these compounds, they are not sufficient and require additional accessory proteins that directly interact with the nitroimidazole. To understand more proximal events in the reductive activation of PA-824, we examined mutants that were wild-type for both FGD1 and F420 and found that, although these mutants had acquired high-level resistance to PA-824 (and another nitroimidazo-oxazine), they retained sensitivity to CGI-17341 (and a related nitroimidazo-oxazole). Microarray-based comparative genome sequencing of these mutants identified lesions in Rv3547, a conserved hypothetical protein with no known function. Complementation with intact Rv3547 fully restored sensitivity to nitroimidazo-oxazines and restored the ability of Mtb to metabolize PA-824. These results suggest that the sensitivity of Mtb to PA-824 and related compounds is mediated by a protein that is highly specific for subtle structural variations in these bicyclic nitroimidazoles.

Bacterial Proteins↗

Systematic analysis of human kinase genes: a large number of genes and alternative splicing events result in functional and structural diversity.

BACKGROUND: Protein kinases are a well defined family of proteins, characterized by the presence of a common kinase catalytic domain and playing a significant role in many important cellular processes, such as proliferation, maintenance of cell shape, apoptosis. In many members of the family, additional non-kinase domains contribute further specialization, resulting in subcellular localization, protein binding and regulation of activity, among others. About 500 genes encode members of the kinase family in the human genome, and although many of them represent well known genes, a larger number of genes code for proteins of more recent identification, or for unknown proteins identified as kinase only after computational studies. RESULTS: A systematic in silico study performed on the human genome, led to the identification of 5 genes, on chromosome 1, 11, 13, 15 and 16 respectively, and 1 pseudogene on chromosome X; some of these genes are reported as kinases from NCBI but are absent in other databases, such as KinBase. Comparative analysis of 483 gene regions and subsequent computational analysis, aimed at identifying unannotated exons, indicates that a large number of kinase may code for alternately spliced forms or be incorrectly annotated. An InterProScan automated analysis was performed to study domain distribution and combination in the various families. At the same time, other structural features were also added to the annotation process, including the putative presence of transmembrane alpha helices, and the cystein propensity to participate into a disulfide bridge. CONCLUSION: The predicted human kinome was extended by identifying both additional genes and potential splice variants, resulting in a varied panorama where functionality may be searched at the gene and protein level. Structural analysis of kinase proteins domains as defined in multiple sources together with transmembrane alpha helices and signal peptide prediction provides hints to function assignment. The results of the human kinome analysis are collected in the KinWeb database, available for browsing and searching over the internet, where all results from the comparative analysis and the gene structure annotation are made available, alongside the domain information. Kinases may be searched by domain combinations and the relative genes may be viewed in a graphic browser at various level of magnification up to gene organization on the full chromosome set.

Algorithms↗

The structure of hepatitis B envelope and molecular variants of hepatitis B virus.

Accumulated evidence in recent years has shown that the variation of hepatitis B virus (HBV) genomes may have profound implications for our understanding of hepatitis B pathogenesis and prevention. Attention has focused on areas of the outer envelope coded by the S gene which are involved in the induction of a protective neutralising antibody response, and mutations which directly affect the production of C gene products, one of which is considered as a target for immune T cells involved in virus clearance. This review highlights recent experimental data which emphasizes the role of such mutations in the establishment and maintenance of chronic HBV infections and focuses attention on the significance of HBV variants with respect to the expanding use of HBV vaccines for mass immunization.

Amino Acid Sequence↗

Human ARHGDIG, a GDP-dissociation inhibitor for Rho proteins: genomic structure, sequence, expression analysis, and mapping to chromosome 16p13.3.

GDP-dissociation inhibitors (GDIs) play a primary role in modulating the activity of GTPases. We recently reported the identification of a new GDI for the Rho-related GTPases named RhoGDIgamma. This gene is now designated ARHGDIG by HUGO. Here, in a detailed analysis of tissue expression of ARHGDIG, we observe high levels in the entire brain, with regional variations. The mRNA is also present at high levels in kidney and pancreas and at moderate levels in spinal cord, stomach, and pituitary gland. In other tissues examined, the mRNA levels are very low (lung, trachea, small intestine, colon, placenta) or undetectable. RT-PCR analysis of total RNA isolated from exocrine pancreas and islets shows that the gene is expressed in both tissues. We also report the genomic structure of ARHGDIG. The gene spans over 4 kb and is organized into six exons and five introns. The upstream region lacks a canonical TATA box and contains several putative binding sites for ubiquitous and tissue-specific factors active in central nervous system development. Using FISH, we have mapped the gene to chromosome band 16p13.3. This band is rich in deletion mutants of genes involved in several human diseases, notably polycystic kidney disease, alpha-thalassemia, tuberous sclerosis, mental retardation, and cancer. The promoter structure and the chromosomal location of RhoGDIgamma suggest its importance and underscore the need for further investigation into its biology.

Base Sequence↗

Structure of the intergenic spacer region from the ribosomal RNA gene family of white spruce (Picea glauca).

Five genomic clones containing ribosomal DNA repeats from the gymnosperm white spruce (Picea glauca) have been isolated and characterized by restriction enzyme analysis. No nucleotide variation or length variation was detected within the region encoding the ribosomal RNAs. Four clones which contained the intergenic spacer (IGS) region from different rDNA repeats were further characterized to reveal the sub-repeat structure within the IGS. The sub-repeats were unusually long, ranging from 540 to 990 bp but in all other respects the structure of the IGS was very similar to the organization of the IGS from wheat, Drosophila and Xenopus.

DNA, Ribosomal↗

Scaling up orphan crop research: genebank genetics highlight geographic structure in cultivated cowpea from 10 617 global accessions.

Vigna unguiculata (L.) Walp. is a dryland legume crop, providing essential food and nutritional security for millions of people across the semi-arid tropics, in Africa, Asia and Latin America. However, as a typical 'orphan crop', cowpea has long remained underrepresented in global genomic research to support crop improvement. Here, we conducted the largest genetic diversity analysis of cowpea to date, comprising 10 617 accessions sourced from seven international collections. Using genotyping-by-sequencing, we characterised the global patterns of genetic diversity, assessed redundancy within and across collections, and examined the geographic structure of the cowpea global allele pool. Our results revealed nine distinct genetic groups with clear geographic associations and fine-scale population differentiation, reflecting dispersal history, regional adaptation and the influence of modern breeding. Duplication across collections was detected, highlighting the need for improved curation and integration of germplasm resources. Landraces from sub-Saharan Africa do not fully capture the genetic diversity present in several other geographic regions, indicating the existence of abundant and untapped genetic resources worldwide. These findings not only provide insights into the genetic structure and evolutionary history of cowpea but also offer a valuable foundation for harnessing global germplasm diversity to enhance breeding potential and accelerate crop improvement.

Vigna↗

Genomics, mutations and the Internet: the naming and use of parts.

Mutations are the source of genetic variation and diversity; by their effect, some are neutral, others are pathogenic. In contemporary genetics, mutations appear at the interface between genomics (structural and functional) and genetics (heredity), where they serve gene discovery and mapping (genomics) and generate challenges to modify their phenotypic effects (medical genetics). Assuming the human genome harbours 80,000 transcribed genes each possessing at least 100 different (germline) alleles in a typical population, how then to record and recover data on at least 8 million human alleles? Bioinformatics is the essential resource to create the corresponding accessible digital libraries (genomic and locus-specific mutation databases) for this purpose, a goal to which The HUGO Mutation Database Initiative (Science 279: 10-11, 1998) aspires. Guidelines now exist for naming alleles (Hum Mutat 11: 1-3, 1998). The principles behind the practice are illustrated by PAHdb (http:/(/)www.mcgill.ca/ pahdb), a prototype locus-specific mutation database (NAR 26: 220-225, 1998), and by prototype genomic mutation databases (HGMD (NAR 26: 285-287, 1998), http:/(/)www.uwcm.ac.uk/uwcm/mg/hgmd0.h tml; the EBI mutation database, http:/(/)www2.ebi.ac.uk/mutations/; and OMIM, http:/(/)www.ncbi.nlm. nih.gov/Omim.html).

Databases, Factual↗

Structural variants deconstruct the genome.

Common genomic structural variants predispose to deleterious de novo genomic rearrangements. Understanding how they do so will require population studies across the continuum of genomic variation and ethical discussion of the nature and uses of human variation.

Chromosomes, Human↗