Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic Structural Variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

Codon usage in nucleopolyhedroviruses.

Phylogenetic analyses based on baculovirus polyhedrin nucleotide and amino acid sequences revealed two major nucleopolyhedrovirus (NPV) clades, designated Group I and Group II. Subsequent phylogenetic analyses have revealed three Group II subclades, designated A, B and C. Variations in amino acid frequencies determine the extent of dissimilarity for divergent but structurally and functionally conserved genes and therefore significantly influence the analysis of phylogenetic relationships. Hence, it is important to consider variations in amino acid codon usage. The Genome Hypothesis postulates that genes in any given genome use the same coding pattern with respect to synonymous codons and that genes in phylogenetically related species generally show the same pattern of codon usage. We have examined codon usage in six genes from six NPVs and found that: (1) there is significant variation in codon use by genes within the same virus genome; (2) there is significant variation in the codon usage of homologous genes encoded by different NPVs; (3) there is no correlation between the level of gene expression and codon bias in NPVs; (4) there is no correlation between gene length and codon bias in NPVs; and (5) that while codon use bias appears to be conserved between viruses that are closely related phylogenetically, the patterns of codon usage also appear to be a direct function of the GC-content of the virus-encoded genes.

Animals↗

Gene families from the Arabidopsis thaliana pollen coat proteome.

The pollen extracellular matrix contains proteins mediating species specificity and components needed for efficient pollination. We identified all proteins >10 kilodaltons in the Arabidopsis pollen coating and showed that most of the corresponding genes reside in two genomic clusters. One cluster encodes six lipases, whereas the other contains six lipid-binding oleosin genes, including GRP17, a gene that promotes efficient pollination. Individual oleosins exhibit extensive divergence between ecotypes, but the entire cluster remains intact. Analysis of the syntenic region in Brassica oleracea revealed even greater divergence, but a similar clustering of the genes. Such allelic flexibility may promote speciation in plants.

Alleles↗

Variations of human adrenomedullin gene and its relation to cardiovascular diseases.

The studies concerning the structure and variations of the human adrenomedullin (AM) gene are reviewed, and their relations to the gene function and genetic predisposition to cardiovascular diseases are discussed. The genomic human AM gene is composed of four exons, and the whole nucleotide sequence corresponding to mature AM resides in the fourth exon. In chromosomal sublocalization, the AM gene is located in the distal portion of the short arm of chromosome 11 (11p15.1-3). Analysis of the promoter region of the AM gene has revealed that two transcription factors, nuclear factor for interleukin-6 expression (NF-IL6) and activator protein 2 (AP-2), participate in the regulation of AM gene expression. It is surmised that NF-IL6 mediates inflammatory stimuli and AP-2 mediates signals of phospholipase C and protein kinase C activation. In addition to these factors, hypoxia induces AM gene expression via the hypoxia inducible factor-1 (HIF-1) binding site. The 3'-end of the AM gene is flanked by a microsatellite marker of cytosine adenine (CA) repeats. In Japanese, there are four types of alleles with different CA-repeat numbers: 11, 13, 14 and 19. It is suggested that existence of the 19-repeat allele is associated with genetic predispositions to develop essential hypertension and diabetic nephropathy.

Adrenomedullin↗

[Heterogeneity of G protein-coupled receptor generated by post-translational mechanisms and its clinical meanings].

G protein-coupled receptors (GPCRs) are the most famous target proteins for medicinal drugs. So far, heterogeneity of GPCRs is mainly focused on genetic variation. However, it has been reported that the structure and function of GPCRs are modified by several mechanisms after translation. RNA editing introduces the amino acid different from that encoded in genome by changing the nucleotide. Dimer formation is another example of how heterogeneity is produced. Many receptors form homo- or hetero-dimers, and obtain different function from original receptors. Receptors are regulated by several means to modulate stimulation strength. Receptor subtype is often differentially regulated by receptor kinases and/or second messenger-regulated kinases. There is a new type of receptor that shows a novel structural feature, a long amino terminal region belonging to class B seven transmembrane receptors. The physiological function of this class of receptor is assumed to play a role in cell-cell communication. This novel structural feature may directly link GPCR to the cytoskeleton. These mechanisms to produce functional and structural heterogeneity may explain how cells evoke different responses in different tissues or cells upon the same stimulation. Thus, the post-translational mechanism to produce heterogeneity provides additional flexibility when cells respond to one extracellular stimulus.

Animals↗

The Celera Discovery System.

The Celera Discovery System (CDS) is a web-accessible research workbench for mining genomic and related biological information. Users have access to the human and mouse genome sequences with annotation presented in summary form in BioMolecule Reports for genes, transcripts and proteins. Over 40 additional databases are available, including sequence, mapping, mutation, genetic variation, mRNA expression, protein structure, motif and classification data. Data are accessible by browsing reports, through a variety of interactive graphical viewers, and by advanced query capability provided by the LION SRS search engine. A growing number of sequence analysis tools are available, including sequence similarity, pattern searching, multiple sequence alignment and Hidden Markov Model search. A user workspace keeps track of queries and analyses. CDS is widely used by the academic research community and requires a subscription for access. The system and academic pricing information are available at http://cds.celera.com.

Animals↗

Comparative genomics of large mitochondria in placozoans.

The first sequenced mitochondrial genome of a placozoan, Trichoplax adhaerens, challenged the conventional wisdom that a compact mitochondrial genome is a common feature among all animals. Three additional placozoan mitochondrial genomes representing highly divergent clades have been sequenced to determine whether the large Trichoplax mtDNA is a shared feature among members of the phylum Placozoa or a uniquely derived condition. All three mitochondrial genomes were found to be very large, 32- to 37-kb, circular molecules, having the typical 12 respiratory chain genes, 24 tRNAs, rnS, and rnL. They share with the Trichoplax mitochondrial genome the absence of atp8, atp9, and all ribosomal protein genes, the presence of several cox1 introns, and a large open reading frame containing an intron group I LAGLIDADG endonuclease domain. The differences in mtDNA size within Placozoa are due to variation in intergenic spacer regions and the presence or absence of long open reading frames of unknown function. Phylogenetic analyses of the 12 respiratory chain genes support the monophyly of Placozoa. The similarities in composition and structure between the three mitochondrial genomes reported here and that of Trichoplax's mtDNA suggest that their uncompacted state is a shared ancestral feature to other nonmetazoans while their gene content is a derived feature shared only among the Metazoa.

Amino Acid Sequence↗

Short sequence repeats in microbial pathogenesis and evolution.

Repetitive DNA is ubiquitous in microbial genomes. Different classes of short sequence repeats (SSRs) have been identified and demonstrated to be generally heterogeneous in a locus-dependent manner, reflected in variation in the number of repeat units present at a given genomic site or by sequence heterogeneity among individual units. Both types of variability can be used to assess intra-species genetic diversity. Repeat variability often affects the coding potential of the region in which the repetitive element is located. This implies that determination of the primary structure of variable numbers of tandem repeats can be used for epidemiological identification purposes, and also for the analysis of gene function. Precise assessment of SSR structure can also generate insight into the regulation of gene expression. Together, DNA repeat analysis in microbial species provides information on both functional and evolutionary aspects of genetic diversity among microbial isolates.

Candida albicans↗

Application of comparative genomics to narrow-leafed lupin (Lupinus angustifolius L.) using sequence information from soybean and Arabidopsis.

The completion of genome-sequencing initiatives for model plants and EST databases for major crop species provides a large resource for gaining fundamental knowledge of complex gene interactions and the functional significance of proteins. There are increasingly numerous opportunities to transfer this information to other plant species with uncharacterized genomes and make advances in genome analysis, gene expression, and predicted protein function. In this study, we have used DNA sequences from soybean and Arabidopsis to determine the feasibility of applying comparative genomics to narrow-leafed lupin. We have used transcribed sequences from soybean and showed that a high proportion cross hybridize to lupin DNA, identifying similar genes and providing landmarks for estimating the degree of chromosomal synteny between species. To further investigate comparative relationships in this study, a detailed analysis of three lupin genes and comparison of orthologs from soybean and Arabidopsis shows that, in some cases, gene structure and expression are highly conserved and their proteins may have similar function. In other cases, genes show variation in expression profiles indicating alternative functions across species. The advantages and limitation of using soybean and Arabidopsis sequences for comparative genomics in lupins are discussed.

Amino Acid Sequence↗

Redefining bacterial populations: a post-genomic reformation.

Sexual reproduction and recombination are essential for the survival of most eukaryotic populations. Until recently, the impact of these processes on the structure of bacterial populations has been largely overlooked. The advent of large-scale whole-genome sequencing and the concomitant development of molecular tools, such as microarray technology, facilitate the sensitive detection of recombination events in bacteria. These techniques are revealing that bacterial populations are comprised of isolates that show a surprisingly wide spectrum of genetic diversity at the DNA level. Our new awareness of this genetic diversity is increasing our understanding of population structures and of how these affect host pathogen relationships.

Bacteria↗

Using natural allelic diversity to evaluate gene function.

Genomics has developed a wide range of tools to identify genes that play roles in specific pathways. However, relating individual genes and alleles to agronomic traits is still quite challenging. We describe how association analysis can be used to relate natural variation at candidate genes with agronomic phenotypes. Association approaches in plants can provide very high resolution and can evaluate a wide range of alleles rapidly. We discuss issues related to experimental design, germplasm sample, molecular assay, population structure, and statistical analysis necessary for association analysis in plants.

Alleles↗

KEGG as a glycome informatics resource.

Bioinformatics approaches to carbohydrate research have recently begun using large amounts of protein and carbohydrate data. In this field called glycome informatics, the foremost necessity is a comprehensive resource for genome-scale bioinformatics analysis of glycan data. Although the accumulation of experimental data may be useful as a reference of biological and biochemical information on carbohydrates, this is insufficient for bioinformatics analysis. Thus, we have developed a glycome informatics resource (http://www.genome.jp/kegg/glycan/) in KEGG (Kyoto Encyclopedia of Genes and Genomes), an integrated knowledge base of protein networks, genomic information, and chemical information. This review describes three noteworthy features: (1) GLYCAN, a database of carbohydrate structures; (2) glycan-related pathways; and (3) Composite Structure Map (CSM), a map illustrating all possible variations of carbohydrate structures within organisms. GLYCAN includes two useful tools: an intuitive drawing tool called KegDraw, and an efficient glycan search and alignment tool called KEGG Carbohydrate Matcher (KCaM). KEGG's glycan biosynthesis and metabolism pathways, integrating carbohydrate structures, proteins, and reactions, are also a pivotal resource. CSM is constructed as a bridge between carbohydrate functions and structures. CSM is able to display, for example, expression data of glycosyltransferases in a compact manner. In all the KEGG resources, various objects including KEGG pathways, chemical compounds, as well as carbohydrate structures are commonly represented as graphs, which are widely studied and utilized in the computer science field.

Carbohydrates↗

Genomic structure and mutations in adipose-specific gene, adiponectin.

BACKGROUND: Adiponectin is a collagen-like plasma protein specifically synthesized in adipose tissue. Plasma adiponectin concentrations are decreased in obesity whereas it is adipose-specific. OBJECTIVE: To clarify the significance of the genetic variations in adiponectin gene on its plasma concentrations and obesity. SUBJECTS: Two hundred and nineteen unrelated adult Japanese subjects (123 men and 96 women, age: 20-83 y, BMI: 16-43 kg/m2) including 77 obese subjects (BMI>26.4 kg/m2). MEASUREMENT: Human adiponectin gene was isolated from PAC DNA pools. Mutations in the adiponectin gene were screened by direct sequencing or restriction-fragment polymorphism. The levels of plasma adiponectin were determined by the enzyme-linked immunosorbent assay (ELISA). RESULTS: Adiponectin gene spanned 17 kb on chromosome 3q27, consisting of three exons and two introns. Within 2.1 kb of the 5'-flanking region, there were two octamer elements present in the promoter of adipsin. Two nucleotide changes were identified. One was a polymorphism (G/T) occurring in exon 2, and the other was a missense mutation (R112C) in exon 3. The mean plasma adiponectin levels of the subjects carrying G allele were low (G/G: 4.5 microg/ml; G/T: 5.9 microg/ml; and T/T: 6.3 microg/ml), but were not statistically significant. The allelic frequency between the obese and the non-obese showed no significant difference. The subject carrying R112C mutation showed markedly low concentration of plasma adiponectin. CONCLUSION: Two nucleotide changes have been identified in the adiponectin gene. G/T polymorphism in exon 2 was associated with neither plasma adiponectin concentrations nor the presence of obesity. A subject carrying missense mutation (R112C) showed markedly low plasma adiponectin concentration.

Adipocytes↗

Comparison of the rate of sequence variation in the hypervariable region of E2/NS1 region of hepatitis C virus in normal and hypogammaglobulinemic patients.

The hypervariable region (HVR) of the E2/NS1 region of hepatitis C virus (HCV) varies greatly between viral isolates with high rates of genomic change reported during the course of chronic infection. The HVR is thought to encode a structurally unconstrained envelope protein containing several linear B cell epitopes recognized by neutralizing antibody. It has been postulated that amino acid changes in the HVR could result from humoral immune pressure leading to the selection of escape mutants. The aim of this study was to compare the rates of nucleotide and amino acid variation in the HVR of control patients to patients with common variable immunodeficiency (CVID) where the effect of the humoral immune system is reduced. Five controls and four patients with CVID were studied. Serum samples were taken over periods of between 1 and 6 years. HCV was detected by polymerase chain reaction (PCR) with primers derived from conserved flanking regions of the HVR. PCR products were cloned into a plasmid vector and recombinant clones identified by restriction enzyme digestion. Purified DNA from at least three individual clones from each time point was sequenced by the dideoxynucleotide chain-termination method. Consensus sequences were extracted from the three clones, and the DNA and deduced protein sequences were compared. Control patients had a mean rate of nucleotide change of 6.954 nucleotide substitutions per year, compared with patients with CVID with a rate of 0.415 nucleotide substitutions per year (P < .02). The corresponding rates for amino acid variation were 3.868 amino acid substitutions per year for the control patients compared with 0.185 amino acid substitutions per year for the patients with CVID. These findings suggest that in the absence of humoral immune selective pressure, the frequency of occurrence of genetic variation in the major viral species is reduced. The mutations occur, but in the absence of immune selection remain as minor species. The evolution of viral mutants capable of evading the host's immune system may contribute to the ability of HCV to establish chronic infection.

Adult↗

Microarray analysis of transposition targets in Escherichia coli: the impact of transcription.

Transposable elements have influenced the genetic and physical composition of all modern organisms. Defining how different transposons select target sites is critical for understanding the biochemical mechanism of this type of recombination and the impact of mobile genes on chromosome structure and function. Phage Mu replicates in Gram-negative bacteria using an extremely efficient transposition reaction. Replicated copies are excised from the chromosome and packaged into virus particles. Each viral genome plus several hundred base pairs of host DNA covalently attached to the prophage right end is packed into a virion. To study Mu transposition preferences, we used DNA microarray technology to measure the abundance of >4,000 Escherichia coli genes in purified Mu phage DNA. Insertion hot- and cold-spot genes were found throughout the genome, reflecting >1,000-fold variation in utilization frequency. A moderate preference was observed for genes near the origin compared to terminus of replication. Large biases were found at hot and cold spots, which often include several consecutive genes. Efficient transcription of genes had a strong negative influence on transposition. Our results indicate that local chromosome structure is more important than DNA sequence in determining Mu target-site selection.

Bacteriophage mu↗

Spectral karyotyping suggests additional subsets of colorectal cancers characterized by pattern of chromosome rearrangement.

The abundant chromosome abnormalities in most carcinomas are probably a reflection of genomic instability present in the tumor, so the pattern and variability of chromosome abnormalities will reflect the mechanism of instability combined with the effects of selection. Chromosome rearrangement was investigated in 17 colorectal carcinoma-derived cell lines. Comparative genomic hybridization showed that the chromosome changes were representative of those found in primary tumors. Spectral karyotyping (SKY) showed that translocations were very varied and mostly unbalanced, with no translocation occurring in more than three lines. At least three karyotype patterns could be distinguished. Some lines had few chromosome abnormalities: they all showed microsatellite instability, the replication error (RER)+ phenotype. Most lines had many chromosome abnormalities: at least seven showed a surprisingly consistent pattern, characterized by multiple unbalanced translocations and intermetaphase variation, with chromosome numbers around triploid, 6-16 structural aberrations, and similarities in gains and losses. Almost all of these were RER-, but one, LS411, was RER+. The line HCA7 showed a novel pattern, suggesting a third kind of genomic instability: multiple reciprocal translocations, with little numerical change or variability. This line was also RER+. The coexistence in one tumor of two kinds of genomic instability is to be expected if the underlying defects are selected for in tumor evolution.

Colorectal Neoplasms↗

Dynamic evolution of plant mitochondrial genomes: mobile genes and introns and highly variable mutation rates.

We summarize our recent studies showing that angiosperm mitochondrial (mt) genomes have experienced remarkably high rates of gene loss and concomitant transfer to the nucleus and of intron acquisition by horizontal transfer. Moreover, we find substantial lineage-specific variation in rates of these structural mutations and also point mutations. These findings mostly arise from a Southern blot survey of gene and intron distribution in 281 diverse angiosperms. These blots reveal numerous losses of mt ribosomal protein genes but, with one exception, only rare loss of respiratory genes. Some lineages of angiosperms have kept all of their mt ribosomal protein genes whereas others have lost most of them. These many losses appear to reflect remarkably high (and variable) rates of functional transfer of mt ribosomal protein genes to the nucleus in angiosperms. The recent transfer of cox2 to the nucleus in legumes provides both an example of interorganellar gene transfer in action and a starting point for discussion of the roles of mechanistic and selective forces in determining the distribution of genetic labor between organellar and nuclear genomes. Plant mt genomes also acquire sequences by horizontal transfer. A striking example of this is a homing group I intron in the mt cox1 gene. This extraordinarily invasive mobile element has probably been acquired over 1,000 times separately during angiosperm evolution via a recent wave of cross-species horizontal transfers. Finally, whereas all previously examined angiosperm mtDNAs have low rates of synonymous substitutions, mtDNAs of two distantly related angiosperms have highly accelerated substitution rates.

Biological Evolution↗

Structural variation and functional importance of a D-loop-T-loop interaction in valine-accepting tRNA-like structures of plant viral RNAs.

Valine-accepting tRNA-like structures (TLSs) are found at the 3' ends of the genomic RNAs of most plant viruses belonging to the genera Tymovirus, Furovirus, Pomovirus and Pecluvirus, and of one Tobamovirus species. Sequence alignment of these TLSs suggests the existence of a tertiary D-loop-T-loop interaction consisting of 2 bp, analogous to those in the elbow region of canonical tRNAs. The conserved G(18).Psi(55) pair of regular tRNAs is found to covary in these TLSs between G.U (possibly also modified to G.Psi) and A.G. We have mutated the relevant bases in turnip yellow mosaic virus (TYMV) and examined the mutants for symptom development on Chinese cabbage plants and for accumulation of genetic reversions. Development of symptoms is shown to rely on the presence of either A.G or G.U in the original mutants or in revertants. This finding supports the existence and functional importance of this tertiary interaction. The fact that only G.U and A.G are accepted at this position appears to result from steric and energetic limitations related to the highly compact nature of the elbow region. We discuss the implications of these findings for the various possible functions of the valine-accepting TLS.

Base Sequence↗

Deep DNA and protein level feature integration for robust clinical variant interpretation using probabilistic gradient boosting.

A major challenge in clinical genomics is to classify genetic variations correctly, since it directly affects disease diagnosis and personal care. The existing methods tend to be based on the combination of different factors, such as protein structure, population frequencies, phenotypic annotations, and sequence conservation. Nevertheless, these methods often cannot be used to achieve the necessary interpretability, quantify uncertainty, and address rare cases. This paper presents a probabilistic gradient boosting model on variant pathogenicity prediction. The suggested framework applies biological characteristics at both level of DNA and protein levels while also scaling the level of uncertainty in clinical decision making. Our machine learning aims to solve the issues of variant interpretation by managing the features and through probability-based pathogenicity prediction. The framework formulation is aimed at generalizing over various datasets and minimizing overfitting. At the same time, it can ensure reasonable performance to facilitate clinical experiments. The model has also been tested on three standard datasets and demonstrated to be more predictive of the pathogenic effect of variants, in comparison with a variety of existing tools. The probabilistic gradient boosting model proposed had ROC AUC values of 0.9293, 0.9610, and 0.9646 on ClinVar variants, GRCh37, and GRCh38 human genome respectively. Furthermore, the dataset was ensured to include both exonic and intronic variants, and Variants of Uncertain Significance were also taken into consideration for Performance Testing. Through this it also aims to provide better clinical significance which will lead to a good interpretable tool for priority of variants for a large variety of disease conditions.

ClinVar↗