Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic Structural Variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

[Increased variability of (TCC)n microsatelline loci in populations of the parthenogenetic lizard Lacerta unisexualis Darevsky].

In four isolated populations of parthenogenetic Caucasian rock lizard Lacerta unisexualis, variability of (TCC)n loci was examined using multilocus DNA fingerprinting. Unexpectedly high variability of (TCC)n microsatellites was found in all four populations. The mean similarity index was 0.825, which is higher than similarity estimates obtained for other mini- and microsatellite loci in L. unisexualis and parthenogenetic species L. dahli and L. armeniaca studied earlier. The high variation level of (TCC)n loci was shown to be at least partially associated with the presence of a diverged (TCC)n sequence fraction in the L. unisexualis genome. Mutations at some other genetically unstable (TCC)n loci may cause their structural diversity in populations of L. unisexualis.

Animals↗

tRNAs are imported into mitochondria of Trypanosoma brucei independently of their genomic context and genetic origin.

The mitochondrial genome of Trypanosoma brucei does not encode any identifiable tRNAs. Instead, mitochondrial tRNAs are synthesized in the nucleus and subsequently imported into mitochondria. In order to analyse the signals which target the tRNAs into the mitochondria, an in vivo import system has been developed: tRNA variants were expressed episomally and their import into mitochondria assessed by purification and nuclease treatment of the mitochondrial fraction. Three tRNA genes were tested in this system: (i) a mutated version of the trypanosomal tRNA(Tyr); (ii) a cytosolic tRNA(His) of yeast; and (iii) a human cytosolic tRNA(Lys). The tRNAs were expressed in their own genomic context, or containing various lengths of the 5'-flanking sequence of the trypanosomal tRNA(Tyr) gene. In all cases efficient import of each of the tRNAs was observed. We independently confirmed the mitochondrial import of the yeast tRNA(His), since in organello [alpha-32P]ATP-labelling of the 3'-end of the tRNA was inhibited by carboxyatractyloside, a highly specific inhibitor of the mitochondrial adenine nucleotide translocator. Import of heterologous tRNAs in their own genomic contexts supports the conclusion that no specific targeting signals are necessary to import tRNAs into mitochondria of T. brucei, but rather that the tRNA structure itself is sufficient to specify import.

Animals↗

Molecular cloning and nucleotide sequence of a pestivirus genome, noncytopathic bovine viral diarrhea virus strain SD-1.

Genomic RNA of noncytopathic (NCP) bovine viral diarrhea virus (BVDV) strain SD-1 was extracted directly from serum obtained from a persistently infected animal. cDNA was synthesized and amplified by polymerase chain reaction (PCR) before cloning. The complete genomic nucleotide sequence was determined by sequencing at least two different clones from independent PCR reactions. The 5' and 3' end sequences of the SD-1 genome was determined from 5'-3' ligation clones. The complete genome sequence was comprised of 12,308 nucleotides containing one large open reading frame which encodes an amino acid sequence of 3898 residues with a calculated molecular weight of 438 kDa. In contrast to cytopathic (CP) BVDV strain NADL, which contains a cellular RNA insert of 270 nucleotides and CP BVDV strain Osloss, which has an inserted ubiquitin RNA sequence of 228 nucleotides, the NCP strain SD-1 had no insertion along the genome. Sequence comparison with other pestiviruses revealed that the overall nucleotide sequence homologies of SD-1 are 88.6% with NADL, 78.3% with Osloss, 67.1% with HoCV Alfort, and 67.2% with HoCV Brescia. The overall deduced amino acid sequence homologies of SD-1 are 92.7% with NADL, 86.2% with Osloss, 72.5% with HoCV Alfort, and 71.2% with HoCV Brescia. The most conserved nucleotide and amino acid sequences are located in the 5' untranslated region (5'UTR) and nonstructural protein p80 region, respectively. The viral glycoproteins, particularly gp53, and nonstructural proteins p54 and p58 have the lowest homology comparing both nucleotide and amino acid sequences between SD-1 and other pestiviruses. Extensive analyses of amino acid sequences for the viral structural proteins and nonstructural protein p54 regions from five pestiviruses led to the identification of four conserved domains (designated as C1, C2, C3, C4) and three highly variable domains (designated as V1, V2, V3) within this region. The C1, C2, and C3 domains are located in the capsid protein p14, glycoprotein gp48, and gp25, respectively. The C4 domain is located in the junction between gp53 and p54. Interestingly, out of three variable domains, two (V1, V2) are located in the same glycoprotein gp53. The third variable domain is located in the nonstructural protein p54.

Amino Acid Sequence↗

Sequence analyses of human tumor-associated SV40 DNAs and SV40 viral isolates from monkeys and humans.

SV40 DNA has been found associated with several types of human tumors. We now report a sequence comparison of SV40 DNAs from pediatric brain tumors and from osteosarcomas with viral isolates from monkeys and from humans. We analyzed the entire genomic sequences of five isolates, Baylor and VA45-54 strains from monkeys and SVCPC, SVMEN, and SVPML-1 recovered from humans, and compared them to the reference virus SV40-776. The viral sequences were highly conserved, but isolates could be distinguished by variations in the structure of the viral regulatory region and in the nucleotide sequence of the variable domain at the C-terminus of the large T-antigen gene. We conclude that multiple strains of SV40 exist that can be identified on the basis of sequences in these regions of the viral genome. The isolates were more similar to each other and to the Baylor strain than to the reference strain SV40-776. Human isolates SVCPC and SVMEN were found to be identical. The DNAs present in some human brain and bone tumors were authentic SV40 sequences. Many of the C-terminal T-ag sequences associated with human tumors were unique, but some sequences were shared by independent sources. There was no compelling evidence for human-specific strains of SV40 or for tumor type-specific associations, suggesting that SV40 has a relatively broad host range. The source of the viral DNA found in human tumors remains unknown.

Animals↗

Assessment of concordance among genealogical reconstructions from various mtDNA segments in three species of Pacific salmon (genus Oncorhynchus).

Seven segments of mitochondrial DNA (mtDNA), comprising 97% of the mitochondrial genome, were amplified by polymerase chain reaction (PCR) and examined for restriction site variation using 13 restriction endonucleases in three species of Pacific salmon: pink (Oncorhynchus gorbuscha), chum (O. keta) and sockeye (O. nerka) salmon. The distribution of variability across the seven mtDNA segments differed substantially among species. Little similarity in the distribution of variable restriction sites was found even between the mitochondrial genomes of the even- and odd-year broodlines of pink salmon. Significantly different levels of nucleotide diversity were detected among three groups of genes: six NADH-dehydrogenase genes had the highest; two rRNA genes had the lowest; and a group that included genes for ATPase and cytochrome oxidase subunits, the cytochrome b gene, and the control region had intermediate levels of nucleotide diversity. Genealogies of mtDNA haplotypes were reconstructed for each species, based on the variation in all mtDNA segments. The contributions of variation within different segments to resolution of the genealogical trees were compared within each species. With the exception of sockeye salmon, restriction site data from different genome segments tended to produce rather different trees (and hence rather different genealogies). In the majority of cases, genealogical information in different segments of mitochondrial genome was additive rather than congruent. This finding has a relevance to phylogeographic studies of other organisms and emphasizes the importance of not relying on a limited segment of the mtDNA genome to derive a phylogeographic structure.

Animals↗

Phylogenetic comparison of the DEN-2 Mexican isolate with other flaviviruses.

Recent attention has focused on the geographic variation of dengue viruses, since major epidemies may follow introduction of a new virus strain into susceptible populations. We cloned and sequenced a very interesting Mexican isolate (200787/1983) which is antigenically unique by signature analysis with respect to all other dengue-2 topotype viruses. This strain is also unique in biological behavior (neurotropism) and is of epidemiological significance in Mexico. The dengue-2 Mexican isolate sequence information was compared with that of other flaviviruses, analyzing the branching structure of the phylogenetic tree reconstructed from the E gene amino acid sequences. The E glycoprotein, is target for neutralizing antibodies and T-cell responses, and defines the tropism and virulence of flaviviruses. In the phylogram, our strain was located in the position of greatest dissimilarity within serotype-2. Also, frequency analysis of amino acids revealed a very different signature pattern from that found in viral serotype-2.

Amino Acid Sequence↗

[Features of the structure and evolution of complex tandemly organized Bsp-repeats in the fox genome. II. Tissue-specific and recombinant BamHI-dimer sites].

The complex structure of the clustered Bsp-repeats in fox genome seems to have evolved throughout a long period of time as a result of multiplication, recombination and divergence events. The sequence of the subrepeat (SR) approximately 245 b.p long is the basic substructure for the hierarchically arranged Bam HI-repeat 1468 b.p. long. The monomer consists of 3 SRs with a 43-59% homology. A dimer is composed of 2 monomers with a 93% homology. Amplification of the Bsp-repeats during evolution seems to have occurred at least twice: first--on the SR ancestral form level, second--on the monomer level. Despite profound divergence, there are still conservative regions in SRs with sequences homologous to known functional sites in eukaryotes. However qualitative and quantitative composition of most functional motifs is stringently individual in every SR. The performed analysis revealed that throughout evolution SRs acquired significant amount of motifs homologous to promoter and enhancer regions in tissue-specific genes and virus regulatory regions. Functional motifs in separate SRs are being differently grouped. Most inducible motifs are located in the III and II subrepeats, putative promoters--in the II one; elements participating both in transcriptional and replicational processes--mainly in the I subrepeat. A few ensembles of functional motifs remotely resemble extended regulatory regions of some tissue-specific genes. The monomers are potentially capable of ensuring diverse aspects of transcriptional regulation. As a whole, motifs of the 3 SRs are potentially capable of regulating the RNA synthesis periodicity with respect to the cellular cycle, activation and repression of genetical material in response to signals from the environment (AP-1, AP-2, AP-4, T-antigen, etc) and temporal ("octamers") etc. Apart from the BamHI-dimer, a few homologues fragments were isolated from fox genome and sequenced. Some of them were rearranged with respect to the BamHI-dimer. Inversion locally alters the composition of motifs and the sequence acquires new functional potential. Thus, the analysis of the emergence and development of Bsp-repeat structural variations allows us to consider repetitive DNA sequences as an ideal material in constructing multiprofile regulatory sequences.

Animals↗

Chromosomal location effects on gene sequence evolution in mammals.

BACKGROUND: Nucleotide substitution rates and G + C content vary considerably among mammalian genes. It has been proposed that the mammalian genome comprises a mosaic of regions - termed isochores - with differing G + C content. The regional variation in gene G + C content might therefore be a reflection of the isochore structure of chromosomes, but the factors influencing the variation of nucleotide substitution rate are still open to question. RESULTS: To examine whether nucleotide substitution rates and gene G + C content are influenced by the chromosomal location of genes, we compared human and murid (mouse or rat) orthologues known to belong to one of the chromosomal (autosomal) segments conserved between these species. Multiple members of gene families were excluded from the dataset. Sets of neighbouring genes were defined as those lying within 1 centiMorgan (cM) of each other on the mouse genetic map. For both synonymous substitution rates and G + C content at silent sites, neighbouring genes were found to be significantly more similar to each other than sets of genes randomly drawn from the dataset. Moreover, we demonstrated that the regional similarities in G + C content (isochores) and synonymous substitution rate were independent of each other. CONCLUSIONS: Our results provide the first substantial statistical evidence for the existence of a regional variation in the synonymous substitution rate within the mammalian genome, indicating that different chromosomal regions evolve at different rates. This regional phenomenon which shapes gene evolution could reflect the existence of 'evolutionary rate units' along the chromosome.

Animals↗

Genomic approaches to typing, taxonomy and evolution of bacterial isolates.

The current literature on bacterial taxonomy, typing and evolution will be critically examined from the perspective of whole-genome structure, function and organization. The following three categories of DNA band pattern studies will be reviewed: (i) random whole-genome analysis; (ii) specific gene variation and (iii) mobile genetic elements. (i) The use of RAPD, PFGE and AFLP to analyse the whole genome will provide a skeleton of polymorphic sites with exact genomic positions as whole-genome sequence data become available. (ii) Different genes provide different levels of evolutionary information for determining isolate relatedness depending on whether they are highly variable (prone to recombination events and horizontal transfer), housekeeping genes with only a small number of single nucleotide differences between isolates or part of the rrn multigene family that is prone to intragenomic recombination and concerted evolution. Comparative analyses of these different gene classes can provide enhanced information about isolate relatedness. (iii) Mobile genetic elements such as insertion sequences, transposons, plasmids and bacteriophages integrate into the bacterial genome at specific (e.g. tRNA genes) or non-specific sites to alter band patterns produced by PFGE, RAPD or AFLP. From the literature it is not clear what level of genetic element duplication constitutes non-relatedness of isolates. A model is presented that incorporates all of the above genomic characteristics for the determination of isolate relatedness in taxonomic, typing and evolutionary studies.

Bacteria↗

Neotelomeres and Telomere-Spanning Chromosomal Arm Fusions in Cancer Genomes Revealed by Long-Read Sequencing.

Alterations in the structure and location of telomeres are key events in cancer genome evolution. However, previous genomic approaches, unable to span long telomeric repeat arrays, could not characterize the nature of these alterations. Here, we applied both long-read and short-read genome sequencing to assess telomere repeat-containing structures in cancers and cancer cell lines. Using long-read genome sequences that span telomeric repeat arrays, we defined four types of telomere repeat variations in cancer cells: neotelomeres where telomere addition heals chromosome breaks, chromosomal arm fusions spanning telomere repeats, fusions of neotelomeres, and peri-centromeric fusions with adjoined telomere and centromere repeats. Analysis of lung adenocarcinoma genome sequences identified somatic neotelomere and telomere-spanning fusion alterations. These results provide a framework for systematic study of telomeric repeat arrays in cancer genomes, that could serve as a model for understanding the somatic evolution of other repetitive genomic elements.

Telomere↗

Sequence variation in ligand binding sites in proteins.

BACKGROUND: The recent explosion in the availability of complete genome sequences has led to the cataloging of tens of thousands of new proteins and putative proteins. Many of these proteins can be structurally or functionally categorized from sequence conservation alone. In contrast, little attention has been given to the meaning of poorly-conserved sites in families of proteins, which are typically assumed to be of little structural or functional importance. RESULTS: Recently, using statistical free energy analysis of tetratricopeptide repeat (TPR) domains, we observed that positions in contact with peptide ligands are more variable than surface positions in general. Here we show that statistical analysis of TPRs, ankyrin repeats, Cys2His2 zinc fingers and PDZ domains accurately identifies specificity-determining positions by their sequence variation. Sequence variation is measured as deviation from a neutral reference state, and we present probabilistic and information theory formalisms that improve upon recently suggested methods such as statistical free energies and sequence entropies. CONCLUSION: Sequence variation has been used to identify functionally-important residues in four selected protein families. With TPRs and ankyrin repeats, protein families that bind highly diverse ligands, the effect is so pronounced that sequence "hypervariation" alone can be used to predict ligand binding sites.

Binding Sites↗

Molecular scanning of the human PPARa gene: association of the L162v mutation with hyperapobetalipoproteinemia.

Peroxisome proliferator-activated receptor alpha (PPARalpha) is a member of the steroid hormone receptor super family involved in the control of cellular lipid utilization. This makes PPARalpha a candidate gene for type 2 diabetes and dyslipidemia. The aim of this study was to investigate whether genetic variation in the human PPARalpha gene can influence the risk of type 2 diabetes and dyslipidemia among French Canadians. We therefore first determined the genomic structure of human PPARalpha, and then designed intronic primers to sequence the coding region and the exon-intron boundaries of the gene in 12 patients with type 2 diabetes and in 2 nondiabetic subjects. Sequence analysis revealed the presence of a L162V missense mutation in exon 5 of one diabetic patient. Leucine 162 is contained within the DNA binding domain of the human PPARalpha gene, and is conserved among humans, mice, rats, and guinea pigs. We subsequently screened a sample of 121 patients newly diagnosed with type 2 diabetes and their age and sex-matched nondiabetic controls, recruited from the Saguenay-Lac-St-Jean region of Northeastern Quebec, for the presence of the L162V mutation by a PCR-RFLP based method. There was no difference in L162 homozygote or V162 carrier frequencies between diabetics and nondiabetics. However, whether diabetic or not, carriers of the V162 allele had higher plasma apolipoprotein B levels compared to noncarriers (P 5 0.05). To further this association, we screened another sample of 193 nondiabetic subjects recruited in the greater Quebec City area. Carriers of the V162 allele compared with homozygotes of the L162 allele had significantly higher concentrations of plasma total and LDL-apolipoprotein B as well as LDL cholesterol (P </= 0.02). These results suggest an association between the PPARalpha V162 allele and the atherogenic/hyperapolipoprotein B dyslipidemia.

Animals↗

Genome-wide SNP-based genomic diversity and population structure analysis in alpaca populations from Europe and Peru.

This study aimed to analyze the genetic diversity and population structure of alpacas in Germany, Switzerland, and Austria (German-speaking regions, GSR) and to compare with that of the country of origin of the species (Peru). A total of 179 animals from GSR and 151 from Peru were genotyped with a species-specific 76k SNP array. The observed and expected heterozygosity was 0.305 and 0.311 for GSR and 0.310 and 0.312 for Peru. The mean FROH values were 0.029 for GSR and 0.023 for Peru. In general, results show that breeders in both analyzed regions efficiently maintain genetic diversity. Principal component analysis identified the GSR and Peru populations as separate from each other, but the relative proximity of both clusters indicates the shared genetic heritage. FST and XPEHH methods identified genomic regions under selection for traits such as coat color and adaptation. Genome-wide association studies comparing black and brown with white or gray alpacas identified associated genome regions containing the ASIP and KIT genes, respectively. The association of a recently identified keratin locus on chromosome 16 with differences in fleece type in alpacas was confirmed, while the putative causality of a TRPV3 variant was rejected.

Animals↗

The effects of alternative splicing on transmembrane proteins in the mouse genome.

Alternative splicing is a major source of variety in mammalian mRNAs, yet many questions remain on its downstream effects on protein function. To this end, we assessed the impact of gene structure and splice variation on signal peptide and transmembrane regions in proteins. Transmembrane proteins perform several key functions in cell signaling and transport, with their function tied closely to their transmembrane architecture. Signal peptides and transmembrane regions both provide key information on protein localization. Thus, any modification to such regions will likely alter protein destination and function. We applied TMHMM and SignalP to a nonredundant set of proteins, and assessed the effects of gene structure and alternative splicing on predicted transmembrane and signal peptide regions. These regions were altered by alternative splicing in roughly half of the cases studied. Transmembrane regions are divided by introns slightly less often than expected given gene structure and transmembrane region size. However, the transmembrane regions in single-pass transmembranes are divided substantially less often than expected. This suggests that intron placement might be subject to some evolutionary pressure to preserve function in these signaling proteins. The data described in this paper is available online at http://www.affymetrix.com/community/publications/affymetrix/tmsplice/.

Alternative Splicing↗

Genomic structure and organization of kringles type 3 to 10 of the apolipoprotein(a) gene in 6q26-27.

Apolipoprotein(a) [apo(a)] is a highly polymorphic glycoprotein covalently linked to the apolipoprotein B-100 of LDL in a particle called lipoprotein(a) [Lp(a)]. High plasma levels of Lp(a) are associated with coronary as well as peripheral atherosclerosis. Plasma levels of Lp(a) show a remarkable variation ranging from 0.1 mg/dl to over 100 mg/dl. The apo(a) gene shows a size polymorphism which resides in the variable number of kringle domains which resemble plasminogen kringle IV. Ten different types of kringle IV repeats have been described, nine of which (kringle IV type 1 and type 3-10) are each supposed to be present in a single copy. The other kringles, namely kringle IV type 2 repeats, vary in number from 3 to 42 between apo(a) alleles and form the basis for the apo(a) size polymorphism. Although an inverse relationship has been observed between the number of kringle type 2 repeats and plasma levels of Lp(a), there are exceptions to this general finding. Indeed, several individuals have been described with similar apo(a) size alleles but very different plasma levels of Lp(a). Genetic studies have linked these differences to the apo(a) locus on 6q26-27, outlining the importance, besides the kringle type 2 repeats, of other regions of the apo(a) gene in contributing to the interindividual differences in the plasma concentration of Lp(a). One of the candidate regions is represented by the non-repeated type-3 to type-10 kringles which are invariably present in each apo(a) allele and whose structural integrity is playing a critical role in the correct assembly of the Lp(a) particle. Biochemical studies with recombinant wild type and mutagenized apo(a) cDNAs with several alterations of the non-repeated kringles have well documented this latter point. As a starting point to search for genetic variations in these kringles associated with different levels of Lp(a), we are presenting the genome organization of type-3 to 10 kringle along with specific PCR primers for easy analysis from genomic DNA. Restriction as well as partial sequencing analyses of the type-3 to 10 kringles region has also provided interesting clues as to the different evolutionary origin of these types of kringle with respect to the polymorphic type-2 kringles.

Apolipoproteins A↗

Structure and functional genomics of lipopolysaccharide expression in Haemophilus influenzae.

The involvement of genes in the lic loci in H. influenzae LPS expression has been known for some time. However, it was not until recently that it was shown that the lic1 locus contains genes required for phase variable expression of phosphocholine substituents, while genes in the lic2 locus and lgtC are required for expression of the globoside trisaccharide, alpha-D-Galp-(1 --> 4)-beta-D-Galp-(1 --> 4)-beta-D-Glcp (i.e., the pK blood group epitope). The availability of the complete sequence of the H. influenzae strain Rd genome has facilitated significant progress in understanding the role of these and other genes in the expression and biosynthesis of LPS. We have employed a comparative structural fingerprinting strategy to establish the structural relationships among LPS from H. influenzae mutant strains in which putative biosynthesis genes were inactivated. Using this functional genomics approach, we have gained considerable insight into the genetic basis for intra-strain and strain-to-strain variation in epitope expression.

Base Sequence↗

DNA structure constraint is probably a fundamental factor inducing CpG deficiency in bacteria.

MOTIVATION: It has been speculated that CpG dinucleotide deficiency in genomes is a consequence of DNA methylation. However, this hypothesis does not adequately explain CpG deficiency in bacteria. The hypothesis based on DNA structure constraint as an alternative explanation was therefore examined. RESULTS: By comparing real bacterial genomes and Markov artificial genomes in the second order, we found that the core structure of a restricted pattern, the TTCGAA pattern, was under represented in low GC content bacterial genomes regardless of CpG dinucleotide level. This is in contrast to the AACGTT pattern, indicating that the counterselection is context-dependent. Further study discovered nine underrepresented patterns that were supposed to be capable of inducing DNA structure constraint. In summary, most of them are in TTCGNA and TTCGAN patterns in both DNA strands. An explanation is also proposed for the strong correlation between GC content and CpG deficiency. The result of random sequence simulation showed that the occurrences of these patterns were correlated with GC content, as well as the percentage of CpG dinucleotides being trapped in these patterns. Finally, we suggest that the degree of counter-selection against these restricted patterns could be influenced by global GC content of a genome.

Base Sequence↗

Evolution of the cetacean mitochondrial D-loop region.

We sequenced the mitochondrial DNA D-loop regions from two cetacean species and compared these with the published D-loop sequences of several other mammalian species, including one other cetacean. Nucleotide substitution rates, DNA sequence simplicity, possible open reading frames (ORFs), and potential RNA secondary structure were investigated. The substitution rate is an order of magnitude lower than would be expected on the basis of reports on human sequence variation in this region but are consistent with interspecific primate and rodent D-loop sequence variation and with estimates of substitution rates from whole mitochondrial genomes. Deletions/insertions are less common in the cetacean D-loop than in other vertebrate species. Areas of high sequence simplicity (clusters of short repetitive motifs) across the region correspond to areas of high sequence divergence. Three regions predicted to form secondary structures are homologous to such putative structures in other species; however, the presumptive structures most conserved in cetaceans are different from those reported for other taxa. While all three species have possible long ORFs, only a short sequence of seven amino acids is shared with other mammalian species, and those changes that had occurred within it are all nonsynonymous. We conclude that DNA slippage, in addition to point mutation, contributes to the evolution of the D-loop and that regions of conserved secondary structure in cetaceans and an ORF are unlikely to contribute significantly to the conservation of the central region.

Amino Acid Sequence↗