Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic Structural Variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40Linked to original sources

Unexpected variation in unique features of the lens-specific type I cytokeratin CP49.

PURPOSE: CP49 is a fiber cell-specific type I cytokeratin, but its function as part of the fiber cell-beaded filament remains unknown. To provide a rational basis for mutational studies that would contribute to an elucidation of function, the study was designed to define elements of CP49s that are highly conserved, discriminate conserved features from species-specific variations, and identify where CP49s have diverged from consensus type I features in their adaptation to selective pressures in the lens. METHODS: The primary sequence and gene structure of CP49 from a third vertebrate order was determined from a combination of cDNA and genomic sequencing. Protein product was characterized by SDS-PAGE and Western blot analysis. Consensus features and phylogenetic relationships were identified by multiple alignment. Coiled-coil analysis was conducted to define central rod domains. RESULTS: Trout CP49 is unique among CP49s in having a 39-amino-acid tail domain and shows both unique sequence and allelic variation at the LNDR motif. Comparison of consensus sequences identified unprecedented divergence between CP49s and other type I cytokeratins, including a shortened central rod domain that is conserved among CP49s, but distinct from type I cytokeratins. CONCLUSIONS: The considerable differences that have emerged between the consensus features of the type I cytokeratins and the CP49s suggest that the beaded filament serves a significantly different function from intermediate filaments in other epithelia and that type I cytokeratins may have limited utility as a model for studies on lens beaded filaments. These differences, in concert with consensus features identified among CP49s, suggest sites that are probably critical to CP49 function in the lens fiber cell.

Amino Acid Sequence↗

A fundamental division in the Alu family of repeated sequences.

The Alu family of repeated sequences from the human genome contains two distinct subfamilies. This division is based on different base preferences in a number of diagnostic sequence positions. One subfamily of the sequences, referred to as the Alu-J subfamily, is very similar to 7SL DNA in these positions. The other subfamily, Alu-S, can be divided further into well-defined branches of sequences. These findings revise the previous picture of the Alu family and expose their complex evolutionary dynamics. They reveal sequence variations of potential importance for the proliferation of Alu repeats and relate them to their structural features. In addition, they open the possibility of using different types of Alu sequences as natural markers for studying genetic rearrangements in the genome.

Base Sequence↗

Characterization of the c-type lysozyme gene family in Anopheles gambiae.

Seven new c-type lysozyme genes were found using the Anopheles gambiae genome sequence, increasing to eight the total number of genes in this family identified in this species. The eight lysozymes in An. gambiae have considerable variation in gene structure and expression patterns. Lys c-6 has the most unusual primary amino acid structure as the predicted protein consists of five lysozyme-like domains. Transcript abundance of each c-type lysozyme was determined by semiquantitative RT-PCR. Lys c-1, c-6 and c-7 are expressed constitutively in all developmental stages from egg to adult. Lys c-2 and c-4 also are found in all stages, but with relatively much higher levels in adults. Conversely, Lys c-3 and c-8 transcripts are highest in larvae. Lys c-1, c-6 and c-7 transcripts are found in nearly all the adult tissue samples examined while Lys c-2 and Lys c-4 are more restricted in their expression. Lys c-1 and c-2 transcripts are clearly immune responsive and are increased significantly 6-12 h post challenge with bacteria. The functional adaptive changes that may have evolved during the expansion of this gene family are briefly discussed in terms of the expression patterns, gene and protein structures.

Amino Acid Sequence↗

Functional recycling of C2 domains throughout evolution: a comparative study of synaptotagmin, protein kinase C and phospholipase C by sequence, structural and modelling approaches.

The C2 domain is one of the most frequent and widely distributed calcium-binding motifs. Its structure comprises an eight-stranded beta-sandwich with two structural types as if the result of a circular permutation. Combining sequence, structural and modelling information, we have explored, at different levels of granularity, the functional characteristics of several families of C2 domains. At the coarsest level, the similarity correlates with key structural determinants of the C2 domain fold and, at the finest level, with the domain architecture of the proteins containing them, highlighting the functional diversity between the various sub-families. The functional diversity appears as different conserved surface patches throughout this common fold. In some cases, these patches are related to substrate-binding sites whereas in others they correspond to interfaces of presumably permanent interaction between other domains within the same polypeptide chain. For those related to substrate-binding sites, the predictions overlap with biochemical data in addition to providing some novel observations. For those acting as protein-protein interfaces, our modelling analysis suggests that slight variations between families are a result of not only complementary adaptations in the interfaces involved but also different domain architecture. In the light of the sequence and structural genomic projects, the work presented here shows that modelling approaches along with careful sub-typing of protein families will be a powerful combination for a broader coverage in proteomics.

Amino Acid Sequence↗

Ribosomes unraveled: The path from variant to impact.

In this issue of Cell Genomics, Rothschild et al.1 reveal how ribosomal RNA diversity impacts ribosome structure and its implications for health and disease. Their innovative methodologies uncover distinct ribosome subtypes with significant structural variations and expression patterns. This work reveals connections to tissue-specific biology and cancer, positing new research avenues.

Ribosomes↗

cDNA sequence diversity and genomic clusters of major surface glycoprotein genes of Pneumocystis carinii.

The major surface glycoprotein (MSG) of Pneumocystis carinii plays a crucial role in the pathobiology of P. carinii, which often causes fatal pneumonia in AIDS patients. The cDNAs encoding MSG antigens were cloned from a lambda gt11 expression library of rat-derived P. carinii by immunoscreening. The cloned cDNAs constituted a gene family containing approximately 70% amino acid identity between subtypes. The diversity of MSG cDNAs was high and reflected the genomic structure of MSG genes clustered in the P. carinii chromosomes. These multiple genes may account for the high-level expression of MSG that could generate potential variations in the cell surface. Moreover, the MSG sequences have significant sequence homology to tropomyosins and myosins, suggesting physical or functional association with the membrane cytoskeleton.

Amino Acid Sequence↗

A vision of how low-coverage sequence data should contribute to genetic evaluation in the future.

Low-coverage sequencing refers to sequencing DNA of individuals to a low depth of coverage (e.g., 0.5X) and imputing that sequence to a genomic sequence based on reference haplotypes from individuals sequenced to a high depth of coverage (e.g., ≥10X). It has been proposed as an alternative to genotyping by Single-nucleotide polymorphisms (SNP) arrays. At least one commercial product based on it is available for agricultural species. Concerns limiting adoption in its current form are: 1) the cost of storing the huge volume of data it generates and 2) whether that additional data will result in improved accuracy of genetic evaluation. This work envisions future implementation of low-coverage sequencing to reduce storage costs and enhance genetic evaluations by leveraging the additional information in the full sequence of the pangenome to account for more genetic variation. We propose addressing the storage issue by representing genomic sequence of an individual in a pair of haplotype arrays with each element pointing to an enumerated haplotype of the sequence within one of approximately 50,000 defined genome segments. Assuming 60 million genomic variants, the infrastructure required to translate the identifier of any enumerated haplotype into its genomic sequence would require less than 10 gigabytes of binary storage. Each haplotype array element would require 2 bytes, so the marginal binary storage required to represent the genomic sequence of an individual would be about 200 kilobytes (KB), similar to the genotypes from a SNP array with 200,000 markers. This assumes no pedigree and no ambiguity of the imputation, though the latter is unrealistic. Strategies to minimize, and when necessary, to manage and efficiently represent ambiguity are proposed. The genomic sequence of an individual could be stored in about 1 KB (binary) if both parents have unambiguous sequences stored as described above. The proposed system for representing the pangenome includes algorithms for read mapping and imputation intended to leverage all known genetic variation in the target population. It is also designed to use sequencing reads generated for imputing the genomic sequence of new individuals to identify unrecognized mutations, crossovers, and structural variants, thus continuously improving the genome representation, especially if widespread use of low-coverage sequencing in livestock industries is realized. This could make improved genetic merit and management of livestock feasible without computational burden.

Animals↗

Genome-wide diversity of chromosomal inversions and their disease relationships.

Chromosomal inversions shape evolution and are implicated in human disease, yet their effects on genomic variation and health outcomes remain poorly understood. We analyze genome-wide human inversion polymorphisms, contrasting single-event and recurrent loci. Inversion recurrence is validated using structured-coalescent simulations. We show that single-event inversions evolve in near-complete isolation: inverted haplotypes show ~16-fold lower diversity and strong differentiation from direct haplotypes (median FST = 0.33). By contrast, recurrent inversions maintain gene flow, resulting in similar diversity across orientations and ~4-fold lower differentiation. We further find marked differences in coding sequence conservation between single-event and recurrent inversions. Using the NIH All of Us biobank, we impute inversions and identify four inversions with significant disease associations. Notably, the 17q21 inversion is associated with reduced risk of cognitive decline (OR=0.919) and breast cancer (OR=0.910) but with increased obesity risk (OR=1.097), consistent with pleiotropic selection. These findings establish inversions as major drivers of human genetic diversity and disease, with evolutionary outcomes critically dependent on recurrence.

Evolution↗

Low frequency of myocilin mutations in Indian primary open-angle glaucoma patients.

Glaucoma is one of the major causes of blindness in the Indian population. Mutations in the myocilin (MYOC) gene have been reported in different populations. However, reports on MYOC mutations in Indian primary open-angle glaucoma (POAG) patients and juvenile open-angle glaucoma (JOAG) patients are sparse. We therefore screened 100 unrelated POAG/JOAG patients for MYOC mutations. Patients with POAG/JOAG were clinically diagnosed. Genomic DNA from such patients was collected and studied for MYOC mutations by direct sequencing. Nucleotide variations were compared with unrelated healthy controls by restriction enzyme digestion. Secondary structure prediction for the sequence variants was performed by Chou-Fasman method. A novel mutation in exon 1 (144 G-->Alpha) resulting in Gln48His substitution was observed in 2% of the patients. Four other polymorphisms were also observed. The novel mutation was seen in four other affected family members of a JOAG patient. The novel mutation was found to alter the secondary structure in the glycosaminoglycan initiation site of the protein. MYOC mutations were found in 2% of the population studied. MYOC gene may not be playing a significant role in causing POAG in the Indian population.

Adolescent↗

Variability and genetic structure of plant virus populations.

Populations of plant viruses, like all other living beings, are genetically heterogeneous, a property long recognized in plant virology. Only recently have the processes resulting in genetic variation and diversity in virus populations and genetic structure been analyzed quantitatively. The subject of this review is the analysis of genetic variation, its quantification in plant virus populations, and what factors and processes determine the genetic structure of these populations and its temporal change. The high potential for genetic variation in plant viruses, through either mutation or genetic exchange by recombination or reassortment of genomic segments, need not necessarily result in high diversity of virus populations. Selection by factors such as the interaction of the virus with host plants and vectors and random genetic drift may in fact reduce genetic diversity in populations. There is evidence that negative selection results in virus-encoded proteins being not more variable than those of their hosts and vectors. Evidence suggests that small population diversity, and genetic stability, is the rule. Populations of plant viruses often consist of a few genetic variants and many infrequent variants. Their distribution may provide evidence of a population that is undifferentiated, differentiated by factors such as location, host plant, or time, or that fluctuates randomly in composition, depending on the virus.

Gene Frequency↗

Highly repetitive elements from Chinese bitterlings (genus Rhodeus, Cyprinidae).

We have isolated and characterized several highly repetitive DNA elements from two species of Chinese bitterlings, Rhodeus atremius suigensis and R. ocellatus ocellatus. They comprise a partly interspersed and partly tandem repetitive family of about 1.0 to 1.3 kb in length. Individual elements showed considerable length variation, but genomic Southern blotting revealed two major length groups. Their restricted presence of these elements among related species and relative copy number differences indicated rapid change of genome structure in this group of fish. The isolated elements may be useful landmarks for further chromosomal studies.

Animals↗

Progress in the development of Helicobacter pylori strain typing methods.

Helicobacter pylori is very different from other Gram negative bacteria that inhabit the human gastroduodenal tract. Its success in adapting to colonise and persist in the stomach is reflected in key features such as unique chemical structure and architecture of lipopolysaccharide, sheathed flagella, genomic diversity, and potent urease activity. Strain diversity within the species is well established and so the challenge is to exploit variations in these features for developing relevant epidemiological typing methods.

DNA, Bacterial↗

Structural variation of type XII collagen at its carboxyl-terminal NC1 domain generated by tissue-specific alternative splicing.

This paper reports the identification of two structural variations in the NC1 domain of rat and mouse type XII collagen. The long NC1 domain encoding 74 amino acids showed homology to chicken type XII and XIV collagens. The short NC1 domain was composed of 19 amino acids. Through genomic DNA analyses, two alternative exons were identified, each of which contained the variable NC1 sequence. With the amino-terminal NC3 splicing alternatives, we propose here a new descriptive nomenclature: types XIIA-1 and XIIB-1 which include a long NC1 sequence encoded by exon 1 (from the 3'-end), and types XIIA-2 and XIIB-2 which include a short NC1 sequence encoded by exon 2. Types XIIA-1 and XIIB-1, the predominant transcripts in 15-day old mouse embryos, showed decreased expression in 17-day old embryos when type XIIB-2 expression was sustained at constant levels. In adult mice, type XIIB-1 associates with ligament and tendon, whereas type XIIB-2 is expressed in various other tissues. The long NC1 domain contains an extended acidic region (pI = 3.4) followed by a terminal basic region (pI = 13.8). Because the short NC1 domain lacks these features, structural variations in the type XII collagen NC1 domain suggests different functional roles in a tissue-specific fashion.

Alternative Splicing↗

The retroviral RNA dimer linkage: different structures may reflect different roles.

Retroviruses are unique among virus families in having dimeric genomes. The RNA sequences and structures that link the two RNA molecules vary, and these differences provide clues as to the role of this feature in the viral lifecycles. This review draws upon examples from different retroviral families. Differences and similarities in both secondary and tertiary structure are discussed. The implication of varying roles for the dimer linkage in related viruses is considered.

Animals↗

Ten families of variant genes encoded in subtelomeric regions of multiple chromosomes of Plasmodium chabaudi, a malaria species that undergoes antigenic variation in the laboratory mouse.

The chromosome ends of human malaria parasites harbour many genes encoding proteins that are exported to the surface of infected red cells, often being involved in host-parasite interactions and immune evasion. Unlike other murine malaria parasites Plasmodium chabaudi undergoes antigenic variation during passage in the laboratory mouse and hence is a model suitable for investigation of switching mechanisms. However, little is known about the subtelomeric regions of P. chabaudi chromosomes and its variable antigens. Here we report 80 kb of sequence from an end of one P. chabaudi chromosome. Hybridization of probes spanning this region to two dimensional pulsed field gels of the genome revealed 10 multicopy gene families located exclusively in subtelomeric regions of multiple P. chabaudi chromosomes, interspersed amongst multicopy intergenic regions. Hence all chromosomes share a common subtelomeric structure, presumably playing a similar role in spatial positioning as the P. falciparum Rep20 sequence. Expression in blood stages, domains characteristic of surface antigens and copy numbers between four and several hundred per genome, indicate a functional role in antigenic variation for some of these families. We identify members of the cir family, as well as novel genes, that although clearly homologous to cir have large low complexity regions in the predicted extracellular domains. Although all families have homologues in other rodent Plasmodium species, four were previously not known to be subtelomeric. Six have homologues in human and simian malarias.

Animals↗

Canine genomics and genetics: running with the pack.

The domestication of the dog from its wolf ancestors is perhaps the most complex genetic experiment in history, and certainly the most extensive. Beginning with the wolf, man has created dog breeds that are hunters or herders, big or small, lean or squat, and independent or loyal. Most breeds were established in the 1800s by dog fanciers, using a small number of founders that featured traits of particular interest. Popular sire effects, population bottlenecks, and strict breeding programs designed to expand populations with desirable traits led to the development of what are now closed breeding populations, with limited phenotypic and genetic heterogeneity, but which are ideal for genetic dissection of complex traits. In this review, we first discuss the advances in mapping and sequencing that accelerated the field in recent years. We then highlight findings of interest related to disease gene mapping and population structure. Finally, we summarize novel results on the genetics of morphologic variation.

Animals↗

Complex trait analysis of the hippocampus: mapping and biometric analysis of two novel gene loci with specific effects on hippocampal structure in mice.

Notable differences in hippocampal structure are associated with intriguing differences in development and behavioral capabilities. We explored genetic and environmental factors that modulate hippocampal size, structure, and cell number using sets of C57BL/6J (B6) and DBA/2J (D2) mice; their F1 and F2 intercrosses (n = 180); and 35 lines of BXD recombinant inbred (RI) strains. Hippocampal weights of the parental strains differ by 20%. Estimates of granule cell number also differ by approximately 20%. Hippocampal weights of RI strains range from 21 to 31 mg, and those of individual F2 mice range from 23 to 36 mg (bilateral weights). Volume and granule cell number are well correlated (r = 0.7-0.8). Significant variation is associated with differences in age and sex. The hippocampus increases in weight by 0.24 mg per month, and those of males are 0.55 mg heavier (bilateral) than those of females. Heritability of variation is approximately 50%, and half of this genetic variation is generated by two quantitative trait loci that map to chromosome 1 (Hipp1a: genome-wide p < 0.005, between 65 and 100 cM) and to chromosome 5 (Hipp5a, p < 0.05, between 15 and 40 cM). These are among the first gene loci known to produce normal variation in forebrain structure. Hipp1a and Hipp5a individually modulate hippocampal weight by 1.0-2.0 mg, an effect size greater than that generated by age or sex. The Hipp gene loci modulate neuron number in the dentate gyrus, collectively shifting the population up or down by as much as 200,000 cells. Candidate genes for the Hipp loci include Rxrg and Fgfr3.

Aging↗