Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequence diversity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

The left end of IS2: a compromise between transpositional activity and an essential promoter function that regulates the transposition pathway.

Cut-and-paste (simple insertion) and replicative transposition pathways are the two classical paradigms by which transposable elements are mobilized. A novel variation of cut and paste, a two-step transposition cycle, has recently been proposed for insertion sequences of the IS3 family. In IS2 this variation involves the formation of a circular, putative transposition intermediate (the minicircle) in the first step. Two aspects of the minicircle may involve its proposed role in the second step (integration into the target). The first is the presence of a highly reactive junction formed by the two abutted ends of the element. The second is the assembly at the minicircle junction of a strong hybrid promoter which generates higher levels of transposase. In this report we show that IS2 possesses a highly reactive minicircle junction at which a strong promoter is assembled and that the promoter is needed for the efficient completion of the pathway. We show that the sequence diversions which characterize the imperfect inverted repeats or ends of this element have evolved specifically to permit the formation and optimal function of this promoter. While these sequence diversions eliminate catalytic activity of the left end (IRL) in the linear element, sufficient sequence information essential for catalysis is retained by the IRL in the context of the minicircle junction. These data confirm that the minicircle is an essential intermediate in the two-step transposition pathway of IS2.

Base Sequence↗

Genetic and linguistic differentiation in the Americas.

The relationship between linguistic differentiation and evolutionary affinities was evaluated in three tribes of the Pacific Northwest. Two tribes (Nuu-Chah-Nulth and Bella Coola) speak Amerind languages, while the language of the third (Haida) belongs to a different linguistic phylum--Na-Dene. Construction of a molecular phylogeny gave no evidence of clustering by linguistic affiliation, suggesting a relatively recent ancestry of these linguistically divergent populations. When the evolutionary affinities of the tribes were evaluated in terms of mitochondrial sequence diversity, the Na-Dene-speaking Haida had a reduced amount of diversity compared to the two Amerind tribes and thus appear to be a biologically younger population. Further, since the sequence diversity between the two Amerind-speaking tribes is comparable to the diversity between the Amerind tribes and the Na-Dene Haida, the evolutionary divergence within the Amerind linguistic phylum may be as great as the evolutionary divergence between the Amerind and Na-Dene phyla. Hence, in the New World, rates of linguistic differentiation appear to be markedly faster than rates of biological differentiation, with little congruence between linguistic hierarchy and the pattern of evolutionary relationships.

Alaska↗

Sequence polymorphisms of the mtDNA control region in a human isolate: the Georgians from Swanetia.

In this work, we analyzed the sequence diversity of the mtDNA control region (HVI and HVII) in a sample of 48 individuals from Swanetia (Georgia), using direct fluorescent-based sequencing methods. We identified 43 different mtDNA haplotypes resulting from 78 polymorphic sites (46 in HVI and 32 in HVII). Most of the variable positions identified in both HVI and HVII were transitions (82.6 and 71.9%, respectively). The frequency of length heteroplasmy in the homopolymeric C-stretch regions was the same for both segments (10.4%). The sequence diversity increased markedly when both hypervariable regions were analyzed jointly (HVI: 0.985, HVII: 0.975, HVI+HVII: 0.994). Accordingly, the probability of two randomly selected sequences matching (random match probability, RMP) decreased from 3.4% (HVI) to 2.6% (HVI+HVII), despite which the RMP values in Georgians remained higher than estimated in most Europeans. This suggests that the variability of maternal lineages tends to be lower in traditional human isolates and, therefore, the potential of discrimination of mtDNA in forensic analysis is more limited in this type of population. The incorporation of HVII data also contributed to the refinement of results regarding the genetic relationships among the samples included in the analyses, which stress the importance of considering HVII in both population and forensic genetics.

Base Sequence↗

Cloning and characterization of receptor kinase class disease resistance gene candidates in Citrus.

The rice gene Xa21 represents a unique class of plant disease resistance ( R) genes with distinct protein structure and broad-spectrum specificity; few sequences or genes of this class have been cloned and characterized in other plant species. Degenerate primers were designed from the conserved motifs in the kinase domains of Xa21 and tomato Pto, and used in PCR amplification to identify this class of resistance gene candidate (RGC) sequences from citrus for future evaluation of possible association with citrus canker resistance. Twenty-nine RGC sequences highly similar to the kinase domain of Xa21 (55%-60% amino-acid identity) were cloned and characterized. To facilitate recovery of full-length gene structures and to overcome RGC mapping limitations, large-insert genomic clones (BACs) were identified, fingerprinted and assembled into contigs. Southern hybridization revealed the presence of 1-3 copies of receptor-like kinase sequences (i.e., clustering) in each BAC. Some of these sequences were sampled by PCR amplification and direct sequencing. Twenty-three sequences were thus obtained and classified into five groups and eight subgroups, which indicates the possibility of enhancing RGC sequence diversity from BACs. A primer-walking strategy was employed to derive full-length gene structures from two BAC clones; both sequences 17o6RLK and 26m19RLK contained all the features of the rice Xa21 protein, including a signal peptide, the same number of leucine-rich-repeats, and transmembrane and kinase domains. These results demonstrate that PCR amplification with appropriately designed degenerate primers is an efficient approach for cloning receptor-like kinase class RGCs. Utilization of BAC clones can facilitate this approach in multiple ways by improving sequence diversity, providing full-length genes, and assisting in understanding gene structures and distribution.

Amino Acid Sequence↗

Signature sequences in diverse proteins provide evidence of a close evolutionary relationship between the Deinococcus-thermus group and cyanobacteria.

A number of proteins have been identified that contain prominent sequence signatures that are uniquely shared by the members of the Deinococcus-Thermus genera and the cyanobacterial species but which are not found in any of the other eubacterial or archaebacterial homologs. The proteins containing such sequence signatures include (1) the DnaJ/Hsp40 family of proteins, (2) DNA polymerase I, (3) the protein synthesis elongation factor EF-Tu, and (4) the elongation factor EF-Ts. A strong affinity of the Deinococcus-Thermus species to cyanobacteria is also seen in the phylogenetic trees based on Hsp70 and DnaJ sequences. These results provide strong evidence of a close and specific evolutionary relationship between species belonging to these two eubacterial divisions.

Amino Acid Sequence↗

Duplication and diversification of the apolipoprotein CI (APOCI) genomic segment in association with retroelements.

We have previously shown that several multicopy gene families within the major histocompatibility complex (MHC) arose from a process of segmental duplication. It has also been observed that retroelements play a role in generating diversity within these duplicated segments. The objective of this study was to compare the genomic organization of a gene duplication within another multicopy gene family outside the MHC. Using new continuous genomic sequence encompassing the APOE-CII gene cluster, we show that APOCI and its pseudogene, APOCI', are contained within large duplicated segments which include sequences from the hepatic control region (HCR). Flanking Alu sequences are observed at both ends of the duplicated unit, suggesting a possible role in the integration of these segments. As observed previously within the MHC, the major differences between the segments are the insertion of sequences (approximately 200-1000 bp in length), consisting predominantly of Alu sequences. Ancestral retroelements also contribute to the generation of sequence diversity between the segments, especially within the 3' poly(A) tract of Alu sequences. The exonic and regulatory sequences of the APOCI and HCR loci show limited sequence diversity, with exon 3 being an exception. Finally, the typing of pre- and postduplication Alus from both segments indicates an estimated time of duplication of approximately 37 million years ago (mya), some time prior to the separation of Old and New World monkeys.

Alu Elements↗

Application of protein structure alignments to iterated hidden Markov model protocols for structure prediction.

BACKGROUND: One of the most powerful methods for the prediction of protein structure from sequence information alone is the iterative construction of profile-type models. Because profiles are built from sequence alignments, the sequences included in the alignment and the method used to align them will be important to the sensitivity of the resulting profile. The inclusion of highly diverse sequences will presumably produce a more powerful profile, but distantly related sequences can be difficult to align accurately using only sequence information. Therefore, it would be expected that the use of protein structure alignments to improve the selection and alignment of diverse sequence homologs might yield improved profiles. However, the actual utility of such an approach has remained unclear. RESULTS: We explored several iterative protocols for the generation of profile hidden Markov models. These protocols were tailored to allow the inclusion of protein structure alignments in the process, and were used for large-scale creation and benchmarking of structure alignment-enhanced models. We found that models using structure alignments did not provide an overall improvement over sequence-only models for superfamily-level structure predictions. However, the results also revealed that the structure alignment-enhanced models were complimentary to the sequence-only models, particularly at the edge of the "twilight zone". When the two sets of models were combined, they provided improved results over sequence-only models alone. In addition, we found that the beneficial effects of the structure alignment-enhanced models could not be realized if the structure-based alignments were replaced with sequence-based alignments. Our experiments with different iterative protocols for sequence-only models also suggested that simple protocol modifications were unable to yield equivalent improvements to those provided by the structure alignment-enhanced models. Finally, we found that models using structure alignments provided fold-level structure assignments that were superior to those produced by sequence-only models. CONCLUSION: When attempting to predict the structure of remote homologs, we advocate a combined approach in which both traditional models and models incorporating structure alignments are used.

Algorithms↗

Characterization of an endogenous retrovirus class in elephants and their relatives.

BACKGROUND: Endogenous retrovirus-like elements (ERV-Ls, primed with tRNA leucine) are a diverse group of reiterated sequences related to foamy viruses and widely distributed among mammals. As shown in previous investigations, in many primates and rodents this class of elements has remained transpositionally active, as reflected by increased copy number and high sequence diversity within and among taxa. RESULTS: Here we examine whether proviral-like sequences may be suitable molecular probes for investigating the phylogeny of groups known to have high element diversity. As a test we characterized ERV-Ls occurring in a sample of extant members of superorder Uranotheria (Asian and African elephants, manatees, and hyraxes). The ERV-L complement in this group is even more diverse than previously suspected, and there is sequence evidence for active expansion, particularly in elephantids. Many of the elements characterized have protein coding potential suggestive of activity. CONCLUSIONS: In general, the evidence supports the hypothesis that the complement had a single origin within basal Uranotheria.

Africa↗

Hungarian mtDNA population databases from Budapest and the Baranya county Roma.

To facilitate forensic mtDNA testing in Hungary, we have generated control region databases for two Hungarian populations: 211 individuals were sampled from the urban Budapest population and 208 individuals were sampled from a Romani ("gypsy") population in Baranya county. Sequences were generated using a highly redundant approach to minimize potential database errors. The Budapest population had high sequence diversity with 180 lineages, 183 polymorphic positions, and a random match probability of 1%. In contrast, the Romani population exhibited low sequence diversity, with only 56 lineages, 109 segregating sites, and a random match probability of 8.8%. The mtDNA haplogroup compositions of the two populations were also distinct, with the large proportion of haplogroup M samples (35%) in the Roma the most obvious difference between the two populations. These factors highlight the importance of considering population structure when generating reference databases for forensic testing purposes. Comparisons between our Romani population sample and other published data indicate the need for heightened caution when sampling and using mtDNA databases of small endogamous populations. The Romani populations that we compared showed significant departures from genetic uniformity.

DNA, Mitochondrial↗

Molecular diversity at the self-incompatibility locus is a salient feature in natural populations of wild tomato (Lycopersicon peruvianum).

A cDNA encoding a stylar protein was cloned from flowers of self-incompatible wild tomato (Lycopersicon peruvianum). The corresponding gene was mapped to the S locus, which is responsible for self-incompatibility. The nucleotide sequence was determined for this allele, and compared to other S-related sequences in the Solanaceae. The S allele was used to probe DNA from 92 plants comprising 10 natural populations of Lycopersicon peruvianum. Hybridization was conducted under moderate and permissive stringencies in order to detect homologous sequences. Few alleles were detected, even under permissive conditions, underscoring the great sequence diversity at this locus. Those alleles that were detected are highly homologous. Sequences could not be detected in self-incompatible Nicotiana alata, self-compatible L. esculentum (cultivated tomato) or self-compatible L. hirsutum. However, hybridization to an individual of self-incompatible L. hirsutum revealed a closely related sequence that maps to the S locus in this reproductively isolated species. This supports the finding that S locus polymorphism predates speciation. The extraordinarily high degree of sequence diversity present in the gametophytic self-incompatibility system is discussed in the context of other highly divergent systems representing several kingdoms.

Alleles↗

Comparative genotyping of Campylobacter jejuni by amplified fragment length polymorphism, multilocus sequence typing, and short repeat sequencing: strain diversity, host range, and recombination.

Three molecular typing methods were used to study the relationships among 184 Campylobacter strains isolated from humans, cattle, and chickens. All strains were genotyped by amplified fragment length polymorphism (AFLP) analysis, multilocus sequence typing (MLST), and sequence analysis of a genomic region with short tandem repeats designated clustered regularly interspaced short palindromic repeats (CRISPRs). MLST and AFLP analysis yielded more than 100 different profiles and patterns, respectively. These multiple-locus typing methods resulted in similar genetic clustering, indicating that both are useful in disclosing genetic relationships between Campylobacter jejuni isolates. Group separation analysis of the AFLP analysis and MLST data revealed an unexpected association between cattle and human strains, suggesting a common source of infection. Analysis of the polymorphic CRISPR region carrying short repeats allowed about two-thirds of the typeable strains to be distinguished, similar to AFLP analysis and MLST. The three methods proved to be equally powerful in identifying strains from outbreaks of human campylobacteriosis. Analysis of the MLST data showed that intra- and interspecies recombination occurs frequently and that the role of recombination in sequence variation is 50 times greater than that of mutation. Examination of strains cultured from cecum swabs revealed that individual chickens harbored multiple Campylobacter strain types and that some genotypes were found in more than one chicken. We conclude that typing of Campylobacter strains is useful for identification of outbreaks but is probably not useful for source tracing and global epidemiology because of carriage of strains of multiple types and an extremely high diversity of strains in animals.

Alleles↗

Genomic sequencing in diverse and underserved pediatric populations: Parent perspectives on understanding, uncertainty, psychosocial impact, and personal utility of results.

PURPOSE: Limited evidence evaluates parents' perceptions of their child's clinical genome-scale sequencing (GS) results, particularly among individuals from medically underserved groups. Five Clinical Sequencing Evidence-Generating Research consortium studies performed GS in children with suspected genetic conditions with high proportions of individuals from underserved groups to address this evidence gap. METHODS: Parents completed surveys of perceived understanding, personal utility, and test-related distress after GS result disclosure. We assessed outcomes' associations with child- and parent-related factors: child age; type of GS finding; and parent health literacy, numeracy, and education. RESULTS: A total of 1763 parents completed surveys; 83% met "underserved" criteria based on race, ethnicity, and risk factors for barriers to access. We observed high perceived understanding and personal utility and low test-related distress. Outcomes were associated with the type of GS finding; parents of children with a pathogenic or likely pathogenic finding endorsed higher personal utility and more test-related distress than those whose children had a variant of uncertain significance or normal finding. Personal utility was higher in parents who met the criteria for "underserved." CONCLUSION: Our findings shed light on correlates of parents' cognitive and emotional responses to their child's GS findings and emphasize the need for tailored support in disclosure discussions.

Humans↗

Molecular population genetics of sequence length diversity in the Adh region of Drosophila pseudoobscura.

Positive and negative selection on indel variation may explain the correlation between intron length and recombination levels in natural populations of Drosophila. A nucleotide sequence analysis of the 3.5 kilobase sequence of the alcohol dehydrogenase (Adh) region from 139 Drosophila pseudoobscura strains and one D. miranda strain was used to determine whether positive or negative selection acts on indel variation in a gene that experiences high levels of recombination. A total of 30 deletion and 36 insertion polymorphisms were segregating within D. pseudoobscura populations and no indels were fixed between D. pseudoobscura and its two sibling species D. miranda and D. persimilis. The ratio of Tajima's D to its theoretical minimum value (D(min)) was proposed as a metric to assess the heterogeneity in D among D. pseudoobscura loci when the number of segregating sites differs among loci. The magnitude of the D/D(min) ratio was found to increase as the rate of population expansion increases, allowing one to assess which loci have an excess of rare variants due to population expansion versus purifying selection. D. pseudoobscura populations appear to have had modest increases in size accounting for some of the observed excess of rare variants. The D/D(min) ratio rejected a neutral model for deletion polymorphisms. Linkage disequilibrium among pairs of indels was greater than between pairs of segregating nucleotides. These results suggest that purifying selection removes deletion variation from intron sequences, but not insertion polymorphisms. Genome rearrangement and size-dependent intron evolution are proposed as mechanisms that limit runaway intron expansion.

Alcohol Dehydrogenase↗

Thoroughly sampling sequence space: large-scale protein design of structural ensembles.

Modeling the inherent flexibility of the protein backbone as part of computational protein design is necessary to capture the behavior of real proteins and is a prerequisite for the accurate exploration of protein sequence space. We present the results of a broad exploration of sequence space, with backbone flexibility, through a novel approach: large-scale protein design to structural ensembles. A distributed computing architecture has allowed us to generate hundreds of thousands of diverse sequences for a set of 253 naturally occurring proteins, allowing exciting insights into the nature of protein sequence space. Designing to a structural ensemble produces a much greater diversity of sequences than previous studies have reported, and homology searches using profiles derived from the designed sequences against the Protein Data Bank show that the relevance and quality of the sequences is not diminished. The designed sequences have greater overall diversity than corresponding natural sequence alignments, and no direct correlations are seen between the diversity of natural sequence alignments and the diversity of the corresponding designed sequences. For structures in the same fold, the sequence entropies of the designed sequences cluster together tightly. This tight clustering of sequence entropies within a fold and the separation of sequence entropy distributions for different folds suggest that the diversity of designed sequences is primarily determined by a structure's overall fold, and that the designability principle postulated from studies of simple models holds in real proteins. This has important implications for experimental protein design and engineering, as well as providing insight into protein evolution.

Amino Acid Sequence↗

Genetic mapping and molecular characterization of the self-incompatibility (S) locus in Petunia inflata.

Gametophytic self-incompatibility (SI) possessed by the Solanaceae is controlled by a highly polymorphic locus called the S locus. The S locus contains two linked genes, S-RNase, which determines female specificity, and the as yet unidentified pollen S gene, which determines male specificity in SI interactions. To identify the pollen S gene of Petunia inflata, we had previously used mRNA differential display and subtractive hybridization to identify 13 pollen-expressed genes that showed S -haplotype-specific RFLP. Here, we carried out recombination analysis of 1205 F2 plants to determine the genetic distance between each of these S -linked genes and S-RNase. Recombination was observed between four of the genes (3.16, G211, G212, and G221) and S-RNase, whereas no recombination was observed for the other nine genes (3.2, 3.15, A113, A134, A181, A301, G261, X9, and X11). A genetic map of the S locus was constructed, with 3.16 and G221 delimiting the outer limits. None of the observed crossovers disrupted SI, suggesting that all the genes required for SI are contained in the chromosomal region defined by 3.16 and G221. These results and our preliminary chromosome walking results suggest that the S locus is a huge multi-gene complex. Allelic sequence diversity of G221 and 3.16, as well as of 3.2, 3.15, A113, A134 and G261, was determined by comparing two or three alleles of their cDNA and/or genomic sequences. In contrast to S-RNase, all these genes showed very low degrees of allelic sequence diversity in the coding regions, introns, and flanking regions.

Alleles↗

Initiator and upstream elements in the alpha2-tubulin promoter of Giardia lamblia.

Giardia lamblia, one of the earliest diverging eukaryotes and a major cause of diarrhea world-wide, has unusually short intergenic regions, raising questions concerning its regulation of gene expression. We have approached this issue through examination of the alpha2-tubulin promoter and in particular investigated the function of an AT-rich element surrounding the transcription start site. Its placement and the ability of this sequence to direct transcription initiation in the absence of any other promoter elements is similar to the initiator element in higher eukaryotes. However, the sequence diversity of extremely short (8-10 bp) initiator elements is surprising, as is their ability to independently direct substantial levels of transcription. We also identified a large AT-rich element located between -64 and -29 bp upstream of the transcriptional start site and show using both deletions and site-specific mutations of this region that sequences between -60 and the start of transcription are important for promoter strength; interestingly this AT-rich sequence is not highly conserved among different Giardia promoters. These data suggest that while the overall structure of the core promoter has been conserved throughout eukaryotic evolution, significant variation and flexibility is allowed in element consensus sequences and roles in transcription. In particular, the short and diverse sequences that function in transcription initiation in Giardia suggest the potential for relaxed transcriptional regulation.

Animals↗

In silico detection of control signals: mRNA 3'-end-processing sequences in diverse species.

We have investigated mRNA 3'-end-processing signals in each of six eukaryotic species (yeast, rice, arabidopsis, fruitfly, mouse, and human) through the analysis of more than 20,000 3'-expressed sequence tags. The use and conservation of the canonical AAUAAA element vary widely among the six species and are especially weak in plants and yeast. Even in the animal species, the AAUAAA signal does not appear to be as universal as indicated by previous studies. The abundance of single-base variants of AAUAAA correlates with their measured processing efficiencies. As found previously, the plant polyadenylation signals are more similar to those of yeast than to those of animals, with both common content and arrangement of the signal elements. In all species examined, the complete polyadenylation signal appears to consist of an aggregate of multiple elements. In light of these and previous results, we present a broadened concept of 3'-end-processing signals in which no single exact sequence element is universally required for processing. Rather, the total efficiency is a function of all elements and, importantly, an inefficient word in one element can be compensated for by strong words in other elements. These complex patterns indicate that effective tools to identify 3'-end-processing signals will require more than consensus sequence identification.

Animals↗

Phylogenetic analysis of sequences from diverse bacteria with homology to the Escherichia coli rho gene.

Genes from Pseudomonas fluorescens, Chromatium vinosum, Micrococcus luteus, Deinococcus radiodurans, and Thermotoga maritima with homology to the Escherichia coli rho gene were cloned and sequenced, and their sequences were compared with other available sequences. The species for all of the compared sequences are members of five bacterial phyla, including Thermotogales, the most deeply diverged phylum. This suggests that a rho-like gene is ubiquitous in the Bacteria and was present in their common ancestor. The comparative analysis revealed that the Rho homologs are highly conserved, exhibiting a minimum identity of 50% of their amino acid residues in pairwise comparisons. The ATP-binding domain had a particularly high degree of conservation, consisting of some blocks with sequences of residues that are very similar to segments of the alpha and beta subunits of F1-ATPase and of other blocks with sequences that are unique to Rho. The RNA-binding domain is more diverged than the ATP-binding domain. However, one of its most highly conserved segments includes a RNP1-like sequence, which is known to be involved in RNA binding. Overall, the degree of similarity is lowest in the first 50 residues (the first half of the RNA-binding domain), in the putative connector region between the RNA-binding and the ATP-binding domains, and in the last 50 residues of the polypeptide. Since functionally defective mutants for E. coli Rho exist in all three of these segments, they represent important parts of Rho that have undergone adaptive evolution.

Adenosine Triphosphate↗