Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequence diversity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

The imprint of somatic hypermutation on the repertoire of human germline V genes.

In the human immune system, antibodies with high affinities for antigen are created in two stages. A diverse primary repertoire of antibody structures is produced by the combinatorial rearrangement of germline V gene segments and antibodies are selected from this repertoire by binding to the antigen. Their affinities are then improved by somatic hypermutation and further rounds of selection. We have dissected the sequence diversity created at each stage in response to a wide range of antigens. In the primary repertoire, diversity is focused at the centre of the binding site. With somatic hypermutation, diversity spreads to regions at the periphery of the binding site that are highly conserved in the primary repertoire. We propose that evolution has favoured this complementarity as an efficient strategy for searching sequence space and that the germline V gene families evolved to exploit the diversity created by somatic hypermutation.

Antibody Diversity↗

B cells selected for apoptosis in the sheep ileal Peyer's patch have enhanced mutational diversity in the Ig V lambda light chain.

To investigate the molecular events associated with B cell apoptosis, we analyzed follicular B cells from the large Peyer's patch (PP) in the sheep ileum. Over 95% of B cells generated in the ileal PP are rapidly destroyed by apoptosis. Ig V lambda sequences from apoptotic B cells were compared with sequence from B cells about to emigrate from the PP. The sequences originated from two germline genes, V lambda 5.1 and V lambda 5.3. Only V lambda 5.1 was rearranged in apoptotic cells, whereas both V lambda 5.1 and V lambda 5.3 were rearranged in B cells about to emigrate. Apoptotic B cells had evidence of increased Ig sequence diversity based on: 1) significantly greater replacement to silent mutation ratios in the complementarity determining regions, 2) the more random distribution of mutations, and 3) the lack of mutational specificity compared with the mutational bias favoring transitions and purines in B cells about to emigrate. Based on this analysis, we propose that the continual proliferation of B cells in the PP follicle might increase their affinity to local Ags. Those Ags that are sequestered in this environment might be expected to stimulate the production of B cells with such high-affinity receptors that ligation would trigger apoptosis. This could account for the deletion of B cells with specificity for self-antigens, selecting ligands as well as gut-derived food and microbial Ags. This process could contribute to the elimination of self-reactive B cells, the expansion of the antibody repertoire, and the generation of oral tolerance.

Animals↗

Species Diversity of Uncultured and Cultured Populations of Soil and Marine Ammonia Oxidizing Bacteria.

Although molecular techniques are considered to provide a more comprehensive view of species diversity of natural microbial populations, few studies have compared diversity assessed by molecular and cultivation-based approaches using the same samples. To achieve this, the diversity of natural populations of ammonia oxidising bacteria in arable soil and marine sediments was determined by analysis of 16S rDNA sequences from enrichment cultures, prepared using standard methods for this group, and from 16S rDNA cloned from DNA extracted directly from the same environmental samples. Soil and marine samples yielded 31 and 18 enrichment cultures, respectively, which were compared with 50 and 40 environmental clones. There was no evidence for selection for particular ammonia oxidizer clusters by different procedures employed for enrichment from soil samples, although no culture was obtained in medium at acid pH. In soil enrichment cultures, Nitrosospira cluster 3 sequences were most abundant, whereas clones were distributed more evenly between Nitrosospira clusters 2, 3, and 4. In marine samples, the majority of enrichment cultures contained Nitrosomonas, whereas Nitrosospira sequences were most abundant among environmental clones. Soil enrichments contained a higher proportion of identical sequences than clones, suggesting laboratory selection for particular strains, but the converse was found in marine samples. In addition, 16% of soil enrichment culture sequences were identical to those in environmental clones, but only 1 of 40 marine enrichments was found among clones, indicating poorer culturability of marine strains represented in the clone library, under the conditions employed. The study demonstrates significant differences in species composition assessed by molecular and culture-based approaches but indicates also that, employing only a limited range of cultivation conditions, 7% of the observed sequence diversity in clones of ammonia oxidizers from these environments could be obtained in laboratory enrichment culture. Further studies and experimental approaches are required to determine which approach provides better representation of the natural community.

Journal Article↗

Multi-gene family of major surface glycoproteins of Pneumocystis carinii: full-size cDNA cloning and expression.

The major surface glycoprotein (MSG) of Pneumocystis carinii plays a crucial role in the fatal pneumonia caused by this organism in AIDS patients. A cDNA encoding a full-length MSG polypeptide was isolated from a lambda library of rat-derived P. carinii cDNAs. The deduced MSG, referred to as the MSG5 subtype, is a 120,765-Da protein composed of 1,076 amino acids and contains an anchoring hydrophobic sequence at the C-terminus of the protein. Sequence analyses of cloned MSG-cDNAs revealed an MSG-gene family with approximately 70% protein sequence identity between subtypes. P. carinii karyotype hybridization analyses indicated that the MSG gene family members are scattered throughout most of the P. carinii chromosomes. These recombinant MSG proteins reacted with the antiserum from P. carinii-infected rats, as expected, and antiserum generated against P. carinii-infected mice, indicating the existence of common determinants in MSG polypeptides. The family of MSG proteins is rich in cysteine residues and these cysteines are highly conserved in all MSG subtypes regardless of species specificity, suggesting the structural and/or functional importance of these cysteines. The pathobiological significance of the MSG gene family and its sequence diversity in P. carinii is discussed.

Amino Acid Sequence↗

Selection of novel forms of a functional domain within the Tetrahymena ribozyme.

P5abc is an RNA structure within the self-splicing Tetrahymena group I intron that provides an activation function to the remainder of the ribozyme, either when present in cis or when added in trans. This 69-nucleotide activator domain was replaced with randomized sequence of 20 or 40 nt in length, and individuals among these pools with sequences that could functionally replace P5abc were selected. The basis of selection was a reaction in which two separate halves of the ribozyme became joined; selection was completed by reverse transcription and the polymerase chain reaction, using primers with sequence from either side of the ligation junction. Selectant sequences fell into three families that appear unrelated to P5abc; for example they lack the A-rich bulge thought to be a important feature of P5abc. Thus, rather than defining some consensus sequence for activator domains, this result reveals a certain tolerance in the ribozyme in its ability to derive activation function from diverse sequence types. In the context of splicing precursor RNA, the new sequences supported self-splicing, but failed to activate a related reaction, hydrolysis of the 3' splice site, implying that this region of the intron can differentially control two related reactions.

Animals↗

Increased nucleotide diversity with transient Y linkage in Drosophila americana.

Recombination shapes nucleotide variation within genomes. Patterns are thought to arise from the local recombination landscape, influencing the degree to which neutral variation experiences hitchhiking with selected variation. This study examines DNA polymorphism along Chromosome 4 (element B) of Drosophila americana to identify effects of hitchhiking arising as a consequence of Y-linked transmission. A centromeric fusion between the X and 4(th) chromosomes segregates in natural populations of D. americana. Frequency of the X-4 fusion exhibits a strong positive correlation with latitude, which has explicit consequences for unfused 4(th) chromosomes. Unfused Chromosome 4 exists as a non-recombining Y chromosome or as an autosome proportional to the frequency of the X-4 fusion. Furthermore, Y linkage along the unfused 4 is disrupted as a function of the rate of recombination with the centromere. Inter-population and intra-chromosomal patterns of nucleotide diversity were assayed using six regions distributed along unfused 4(th) chromosomes derived from populations with different frequencies of the X-4 fusion. No difference in overall level of nucleotide diversity was detected among populations, yet variation along the chromosome exhibits a distinct pattern in relation to the X-4 fusion. Sequence diversity is inflated at loci experiencing the strongest Y linkage. These findings are inconsistent with the expected reduction in nucleotide diversity resulting from hitchhiking due to background selection or selective sweeps. In contrast, excessive polymorphism is accruing in association with transient Y linkage, and furthermore, hitchhiking with sexually antagonistic alleles is potentially responsible.

Animals↗

Differential regulation of A gamma and G gamma fetal hemoglobin mRNA levels by hydroxyurea and butyrate.

In clinical studies, both hydroxyurea and butyrate increase fetal hemoglobin expression and ameliorate the symptoms of sickle cell anemia. However, comparative studies of the effects of hydroxyurea and butyrate on the expression of the individual fetal hemoglobin genes, A gamma and G gamma, have not been performed. The present study reports the effects of hydroxyurea and butyrate on steady-state A gamma and G gamma mRNA levels in K562 cells. Because the high degree of homology between the A gamma and G gamma cDNA sequences precludes the use of large cDNA probes for detection of individual fetal hemoglobin gene products, we investigated the specificity of two 20-base oligonucleotide probes synthesized from the region of greatest sequence diversity between these genes. Hybridization experiments demonstrated that the A gamma oligonucleotide probe was specific for A gamma DNA and RNA sequences and the G gamma oligonucleotide probe was specific for G gamma DNA and RNA sequences. These oligonucleotide probes detected both A gamma and G gamma mRNAs in K562 cells. In K562 cells treated with 2 mM sodium butyrate for 168 hours, the G gamma mRNA level increased 3.6-fold, whereas the A gamma mRNA level was not significantly different from untreated cells. Similar results were obtained when K562 cells were treated with 80 microM hydroxyurea. The G gamma mRNA level increased 2.3-fold at 168 hours, whereas the A gamma mRNA level did not change. The above results demonstrate that both butyrate and hydroxyurea selectively increase G gamma expression. Selective regulation of individual fetal hemoglobin genes is also seen in human development, where approximately 70% of the total fetal hemoglobin in the fetus is G gamma. Therefore, understanding the mechanisms by which butyrate and hydroxyurea differentially regulate fetal hemoglobin gene expression may provide insights into the developmental regulation of hemoglobin expression as well as the mechanisms of action of pharmacological agents currently being used to treat sickle cell disease.

Butyrates↗

Fast, accurate construction of multiple sequence alignments from protein language embeddings.

Multiple sequence alignment (MSA) is a foundational task in computational biology, underpinning protein structure prediction, evolutionary analysis, and domain annotation. Traditional MSA algorithms rely on pairwise amino acid substitution matrices derived from conserved protein families. While effective for aligning closely related sequences, these scoring schemes struggle in the low-identity "twilight zone." Here, we present a new approach for constructing MSAs leveraging amino acid embeddings generated by protein language models (PLMs), which capture rich evolutionary and contextual information from massive and diverse sequence datasets. We introduce a windowed reciprocal-weighted embedding similarity metric that is surprisingly effective in identifying corresponding amino acids across sequences. Building on this metric, we develop ARIES (Alignment via RecIprocal Embedding Similarity), an algorithm that constructs a PLM-generated template embedding and aligns each sequence to this template via dynamic time warping in order to build a global MSA. Across diverse benchmark datasets, ARIES achieves higher accuracies than existing state-of-the-art approaches, especially in low-identity regimes where traditional methods degrade, while scaling almost linearly with the number of sequences to be aligned. Together, these results provide the first large-scale demonstration of the power of PLMs for accurate and scalable MSA construction across protein families of varying sizes and levels of similarity, highlighting the potential of PLMs to transform comparative sequence analysis.

Deep Learning↗

The phylogenetics of Desmognathine salamander populations across the southern Appalachians.

Salamanders in the genus Desmognathus (Caudata: Plethodontidae) are distributed along an aquatic to terrestrial habitat gradient in the southern Appalachian Mountains. The spatial distribution of species is believed to have formed as aquatic ancestors displaced lineages by competition and predatory interactions into less optimal terrestrial habitats. Aquatic and terrestrial species may also display different patterns of genetic diversity due to the differing likelihood of gene flow via aquatic corridors. To determine whether phylogenetic patterns were consistent with these hypotheses, we sequenced portions of the cytochrome oxidase I and 12S rRNA genes of the mitochondrial genome from 96 individuals belonging to 10 species in the genus Desmognathus. In addition, we combined our dataset with an earlier published dataset for the 12S rRNA genes. The order of species divergence is consistent with aquatic ancestors having displaced taxa into more terrestrial habitats, but the major lineages within the genus Desmognathus arose suddenly, and therefore, the specific sequence of events is not well resolved. The phylogenetic analyses among species suggest that direct-development and a terrestrial lifestyle are ancestral in the genus Desmognathus, but the degree of adult terrestriallity is labile, with some species having re-invaded terrestrial habitats. We present evidence of a clade of Desmognathus quadramaculatus from North Carolina that is distinct from the D. quadramaculatus/Desmognathus marmoratus clade. Within species, estimates of Tajima's D and Fu and Li's statistics suggest the species experienced population expansions at different times in the past. Current levels of sequence diversity in northern populations, therefore, reflect different arrival times, and hence, differences in the opportunity for among population divergence. The recent arrival of most species over large portions of their geographic ranges suggests that most extant communities have been assembled, a posteriori, by the recent assortment of species along the aquatic to terrestrial gradient according to their ecologies.

Animals↗

Evolution of MHC class II E beta diversity within the genus Peromyscus.

Progress in understanding the evolution of variation at the MHC has been slowed by an inability to assess the relative roles of mutation vs. intragenic recombination in contributing to observed polymorphism. Recent theoretical advances now permit a quantitative treatment of the problem, with the result that the amount of recombination is at least an order of magnitude greater than that of mutation in the history of class II genes. We suggest that this insight allows progress in evaluating the importance of other factors affecting the evolution of the MHC. We investigated the evolution of MHC class II E beta sequence diversity in the genus Peromyscus. We find evidence for extensive recombination in the history of these sequences. Nevertheless, it appears that intragenic recombination alone is insufficient to account for evolution of MHC diversity in Peromyscus. Significant differences in silent variation among subgenera arose over a relatively short period of time, with little subsequent change. We argue that these observations are consistent with the effects of historical population bottleneck(s). Population restrictions may explain general features of MHC evolution, including the large amount of recombination in the history of MHC genes, because intragenic recombination may efficiently regenerate allelic polymorphism following a population constriction.

Amino Acid Sequence↗

Genetic alteration of the hepatitis C virus hypervariable region obtained from an asymptomatic carrier.

Hepatitis C virus (HCV) genome shows extensive sequence diversity at 2 hypervariable regions (HVR1 and HVR2) of the putative envelope glycoprotein (gp70). We recently reported that the amino-acid sequence of HVR1, but not of HVR2, underwent a striking mutation or mutated sequentially over a period of several months in patients with chronic hepatitis (CH). Here, we examined whether these genetic alterations in HVR1 occurred in an asymptomatic HCV carrier. The level of HCV RNA in serum was almost the same throughout the 4 time points sampled over 16 months. However, we found that the amino-acid sequence of the HCV HVR1 from this asymptomatic carrier altered with time, as seen in patients with CH. Alterations of amino acids in the HVR1 were correlated with persistent HCV infection rather than with clinical symptoms. Sequence heterogeneity of HVR1 was not correlated with alanine aminotransferase (ALT) values or liver histological findings. The necessity of clinical follow-up of HCV asymptomatic carriers is discussed.

Adult↗

Asymmetric infectivity of pseudorecombinants of cabbage leaf curl virus and squash leaf curl virus: implications for bipartite geminivirus evolution and movement.

The bipartite geminiviruses squash leaf curl virus (SqLCV) and cabbage leaf curl virus (CLCV) have distinct host ranges. SqLCV infects a broad range of plants within the Cucurbitaceae, including pumpkin and squash, and CLCV has a broad host range within Brassicaceae that includes cabbage and Arabidopsis thaliana. Despite this, the genomic A components of these viruses share a high degree of sequence identity, particularly in the gene encoding the replication protein AL1, and their common regions are 77% identical. However, there is unexpected sequence diversity in the common regions of the two CLCV genomic A and B components, these being only 80% identical. Based on these sequence similarities, we investigated the host range properties of pseudorecombinants of SqLCV and CLCV. We found that in a pseudorecombinant virus consisting of the A component of CLCV and the B component of SqLCV, both components replicated in tobacco protoplasts, and this pseudorecombinant was infectious and caused systemic disease in Nicotiana benthamiana, a common host to all bipartite geminiviruses. However, this pseudorecombinant did not move systemically in pumpkin or Arabidopsis, despite the demonstrated replication compatibility of the genome components. As a result of the greater sequence differences between the common regions, the pseudorecombinant of SqLCV A and CLCV B components neither replicated the CLCV B component nor systemically infected any of the hosts tested. These findings demonstrate that for different geminiviruses with distinct host ranges, the replication origins and AL1 proteins can be sufficiently similar to permit infectious pseudorecombinants, but replication alone is not sufficient to cause systemic disease, and host range may ultimately be limited at the level of movement. The results of this study further suggest that CLCV is an evolving virus that can provide insights into how new bipartite geminiviruses arise from mixed infections.

Base Sequence↗

Characterisation of T cell antigen receptor alpha chain isotypes in the common carp.

T cell receptor alpha (TCRalpha) chain has been characterised in several teleost species to date. Here, a reverse transcription-polymerase chain reaction (RT-PCR) strategy was used to isolate cDNA clones encoding TCRalpha chain from an individual of the common carp (Cyprinus carpio.L.). The Valpha sequences identified were most similar to Valpha of other teleosts, and could be classified into as many as 14 Valpha families. For the Jalpha sequences, diversity comparable to that seen in other teleosts could be identified, and the J-region motif was well conserved. The Calpha sequences demonstrated the highest similarity to zebrafish Calpha and possessed a well-conserved transmembrane (TM) region. Two Calpha isotypes with a complete C region were obtained, designated Calpha1 and Calpha2, with approximately 70% similarity at the amino acid level ( approximately 85% identity at the nucleotide level), and, in addition, Calpha2 contained two unique sequences, designated Calpha2a and Calpha2b, with 93% similarity (96% identity). Therefore, the results obtained using an individual clearly showed that carp possesses at least two Calpha loci, possibly as a result of tetraploidisation.

Amino Acid Sequence↗

Diversification of Drosophila chloride channel gene by multiple posttranscriptional mRNA modifications.

We have identified and analyzed a Drosophila melanogaster gene that encodes a chloride channel subunit (DrosGluCl-alpha) previously shown to function as a glutamate-gated chloride channel in an in vitro expression system. Sequence analysis of several cDNAs corresponding to the gene revealed sequence diversity in their open reading frames at seven specific sites. Site-specific A-to-G variations between cDNA and genomic sequences, consistent with RNA editing, were detected at five nucleotide positions. In addition, sequence variations among cDNA clones consistent with alternative splicing of mRNA were found at two different sites. In the 5' region, two small adjacent exons, containing similar but distinct modular sequences, are alternatively incorporated into the mature mRNA. In the 3' region, alternative splicing generates a variant encoding a protein with four additional amino acids just upstream of the fourth transmembrane domain. Combinations of RNA editing and alternative splicing can lead to extensive diversification of transcripts. These results give the first example of RNA editing in neurotransmitter-gated chloride channel genes or of alternative splicing in a glutamate-gated chloride channel gene of Drosophila.

Alternative Splicing↗

Isolation and analysis of the breakpoint sequences of chromosome inversion In(3L)Payne in Drosophila melanogaster.

Chromosomal rearrangements constitute a significant feature of genome evolution, and inversion polymorphisms in Drosophila have been studied intensely for decades. Population geneticists have long recognized that the sequence features associated with inversion breakpoints would reveal much about the mutational origin, uniqueness, and genealogical history of individual inversion polymorphisms, but the cloning of breakpoint sequences is not trivial. With the aid of a method for rapid recovery of DNA clones spanning rearrangement breakpoints, we recover and examine the DNA sequences spanning the breakpoints of the cosmopolitan inversion In(3L)Payne in Drosophila melanogaster. By examining the sequence diversity associated with six standard and seven inverted chromosomes from natural populations, we find that the inversion is monophyletic in origin, the sequences are genetically isolated from recombination at the breakpoints, and there is no association with features such as transposable elements. The inverted sequences show 17-fold less nucleotide polymorphism, but there are eight fixed differences in the region spanning both breakpoints. This suggests that this inversion is not recently derived. Finally, Northern analysis and transcript mapping find that the distal breakpoint has disrupted three transcripts that are normally expressed in the standard arrangement. Incidentally, the method introduced here can be used to isolate breakpoint sequences of arrangements associated with many human diseases.

Animals↗

Intrapatient variability of HIV type 1 group O ANT70 during a 10-year follow-up.

HIV-1 ANT70 is the first HIV-1 group O virus isolate obtained from a 25-year-old Cameroonian woman, who seroconverted in March 1987. This individual has remained asymptomatic and clinically healthy (clinical stage WHO 1, CDC II) even though she did not receive any antiretroviral therapy for HIV-1 before 97 months post-seroconversion. CD4+ T cell counts declined steadily to 200/microl at 70 months postseroconversion. The HIV-1 ANT70 nucleotide and amino acid sequence diversity of the V3C3-encoding env fragment within this individual was followed over a 10-year period. RT-PCR, cloning, sequencing, and genetic analyses were performed on eight plasma follow-up samples. Extensive increasing intra- and intersample variation was observed. This is the first long-term (>10 years) follow-up of the genetic variability of an HIV-1 group O-infected individual. As the course of the disease in the HIV-1 ANT70-infected woman was similar in many aspects to that of group M-infected individuals, it remains to be elucidated whether the changes observed in the V3 loop are critical for disease progression.

Adult↗

Improved template representation in cpn60 polymerase chain reaction (PCR) product libraries generated from complex templates by application of a specific mixture of PCR primers.

Some classes of high G+C content organisms such as the Actinobacteria, which are known through culture-based studies to be present in large numbers in particular microbial communities, are under-represented or even absent from 16S rRNA or cpn60 polymerase chain reaction (PCR) product libraries derived from these templates. Using reference cpn60 sequence data from organisms with high G+C content genomes, a pair of PCR primers were designed which, when used in combination with the previously developed degenerate, universal cpn60 primers, improve the representation of templates with high G+C content. The primers were validated using a combination of traditional and quantitative real-time PCR on both manufactured template mixtures and biological samples. The development and optimization of this specific primer mixture represents an improvement of established methods and a significant advance in the ability to generate cpn60 PCR product libraries that more closely represent the sequence diversity in complex templates.

Bacteria↗

The salmonid class I MHC: limited diversity in a primitive teleost.

Three MHC class I genes have been characterized in salmonids: A, B, and UA. Levels of polymorphism vary among the genes, but they all share one common feature: a lack of sequence diversity. Although individual species can carry over 30 alleles at a given locus (A), intraspecific diversity is generally less than 5% in Pacific salmon (genus Oncorhynchus), and less than 10% in Atlantic salmon (genus Salmo). These levels of diversity suggest that few ancient allelic lineages have persisted within species, and that most of the allelic radiation has occurred during or since speciation. Also apparent is the greater retention of allelic lineages in Atlantic salmon than Pacific salmon, which reflects historic differences of the two genera. Comparison of the salmonid class I sequences with those of other teleosts reveals two well supported groups: one containing the Cypriniformes and the salmonid UA, and the other containing the neoteleosts and the salmonid A and B. There is no homology between known Cypriniformes and neoteleostean sequences. If this relationship is borne out, it offers strong support for the hypothesis that the higher teleosts diverged more recently from the Salmoniformes than the Cypriniformes. The salmonid MHC may provide a snapshot of the neoteleostean MHC prior to the extensive class I duplication that has taken place in at least some of the more advanced species.

Animals↗