Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequence diversity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Sub-grouping of Plasmodium falciparum 3D7 var genes based on sequence analysis of coding and non-coding regions.

BACKGROUND: The variant surface antigen family Plasmodium falciparum erythrocyte membrane protein-1 (PfEMP1) is an important target for protective immunity and is implicated in the pathology of malaria through its ability to adhere to host endothelial receptors. The sequence diversity and organization of the 3D7 PfEMP1 repertoire was investigated on the basis of the complete genome sequence. METHODS: Using two tree-building methods we analysed the coding and non-coding sequences of 3D7 var and rif genes as well as var genes of other parasite strains. RESULTS: var genes can be sub-grouped into three major groups (group A, B and C) and two intermediate groups B/A and B/C representing transitions between the three major groups. The best defined var group, group A, comprises telomeric genes transcribed towards the telomere encoding PfEMP1s with complex domain structures different from the 4-domain type dominant of groups B and C. Two sequences belonging to the var1 and var2 subfamilies formed independent groups. A rif subgroup transcribed towards the centromere was found neighbouring var genes of group A such that the rif and var 5' regions merged. This organization appeared to be unique for the group A var genes CONCLUSION: The grouping of var genes implies that var gene recombination preferentially occurs within var gene groups and it is speculated that the groups reflect a functional diversification evolved to cope with the varying conditions of transmission and host immune response met by the parasite.

Animals↗

Tracking Cryptosporidium parvum by sequence analysis of small double-stranded RNA.

We sequenced a 173-nucleotide fragment of the small double-stranded viruslike RNA of Cryptosporidium parvum isolates from 23 calves and 38 humans. Sequence diversity was detected at 17 sites. Isolates from the same outbreak had identical double-stranded RNA sequences, suggesting that this technique may be useful for tracking Cryptosporidium infection sources.

Animals↗

Pan-genomics and multi-omics for deciphering genetic variation and accelerating genetic improvement in ruminant livestock.

Livestock reference genomes have transformed the discovery of variants associated with production, reproduction, health, and environmental adaptation. Nevertheless, a single linear reference represents only one mosaic haplotype and incompletely captures sequence diversity within a species, particularly structural variants, copy-number changes, repeat-rich regions, and breed-specific sequences. Pangenomes address this limitation by integrating multiple high-quality assemblies or population-scale variants into a unified sequence or graph representation. Concurrently, multi-omics approaches connect genomic variation with transcriptomic, epigenomic, manuscriptproteomic, metabolomic, and microbiome responses, thereby improving biological interpretation of genotype-phenotype relationships. This review synthesizes recent progress in livestock pangenomics and multi-omics, with emphasis on cattle, goats, sheep, water buffalo, and chickens. It describes advances in long-read and haplotype-resolved sequencing, graph construction, structural-variant discovery and genotyping, functional annotation, and integrative analysis. Recent pangenome studies have uncovered substantial non-reference sequence, reduced reference bias, identified breed- and population-specific structural variants, and resolved candidate variants underlying pigmentation, body size, tail morphology, cashmere production, altitude adaptation, and other economically relevant traits. However, translation into routine breeding remains constrained by uneven population representation, inconsistent structural-variant definitions, limited functional annotation, computational demands, and insufficient validation across environments. Future progress will depend on diverse near-complete assemblies, graph-aware imputation and genomic prediction, long-read transcriptomics, single-cell and spatial omics, rigorous causal validation, and open, interoperable resources. Together, these developments can support more accurate, resilient, and biologically informed livestock improvement. Importantly, current dairy-cattle evidence indicates that pangenome-derived structural variants can substantially improve variant discovery and functional interpretation while yielding only marginal average gains in routine genomic prediction, favoring targeted augmentation rather than wholesale replacement of established SNP-based evaluations.

Animals↗

DNA diversity in Hawaiian endemic plant Schiedea globosa.

This is the first report of a study devoted to the population genetics of speciation in the endemic Hawaiian plant genus Schiedea (Caryophyllaceae). Here, we report the estimates of DNA sequence diversity and divergence in a newly isolated nuclear gene from Maui and Oahu Schiedea globosa populations. Overall, the species-wide average heterozygosity per silent site is pi = 0.3%. The silent DNA diversity on the older island of Oahu (pi = 0.24%) is almost twice as high as on the younger Maui (pi = 0.14%). Consistent with this, the haplotype phylogeny suggests a more recent origin of the Maui populations. There is no significant isolation between the two Maui populations (F(st)=0.027), while isolation between the two islands is high (F(st)=0.57, P<0.0001). Pairwise mismatch distributions suggest population growth approximately 660 and 310 thousand generations ago for the Oahu and the Maui populations, respectively, which may be the minimal age for these populations. This is consistent with a fairly neutral frequency spectrum (Tajima's D is 0.34 and -0.94 for the Oahu and the Maui populations, respectively), suggesting that both populations are sufficiently old to have recovered from any initial founder effects. Relatively high nuclear DNA diversity in the S. globosa populations illustrates the usefulness of a DNA sequence-based approach to the population genetics of island plant populations.

Caryophyllaceae↗

Increased promoter diversity reveals a complex phylogeny of human immunodeficiency virus type 1 subtype C in India.

OBJECTIVE: To evaluate human immunodeficiency virus type 1 (HIV-1) long terminal repeat (LTR) sequence diversity among distinct populations within India and to determine the prevalent subtype. STUDY DESIGN/METHODS: Analysis of the 3'LTR was conducted from 28 HIV-1-positive samples: 1992-1993 (Pune, New Delhi) and 1995-1996 (Pune, Mumbai and Vellore). Genomic DNA was extracted from cocultivated peripheral blood mononuclear cells (PBMCs) and used for polymerase chain reaction (PCR) amplification and sequencing using dye terminator chemistry. Sequences were edited, aligned, and analyzed phylogenetically utilizing gap-stripped and bootstrapping parameters. Mobility shift assays were used to confirm binding activity. RESULTS: All nucleotide sequences were HWV-1 subtype C based on phylogenetic analysis. The isolates from Pune/Delhi formed subclusters when analyzed separately, irrespective of time or sample source. However, no significant subclustering was observed with isolates from Mumbai or Vellore or with the entire sample set when analyzed collectively. Subtype-specific enhancer analysis revealed an expected third NF-kappaB site but also revealed six isolates with insertions and deletions not previously described, one of which resembles an AP-1 binding site. CONCLUSIONS: The results confirm the prevalence of HIV-1C and suggest increasingly complex phylogeny of HIV-1C within India, such that the previously observed subclustering may no longer adequately reflect the diversity of isolates currently circulating throughout India.

Base Sequence↗

Molecular ecology and evolution of Streptococcus thermophilus bacteriophages--a review.

Bacteriophages attacking Streptococcus thermophilus, a lactic acid bacterium used in milk fermentation, are a threat to the dairy industry. These small isometric-headed phages possess double-stranded DNA genomes of 31 to 45 kb. Yoghurt-derived phages exhibit a limited degree of variability, as defined by restriction pattern and host range, while a large diversity of phage types have been isolated from cheese factories. Despite this diversity all S. thermophilus phages, virulent and temperate, belong to a single DNA homology group. Several mechanisms appear to create genetic variability in this phage group. Site-specific deletions, one type possibly mediated by a viral recombinase/integrase, which transformed a temperate into a virulent phage, were observed. Recombination as a result of superinfection of a lysogenic host has been reported. Comparative DNA sequencing identified up to 10% sequence diversity due to point mutations. Genome sequencing of the prototype temperate phage phi Sfi21 revealed many predicted proteins which showed homology with phages from Lactococcus lactis suggesting horizontal gene transfer. Homology with phages from evolutionary unrelated bacteria like E. coli (e.g. lambdoid phage 434 and P1) and Mycobacterium phi L5 was also found. Due to their industrial importance, the existence of large phage collections, and the whole phage genome sequencing projects which are currently underway, the S. thermophilus phages may present an interesting experimental system to study bacteriophage evolution.

Bacteriophages↗

Allelic diversity of the human plasma alpha(1,3)fucosyltransferase gene (FUT6).

The 1080-bp coding region of the human plasma alpha(1,3)fucosyltransferase gene (FUT6) was sequenced in a total of 161 individuals (322 chromosomes) drawn from three populations, involving 56 Africans (Xhosa), 52 European-Africans of South Africa, and 53 Japanese. In addition to six reported base substitutions, eleven new base substitutions and a single base insertion were found in the coding region of the FUT6. Eleven functional and four null alleles were encountered, of which 10 alleles were novel alleles identified in this study. Two null alleles have been identified previously, whereas two novel null alleles, which contained a single base (cytosine) insertion at nucleotide 499, were found in a Xhosa population. The allelic distributions of FUT6 were different among these three populations. The heterozygosity of FUT6 was 0.860, 0.699, and 0.632, in Xhosa, European-African (South Africa), and in Japanese populations, respectively. The extensive DNA sequence diversity of the FUT6 may be suitable for application as a tool in genetic studies for modern human evolution.

Alleles↗

Genes coding for tryptophan-rich proteins are transcribed throughout the asexual cycle of Plasmodium falciparum.

Multigene families are a common feature in Plasmodia spp. and constitute a substantial content of the parasite genome. Here, we analyse the structural organisation and sequence diversity of two further members of the Trp-rich multigene family of P. falciparum. The complete DNA sequence of both genes was determined from a series of laboratory adapted and field isolates. Based on the amino acid sequences, we have termed them tryptophan-rich antigen-3 (TrpA-3) and lysine-tryptophan-rich antigen (LysTrpA). Analysis of the genes using reverse transcriptase-polymerase chain reaction (RT-PCR), showed that both genes are transcribed and that introns are spliced out at predicted positions. Gene expression profiles obtained from microarray analysis indicate that both genes are expressed in the mid-stages of the asexual cycle. In-frame stop codons were detected which interrupted the reading frame of LysTrpA. Whereas the number of the Trp-rich proteins is rather low in P. falciparum, P. chabaudi, P. berghei and P. yoelii, this family seems to have 15 or more members in P. knowlesi and P. vivax.

Amino Acid Sequence↗

Selection of linkers for a catalytic single-chain antibody using phage display technology.

Phage display has been evaluated as a means of rapidly selecting tailored linkers for single-chain antibodies (scFvs) from protein linker libraries. Preliminary experiments with a conventional linker failed to yield a functional single-chain version of a catalytic antibody with chorismate mutase activity. A random linker library was therefore constructed in which the genes for the heavy and light chain variable domains were linked by a segment encoding an 18-amino acid polypeptide of variable composition. The scFv repertoire ( approximately 5 x 10(6) different members) was displayed on filamentous phage and subjected to affinity selection with hapten. The population of selected variants exhibited significant increases in binding activity but retained considerable sequence diversity. Screening 1054 individual variants subsequently yielded a catalytically active scFv that was produced efficiently in soluble form. Sequence analysis revealed a conserved proline in the linker two residues after the VH C terminus and an abundance of arginines and prolines at other positions as the only common features of the selected tethers. There are apparently many viable solutions to the problem of linking individual VH and VL domains, but subtle differences in sequence dramatically influence the production, stability, and recognition properties of the scFv. The success of these experiments suggests that phage display will be generally useful for identifying peptide sequences for covalently linking any two protein domains.

Amino Acid Sequence↗

Evolutionary analysis of variants of hepatitis C virus found in South-East Asia: comparison with classifications based upon sequence similarity.

Variants of hepatitis C virus (HCV) have been classified by nucleotide sequence comparisons in different regions of the genome. Many investigators have defined the ranges of sequence similarity values or evolutionary distances corresponding to divisions of HCV into types, subtypes and isolates. Using these criteria, novel variants of HCV from Vietnam, Thailand and Indonesia have been classified as types 7, 8, 9, 10 and 11, many of which can be further subdivided into between two to four subtypes. In this study, this distance-based method of virus classification was compared with phylogenetic analysis and statistical measures to establish the confidence of the groupings. Using bootstrap resampling of phylogenetic trees in several subgenomic regions (core, E1, NS5) and with complete genomic sequences, we found that one set of novel HCV variants ('types 7, 8, 9 and 11') consistently grouped together into a single clade that also contained type 6a, while 'type 10a' grouped with type 3. In contrast, no robust higher-order groupings were observed between any of the other five previously described HCV genotypes (types 1-5). In each subgenomic region, the distribution of pairwise distances between members of the type 6 clade were consistently bi-modal and therefore provided no justification for classification of these variants into the three proposed categories (type, subtype, isolate). Based on these results, we propose that a more useful classification would regard all these variants as subtypes of type 6 or type 3, even though the level of sequence diversity within the clade was greater than observed for other genotypes. Classification by phylogenetic relatedness rules out simple sequence similarity measurements as a method for assigning HCV genotypes, but provides a more appropriate description of the evolutionary and epidemiological history of a virus.

Asia, Southeastern↗

The structure of Ski8p, a protein regulating mRNA degradation: Implications for WD protein structure.

Ski8p is a 44-kD protein that primarily functions in the regulation of exosome-mediated, 3'--> 5' degradation of damaged mRNA. It does so by forming a complex with two partner proteins, Ski2p and Ski3p, which complete a complex that is capable of recruiting and activating the exosome/Ski7p complex that functions in RNA degradation. Ski8p also functions in meiotic recombination in complex with Spo11 in yeast. It is one of the many hundreds of primarily eukaryotic proteins containing tandem copies of WD repeats (also known as WD40 or beta-transducin repeats), which are short ~40 amino acid motifs, often terminating in a Trp-Asp dipeptide. Genomic analyses have demonstrated that WD repeats are found in 1%-2% of proteins in a typical eukaryote, but are extremely rare in prokaryotes. Almost all structurally characterized WD-repeat proteins are composed of seven such repeats and fold into seven-bladed beta propellers. Ski8p was thought to contain five WD repeats on the basis of primary sequence analysis implying a five-bladed propeller. The 1.9 A crystal structure unexpectedly exhibits a seven-bladed propeller fold with seven structurally authentic WD repeats. Structure-based sequence alignments show additional sequence diversity in the two undetected repeats. This demonstrates that many WD repeats have not yet been identified in sequences and also raises the possibility that the seven-bladed propeller may be the predominant fold for this family of proteins.

Amino Acid Motifs↗

Immunity and vaccine control of Echinococcus granulosus infection in animal intermediate hosts.

Much progress has been made with characterisation of the EG95 vaccine which can be used to prevent hydatid infection in animal intermediate hosts of Echinococcus granulosus. The vaccine comprises a single recombinant oncosphere antigen and the adjuvant Quil A. It induces complement-fixing antibodies that kill the invading oncosphere early in an infection. In the majority of vaccinated animals, no hydatid cysts occur following a challenge infection. However, a small number of viable cysts may occur in some vaccinated animals. The vaccine has proved effective in vaccine trials carried out in sheep in New Zealand, Australia, Argentina, Chile and China as well as in goats and cattle. Investigations of the genetic diversity of the gene encoding EG95 have identified no unequivocal variation within the G1 strain parasites; however DNA sequence diversity within the EG95 family of genes has been found in G6/G7 parasites. GMP production scale-up of the vaccine has been undertaken in New Zealand and China and it is expected that the vaccine will be become available through these sources for implementation as part of hydatid control programs worldwide.

Animals↗

Template selection during manipulation of complex mixtures by PCR.

PCR is ubiquitous in molecular biology. It is used to amplify single sequences from large genomes, or populations of sequences from complex mixtures such as cDNA libraries in mammalian cells. These cDNA libraries are often employed in subsequent labor-intensive experiments such as genetic screens, the outcome of which depends on library quality. The use of PCR to amplify diverse sequence populations raises important technical issues. One question is whether or not PCR is capable of maintaining population diversity, specifically with respect to template selection in the first rounds of the amplification process (i.e., the possibility that rare sequences in a complex mixture are lost because of amplification failure at the outset of the PCR). Here, we analyze the properties of PCR in the context of template selection in complex mixtures and show that it is an excellent method for preserving diversity.

Animals↗

Conservation of structural motifs and antigenic diversity in the Plasmodium falciparum merozoite surface protein-3 (MSP-3).

Merozoite surface protein-3 (MSP-3) is a secreted polymorphic antigen associated with erythrocytic schizonts and merozoites of Plasmodium falciparum asexual blood-stages. A prominent structural feature of MSP-3 is a domain composed of three blocks of tandemly-repeated heptads with the consensus sequence AXXAXXX. The three blocks of four alanine heptad-repeats are separated by short stretches of non-repetitive sequence unrelated to the heptad-repeat. C-terminal to the heptad-repeats, MSP-3 contains a glutamic acid-rich domain followed by another heptad-repeat similar to a leucine-zipper motif. An analysis of the msp-3 gene from four P. falciparum isolates shows that polymorphism in MSP-3 is predominantly due to sequence diversity in the N-terminal half of the predicted polypeptide within and flanking the heptad-repeats. Mutations in the region of the gene that encodes the alanine heptad-repeats appear to be of two types. Unique mutations in non-repetitive sequence have generated amino acid substitutions and deletions that result in unique sequences among MSP-3 variants. In contrast, mutations in the heptad-coding sequence are largely dimorphic and are clustered in one or two heptads in each of the three blocks of heptads. Despite the diversity within and flanking the heptad domain the AXXAXXX motif is highly conserved as are other features of the sequence that predict the formation of alpha-helical secondary structure. Recombinant proteins and a synthetic peptide were used to raise antisera to conserved and variable regions of MSP-3. Differential reactivity of these reagents with the parasite antigen identified the alanine heptad-repeat domain as a site of antigenic diversity among MSP-3 polypeptides.

Amino Acid Sequence↗

Genetic diversity within populations of cyanobacteria assessed by analysis of single filaments.

We have developed a technique for determining the genetic structure of populations of filamentous cyanobacteria. The sequence diversity at specific gene loci is first characterised in a range of clonal cultures; subsequent analysis involves individual trichomes collected directly from natural populations. This technique has been used to examine the population genetic structure of Nodularia in the Baltic Sea and Planktothrix in Lake Zürich. For Nodularia, studies utilising four polymorphic loci reveal that even though there is a degree of linkage disequilibrium, horizontal transfer of genetic information has been sufficient to generate many of the possible allelic combinations. Analyses reveal both spatial and temporal variation in population genetic structure. Other studies of both Nodularia and Planktothrir have shown a correlation between particular alleles at the gvpC locus and the critical pressure of the gas vesicles that accumulate within the cell. We are now investigating how the natural selection of different gas vesicle phenotypes, imposed by changes in the depth of the upper mixed layer of the water column, affects the relative success of individual cyanobacteria possessing different gvpC alleles.

Archaeal Proteins↗

Caenorhabditis elegans has two isozymic forms, CE-1 and CE-2, of fructose-1,6-bisphosphate aldolase which are encoded by different genes.

Two distinct types of cDNAs for fructose-1,6-bisphosphate (FBP) aldolase, Ce-1 and Ce-2, have been isolated from nematode Caenorhabditis elegans, and the respective recombinant aldolase isozymes, CE-1 and CE-2, have been purified and characterized. The Ce-1 and Ce-2 are 1282 and 1248 bp in total length, respectively, and both have an open reading frame of 1098 bp, which encodes 366 amino acid residues. The entire amino acid sequences deduced from Ce-1 and Ce-2 show a high degree of identity to one another and to those of vertebrate and invertebrate aldolases. The highest sequence diversity was found in the carboxyl-terminal region that corresponds to one of the isozyme group-specific sequences of vertebrate aldolase isozymes that play a role in determining isozyme-specific functions. Southern blot analysis suggests that CE-1 and CE-2 are encoded by different genes. Concerning general or kinetic properties, CE-2 is quite different from CE-1. CE-1 exhibits unique characteristics which are not identical to any aldolase isozymes previously reported, whereas CE-2 is similar to vertebrate aldolase C. These results suggest that CE-2 might preserve the properties of a progenitor aldolase with a moderate preference for FBP over fructose 1-phosphate (F1P) as a substrate, whereas CE-1 evolved to act as an intrinsic enzyme that exhibits a much broader substrate specificity than dose CE-2.

Amino Acid Sequence↗

The salmonid MHC class I: more ancient loci uncovered.

An unprecedented level of sequence diversity has been maintained in the salmonid major histocompatibility complex (MHC) class I UBA gene, with between lineage AA sequence identities as low as 34%. The derivation of deep allelic lineages may have occurred through interlocus exon shuffling or convergence of ancient loci with the UBA locus, but until recently, no such ancient loci were uncovered. Herein, we document the existence of eight additional MHC class I loci in salmon (UCA, UDA, UEA, UFA, UGA, UHA, ULA, and ZE), six of which share exon 2 and 3 lineages with UBA, and three of which have not been described elsewhere. Half of the UBA exon 2 lineages and all UBA exon 3 lineages are shared with other loci. Two loci, UGA and UEA, share only a single exon lineage with UBA, likely generated through exon shuffling. Based on sequence homologies, we hypothesize that most exchanges and duplications occurred before or during tetraploidization (50 to 100 Ma). Novel loci that share no relationship with other salmonid loci are also identified (UHA and ZE). Each locus is evaluated for its potential to function as a class Ia gene based on gene expression, conserved residues and polymorphism. UBA is the only locus that can indisputably be classified as a class Ia gene, although three of the eight loci (ZE, UCA, and ULA) conform in three out of four measures. We hypothesize that these additional loci are in varying states of degradation to class Ib genes.

Amino Acid Sequence↗

Gene conversion between murine class II major histocompatibility complex loci. Functional and molecular evidence from the bm 12 mutant.

The experiments presented in this study define the molecular basis of the bm 12 mutation. Initial characterization of an alloreactive T cell clone, 4.1.4, showed this clone to recognize an allodeterminant present on the E beta b and A beta bm12 chains, but not on the bm 12 parent A beta b chain. To define the extent of sequence shared between the I-E beta product and the mutant I-A beta product, we isolated a cDNA clone of the E beta b gene and determined its nucleotide sequence. Comparison of the nucleotide sequences of E beta b, A beta b, and A beta bm12 shows the the A beta bm12 gene to be identical to the E beta b gene in the region where it differs from its A beta b parent. We predict that the bm 12 mutation arose by gene conversion of this region, which spans 14 nucleotides between amino acid residues 67-71 of the mature A beta chain, from the E beta b locus to the corresponding position at the A beta b locus. Recognition of this region, which spans one of the previously defined E beta allelic "hypervariable" regions, by an alloreactive T cell clone provides the first direct evidence of the functional importance of these hypervariable regions in T cell stimulation. The identification of a gene conversion event involving one of these allelic variable regions implicates conversion as a mechanism that acts on class II beta genes to create sequence diversity in regions of Ia molecules that interact with foreign antigen or a T cell receptor, regions where protein sequence polymorphism would presumably be selected for by the expanded ability it affords the organism to mount effective immune responses against a wider variety of foreign antigens.

Animals↗