Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequence diversity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

The systematics of North American Daphnia (Crustacea: Anomopoda): a molecular phylogenetic approach.

Despite extensive studies on the ecology and evolution of the freshwater microcrustacean Daphnia, there is little understanding of the evolutionary history of the genus. Past attempts at reconstructing phylogenetic relationships among Daphnia species have been highly controversial, mainly because of the poor taxonomy of the genus. However, following a revised taxonomy of the daphniid fauna of North America, we conducted a comprehensive appraisal of systematic relationships within the genus through the analysis of sequence diversity in 503 b.p. of the 12S rRNA gene of the mtDNA. The large sequence divergence among its 34 North American members indicates that the genus Daphnia originated during the Mesozoic, even though many lineages exhibit extreme morphological stasis. Results from both cladistic and phenetic analyses indicate the presence of three subgenera comprised of 15 species complexes. Only four of these lineages have shown active speciation over the past 3 Ma, suggesting that cladogenesis in the genus has been constrained. Our study also reveals that interspecific hybridization occurs between taxa which show very large sequence divergence (up to 14%), suggesting that reproductive isolation within the genus evolves slowly.

Animals↗

A superfamily of S locus-related sequences in Arabidopsis: diverse structures and expression patterns.

Six sequences that are closely related to the S gene family of the largely self-incompatible Brassica species have been identified in self-fertilizing Arabidopsis. The sequences define four genomic regions that map to chromosomes 1 and 3. Of the four functional genes identified, only the previously reported Arabidopsis AtS1 gene was expressed specifically in papillar cells and may function in pollination. The remaining three genes, including two novel genes designated ARK2 and ARK3, encode putative receptor-like serine/threonine protein kinases that are expressed predominantly in vegetative tissues. ARK2 promoter activity was detected exclusively in above-ground tissues, specifically in cotyledons, leaves, and sepals, in correlation with the maturation of these structures. ARK3 promoter activity was detected in roots as well as above-ground tissues but was limited to small groups of cells in the root-hypocotyl transition zone and at the base of lateral roots, axillary buds, and pedicels. The nonoverlapping patterns of expression of the ARK genes and the divergence of their sequences, particularly in their predicted extracellular domains, suggest that these genes perform nonredundant functions in specific aspects of development or growth of the plant body.

Amino Acid Sequence↗

Evolutionary genetics of the capsular locus of serogroup 6 pneumococci.

The evolution of the capsular biosynthetic (cps) locus of serogroup 6 Streptococcus pneumoniae was investigated by analyzing sequence variation within three serotype-specific cps genes from 102 serotype 6A and 6B isolates. Sequence variation within these cps genes was related to the genetic relatedness of the isolates, determined by multilocus sequence typing, and to the inferred patterns of recent evolutionary descent, explored using the eBURST algorithm. The serotype-specific cps genes had a low percent G+C, and there was a low level of sequence diversity in this region among serotype 6A and 6B isolates. There was also little sequence divergence between these serotypes, suggesting a single introduction of an ancestral cps sequence, followed by slight divergence to create serotypes 6A and 6B. A minority of serotype 6B isolates had cps sequences (class 2 sequences) that were approximately 5% divergent from those of other serotype 6B isolates (class 1 sequences) and which may have arisen by a second, more recent introduction from a related but distinct source. Expression of a serotype 6A or 6B capsule correlated perfectly with a single nonsynonymous polymorphism within wciP, the rhamnosyl transferase gene. In addition to ample evidence of the horizontal transfer of the serotype 6A and 6B cps locus into unrelated lineages, there was evidence for relatively frequent changes from serotype 6A to 6B, and vice versa, among very closely related isolates and examples of recent recombinational events between class 1 and 2 cps serogroup 6 sequences.

Amino Acid Sequence↗

SeqVISTA: a graphical tool for sequence feature visualization and comparison.

BACKGROUND: Many readers will sympathize with the following story. You are viewing a gene sequence in Entrez, and you want to find whether it contains a particular sequence motif. You reach for the browser's "find in page" button, but those darn spaces every 10 bp get in the way. And what if the motif is on the opposite strand? Subsequently, your favorite sequence analysis software informs you that there is an interesting feature at position 13982-14013. By painstakingly counting the 10 bp blocks, you are able to examine the sequence at this location. But now you want to see what other features have been annotated close by, and this information is buried several screenfuls higher up the web page. RESULTS: SeqVISTA presents a holistic, graphical view of features annotated on nucleotide or protein sequences. This interactive tool highlights the residues in the sequence that correspond to features chosen by the user, and allows easy searching for sequence motifs or extraction of particular subsequences. SeqVISTA is able to display results from diverse sequence analysis tools in an integrated fashion, and aims to provide much-needed unity to the bioinformatics resources scattered around the Internet. Our viewer may be launched on a GenBank record by a single click of a button installed in the web browser. CONCLUSION: SeqVISTA allows insights to be gained by viewing the totality of sequence annotations and predictions, which may be more revealing than the sum of their parts. SeqVISTA runs on any operating system with a Java 1.4 virtual machine. It is freely available to academic users at http://zlab.bu.edu/SeqVISTA.

Amino Acid Sequence↗

Genetic evidence for the origins of Venezuelan equine encephalitis virus subtype IAB outbreaks.

Epizootics of Venezuelan equine encephalitis (VEE) involving subtype IAB viruses occurred sporadically in South, Central and North America from 1938 to 1973. Incompletely inactivated vaccines have long been suspected as a source of the later epizootics. We tested this hypothesis by sequencing the PE2 glycoprotein precursor (1,677 nucleotides) or 26S/nonstructural protein 4 (nsP4) genome regions (4,490 nucleotides) for isolates representing most major outbreaks. Two distinct IAB genotypes were identified: 1) 1940s Peruvian strains and 2) 1938-1973 isolates from South, Central, and North America. Nucleotide sequences of these two genotypes differed by 1.1%, while the latter group showed only 0.6% sequence diversity. Early VEE virus IAB strains that were used for inactivated vaccine preparation had sequences identical to those predicted by phylogenetic analyses to be ancestors of the 1960s-1970s outbreaks. These data support the hypothesis of a vaccine origin for many VEE outbreaks. However, continuous, cryptic circulation of IAB viruses cannot be ruled out as a source of epizootic emergence.

Amino Acid Sequence↗

Studies on the diversity of the distinct phylogenetic lineage encompassing Glomus claroideum and Glomus etunicatum.

Morphological and molecular characters were analysed to investigate diversity within isolates of the Glomus claroideum/Glomus etunicatum species group in the genus Glomus. The inter- and intra-isolate sequence diversity of the large subunit (LSU) rRNA gene D2 region of eight isolates of G. claroideum and G. etunicatum was studied using PCR-single strand conformational polymorphism (SSCP)-sequencing. In addition, two isolates recently obtained from Southern China were included in the analysis to allow for a wider geographic screening. Single spore DNA isolation confirmed the magnitude of gene diversity found in multispore DNA extractions. An apparent overlap of spore morphological characters was found between G. claroideum and G. etunicatum in some isolates. Analysis of the sequence frequencies in all G. etunicatum and G. claroideum isolates (ten) showed that four LSU D2 sequences, representing 32.1% of the clones analysed for multispore extraction (564) were found to be common to both species, and those sequences were the most abundant in four of the ten isolates analysed. The frequency of these sequences ranged between 23.2% and 87.5% of the clones analysed in each isolate. The implications for the use of phenotypic characters to define species in arbuscular mycorrhizal fungi are discussed. The current position of G. claroideum/G.etunicatum in the taxonomy of the Glomeromycota is also discussed.

Biodiversity↗

Phylogenetic comparison and molecular epidemiology of classical swine fever virus.

The genetic diversity of classical swine fever virus (CSFV) was studied by RT-PCR amplification and sequencing of a 409 bp fragment of the NS5B polymerase region. A total of 106 viruses isolated from 20 countries over a period of 52 years (1945-1997) were included in the phylogenetic study. The results showed that the viruses could be divided into two main groups. Group 1 consisted of Asian and South American isolates from the 1980s, as well as of old European and American isolates. Group 2 consisted mostly of recent European viruses from the 1980s and 1990s, and was further divided into three subgroups largely according to geographic origin and/or year of isolation. Five 1997 CSFV isolates from Germany, Netherlands and Italy clustered together indicating a common origin for these outbreaks, but two other 1997 isolations in different regions of Germany are likely due to different epidemiological events. The results show that the NSSB region of the genome gives a good resolution for phylogenetic studies of CSFV. Molecular epidemiology based on nucleotide sequence diversity is a useful tool for tracing virus spread and for developing disease control strategies.

Animals↗

Selection for antigenic diversity of Tams1, the major merozoite antigen of Theileria annulata.

Tams1, the major merozoite/piroplasm surface antigen of Theileria annulata has the potential to be a component of a diagnostic ELISA test and be included in a recombinant subunit vaccine. However, the observation that this antigen displays diversity could constrain these applications. In this paper we have extensively characterized Tams1 diversity at the DNA level, using a PCR/sequencing strategy. Up to 44 alleles have been cloned and sequenced. The comparison of these alleles has identified regions of sequence conservation, variability and hyper-variability. Computer analysis of these alleles has indicated that positive selection may operate on certain regions of Tams1. Expression and Western blot analysis of selected alleles has indicated that sequence diversity is reflected in altered antigenicity and a continuum of relatedness and antibody cross recognition may exist. The possible function of the sequence conservation and polymorphism within Tams1 is discussed in relation to protein structure, host cell invasion and immune evasion.

Alleles↗

Detection of genetic diversity in linear plasmids 28-3 and 36 in Borrelia burgdorferi sensu stricto isolates by subtractive hybridization.

Recent studies based on sequence divergence in the ospC gene have identified limited subpopulations of B. burgdorferi associated with invasive human disease. Spirochetes with certain OspC types never cause human disease, while some others cause local infection at the primary skin site but do not hematogenously disseminate. Only four OspC genotypes (A, B, I and K) are responsible for disseminated disease and are found in the blood and cerebrospinal fluid, and hence are termed invasive strains. Subtractive hybridization was carried out between a prototype of a low passage invasive type, strain B31, and a strain associated only with local infection, group E, to identify genes associated with hematogenous dissemination. Two clones isolated from the subtraction library were unique to the B31 genome and mapped to locus BBH26 located on linear plasmid 28-3 (lp28-3) and to locus BBK48 located on linear plasmid 36 (lp36). Sequence analysis of the BBH26 locus revealed an amino acid repeat motif in the group E DNA that was absent in the B31 genome. This in-frame repeat motif was present yet variable in DNA isolated from several major OspC groups. However, no consistent sequence diversity was noted when other invasive and non-invasive strains were compared. In contrast, analysis of the BBK48 locus revealed a striking distinction between invasive and non-invasive spirochetes. PCR and Southern blot analysis indicated this locus was only present in invasive groups A, B, I, and K. BBK48 is a member of a gene family clustered on lp36. Therefore, these findings indicate that this genetic loci may participate in differentiating pathogens from non-pathogens and that its presence, which is correlated with ospC type, may play a role determining infectivity in humans.

Amino Acid Sequence↗

Interhomologue sequence variation of alpha satellite DNA from human chromosome 17: evidence for concerted evolution along haplotypic lineages.

Alpha satellite DNA is a family of tandemly repeated DNA found at the centromeres of all primate chromosomes. Different human chromosomes 17 in the population are characterized by distinct alpha satellite haplotypes, distinguished by the presence of variant repeat forms that have precise monomeric deletions. Pair-wise comparisons of sequence diversity between variant repeat units from each haplotype show that they are closely related in sequence. Direct sequencing of PCR-amplified alpha satellite reveals heterogeneous positions between the repeat units on a chromosome as two bands at the same position on a sequencing ladder. No variation was detected in the sequence and location of these heterogeneous positions between chromosomes 17 from the same haplotype, but distinct patterns of variation were detected between chromosomes from different haplotypes. Subsequent sequence analysis of individual repeats from each haplotype confirmed the presence of extensive haplotype-specific sequence variation. Phylogenetic inference yielded a tree that suggests these chromosome 17 repeat units evolve principally along haplotypic lineages. These studies allow insight into the relative rates and/or timing of genetic turnover processes that lead to the homogenization of tandem DNA families.

Base Sequence↗

The dual origin of the Malagasy in Island Southeast Asia and East Africa: evidence from maternal and paternal lineages.

Linguistic and archaeological evidence about the origins of the Malagasy, the indigenous peoples of Madagascar, points to mixed African and Indonesian ancestry. By contrast, genetic evidence about the origins of the Malagasy has hitherto remained partial and imprecise. We defined 26 Y-chromosomal lineages by typing 44 Y-chromosomal polymorphisms in 362 males from four different ethnic groups from Madagascar and 10 potential ancestral populations in Island Southeast Asia and the Pacific. We also compared mitochondrial sequence diversity in the Malagasy with a manually curated database of 19,371 hypervariable segment I sequences, incorporating both published and unpublished data. We could attribute every maternal and paternal lineage found in the Malagasy to a likely geographic origin. Here, we demonstrate approximately equal African and Indonesian contributions to both paternal and maternal Malagasy lineages. The most likely origin of the Asia-derived paternal lineages found in the Malagasy is Borneo. This agrees strikingly with the linguistic evidence that the languages spoken around the Barito River in southern Borneo are the closest extant relatives of Malagasy languages. As a result of their equally balanced admixed ancestry, the Malagasy may represent an ideal population in which to identify loci underlying complex traits of both anthropological and medical interest.

Africa↗

Classification of hepatitis C virus into six major genotypes and a series of subtypes by phylogenetic analysis of the NS-5 region.

Hepatitis C virus (HCV) showed substantial nucleotide sequence diversity distributed throughout the viral genome, with many variants showing only 68 to 79% overall sequence similarity to one another. Phylogenetic analysis of nucleotide sequences derived from part of the gene encoding a non-structural protein (NS-5) has provided evidence for six major genotypes of HCV amongst a worldwide collection of 76 samples from HCV-infected blood donors and patients with chronic hepatitis. Many of these HCV types comprised a number of more closely related subtypes, leading to a current total of 11 genetically distinct viral populations. Phylogenetic analysis of other regions of the viral genome produced relationships between published sequences equivalent to those found in NS-5, apart from the more highly conserved 5' non-coding region in which only the six major HCV types, but not subtypes, could be differentiated. A new nomenclature for HCV variants is proposed in this communication that reflects the two-tiered nature of sequence differences between different viral isolates. The scheme classifies all known HCV variants to date, and describes criteria that would enable new variants to be assigned within the classification as they are discovered.

Base Sequence↗

Molecular organization of the class I genes of human major histocompatibility complex.

In this brief review, our main emphasis has been on the analysis of the sequence diversity among various class I genes and their functional implications. The availability of complete nucleotide sequences of 7 different genes representing different loci allowed us to derive a consensus sequence. One mouse MHC Class I gene was included in these comparisons as a representative of H2 genes Evolutionary patterns can be seen on the basis of divergence of various genes from the derived consensus sequence. At least 1 human gene which has a promoter similar to that of H2 genes and which contains a single initiation codon following this promoter (unlike all other human genes and like all the H2 genes) has been identified. Both variable and homology regions can be identified in the entire length of the gene. While exons show relatively strong conservation of sequences, the introns have many variable regions, introns 6 and 7 being the most heterogeneous. Stretches of conserved nucleotide sequences are noticed at the 3' regions of most introns. Estimation of total number of class I genes is presented on the basis of cloning experiments, and the abundance of 1 particular pseudogene is discussed.

Alleles↗

Prediction of the functional class of lipid binding proteins from sequence-derived properties irrespective of sequence similarity.

Lipid binding proteins play important roles in signaling, regulation, membrane trafficking, immune response, lipid metabolism, and transport. Because of their functional and sequence diversity, it is desirable to explore additional methods for predicting lipid binding proteins irrespective of sequence similarity. This work explores the use of support vector machines (SVMs) as such a method. SVM prediction systems are developed using 14,776 lipid binding and 133,441 nonlipid binding proteins and are evaluated by an independent set of 6,768 lipid binding and 64,761 nonlipid binding proteins. The computed prediction accuracy is 78.9, 79.5, 82.2, 79.5, 84.4, 76.6, 90.6, 79.0, and 89.9% for lipid degradation, lipid metabolism, lipid synthesis, lipid transport, lipid binding, lipopolysaccharide biosynthesis, lipoprotein, lipoyl, and all lipid binding proteins, respectively. The accuracy for the nonmember proteins of each class is 99.9, 99.2, 99.6, 99.8, 99.9, 99.8, 98.5, 99.9, and 97.0%, respectively. Comparable accuracies are obtained when homologous proteins are considered as one, or by using a different SVM kernel function. Our method predicts 86.8% of the 76 lipid binding proteins nonhomologous to any protein in the Swiss-Prot database and 89.0% of the 73 known lipid binding domains as lipid binding. These findings suggest the usefulness of SVMs for facilitating the prediction of lipid binding proteins. Our software can be accessed at the SVMProt server (http://jing.cz3.nus.edu.sg/cgi-bin/svmprot.cgi).

Algorithms↗

Structural diversity of streptokinase and activation of human plasminogen.

The beta domain of streptokinase is required for plasminogen activation and contains a region of sequence diversity associated with infection and disease in group A streptococci. We report that mutagenesis of this polymorphic region does not alter plasminogen activation, which suggests an alternative function for this molecular motif in streptococcal disease.

Amino Acid Sequence↗

Molecular cloning and characterization of Dr-II, a nonfimbrial adhesin-I-like adhesin isolated from gestational pyelonephritis-associated Escherichia coli that binds to decay-accelerating factor.

Bacterial adhesins play an important role in the colonization of the human urogenital tract. Escherichia coli Dr family adhesins have been found to be frequently expressed in strains associated with pyelonephritis in pregnant females. The tissue receptor for known Dr adhesins has been localized to the short consensus repeat-3 (SCR-3) domain of decay accelerating factor (DAF), a complement regulatory protein. In this report, we identified and cloned draE2, a gene encoding a novel 17-kDa DAF-binding adhesin, Dr-II, from a strain of E. coli associated with acute gestational pyelonephritis. Despite the significant sequence diversity between Dr-II and Dr family adhesins, the receptor of Dr-II was found to be the SCR-3 domain of DAF. Sequence analysis of the 186-amino-acid Dr-II open reading frame revealed significant diversity from other members of the Dr adhesin family, including Dr, AFA-I, AFA-III, and F1845, but only an 8-amino-acid difference in sequence from that of the 17-kDa nonfimbrial adhesin NFA-I of unknown receptor specificity. N-terminal peptide sequencing of the purified adhesin confirmed the identity of the open reading frame and indicated cleavage of a 28-amino-acid signal peptide. Antibodies raised against purified Dr-II adhesin exhibited little or no cross-reactivity to Dr adhesin. Characterization of the biological properties demonstrated that like the Dr adhesins, Dr-II was associated with the ability of E. coli to bind to tubular basement membranes and Bowman's capsule and to be internalized into HeLa cells.

Adhesins, Escherichia coli↗

Distribution of genetic diversity in relation to chromosomal inversions in the malaria mosquito Anopheles gambiae.

The epidemiology of malaria in Africa is complicated by the fact that its principal vector, the mosquito Anopheles gambiae, constitutes a complex of six sibling species. Each species is characterized by a unique array of paracentric inversions, as deduced by karyotypic analysis. In addition, most of the species carry a number of polymorphic inversions. In order to develop an understanding of the evolutionary histories of different parts of the genome, we compared the genetic variation of areas inside and outside inversions in two distinct inversion karyotypes of A. gambiae. Thirty-five cDNA clones were mapped on the five arms of the A. gambiae chromosomes with divisional probes. Sixteen of these clones, localized both inside and outside inversions of chromosome 2, were used as probes in order to determine the nucleotide diversity of different parts of the genome in the two inversion karyotypes. We observed that the sequence diversity inside the inversion is more than three-fold lower than in areas outside the inversion and that the degree of divergence increases gradually at loci at increasing distance from the inversion. To interpret the data we present a selectionist and a stochastic model, both of which point to a relatively recent origin of the studied inversion and may suggest differences between the evolutionary history of inversions in Anopheles and Drosophila species.

Animals↗

Detection of identical JC virus DNA sequences in both human kidneys.

We studied JC virus (JCV) DNA sequence diversity among kidneys derived from cadavers with various causes of death. The 610-bp JCV DNA sequences we evaluated were identical not only among specimens derived from the same kidney but also among those derived from both kidneys of the same cadaver. Because the left and right kidneys are anatomically independent, our findings suggest that the viremia that has been proposed to occur after primary infection distributes the same JCV strain to both kidneys.

Adult↗