Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,711 records · Page 95Linked to original sources

A method for detecting positive selection at single amino acid sites.

A method was developed for detecting the selective force at single amino acid sites given a multiple alignment of protein-coding sequences. The phylogenetic tree was reconstructed using the number of synonymous substitutions. Then, the neutrality was tested for each codon site using the numbers of synonymous and nonsynonymous changes throughout the phylogenetic tree. Computer simulation showed that this method accurately estimated the numbers of synonymous and nonsynonymous substitutions per site, as long as the substitution number on each branch was relatively small. The false-positive rate for detecting the selective force was generally low. On the other hand, the true-positive rate for detecting the selective force depended on the parameter values. Within the range of parameter values used in the simulation, the true-positive rate increased as the strength of the selective force and the total branch length (namely the total number of synonymous substitutions per site) in the phylogenetic tree increased. In particular, with the relative rate of nonsynonymous substitutions to synonymous substitutions being 5.0, most of the positively selected codon sites were correctly detected when the total branch length in the phylogenetic tree was > or = 2.5. When this method was applied to the human leukocyte antigen (HLA) gene, which included antigen recognition sites (ARSs), positive selection was detected mainly on ARSs. This finding confirmed the effectiveness of the present method with actual data. Moreover, two amino acid sites were newly identified as positively selected in non-ARSs. The three-dimensional structure of the HLA molecule indicated that these sites might be involved in antigen recognition. Positively selected amino acid sites were also identified in the envelope protein of human immunodeficiency virus and the influenza virus hemagglutinin protein. This method may be helpful for predicting functions of amino acid sites in proteins, especially in the present situation, in which sequence data are accumulating at an enormous speed.

Amino Acid Substitution↗

Analysis of the levels of conservation of the J domain among the various types of DnaJ-like proteins.

DnaJ-like proteins are defined by the presence of an approximately 73 amino acid region termed the J domain. This region bears similarity to the initial 73 amino acids of the Escherichia coli protein DnaJ. Although the structures of the J domains of E coli DnaJ and human heat shock protein 40 have been solved using nuclear magnetic resonance, no detailed analysis of the amino acid conservation among the J domains of the various DnaJ-like proteins has yet been attempted. A multiple alignment of 223 J domain sequences was performed, and the levels of amino acid conservation at each position were established. It was found that the levels of sequence conservation were particularly high in 'true' DnaJ homologues (ie, those that share full domain conservation with DnaJ) and decreased substantially in those J domains in DnaJ-like proteins that contained no additional similarity to DnaJ outside their J domain. Residues were also identified that could be important for stabilizing the J domain and for mediating the interaction with heat shock protein 70.

Amino Acid Sequence↗

Combining bioinformatics and phylogenetics to identify large sets of single-copy orthologous genes (COSII) for comparative, evolutionary and systematic studies: a test case in the euasterid plant clade.

We report herein the application of a set of algorithms to identify a large number (2869) of single-copy orthologs (COSII), which are shared by most, if not all, euasterid plant species as well as the model species Arabidopsis. Alignments of the orthologous sequences across multiple species enabled the design of "universal PCR primers," which can be used to amplify the corresponding orthologs from a broad range of taxa, including those lacking any sequence databases. Functional annotation revealed that these conserved, single-copy orthologs encode a higher-than-expected frequency of proteins transported and utilized in organelles and a paucity of proteins associated with cell walls, protein kinases, transcription factors, and signal transduction. The enabling power of this new ortholog resource was demonstrated in phylogenetic studies, as well as in comparative mapping across the plant families tomato (family Solanaceae) and coffee (family Rubiaceae). The combined results of these studies provide compelling evidence that (1) the ancestral species that gave rise to the core euasterid families Solanaceae and Rubiaceae had a basic chromosome number of x=11 or 12.2) No whole-genome duplication event (i.e., polyploidization) occurred immediately prior to or after the radiation of either Solanaceae or Rubiaceae as has been recently suggested.

Algorithms↗

[Cloning and identification of the priming glycosyltransferase gene involved in exopolysaccharide 139A biosynthesis in Streptomyces].

Recently in our laboratory, Streptomyces sp. 139 has been identified to produce a new exopolysaccharide designated EPS 139A that shows anti-rheumatic arthritis activity. The strategy of studying EPS 139A biosynthesis is to clone the key gene in the EPS biosynthesis pathway, i.e. the priming glycosyltransferase gene catalyzing the first step of nucleotide sugar transfer. Degenerate primers-based PCR approach was adopted to isolate the putative priming glycosyltransferase gene in Streptomyces sp. 139. According to the genes encoding the priming glycosyltransferases that have been identified in several microorganisms, a multiple alignment of the amino acid sequences of these genes was used to identify regions conserved between all genes. To clone the priming glycosyltransferase gene in Streptomyces sp. 139, degenerate primers were designed from these conserved regions taking into account information on Streptomyces codon usage to amplify an internal DNA fragment of this gene. A distinctive PCR product with the expected size of 0.3 kb was amplified from Streptomyces sp. 139 total genomic DNA. Sequence analysis showed that it is part of a putative priming glycosyltransferase gene and contains the predicted conserved domain B. To isolate the complete priming glycosyltransferase gene, a Streptomyces sp. 139 genomic library was constructed in the E. coli--Streptomyces shuttle vector pOJ446. Using the 0.3 kb PCR product of priming glycosyltransferase gene as a probe, 17 positive colonies were isolated by colony hybridization. A 4.0 kb BamHI fragment from all positive cosmids that hybridized to this probe was sequenced, which revealed the complete priming glycosyltransferase gene. The priming glycosyltransferase gene ste5 (GenBank under accession number AY131229) most likely begins with GTG, preceded by a probable ribosome binding site (RBS), GGGGA. It encodes a 492-amino-acid protein with molecular weight of 54 kDa and isoelectric point of 10.6. The G + C content of ste5 is 73%, close to the average of G + C content (74%) for Streptomyces. Moreover, the preference usage of G or C as third base of codons are found in the ste5, which is in accordance with the Streptomyces codon usage. A BlastP search showed that the C-terminal region of Ste5 shows highly homology with a number of priming glycosyltransferases from many different organisms. Ste5 contains two putative catalytic residues, Glu and Asp (residues 423 and 474) with a spacing of approximately 50 amino acids that conserved in various beta-glycosyltransferases. Moreover, the C-terminal one third of Ste5 contains three domains, A, B and C that is reported to be common to glycosyltransferases. By hydrophilicity plot prediction, the N-terminal two thirds of Ste5 exhibits 5 putative transmembrane domains. To investigate the involvement of the identified polysaccharide gene cluster in EPS 139A biosynthesis, the gene ste5 encoding priming glycosyltransferase was insertionally disrupted by a single-crossover homologous recombination event. A 0.85 kb internal fragment of ste5 was cloned into vector pKC1139 to yield pLY5015 that was transduced into Streptomyces sp. 139. Correct integration in Streptomyces LY1001 ste5- mutant strain was confirmed by Southern hybridization. After fermentation, no EPS 139A could be detected in the cultures of ste5- mutant strain Streptomyces LY1001. Therefore, the gene ste5 identified in this work is involved in the synthesis of the Streptomyces sp. 139 EPS.

Amino Acid Sequence↗

FIE2: A program for the extraction of genomic DNA sequences around the start and translation initiation site of human genes.

FIE2 (5' end Information Extraction v2) is a web-based program for easy identification and extraction of nucleotide sequence around the start of genes (promoter region) and their translation initiation site (TIS). Using information provided by the National Center for Biotechnology Information's (NCBI's) LocusLink, FIE2 identifies the 5'-most end of a gene on its respective chromosome based on alignment of a selected set of mRNAs representative of the gene. FIE2 then uses currently available human genome sequence information to extract the desired sequences. The accuracy of the information extracted is therefore limited by the accuracy and completeness of the sequence annotation and sequence alignment provided by LocusLink. In addition, multiple TIS positions are also occasionally presented, for example, as a result of multiple alignments of transcript variants. One of the key criteria of FIE2 is that it should extract only the correct information or attempt no extraction at all. To date, the authors are not aware of any publicly available web-based tool that uses the human genomic sequence to extract pertinent promoter- and TIS-region information in this fashion. FIE2 is freely available at http://sdmc.lit.org.sg/FIE2.0.

Base Sequence↗

Homology modeling of a human glycine alpha 1 receptor reveals a plausible anesthetic binding site.

The superfamily of ligand-gated ion channels (LGICs) has been implicated in anesthetic and alcohol responses. Mutations within glycine and GABA receptors have demonstrated that possible sites of anesthetic action exist within the transmembrane subunits of these receptors. The exact molecular arrangement of this transmembrane region remains at intermediate resolution with current experimental techniques. Homology modeling methods were therefore combined with experimental data to produce a more exact model of this region. A consensus from multiple bioinformatics techniques predicted the topology within the transmembrane domain of a glycine alpha one receptor (GlyRa1) to be alpha helical. This fold information was combined with sequence information using the SeqFold algorithm to search for modeling templates. Independently, the FoldMiner algorithm was used to search for templates that had structural folds similar to published coordinates of the homologous nAChR (1OED). Both SeqFold and Foldminer identified the same modeling template. The GlyRa1 sequence was aligned with this template using multiple scoring criteria. Refinement of the alignment closed gaps to produce agreement with labeling studies carried out on the homologous receptors of the superfamily. Structural assignment and refinement was achieved using Modeler. The final structure demonstrated a cavity within the core of a four-helix bundle. Residues known to be involved in modulating anesthetic potency converge on and line this cavity. This suggests that the binding sites for volatile anesthetics in the LGICs are the cavities formed within the core of transmembrane four-helix bundles.

Amino Acid Sequence↗

Expression of Anaplasma marginale major surface protein 2 variants in persistently infected ticks.

Anaplasma marginale, an intraerythrocytic ehrlichial pathogen of cattle, establishes persistent infections in both vertebrate (cattle) and invertebrate (tick) hosts. The ability of A. marginale to persist in cattle has been shown to be due, in part, to major surface protein 2 (MSP2) variants which are hypothesized to emerge in response to the bovine immune response. MSP2 antigenic variation has not been studied in persistently infected ticks. In this study we analyzed MSP2 in A. marginale populations from the salivary glands of male Dermacentor variabilis persistently infected with A. marginale after feeding successively on one susceptible bovine and three sheep. New MSP2 variants appeared in each A. marginale population, and sequence alignment of the MSP2 variants revealed multiple amino acid substitutions, insertions, and deletions. These results suggest that selection pressure on MSP2 occurred in tick salivary glands independent of the bovine immune response.

Amino Acid Sequence↗

Multiple model approach--dealing with alignment ambiguities in protein modeling.

Sequence alignments for distantly homologous proteins are often ambiguous, which creates a weak link in structure prediction by homology. We address this problem by using several plausible alignments in a modeling procedure, obtaining many models of the target. All are subsequently evaluated by a threading algorithm. It is shown that this approach can identify best alignments and produce reasonable models, whose quality is now limited only by the extent of the structural similarity between the known and predicted protein. Using a similar approach structure prediction for the oxidized dimer of S100A1 protein, for which the structure is not known, is presented.

Amino Acid Sequence↗

Integrated graphical analysis of protein sequence features predicted from sequence composition.

Several protein sequence analysis algorithms are based on properties of amino acid composition and repetitiveness. These include methods for prediction of secondary structure elements, coiled-coils, transmembrane segments or signal peptides, and for assignment of low-complexity, nonglobular, or intrinsically unstructured regions. The quality of such analyses can be greatly enhanced by graphical software tools that present predicted sequence features together in context and allow judgment to be focused simultaneously on several different types of supporting information. For these purposes, we describe the SFINX package, which allows many different sets of segmental or continuous-curve sequence feature data, generated by individual external programs, to be viewed in combination alongside a sequence dot-plot or a multiple alignment of database matches. The implementation is currently based on extensions to the graphical viewers Dotter and Blixem and scripts that convert data from external programs to a simple generic data definition format called SFS. We describe applications in which dot-plots and flanking database matches provide valuable contextual information for analyses based on compositional and repetitive sequence features. The system is also useful for comparing results from algorithms run with a range of parameters to determine appropriate values for defaults or cutoffs for large-scale genomic analyses.

Amino Acid Motifs↗

Identification, cloning and characterization of a new DNA-binding protein from the hyperthermophilic methanogen Methanopyrus kandleri.

Three novel DNA-binding proteins with apparent molecular masses of 7, 10 and 30 kDa have been isolated from the hyperthermophilic methanogen Methanopyrus kandleri. The proteins were identified using a blot overlay assay that was modified to emulate the high ionic strength intracellular environment of M.kandleri proteins. A 7 kDa protein, named 7kMk, was cloned and expressed in Escherichia coli. As indicated by CD spectroscopy and computer-assisted structure prediction methods, 7kMk is a substantially alpha-helical protein possibly containing a short N-terminal beta-strand. According to analytical gel filtration chromatography and chemical crosslinking, 7kMk exists as a stable dimer, susceptible to further oligomerization. Electron microscopy showed that 7kMk bends DNA and also leads to the formation of loop-like structures of approximately 43.5 +/- 3.5 nm (136 +/- 11 bp for B-form DNA) circumference. A topoisomerase relaxation assay demonstrated that looped DNA is negatively supercoiled under physiologically relevant conditions (high salt and temperature). A BLAST search did not yield 7kMk homologs at the amino acid sequence level, but based on a multiple alignment with ribbon-helix-helix (RHH) transcriptional regulators, fold features and self-association properties of 7kMk we hypothesize that it could be related to RHH proteins.

Amino Acid Sequence↗

A method of estimating from two aligned present-day DNA sequences their ancestral composition and subsequent rates of substitution, possibly different in the two lineages, corrected for multiple and parallel substitutions at the same site.

The course of evolutionary change in DNA sequences has been modeled as a Markov process. The Markov process was represented by discrete time matrix methods. The parameters of the Markov transition matrices were estimated by least-squares direct-search optimization of the fit of the calculated divergence matrix to that observed for two aligned sequences. The Markov process corrected for multiple and parallel substitutions of bases at the same site. The method avoided the incorrect assumption of all previously described methods that the divergence between two present-day sequences is twice the divergence of either from the common and unknown ancestral sequence. The three previous methods were shown to be equivalent. The present method also avoided the undesirable assumptions that sequence composition has not changed with time and that the substitution rates in the two descendant lineages were the same. It permitted simultaneous estimation of ancestral sequence composition and, if applicable, of different substitution rates for the two descendant lineages, provided the total number of estimated parameters was less than 16. Properties of the Markov chain were discussed. It was proved for symmetric substitution matrices that all elements of the equilibrium divergence matrix equal 1/16, and that the total difference in the divergence matrix at epoch k equals the total change in the common substitution matrix at epoch 2k for all values of k. It was shown how to resolve an ambiguity in the assignment of two different substitution rates to the two descendant lineages when four or more similar sequences are available. The method was applied to the divergence matrix for codon site 3 for the mouse and rabbit beta-globins. This observed divergence matrix was significantly asymmetric and required at least two different substitution rates. This result could be achieved only by using different asymmetric substitution matrices for the two lineages.

Animals↗

Neighboring base composition and transversion/transition bias in a comparison of rice and maize chloroplast noncoding regions.

The correspondence between the transversion/transition ratio and the neighboring base composition in chloroplast DNA is examined. For 18 noncoding regions of the chloroplast genome, alignments between rice (Oryza sativa) and maize (Zea mays) were generated by two different methods. Difficulties of aligning noncoding DNA are discussed, and the alignments are analyzed in a manner that reduces alignment artifacts. Sequence divergence is < 10%, so multiple substitutions at a site are assumed to be rare. Observed substitutions were analyzed with respect to the A+T content of the two immediately flanking bases. It is shown that as this content increases, the proportion of transversions also increases. When both the 5'- and 3'-flanking nucleotides are G or C (A+T content of 0), only 25% of the observed substitutions are transversions. However, when both the 5'- and 3'-flanking nucleotides are A or T (A+T content of 2), 57% of the observed substitutions are transversions. Therefore, the influence of flanking base composition on substitutions, previously reported for a single noncoding region, is a general feature of the chloroplast genome.

Algorithms↗

Extensive sequence conservation among insect, nematode, and vertebrate vitellogenins reveals ancient common ancestry.

The eggs of most oviparous animals are provisioned with a class of protein called vitellogenin (Vg) which is stored as the major component of yolk. Until recently, deduced amino acid sequences were available only from vertebrate and nematode Vgs, which proved to be homologous. The sequences of several insect Vgs are now known, but early attempts at pairwise alignments with vertebrate and nematode Vgs have been problematic, leading to conflicting conclusions about how closely insect Vgs are related to the others. In this paper we demonstrate that insect Vg sequences can be confidently aligned with one another along their entire lengths and with multiple vertebrate and nematode Vg sequences along most of their spans. Although divergence is high, conservation among insect, vertebrate, and nematode Vg sequences is widespread with a preponderance of glycine, proline, and cysteine residues among strictly conserved amino acids, establishing conclusively that Vgs from the three phyla are homologous. Areas of least-certain alignment are primarily in and around insect and vertebrate polyserine domains which are not homologous. Phylogenetic reconstructions of Vgs based on sequence identities indicate that the insect lineage is the most diverged and that the mammalian serum protein, apolipoprotein B-100, arose from a Vg ancestor after the nematode/vertebrate divergence.

Amino Acid Sequence↗

T-cell-epitope mapping of the idiotypic monoclonal IgG heavy and light chains in multiple myeloma.

The idiotypic structures of the myeloma protein might be regarded as tumor-specific antigens. The present study was designed to map T-cell epitopes of the idiotypic myeloma protein to prove the existence of naturally occurring major-histocompatibility-complex-dependent idiotype (peptide)-specific T cells in multiple myeloma. The fine specificity of idiotype-reactive, interferon-gamma-producing blood T cells of a patient with multiple myeloma stage I was characterized by identification of idiotype (heavy and light chains)-derived MHC-restricted T-cell epitopes. T cells specifically reacting with peptides corresponding to each of the 3 complementarity-determining regions (CDRs) of the heavy-chain variable part (V(H)) of the autologous idiotype were found. In contrast, none of the peptides corresponding to the 3 CDRs of the light chain (V(L)) induced a specific T-cell response. The idiotype amino-acid sequence corresponding to the junction of the V(H), diversity (D), and joining (J) gene segments of the VH appeared to be an important target for T cells, since the sequence expressed MHC-class-I- as well as MHC-class-II-restricted epitopes. The study provides further support for the existence of MHC-restricted idiotype-specific T cells, which may target immunogenic CDR peptides in multiple myeloma. Such T cells could be an important part of the specific anti-tumor immune responses induced in idiotype vaccination protocols.

Aged↗

Evolutionary correlation between linker histones and microtubular structures.

Histones of the H1 group (linker histones) are abundant components of chromatin in eukaryotes, occurring on average at one molecule per nucleosome. The recent reports on the lack of a clear phenotypic effect of knock-out mutations as well as overexpression of histone H1 genes in different organisms have seriously undermined the long-held view that linker histones are essential for the basic functions of eukaryotic cells. In an attempt to resolve the paradox of an abundant conserved protein without a clear function, we re-examined the molecular and phylogenetic data on linker histones to see if they could reveal any correlation between the features of H1 and the functional or morphological characteristics of cells or organisms. Because of an earlier demonstration that in sea urchin the chromatin-type histone H1 is also found in the flagellar microtubules (Multigner et al. 1992), we focused on the correlation between the features of H1 and those of microtubular structures. A phylogenetic tree based on multiple alignment of over 100 available HI sequences suggests that the first divergence of the globular domain of H1 (GH1) resulted in branching into separate types characteristic for plants/Dictyostelium and for animals/ascomycetes, respectively. The GH1s of these two types differ by a short region (usually 5 amino acids) placed at a specific location within the C-terminal wing subdomain of GH1. Evolutionary analysis of the diversification of H1 mRNA into cell-cycle-dependent (polyA-) and independent (polyA+) forms showed a mosaic occurrence of these two forms in plants and animals, despite the fact that the H1 proteins of plants and animals belong to two well-distinguished groups. However, among organisms from both animal and plant kingdom, only those with H1 mRNA of a polyA- type have flagellated gametes. This correlation as well as the demonstration that in Volvox carteri the accumulation of polyA- mRNA of H1 occurs concurrently with the production of new flagella (Lindauer et al. 1993), suggests a direct link between polyA- phenotype of histone H1 mRNA and flagellogenesis.

Animals↗

The role of N286 and D320 in the reaction mechanism of human dihydrolipoamide dehydrogenase (E3) center domain.

According to the multiple alignment of various dihydrolipoamide dehydrogenases (E3s) sequences, three human mutant E3s of the conserved residues in the center domain, N286D, N286Q, and D320N were created, over-expressed and purified. We characterized these mutants to investigate the reaction mechanism of human dihydrolipoamide dehydrogenases. The specific activities of N286D, N286Q, and D320N are 30.84%, 24.57% and 48.60% to that of the wild-type E3 respectively. The FAD content analysis indicated that these mutant E3s about 96.0%, 99.4% and 82.7% of FAD content compared to that of wild-type E3 respectively. The molecular weight analysis showed that these three mutant proteins form the dimer. Kinetic's data demonstrated that the K(cat) of both forward and reverse reactions of these mutant proteins were decreased. These results suggest that N286 and D320 play a role in the catalytic function of the E3.

Amino Acid Sequence↗

Immune cascade of Spodoptera litura: cloning, expression, and characterization of inducible prophenol oxidase.

Haemolymph associated phenol oxidase is a critical component of invertebrate immune reaction and cuticle sclerotization. Phenol oxidase catalyses the conversion of mono-phenols to diphenols and quinones which finally leads to melanin formation. We have cloned the c-DNA encoding phenol oxidase from the haemocytes of Spodoptera litura and expressed it in Escherichia coli. The encoding gene is 2452bp with an open reading frame of 2091 bp translating into a 697 amino acid protein. Multiple alignment analysis of the predicted protein sequence shows close homology to other lepidopeteran PPOII type genes. The transcription of the gene is induced upon microbial challenge of 6th instar larvae with E. coli and is unresponsive to injury. Cloning of the ORF of SLPPO in-frame in the E. coli expression vector pQE30 resulted in its expression. Enzymatic analysis of the recombinant protein reveals that the recombinant protein is catalytically active on 4-methyl pyrocatechol upon activation by cetyl pyridinium chloride.

Amino Acid Sequence↗

Improving profile HMM discrimination by adapting transition probabilities.

Profile hidden Markov models (HMMs) are used to model protein families and for detecting evolutionary relationships between proteins. Such a profile HMM is typically constructed from a multiple alignment of a set of related sequences. Transition probability parameters in an HMM are used to model insertions and deletions in the alignment. We show here that taking into account unrelated sequences when estimating the transition probability parameters helps to construct more discriminative models for the global/local alignment mode. After normal HMM training, a simple heuristic is employed that adjusts the transition probabilities between match and delete states according to observed transitions in the training set relative to the unrelated (noise) set. The method is called adaptive transition probabilities (ATP) and is based on the HMMER package implementation. It was benchmarked in two remote homology tests based on the Pfam and the SCOP classifications. Compared to the HMMER default procedure, the rate of misclassification was reduced significantly in both tests and across all levels of error rate.

Algorithms↗