Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,801 records · Page 100Linked to original sources

Identification of mycobacterial species by comparative sequence analysis of the RNA polymerase gene (rpoB).

For the differentiation and identification of mycobacterial species, the rpoB gene, encoding the beta subunit of RNA polymerase, was investigated. rpoB DNAs (342 bp) were amplified from 44 reference strains of mycobacteria and clinical isolates (107 strains) by PCR. The nucleotide sequences were directly determined (306 bp) and aligned by using the multiple alignment algorithm in the MegAlign package (DNASTAR) and the MEGA program. A phylogenetic tree was constructed by the neighbor-joining method. Comparative sequence analysis of rpoB DNAs provided the basis for species differentiation within the genus Mycobacterium. Slowly and rapidly growing groups of mycobacteria were clearly separated, and each mycobacterial species was differentiated as a distinct entity in the phylogenetic tree. Pathogenic Mycobacterium kansasii was easily differentiated from nonpathogenic M. gastri; this differentiation cannot be achieved by using 16S rRNA gene (rDNA) sequences. By being grouped into species-specific clusters with low-level sequence divergence among strains of the same species, all of the clinical isolates could be easily identified. These results suggest that comparative sequence analysis of amplified rpoB DNAs can be used efficiently to identify clinical isolates of mycobacteria in parallel with traditional culture methods and as a supplement to 16S rDNA gene analysis. Furthermore, in the case of M. tuberculosis, rifampin resistance can be simultaneously determined.

Amino Acid Sequence↗

WWW access to the SYSTERS protein sequence cluster set.

SUMMARY: We present a Web server where the SYSTERS cluster set of the non-redundant protein database consisting of sequences from SWISS-PROT and PIR is being made available for querying and browsing. The cluster set can be searched with a new sequence using the SSMAL search tool. Additionally, a multiple alignment is generated for each cluster and annotated with domain information from the Pfam protein family database. AVAILABILITY: The server address is http://www.dkfz-heidelberg.de/tbi/services/cluster/ systersform

Algorithms↗

Gibbs motif sampling: detection of bacterial outer membrane protein repeats.

The detection and alignment of locally conserved regions (motifs) in multiple sequences can provide insight into protein structure, function, and evolution. A new Gibbs sampling algorithm is described that detects motif-encoding regions in sequences and optimally partitions them into distinct motif models; this is illustrated using a set of immunoglobulin fold proteins. When applied to sequences sharing a single motif, the sampler can be used to classify motif regions into related submodels, as is illustrated using helix-turn-helix DNA-binding proteins. Other statistically based procedures are described for searching a database for sequences matching motifs found by the sampler. When applied to a set of 32 very distantly related bacterial integral outer membrane proteins, the sampler revealed that they share a subtle, repetitive motif. Although BLAST (Altschul SF et al., 1990, J Mol Biol 215:403-410) fails to detect significant pairwise similarity between any of the sequences, the repeats present in these outer membrane proteins, taken as a whole, are highly significant (based on a generally applicable statistical test for motifs described here). Analysis of bacterial porins with known trimeric beta-barrel structure and related proteins reveals a similar repetitive motif corresponding to alternating membrane-spanning beta-strands. These beta-strands occur on the membrane interface (as opposed to the trimeric interface) of the beta-barrel. The broad conservation and structural location of these repeats suggests that they play important functional roles.

Algorithms↗

Molecular evolution of the genes encoding receptor tyrosine kinase with immunoglobulinlike domains.

Receptor tyrosine kinases (RTK) with five, three, or seven immunoglobulinlike domains in their extracellular regions are classified as subclasses III, IV, and V, respectively. Conservation of the exon/intron structure of the downstream part of the human KIT, FMS, and FLT3 genes that encode RTK of subclass III together with the particular chromosomal localization of these genes suggests that RTKIII genes have evolved from a common ancestor by cis and trans duplications. To strengthen this model of evolution and to determine if it can be extended to RTKIV and V genes, we constructed a phylogenetic tree of RTKIII, IV, and V on the basis of a multiple alignment of their catalytic tyrosine kinase domain sequences and determined the exon/intron structure of PDGFRA (subclass III), FGFR4 (subclass IV), and FLT4 (subclass V) genes in their downstream part. Phylogenetic analyses with amino acid or nucleotide sequences both resulted in one most parsimonious tree. The phylogenetic trees obtained indicate that all three subclasses are well individuated and that RTKIII and RTKV are closer to each other than RTKIV. Furthermore, RTKIII and FLT4 (subclass V) genes possess the same exon/intron structure in their downstream part while the structure of the RTKIV genes is very similar to that of RTKIII and FLT4. Both approaches are in complete agreement and indicate that RTKIII, IV, and V genes most probably evolved from a common ancestor already "in pieces" by successive duplications involving entire genes.

Alternative Splicing↗

Strategies for comparing gene expression profiles from different microarray platforms: application to a case-control experiment.

Meta-analysis of microarray data is increasingly important, considering both the availability of multiple platforms using disparate technologies and the accumulation in public repositories of data sets from different laboratories. We addressed the issue of comparing gene expression profiles from two microarray platforms by devising a standardized investigative strategy. We tested this procedure by studying MDA-MB-231 cells, which undergo apoptosis on treatment with resveratrol. Gene expression profiles were obtained using high-density, short-oligonucleotide, single-color microarray platforms: GeneChip (Affymetrix) and CodeLink (Amersham). Interplatform analyses were carried out on 8414 common transcripts represented on both platforms, as identified by LocusLink ID, representing 70.8% and 88.6% of annotated GeneChip and CodeLink features, respectively. We identified 105 differentially expressed genes (DEGs) on CodeLink and 42 DEGs on GeneChip. Among them, only 9 DEGs were commonly identified by both platforms. Multiple analyses (BLAST alignment of probes with target sequences, gene ontology, literature mining, and quantitative real-time PCR) permitted us to investigate the factors contributing to the generation of platform-dependent results in single-color microarray experiments. An effective approach to cross-platform comparison involves microarrays of similar technologies, samples prepared by identical methods, and a standardized battery of bioinformatic and statistical analyses.

Breast Neoplasms↗

Structural and functional differences between 3-repeat and 4-repeat tau isoforms. Implications for normal tau function and the onset of neurodegenetative disease.

Tau, MAP2, and MAP4 are members of a microtubule-associated protein (MAP) family that are each expressed as "3-repeat" and "4-repeat" isoforms. These isoforms arise from tightly controlled tissue-specific and/or developmentally regulated alternative splicing of a 31-amino acid long "inter-repeat:repeat module," raising the possibility that different MAP isoforms may possess some distinct functional capabilities. Consistent with this hypothesis, regulatory mutations in the human tau gene that disrupt the normal balance between 3-repeat and 4-repeat tau isoform expression lead to a collection of neurodegenerative diseases known as FTDP-17 (fronto-temporal dementias and Parkinsonism linked to chromosome 17), which are characterized by the formation of pathological tau filaments and neuronal cell death. Unfortunately, very little is known regarding structural and functional differences between the isoforms. In our previous analyses, we focused on 4-repeat tau structure and function. Here, we investigate 3-repeat tau, generating a series of truncations, amino acid substitutions, and internal deletions and examining the functional consequences. 3-Repeat tau possesses a "core microtubule binding domain" composed of its first two repeats and the intervening inter-repeat. This observation is in marked contrast to the widely held notion that tau possesses multiple independent tubulin-binding sites aligned in sequence along the length of the protein. In addition, we observed that the carboxyl-terminal sequences downstream of the repeat region make a strong but indirect contribution to microtubule binding activity in 3-repeat tau, which is in contrast to the negligible effect of these same sequences in 4-repeat tau. Taken together with previous work, these data suggest that 3-repeat and 4-repeat tau assume complex and distinct structures that are regulated differentially, which in turn suggests that they may possess isoform-specific functional capabilities. The relevance of isoform-specific structure and function to normal tau action and the onset of neurodegenerative disease are discussed.

Alternative Splicing↗

PairWise and SearchWise: finding the optimal alignment in a simultaneous comparison of a protein profile against all DNA translation frames.

DNA translation frames can be disrupted for several reasons, including: (i) errors in sequence determination; (ii) RNA processing, such as intron removal and guide RNA editing; (iii) less commonly, polymerase frameshifting during transcription or ribosomal frameshifting during translation. Frameshifts frequently confound computational activities involving homologous sequences, such as database searches and inferences on structure, function or phylogeny made from multiple alignments. A dynamic alignment algorithm is reported here which compares a protein profile (a residue scoring matrix for one or more aligned sequences) against the three translation frames of a DNA strand, allowing frameshifting. The algorithm has been incorporated into a new package, WiseTools, for comparison of biological sequences. A protein profile can be compared against either a DNA sequence or a protein sequence. The program PairWise may be used interactively for alignment of any two sequence inputs. SearchWise can perform combinations of searches through DNA or protein databases by a protein profile or DNA sequence. Routine application of the programs has revealed a set of database entries with frameshifts caused by errors in sequence determination.

Algorithms↗

Alignment and structure prediction of divergent protein families: periplasmic and outer membrane proteins of bacterial efflux pumps.

Broad-specificity efflux pumps have been implicated in multidrug-resistant strains of Pseudomonas aeruginosa and other Gram-negative bacteria. Most Gram-negative pumps of clinical relevance have three components, an inner membrane transporter, an outer membrane channel protein, and a periplasmic protein, which together coordinate efflux from the cytoplasmic membrane across the outer membrane through an unknown mechanism. The periplasmic efflux proteins (PEPs) and outer membrane efflux proteins (OEPs) are not obviously related to proteins of known structure, and understanding the structure and function of these proteins has been hindered by the difficulty of obtaining reasonable multiple alignments. We present a general strategy for the alignment and structure prediction of protein families with low mutual sequence similarity using the PEP and OEP families as detailed examples. Gibbs sampling, hidden Markov models, and other analysis techniques were used to locate motifs, generate multiple alignments, and assign PEP or OEP function to hypothetical proteins in several species. We also developed an automated procedure which combines multiple alignments with structure prediction algorithms in order to identify conserved structural features in protein families. This process was used to identify a probable alpha-helical hairpin in the PEP family and was applied to the detection of transmembrane beta-strands in OEPs. We also show that all OEPs contain a large tandem duplication, and demonstrate that the OEP family is unlikely to adopt a porin fold, in contrast to previous predictions.

Amino Acid Sequence↗

Phylogenetic internal control for HIV-1 genotypic antiretroviral testing.

Genotypic testing includes several steps (RNA purification, RT-PCR amplification, DNA sequencing, sequence editing and analysis) that should be individually controlled. In our laboratory, we have added to this step-by-step internal control a final phylogenetic quality control: this is performed every time a sequence is obtained from a patient previously subjected to the same test. Each sequence with this characteristic is routinely compared with sequences from previous samples of the same patient by multiple alignment and a neighbor-joining tree by using Kimura two-parameter method is constructed. To validate the quality control procedure, we have aligned and calculated the mean similarity of the reverse transcriptase (first 984 nucleotides) and protease (whole gene) sequences from 30 patients whose virus was completely wild-type for both reverse transcriptase and protease. In the same tree, we have added the sequences obtained from 5 out of the 30 patients, tested at a second time point. The wild type sequences have shown a mean inter-sample divergence of 2.9%, and all the sequence pairs from individual patients clustered together in the tree constructed with the nucleotide sequences, while the tree constructed with the inferred aminoacid sequences did not always permit to cluster the sequences from the same patients. This indicates that: 1) the phylogenetic analysis of nucleic acid sequences can be useful to rule out sample mix-up; 2) the belonging of a sequence to each individual patient can efficiently be assessed also in the cases of extreme divergence in terms of drug resistance mutations.

Amino Acid Sequence↗

A new distance measure for comparing sequence profiles based on path lengths along an entropy surface.

We describe a new distance measure for comparing DNA sequence profiles. For this measure, columns in a multiple alignment are treated as character frequency vectors (sum of the frequencies equal to one). The distance between two vectors is based on minimum path length along an entropy surface. Path length is estimated using a random graph generated on the entropy surface and Dijkstra's algorithm for all shortest paths to a source. We use the new distance measure to analyze similarities within familes of tandem repeats in the C. elegans genome and show that this new measure gives more accurate refinement of family relationships than a method based on comparing consensus sequences.

Algorithms↗

Scoredist: a simple and robust protein sequence distance estimator.

BACKGROUND: Distance-based methods are popular for reconstructing evolutionary trees thanks to their speed and generality. A number of methods exist for estimating distances from sequence alignments, which often involves some sort of correction for multiple substitutions. The problem is to accurately estimate the number of true substitutions given an observed alignment. So far, the most accurate protein distance estimators have looked for the optimal matrix in a series of transition probability matrices, e.g. the Dayhoff series. The evolutionary distance between two aligned sequences is here estimated as the evolutionary distance of the optimal matrix. The optimal matrix can be found either by an iterative search for the Maximum Likelihood matrix, or by integration to find the Expected Distance. As a consequence, these methods are more complex to implement and computationally heavier than correction-based methods. Another problem is that the result may vary substantially depending on the evolutionary model used for the matrices. An ideal distance estimator should produce consistent and accurate distances independent of the evolutionary model used. RESULTS: We propose a correction-based protein sequence estimator called Scoredist. It uses a logarithmic correction of observed divergence based on the alignment score according to the BLOSUM62 score matrix. We evaluated Scoredist and a number of optimal matrix methods using three evolutionary models for both training and testing Dayhoff, Jones-Taylor-Thornton, and Muller-Vingron, as well as Whelan and Goldman solely for testing. Test alignments with known distances between 0.01 and 2 substitutions per position (1-200 PAM) were simulated using ROSE. Scoredist proved as accurate as the optimal matrix methods, yet substantially more robust. When trained on one model but tested on another one, Scoredist was nearly always more accurate. The Jukes-Cantor and Kimura correction methods were also tested, but were substantially less accurate. CONCLUSION: The Scoredist distance estimator is fast to implement and run, and combines robustness with accuracy. Scoredist has been incorporated into the Belvu alignment viewer, which is available at ftp://ftp.cgb.ki.se/pub/prog/belvu/.

Algorithms↗

PoInTree: a polar and interactive phylogenetic tree.

PoInTree (Polar and Interactive Tree) is an application that allows to build, visualize and customize phylogenetic trees in a polar interactive and highly flexible view. It takes as input a FASTA file or multiple alignment formats. Phylogenetic tree calculation is based on a sequence distance method and utilizes the Neighbor Joining (NJ) algorithm. It also allows displaying precalculated trees of the major protein families based on Pfam classification. In PoInTree, nodes can be dynamically opened and closed and distances between genes are graphically represented. Tree root can be centered on a selected leaf. Text search mechanism, color-coding and labeling display are integrated. The visualizer can be connected to an Oracle database containing information on sequences and other biological data, helping to guide their interpretation within a given protein family across multiple species. The application is written in Borland Delphi and based on VCL Teechart Pro 6 graphical component (Steema software).

Algorithms↗

Constructing aligned sequence blocks.

This paper presents an efficient method for constructing aligned blocks (i.e., gap-free multiple alignments) from a set of pairwise alignments. The method is more sensitive than some earlier block-constructing methods for detecting conserved sequence regions. The technique is applied to analyze conserved regions in protein prenyltransferases and to detect regulatory elements in the 5' flank of the beta-globin gene.

Algorithms↗

FFAS03: a server for profile--profile sequence alignments.

The FFAS03 server provides a web interface to the third generation of the profile-profile alignment and fold-recognition algorithm of fold and function assignment system (FFAS) [L. Rychlewski, L. Jaroszewski, W. Li and A. Godzik (2000), Protein Sci., 9, 232-241]. Profile-profile algorithms use information present in sequences of homologous proteins to amplify the patterns defining the family. As a result, they enable detection of remote homologies beyond the reach of other methods. FFAS, initially developed in 2000, is consistently one of the best ranked fold prediction methods in the CAFASP and LiveBench competitions. It is also used by several fold-recognition consensus methods and meta-servers. The FFAS03 server accepts a user supplied protein sequence and automatically generates a profile, which is then compared with several sets of sequence profiles of proteins from PDB, COG, PFAM and SCOP. The profile databases used by the server are automatically updated with the latest structural and sequence information. The server provides access to the alignment analysis, multiple alignment, and comparative modeling tools. Access to the server is open for both academic and commercial researchers. The FFAS03 server is available at http://ffas.burnham.org.

Algorithms↗

Comparative analysis of sequences encoding ABC systems in the genome of the microsporidian Encephalitozoon cuniculi.

Microsporidia are amitochondriate eukaryotic microbes with fungal affinities and a common status of obligate intracellular parasites. A set of 13 potential genes encoding ATP-binding cassette (ABC) systems was identified in the fully sequenced genome of Encephalitozoon cuniculi. Our analyses of multiple alignments, phylogenetic trees and conserved motifs support a distribution of E. cuniculi ABC systems within only four subfamilies. Six half transporters are homologous to the yeast ATM1 mitochondrial protein, a finding which is in agreement with the hypothesis of a cryptic mitochondrion-derived compartment playing a role in the synthesis and transport of Fe-S clusters. Five half transporters are similar to the human ABCG1 and ABCG2 proteins, involved in regulation of lipid trafficking and anthracyclin resistance respectively. Two proteins with duplicated ABC domains are clearly candidate to non-transport ABC systems: the first is homologous to mammalian RNase L inhibitor and the second to the yeast translation initiation regulator GCN20. An unusual feature of ABC systems in E. cuniculi is the lack of homologs of P-glycoprotein and other ABC transporters which are involved in multiple drug resistance in a large number of eukaryotic microorganisms.

ATP-Binding Cassette Transporters↗

Use of a database of structural alignments and phylogenetic trees in investigating the relationship between sequence and structural variability among homologous proteins.

The database PALI (Phylogeny and ALIgnment of homologous protein structures) consists of families of protein domains of known three-dimensional (3D) structure. In a PALI family, every member has been structurally aligned with every other member (pairwise) and also simultaneous superposition (multiple) of all the members has been performed. The database also contains 3D structure-based and structure-dependent sequence similarity-based phylogenetic dendrograms for all the families. The PALI release used in the present analysis comprises 225 families derived largely from the HOMSTRAD and SCOP databases. The quality of the multiple rigid-body structural alignments in PALI was compared with that obtained from COMPARER, which encodes a procedure based on properties and relationships. The alignments from the two procedures agreed very well and variations are seen only in the low sequence similarity cases often in the loop regions. A validation of Direct Pairwise Alignment (DPA) between two proteins is provided by comparing it with Pairwise alignment extracted from Multiple Alignment of all the members in the family (PMA). In general, DPA and PMA are found to vary rarely. The ready availability of pairwise alignments allows the analysis of variations in structural distances as a function of sequence similarities and number of topologically equivalent Calpha atoms. The structural distance metric used in the analysis combines root mean square deviation (r.m.s.d.) and number of equivalences, and is shown to vary similarly to r.m.s.d. The correlation between sequence similarity and structural similarity is poor in pairs with low sequence similarities. A comparison of sequence and 3D structure-based phylogenies for all the families suggests that only a few families have a radical difference in the two kinds of dendrograms. The difference could occur when the sequence similarity among the homologues is low or when the structures are subjected to evolutionary pressure for the retention of function. The PALI database is expected to be useful in furthering our understanding of the relationship between sequences and structures of homologous proteins and their evolution.

Algorithms↗

Genomic organization of Trypanosoma brucei kinetoplast DNA minicircles.

The sequences of seven new Trypanosoma brucei kinetoplast DNA minicircles were obtained. A detailed comparative analysis of these sequences and those of the 18 complete kDNA minicircle sequences from T. brucei available in the database was performed. These 25 different minicircles contain 86 putative gRNA genes. The number of gRNA genes per minicircle varies from 2 to 5. In most cases, the genes are located between short imperfect inverted repeats, but in several minicircles there are inverted repeat cassettes that did not contain identifiable gRNA genes. Five minicircles contain single gRNA genes not surrounded by identifiable repeats. Two pairs of closely related minicircles may have recently evolved from common ancestors: KTMH1 and KTMH3 contained the same gRNA genes in the same order, whereas KTCSGRA and KTCSGRB contained two gRNA genes in the same order and one gRNA gene specific to each. All minicircles could be classified into two classes on the basis of a short substitution within the highly conserved region, but the minicircles in these two classes did not appear to differ in terms of gRNA content or gene organization. A number of redundant gRNAs containing identical editing information but different sequences were present. The alignments of the predicted gRNAs with the edited mRNA sequences varied from a perfect alignment without gaps to alignments with multiple mismatches. Multiple gRNAs overlapped with upstream gRNAs, but in no case was a complete set of overlapping gRNAs covering an entire editing domain obtained. We estimate that a minimum set of approximately 65 additional gRNAs would be required for complete overlapping sets. This analysis should provide a basis for detailed studies of the evolution and role in RNA editing of kDNA minicircles in this species.

Animals↗

Refine your search to explore more results.