Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,621 records · Page 90Linked to original sources

A sliding window-based method to detect selective constraints in protein-coding genes and its application to RNA viruses.

Here we present a new sliding window-based method specially designed to detect selective constraints in specific regions of a multiple protein-coding sequence alignment. In contrast to previous window-based procedures, our method is based on a nonarbitrary statistical approach to find the appropriate codon-window size to test deviations of synonymous (d(S)) and nonsynonymous (d(N)) nucleotide substitutions from the expectation. The probabilities of d(N) and d(S) are obtained from simulated data and used to detect significant deviations of d(N) and d(S) in a specific window region of the real sequence alignment. The nonsynonymous-to-synonymous rate ratio (w = d(N)/d(S)) was used to highlight selective constraints in any window wherein d(S) or d(N) was significantly different from the expectation. In these significant windows, w and its variance [V(w)] were calculated and used to test the neutral hypothesis. Computer simulations showed that the method is accurate even for highly divergent sequences. The main advantages of the new method are that it (i) uses a statistically appropriate window size to detect different selective patterns, (ii) is computationally less intensive than maximum likelihood methods, and (iii) detects saturation of synonymous sites, which can give deviations from neutrality. Hence, it allows the analysis of highly divergent sequences and the test of different alternative hypothesis as well. The application of the method to different human immunodeficiency virus type 1 and to foot-and-mouth disease virus genes confirms the action of positive selection on previously described regions as well as on new regions.

Base Sequence↗

Molecular mechanism of ferricsiderophore passage through the outer membrane receptor proteins of Escherichia coli.

Iron is an essential nutrient for all microorganisms with a few exceptions. Microorganisms use a variety of systems to acquire iron from the surrounding environment. One such system includes production of an organic molecule known as a siderophore by many bacteria and fungi. Siderophores have the capacity to specifically chelate ferric ions. The ferricsiderophore complex is then transported into the cell via a specific receptor protein located in the outer membrane. This is an energy dependent process and is the subject of investigation in many research laboratories. The crystal structures of three outer membrane ferricsiderophore receptor proteins FepA, FhuA and FecA from Escherichia coli and two FpvA and FptA from Pseudomonas aeruginosa have recently been solved. Four of them, FhuA, FecA, FpvA and FptA have been solved in ligand-bound forms, which gave insight into the residues involved in ligand binding. The structures are similar and show the presence of similar domains; for example, all of them consist of a 22 strand-beta-barrel formed by approximately 600 C-terminal residues while approximately 150 N-terminal residues fold inside the barrel to form a plug domain. The plug domain obstructs the passage through the barrel; therefore our research focuses on the mechanism through which the ferricsiderophore complex is transported across the receptor into the periplasm. There are two possibilities, one in which the plug domain is expelled into the periplasm making way for the ferricsiderophore complex and the second in which the plug domain undergoes structural rearrangement to form a channel through which the complex slides into the periplasm. Multiple alignment studies involving protein sequences of a large number of outer membrane receptor proteins that transport ferricsiderophores have identified several conserved residues. All of the conserved residues are located within the plug and barrel domain below the ligand binding site. We have substituted a number of these residues in FepA and FhuA with either alanine or glutamine resulting in substantial changes in the chemical properties of the residues. This was done to study the effect of the substitutions on the transport of ferricsiderophores. Another strategy used was to create a disulfide bond between the residues located on two adjacent beta-strands of the plug domain or between the residues of the plug domain and the beta-barrel in FhuA by substituting appropriate residues with cysteine. We have looked for the variants where the transport is affected without altering the binding. The data suggest a distinct role of these residues in the mechanism of transport. Our data also indicate that these transporters share a common mechanism of transport and that the plug remains within the barrel and possibly undergoes rearrangement to form a channel to transport the ferricsiderophore from the binding site to the periplasm.

Bacterial Outer Membrane Proteins↗

Crystal structure of an ADP-dependent glucokinase from Pyrococcus furiosus: implications for a sugar-induced conformational change in ADP-dependent kinase.

ADP-dependent kinases are used in the modified Embden-Meyerhoff pathway of certain archaea. Our previous study has revealed a mechanism for ADP-dependent phosphoryl transfer by Thermococcus litoralis glucokinase (tlGK), and its evolutionary relationship with ATP-dependent ribokinases and adenosine kinases (PFKB carbohydrate kinase family members). Here, we report the crystal structure of glucokinase from Pyrococcus furiosus (pfGK) in a closed conformation complexed with glucose and AMP at 1.9A resolution. In comparison with the tlGK structure, the pfGK structure shows significant conformational changes in the small domain and a region around the hinge, suggesting glucose-induced domain closing. A part of the large domain next to the hinge is also shifted accompanied with domain closing. In the pfGK structure, glucose binds in a groove between the large and small domains, and the electron density of O1 atoms for both the alpha and beta-anomer configurations was observed. The structural details of the sugar-binding site of ADP-dependent glucokinase were firstly clarified and then site-directed mutagenesis analysis clarified the catalytic residues for ADP-dependent kinase, such as Arg205 and Asp451 of tlGK. Homology search and multiple alignment of amino acid sequences using the information obtained from the structures reveals that eucaryotic hypothetical proteins homologous to ADP-dependent kinases retain the residues for the recognition of a glucose substrate.

Adenosine Diphosphate↗

DNA sequence motifs are associated with aberrant homologous recombination in the mouse macrophage migration inhibitory factor (Mif) locus.

Homologous recombination is a precise genetic event that can introduce specific alteration in the genome. A planned targeted disruption by homologous recombination of the macrophage migration inhibitory factor (Mif) locus in mouse embryonic stem (ES) cells yielded the targeted clones, some of which had genomic rearrangements inconsistent with the expected homologous recombination event. A detailed characterization of the recombination breakpoints in two of these clones revealed several sequence motifs with possible roles in recombination. These motifs included short regions of sequence identity that may promote DNA alignment, multiple 5'-AAGG/TTCC-3' tetrameres, topoisomerase I consensus sites, and AT-rich sequences that can promote DNA cleavage and recombination. A retrovirus-like intracisternal-A particle (IAP) family sequence was also identified upstream of the Mif gene, and the LTR of this IAP was involved in one of the recombinations. Identification and characterization of such sequence motifs will be valuable for the gene targeting experiments.

Adenine Nucleotides↗

Bioinformatics and molecular modeling in chemical enzymology. Active sites of hydrolases.

Comparison and multiple alignments of amino acid sequences of a representative number of related enzymes demonstrate the existence of certain positions of amino acid residues which are permanently reproducible in all members of the whole family. The use of the bioinformatic approach revealed conservative residues in each of the related enzymes and ranked amino acid conservatism for the overall enzymatic catalysis. Glycine and aspartic acid residues were shown to be the most essential for structure and catalytic activity of enzymes. Amino acid residues forming catalytic subsite of the active site of enzymes are always highly conservative. Analysis revealed that aspartic acid carboxyl group is the most frequently employed nucleophilic (in deprotonated form) and electrophilic (in protonated form) agent involved in activation of molecules by the mechanism of general base and acidic catalyses in the catalytic sites of enzymes. Glycine is a unique amino acid possessing the highest possibilities for rotation along C-C and C-N bonds of the polypeptide chain. The conservative fixation of the glycine residue in polypeptide chains of related enzymes provides a possibility for directed assembly of amino acid residues into the catalytic subsite structure. It is possible that the conservative glycines provide known conformational mobility of the protein and the active site. Methods of molecular modeling were used for analysis of structural substitutions of conservative and non-conservative glycines and their effects on geometry of catalytic site of typical hydrolases. The substitution of glycine(s) for alanine significantly altered the catalytic site structures.

Binding Sites↗

Mutagenesis of the fructose-6-phosphate-binding site in the 2-kinase domain of 6-phosphofructo-2-kinase/fructose-2,6-bisphosphatase.

Multiple alignment of several isozyme sequences of the bifunctional enzyme 6-phosphofructo-2-kinase/fructose-2,6-bisphosphatase revealed conserved residues in the 2-kinase domain. Among these residues, three asparagine residues (Asn76, Asn97 and Asn133; numbering refers to the liver isozyme sequence) and three threonine residues (Thr132, Thr134 and Thr135) are located near the fructose 6-phosphate-binding site in the crystal structure of the bifunctional enzyme. The role of these residues in substrate binding and catalysis in the 6-phosphofructo-2-kinase domain has been studied by mutagenesis to alanine. Since the crystal structure of 6-phosphofructo-2-kinase does not contain fructose 6-phosphate, this substrate was docked into the putative binding site by computer modelling, and its interactions with the protein were predicted. Analysis of the mutagenesis-induced changes in kinetic properties and of the substrate-docking model revealed that all these residues are directly or indirectly involved in fructose-6-phosphate binding. All the mutants displayed an increased Km for fructose 6-phosphate (10-200-fold). We propose that Asn133 stabilises Arg138, which itself makes a direct electrostatic bond with the 6-phosphate group of fructose 6-phosphate, that Asn76 interacts with the C3 hydroxyl group of fructose 6-phosphate, that Thr132 makes a hydrogen bond with the C6 oxygen of this substrate, and that Thr134 interacts with two residues involved in fructose-6-phosphate binding, Thr132 and Tyr199. On the other hand, Asn97 and Thr135 play structural roles, by maintaining the structure of the fructose-6-phosphate-binding pocket.

Base Sequence↗

Characterization of recombinant Arabidopsis thaliana threonine synthase.

Threonine synthase (TS) catalyses the last step in the biosynthesis of threonine, the pyridoxal 5'-phosphate dependent conversion of L-homoserine phosphate (HSerP) into L-threonine and inorganic phosphate. Recombinant Arabidopsis thaliana TS (aTS) was characterized to compare a higher plant TS with its counterparts from Escherichia coli and yeast. This comparison revealed several unique properties of aTS: (a) aTS is a regulatory enzyme whose activity was increased up to 85-fold by S-adenosyl-L-methionine (SAM) and specifically inhibited by AMP; (b) HSerP analogues shown previously to be potent inhibitors of E. coli TS failed to inhibit aTS; and (c) aTS was a dimer, while the E. coli and yeast enzymes are monomers. The N-terminal region of aTS is essential for its regulatory properties and protects against inhibition by HSerP analogues, as an aTS devoid of 77 N-terminal residues was neither activated by SAM nor inhibited by AMP, but was inhibited by HSerP analogues. The C-terminal region of aTS seems to be involved in dimer formation, as the N-terminally truncated aTS was also found to be a dimer. These conclusions are supported by a multiple amino-acid sequence alignment, which revealed the existence of two TS subfamilies. aTS was classified as a member of subfamily 1 and its N-terminus is at least 35 residues longer than those of any nonplant TS. Monomeric E. coli and yeast TS are members of subfamily 2, characterized by C-termini extending about 50 residues over those of subfamily 1 members. As a first step towards a better understanding of the properties of aTS, the enzyme was crystallized by the sitting drop vapour diffusion method. The crystals diffracted to beyond 0.28 nm resolution and belonged to the space group P222 (unit cell parameters: a = 6.16 nm, b = 10.54 nm, c = 14.63 nm, alpha = beta = gamma = 90 degrees).

Amino Acid Sequence↗

Development of humanized monoclonal antibody TMA-15 which neutralizes Shiga toxin 2.

A murine monoclonal antibody (MAb), VTm1.1, specifically recognizing and neutralizing Shiga toxin 2 (Stx2), was obtained. To prevent a humoral response against murine antibody when used clinically, a humanized antibody was constructed by combining the complementarity-determining regions of VTm1.1 with human framework and constant regions. In addition, several amino acids in the framework were changed to improve the binding affinity of the antibody and further reduce its potential immunogenicity. The humanized antibody, TMA-15, recognized the B-subunit of Stx2 and had affinity for Stx2 of 3.3 x 10(-9) M, within two-fold of that of the original murine antibody. TMA-15 neutralized the cytotoxicity of Stx2 and several different Stx2 variants in vitro, and it completely protected mice from death in a Stx2-challenged mice model. These results suggest that TMA-15 will have clinical potency in Stx-producing Escherichia coli infections, including E. coli O157 infections.

Amino Acid Sequence↗

3Dee: a database of protein structural domains.

UNLABELLED: The 3Dee database is a repository of protein structural domains. It stores alternative domain definitions for the same protein, organises domains into sequence and structural hierarchies, contains non-redundant set(s) of sequences and structures, multiple structure alignments for families of domains, and allows previous versions of the database to be regenerated. AVAILABILITY: 3Dee is accessible on the World Wide Web at the URL http://barton.ebi.ac.uk/servers/3Dee.html.

Databases, Factual↗

DIVERGE: phylogeny-based analysis for functional-structural divergence of a protein family.

SUMMARY: DetectIng Variability in Evolutionary Rates among GEnes (DIVERGE) is a software system to study functional divergence of a protein family by detecting site-specific change in evolutionary rate using a multiple alignment of amino acid sequences for a given phylogenetic tree. The program first conducts a statistical test for site-specific rate shifts along the tree, and predicting candidate amino acid residues responsible for functional divergence based on posterior analysis. These results can then be mapped on the 3D protein structure if available. AVAILABILITY: DIVERGE is available free of charge from http://xgu1.zool.iastate.edu/. Distribution packages for both Linux and Microsoft Windows operating systems are available, including manual and example files.

Amino Acid Sequence↗

aPhyloGeo: a Python application for correlating genetic and climatic conditions.

MOTIVATION: Environmental variation and its influence on genetic diversity is a central topic in evolutionary biology and phylogeography. Accurate correlations between genetic and climatic datasets to understand the genetic adaptations of different species to specific environments. It requires integrated and reproducible workflows. RESULTS: We developed aPhyloGeo, an open-source and multiplatform application implemented in Python, for investigating correlations between genetic variation and environmental data within a phylogenetic framework. The workflow integrates multiple analytical steps, including sequence alignment, sliding window phylogenetic inference, and statistical approaches such as the Mantel test and the Procrustean randomization test. These analyses enable the identification of mutation hotspots that exhibit strong associations with environmental variables. In addition, aPhyloGeo supports multicore data processing and provides a fully reproducible pipeline for evaluating localized relationships between genomic variation and climatic distributions. AVAILABILITY AND IMPLEMENTATION: aPhyloGeo is freely available on GitHub at: https://github.com/tahiri-lab/aPhyloGeo, as both a PyPI package and as Python scripts for Linux, macOS, and Windows.

Software↗

Distribution and molecular characterization of integron classes from Escherichia coli and Klebsiella pneumoniae isolates in Sulaymaniyah province of Iraq.

UNLABELLED: The environmental pollution from the misuse of antimicrobial drugs is fueling selection pressure in bacteria, thereby exacerbating the threat to global health. In Iraq, the situation is made worse by the poor implementation of the World Health Organization's Global Antimicrobial Resistance and Use Surveillance System (WHO-GLASS). Consequently, this study aimed to increase surveillance of the spread of antimicrobial resistance in Sulaymaniyah, Iraq. A total of 296 Enterobacteriaceae comprising 147 Klebsiella pneumoniae and 149 Escherichia coli were isolated from humans, poultry, and dairy farms. The isolates were screened using multiplex PCR to assess the prevalence of the clinically important integron integrase (intI) classes and antimicrobial resistance genes (ARGs) of commonly used antibiotics. Remarkably, 81.14% of the isolates carried at least 2 ARGs, 10.47% intI1, and 3.72% intI2. No intI3 was detected. A total of 663 ARGs were identified using multiplex PCR in the two Enterobacteriaceae: beta-lactamase genes were 43%, tetracycline resistance genes 25.20%, sulfonamide resistance gene 16.10%, quinolone resistance gene 10.2%, and aminoglycoside resistance genes 5.7%. K. pneumoniae harbored more integrons and ARGs than E. coli, thus posing a higher antimicrobial resistance threat in this province. This study underscores the importance of implementing more stringent WHO-GLASS and antibiotic stewardship to end the multidrug resistance crisis in Iraq. IMPORTANCE: These data are about the prevalence of integrons and resistance genes, helping to fill a significant gap in global surveillance efforts. Results can be used by global health authorities and the World Health Organization to develop national and international antimicrobial resistance (AMR) control strategies. The study is important because integrons are key genetic platforms that capture and disseminate antibiotic resistance genes among bacteria. In addition, Escherichia coli and Klebsiella spp. are among the top causes of hospital- and community-acquired infections, especially urinary tract infections, bloodstream infections, and pneumonia. Therefore, it will be riskier when these bacteria have a high rate of integrons and resistance genes because it impedes treatments during infection. Another importance of this study is that the study was carried out in Iraq. Iraq, like many low- and middle-income countries, faces challenges with unregulated antibiotic use, leading to high rates of AMR.

Escherichia coli↗

On the inference of parsimonious indel evolutionary scenarios.

Given a multiple alignment of orthologous DNA sequences and a phylogenetic tree for these sequences, we investigate the problem of reconstructing a most parsimonious scenario of insertions and deletions capable of explaining the gaps observed in the alignment. This problem, called the Indel Parsimony Problem, is a crucial component of the problem of ancestral genome reconstruction, and its solution provides valuable information to many genome functional annotation approaches. We first show that the problem is NP-complete. Second, we provide an algorithm, based on the fractional relaxation of an integer linear programming formulation. The algorithm is fast in practice, and the solutions it produces are, in most cases, provably optimal. We describe a divide-and-conquer approach that makes it possible to solve very large instances on a simple desktop machine, while retaining guaranteed optimality. Our algorithms are tested and shown efficient and accurate on a set of 1.8 Mb mammalian orthologous sequences in the CFTR region.

Algorithms↗

Subsets with restricted immunoglobulin gene rearrangement features indicate a role for antigen selection in the development of chronic lymphocytic leukemia.

We recently identified a chronic lymphocytic leukemia (CLL) subgroup using the immunoglobulin variable heavy-chain (V(H)) gene V(H)3-21 with almost identical heavy-chain complementarity determining region 3s (HCDR3s) and preferential variable light-chain (V(L)) gene usage, suggesting recognition of a common antigen epitope in this subset. To further explore the B-cell receptors (BCRs) in CLL, we characterized 407 V(H) rearrangements amplified from 346 CLLs regarding V(H), diversity (D), and joining (J(H)) gene usage and performed multiple alignment of the HCDR3 sequences. These analyses revealed 3 small subsets (2 V(H)1-69 groups, 7 cases; and 1 V(H)1-2 group, 5 cases) with highly restricted HCDR3 features including identical V(H)/D/J(H) usage, HCDR3 lengths, and shared N-sequences, in addition to the V(H)3-21 group (22 cases). Furthermore, another 3 groups (9 V(H)1-3(+) cases, 3 V(H)1-18(+) cases, and 5 V(H)4-39(+) cases) had essentially identical V(H)/D/J(H) use and similar HCDR3 lengths but less conserved N-regions. Analysis in all 6 of these subgroups showed restriction in V(L) gene use, whereas no association between V(H) and V(L) usage was found in cases without HCDR3 similarities. Altogether, structurally similar HCDR3s associated with preferential V(L) gene usage implies selection of BCRs, especially in subsets showing high HCDR3 similarities, thus pointing to restricted antigen recognition sites and possibly involvement of specific antigens in CLL development.

Adult↗

A web-based program for the prediction of average hydropathy, average amphipathicity and average similarity of multiply aligned homologous proteins.

We designed a web-based program, AveHAS, to determine and plot the average hydropathy, average amphipathicity and average similarity for a clustal X-derived multiple alignment of homologous protein sequences. This method is based on the TREEMOMENT and Hydro programs. It has a user-friendly interface, a convenient input format and an improved algorithm.

Algorithms↗

[Development of "Amplisens-HCV-genotype" reagent set for identification of hepatitis C virus genotypes 1a, 1b, 2a and 3a].

Multiple alignments of 119 nucleotide sequences of isolates of hepatitis C virus (HCV) were carried out to choose the type-specific primers for the 5'-ultra-core fragment of viral genome for the purpose of detecting the HCV 1a, 1b, 2a, and 3a subtypes. A PCR kit of reagents was designed for the amplification of cDNA HCV with selected type-specific primers and for making the electrophoresis in agarous gel. The kit comprises the positive control samples, i.e. HCV genome fragments, subtypes 1a, 1b, 2a and 3a, cloned in the plasmid vector. 440 cDNAHCV samples were simultaneously tested by using the worked out reagents' set and according to the method of Ohno et al. The results were found to be concordant in 336 cases, and were discordant in 4 samples. A sequencing of the PCR products and phylogenetic analysis showed that 1 sample belonged to subtype 4a, 2 samples belonged to subtypes 2k and 1 sample--to subtype 31.

DNA Primers↗

CBCAnalyzer: inferring phylogenies based on compensatory base changes in RNA secondary structures.

The CBCAnalyzer (CBC=compensatory base change) is a custom written software toolbox consisting of three parts, CTTransform, CBCDetect, and CBCTree. CTTransform reads several ct-file formats, and generates a so called "bracket-dot-bracket" format that typically is used as input for other tools such as RNAforester, RNAmovie or MARNA. The latter one creates a multiple alignment based on primary sequences and secondary structures that now can be used as input for CBCDetect. CBCDetect counts CBCs in all against all of the aligned sequences. This is important in detecting species that are discriminated by their sexual incompatibility. The count (distance) matrix obtained by CBCDetect is used as input for CBCTree that reconstructs a phylogram by using the algorithm of BIONJ. In this note we describe the features of the toolbox as well as application examples. The toolbox provides a graphical user interface. It is written in C++ and freely available at: http://cbcanalyzer.bioapps.biozentrum.uni-wuerzburg.de.

Algorithms↗