Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,603 records · Page 89Linked to original sources

Predicting DNA-binding sites of proteins from amino acid sequence.

BACKGROUND: Understanding the molecular details of protein-DNA interactions is critical for deciphering the mechanisms of gene regulation. We present a machine learning approach for the identification of amino acid residues involved in protein-DNA interactions. RESULTS: We start with a Naïve Bayes classifier trained to predict whether a given amino acid residue is a DNA-binding residue based on its identity and the identities of its sequence neighbors. The input to the classifier consists of the identities of the target residue and 4 sequence neighbors on each side of the target residue. The classifier is trained and evaluated (using leave-one-out cross-validation) on a non-redundant set of 171 proteins. Our results indicate the feasibility of identifying interface residues based on local sequence information. The classifier achieves 71% overall accuracy with a correlation coefficient of 0.24, 35% specificity and 53% sensitivity in identifying interface residues as evaluated by leave-one-out cross-validation. We show that the performance of the classifier is improved by using sequence entropy of the target residue (the entropy of the corresponding column in multiple alignment obtained by aligning the target sequence with its sequence homologs) as additional input. The classifier achieves 78% overall accuracy with a correlation coefficient of 0.28, 44% specificity and 41% sensitivity in identifying interface residues. Examination of the predictions in the context of 3-dimensional structures of proteins demonstrates the effectiveness of this method in identifying DNA-binding sites from sequence information. In 33% (56 out of 171) of the proteins, the classifier identifies the interaction sites by correctly recognizing at least half of the interface residues. In 87% (149 out of 171) of the proteins, the classifier correctly identifies at least 20% of the interface residues. This suggests the possibility of using such classifiers to identify potential DNA-binding motifs and to gain potentially useful insights into sequence correlates of protein-DNA interactions. CONCLUSION: Naïve Bayes classifiers trained to identify DNA-binding residues using sequence information offer a computationally efficient approach to identifying putative DNA-binding sites in DNA-binding proteins and recognizing potential DNA-binding motifs.

Algorithms↗

A probabilistic model for the evolution of RNA structure.

BACKGROUND: For the purposes of finding and aligning noncoding RNA gene- and cis-regulatory elements in multiple-genome datasets, it is useful to be able to derive multi-sequence stochastic grammars (and hence multiple alignment algorithms) systematically, starting from hypotheses about the various kinds of random mutation event and their rates. RESULTS: Here, we consider a highly simplified evolutionary model for RNA, called "The TKF91 Structure Tree" (following Thorne, Kishino and Felsenstein's 1991 model of sequence evolution with indels), which we have implemented for pairwise alignment as proof of principle for such an approach. The model, its strengths and its weaknesses are discussed with reference to four examples of functional ncRNA sequences: a riboswitch (guanine), a zipcode (nanos), a splicing factor (U4) and a ribozyme (RNase P). As shown by our visualisations of posterior probability matrices, the selected examples illustrate three different signatures of natural selection that are highly characteristic of ncRNA: (i) co-ordinated basepair substitutions, (ii) co-ordinated basepair indels and (iii) whole-stem indels. CONCLUSIONS: Although all three types of mutation "event" are built into our model, events of type (i) and (ii) are found to be better modeled than events of type (iii). Nevertheless, we hypothesise from the model's performance on pairwise alignments that it would form an adequate basis for a prototype multiple alignment and genefinding tool.

Evolution, Molecular↗

Acute hemorrhagic conjunctivitis epidemic caused by coxsackievirus A24 variants in Korea during 2002-2003.

A variant of coxsackievirus A24 (CA24v) is one of the agents causing acute hemorrhagic conjunctivitis. There was an epidemic of acute hemorrhagic conjunctivitis caused by CA24v in Korea from 2002 to 2003. Seventy-one strains of CA24v were isolated from 159 conjunctival specimens (45%). Most of the patients were school children under the age of 20. The epidemic began in the first week of August in 2002, and spread extensively, with a peak in the third week of September. CA24v strains were also isolated from conjunctival specimens in 2003. Reverse transcription polymerase chain reaction (RT-PCR) was performed and sequencing of the 340 bp fragment of the VP1 region of the viruses. Sequencing data were multiple-aligned using CLUSTAL W (version 1.81). Phylogenetic trees were plotted using TreeView (version 1.6.6). Homologies ranged from 97.7%-100%, depending on geographical regions: from 99.4%-100% in 2002 and 98.4%-100% in 2003. A phylogenetic tree based on the nucleotide sequence homologies formed clusters depending on years rather than on geographical regions. Identities (98%-100%) were found among the Korean CA24v strains, and there was 85%-90% homology between these and the prototype strain.

Adolescent↗

DoOP: Databases of Orthologous Promoters, collections of clusters of orthologous upstream sequences from chordates and plants.

DoOP (http://doop.abc.hu/) is a database of eukaryotic promoter sequences (upstream regions) aiming to facilitate the recognition of regulatory sites conserved between species. The annotated first exons of human and Arabidopsis thaliana genes were used as queries in BLAST searches to collect the most closely related orthologous first exon sequences from Chordata and Viridiplantae species. Up to 3000 bp DNA segments upstream from these first exons constitute the clusters in the chordate and plant sections of the Database of Orthologous Promoters. Release 1.0 of DoOP contains 21,061 chordate clusters from 284 different species and 7548 plant clusters from 269 different species. The database can be used to find and retrieve promoter sequences of a given gene from various species and it is also suitable to see the most trivial conserved sequence blocks in the orthologous upstream regions. Users can search DoOP with either sequence or text (annotation) to find promoter clusters of various genes. In addition to the sequence data, the positions of the conserved sequence blocks derived from multiple alignments, the positions of repetitive elements and the positions of transcription start sites known from the Eukaryotic Promoter Database (EPD) can be viewed graphically.

Animals↗

On the role of structural information in remote homology detection and sequence alignment: new methods using hybrid sequence profiles.

Structural alignments often reveal relationships between proteins that cannot be detected using sequence alignment alone. However, profile search methods based entirely on structural alignments alone have not been found to be effective in finding remote homologs. Here, we explore the role of structural information in remote homolog detection and sequence alignment. To this end, we develop a series of hybrid multidimensional alignment profiles that combine sequence, secondary and tertiary structure information into hybrid profiles. Sequence-based profiles are profiles whose position-specific scoring matrix is derived from sequence alignment alone; structure-based profiles are those derived from multiple structure alignments. We compare pure sequence-based profiles to pure structure-based profiles, as well as to hybrid profiles that use combined sequence-and-structure-based profiles, where sequence-based profiles are used in loop/motif regions and structural information is used in core structural regions. All of the hybrid methods offer significant improvement over simple profile-to-profile alignment. We demonstrate that both sequence-based and structure-based profiles contribute to remote homology detection and alignment accuracy, and that each contains some unique information. We discuss the implications of these results for further improvements in amino acid sequence and structural analysis.

Amino Acid Sequence↗

PHOG: a database of supergenomes built from proteome complements.

BACKGROUND: Orthologs and paralogs are widely used terms in modern comparative genomics. Existing procedures for resolving orthologous/paralogous relationships are often based on manual revision of clusters of orthologous groups and/or lack any rigorous evolutionary base. DESCRIPTION: We developed a completely automated procedure that creates clusters of orthologous groups at each node of the taxonomy tree (PHOGs--Phylogenetic Orthologous Groups). As a result of this procedure, a tree of orthologous groups was obtained. Each cluster is a "supergene" and it is represented by an "ancestral" sequence obtained from the multiple alignment of orthologous and paralogous genes. The procedure has been applied to the taxonomy tree of organisms from all three domains of life. Protein complements from 50 bacterial, archaeal and eukaryotic species were used to create PHOGs at all tree nodes. 51367 PHOGs were obtained at the root node. CONCLUSION: The PHOG database demonstrates that it is possible to automatically process any number of sequenced genomes and to reconstruct orthologous and paralogous relationships between genomes using a rigorous evolutionary approach. This database can become a very useful tool in various areas of comparative genomics.

Databases, Genetic↗

Diversity and evolution of blaZ from Staphylococcus aureus and coagulase-negative staphylococci.

OBJECTIVES: To elucidate the diversity and evolutionary history of plasmid- and chromosomally-located blaZ, to detect indications of frequent exchange of blaZ between human and bovine staphylococci and to estimate the frequency of transfer of blaZ between coagulase-negative staphylococci (CoNS) and Staphylococcus aureus of bovine origin. METHODS: blaZ was detected in 143 strains of penicillin-resistant S. aureus and CoNS from five Danish cattle herds (n = 25/23), random CoNS isolates from Denmark (n = 37), a collection of S. aureus from six different countries (n = 52), humans in Denmark (n = 3) and beta-lactamase control strains (n = 3). The sequence was determined in 105 strains and compared to published sequences by pairwise and multiple alignments. Maximum likelihood analysis was performed including bootstrap analysis. Parsimony, neighbour joining and consensus comparisons were performed for recombination. The localization of blaZ was determined by Southern blotting in 108 isolates. RESULTS: All penicillin-resistant strains carried blaZ and showed a similar organization of blaR1 and blaZ. The blaZ gene was localized to a plasmid in only 16 of the resistant strains. Sixty-nine sequences representing 105 isolates and sequences retrieved from public databases were compared. A phylogenetic tree showed that blaZ exists in three evolutionary lines: one group was of plasmid origin, one group was of chromosomal origin and one intermediate group. Sixty-nine sequence types were demonstrated. They translated into 11 BlaZ protein types. The major types all contained strains of both human and bovine origin, and more than one Staphylococcus species, demonstrating a shared gene pool. In a comparison of S. aureus and CoNS obtained from five Danish cattle herds, the same type of blaZ was only detected in one case. CONCLUSIONS: Results indicated a separate evolution for plasmid- and chromosomally-encoded blaZ. Although a common gene pool seems to exist among staphylococci, exchange of blaZ between strains and species is judged to be an extremely rare event.

Alleles↗

A-liner: linear alignment visualizer for genome comparisons.

SUMMARY: A-liner is a flexible command-line tool for linear visualization of genome-scale sequence alignments, supporting outputs from multiple aligners and integrated visualization of annotations, highlights, quantitative tracks, and coordinate scales. It is applicable to a wide range of organisms, from bacteria to large eukaryotic genomes, and facilitates efficient generation of publication-ready comparative genome visualizations. AVAILABILITY AND IMPLEMENTATION: The source code and example output files for a-liner are available in the GitHub repository: https://github.com/mokuno3430/a-liner. A-liner v1.1.0 has been archived on Zenodo at https://doi.org/10.5281/zenodo.19702001.

Software↗

ANTHEPROT 2.0: a three-dimensional module fully coupled with protein sequence analysis methods.

ANTHEPROT is a fully interactive graphics program devoted to the analysis of the sequences and structures of proteins. This program, originally developed to facilitate the protein sequence analysis coupled with multiple alignments and predicted secondary structures of proteins, now comprises a powerful 3D module to display and handle macromolecular structures. All the methods that were previously integrated into ANTHEPROT are now directly coupled with a 3D window that provides the user all the classic features of a molecular modeling package. Indeed, it allows real-time rotation and translation of 3D structures with many kinds of models in depth-cueing mode (space filling, backbone, wire models, main chain, and ribbons), selections (atom type, residue type, segments, and chain), color-coding systems (amino acid properties, predicted or observed secondary structures, temperature B factor, and subunits), geometric calculations (Ramachandran plot, distances, and angles), and fitting molecules. Stereo views are possible as well as HPGL standard files. A module specifically devoted to the determination of 3D structures using nuclear magnetic resonance is also available. This major release of our program for IBM rs6000 workstations is available by anonymous ftp to ibcp.fr for academic institutions.

Antigens↗

Pathogenic potentials of glycoprotein C-negative syncytial mutants from rabbit T cells infected persistently with herpes simplex virus type 1.

Human T cell lymphotropic virus type I (HTLV-I)-transformed T cells of rabbits were infected persistently with Herpes simplex virus type 1 (HSV-1) strain KOS. These infected cells yielded syncytial mutants, either glycoprotein C (gC)-negative or -positive, which predominated over and replaced the wild-type virus in a long-term culture for 2 years. An alignment of nucleotide sequences showed multiple mutations in glycoprotein B (gB) and gC genes of these mutants, which are or may be responsible for the mutant phenotypes. One of four mutants analyzed produced extensively large syncytia and possessed point mutations within the cytoplasmic domain of gB. All four mutants possessed multiple point mutations in gC and two possessed single insertions which resulted in a frame shift, leading to the premature termination of the gC polypeptide chain. The supernatant of the 2-year culture of cells infected persistently, containing only gC-negative syncytial mutants, induced encephalitic symptoms in B/Jas inbred rabbits, when injected intravenously. One gC-negative syncytial isolate from an encephalitic lesion, together with those from the culture supernatant, were examined for pathogenic potential in vitro and in vivo. All these mutants were more cytotoxic and more susceptible to complement inactivation than the parental virus, and could infect and replicate in adrenal glands when injected intravenously into rabbits. Invasion into the central nervous system appeared to be blocked at the portal of entry, the adrenal gland, i.e., none exhibited neuroinvasive potential by itself. Syncytial gC-negative mutants could thus be pathogenic in rabbits.

Adrenal Glands↗

Multiple mapping method: a novel approach to the sequence-to-structure alignment problem in comparative protein structure modeling.

A major bottleneck in comparative protein structure modeling is the quality of input alignment between the target sequence and the template structure. A number of alignment methods are available, but none of these techniques produce consistently good solutions for all cases. Alignments produced by alternative methods may be superior in certain segments but inferior in others when compared to each other; therefore, an accurate solution often requires an optimal combination of them. To address this problem, we have developed a new approach, Multiple Mapping Method (MMM). The algorithm first identifies the alternatively aligned regions from a set of input alignments. These alternatively aligned segments are scored using a composite scoring function, which determines their fitness within the structural environment of the template. The best scoring regions from a set of alternative segments are combined with the core part of the alignments to produce the final MMM alignment. The algorithm was tested on a dataset of 1400 protein pairs using 11 combinations of two to four alignment methods. In all cases MMM showed statistically significant improvement by reducing alignment errors in the range of 3 to 17%. MMM also compared favorably over two alignment meta-servers. The algorithm is computationally efficient; therefore, it is a suitable tool for genome scale modeling studies.

Algorithms↗

Role of phosphorylated Thr160 for the activation of the CDK2/Cyclin A complex.

The enzymatic activity of the CDK2/Cyclin A complex increases upon the specific phosphorylation of Thr160@CDK2. In the present study, we have performed a comparative molecular dynamics (MD) study of models of the complex CDK2/Cyclin A/Substrate, which differ for the presence or absence of the phosphate group bound to Thr160. The models are based on two X-ray structures available for CDK2/CyclinA and pCDK2/CyclinA/Substrate complexes. In this way, we analyze the influence of the phosphorylated Thr160 (pThr160) on both the flexibility of CDK2 activation loop (AL) and substrate binding in CDK2. Our calculations point to a decreased flexibility of the AL in the phosphorylated model, in fairly good agreement with experimental data, and to a key role of pThr160 for substrate recognition and stability. Multiple alignments of the CDKs sequences point to the very high conservation of the AL sequence among the CDKs, thus extending our results to all CDKs.

Binding Sites↗

Cloning and characterization of the murine genes for bHLH-ZIP transcription factors TFEC and TFEB reveal a common gene organization for all MiT subfamily members.

The microphthalmia-TFE (MiT) subfamily of basic helix-loop-helix leucine zipper (bHLH-ZIP) transcription factors, including TFE3, TFEB, TFEC, and Mitf, has been implicated in the regulation of tissue-specific gene expression in several cell lineages. In this report, we investigate the genomic organization and structural relatedness of MiT transcription factors. We characterized the gene for mTFEC, which covers a region of more than 50 kb and is composed of seven exons. Further, we cloned a cDNA for the murine TFEB homologue and characterized its genomic structure. The eight coding exons of mTFEB are distributed over a 6-kb region. A multiple alignment of amino acid sequences of known MiT subfamily members indicates undescribed, conserved N-terminal regions and common putative phosphorylation sites for TFE3, TFEB, and Mitf. Also, intron-exon borders for characterized MiT genes appear completely conserved. A new family member and closely related putative transcription factor in Caenorhabditis elegans was identified by database searches that show a similar genomic organization within the bHLH-ZIP region and the acidic domain. Evolutionary aspects and implications for structure-function relationships are discussed.

Amino Acid Sequence↗

A sliding window-based method to detect selective constraints in protein-coding genes and its application to RNA viruses.

Here we present a new sliding window-based method specially designed to detect selective constraints in specific regions of a multiple protein-coding sequence alignment. In contrast to previous window-based procedures, our method is based on a nonarbitrary statistical approach to find the appropriate codon-window size to test deviations of synonymous (d(S)) and nonsynonymous (d(N)) nucleotide substitutions from the expectation. The probabilities of d(N) and d(S) are obtained from simulated data and used to detect significant deviations of d(N) and d(S) in a specific window region of the real sequence alignment. The nonsynonymous-to-synonymous rate ratio (w = d(N)/d(S)) was used to highlight selective constraints in any window wherein d(S) or d(N) was significantly different from the expectation. In these significant windows, w and its variance [V(w)] were calculated and used to test the neutral hypothesis. Computer simulations showed that the method is accurate even for highly divergent sequences. The main advantages of the new method are that it (i) uses a statistically appropriate window size to detect different selective patterns, (ii) is computationally less intensive than maximum likelihood methods, and (iii) detects saturation of synonymous sites, which can give deviations from neutrality. Hence, it allows the analysis of highly divergent sequences and the test of different alternative hypothesis as well. The application of the method to different human immunodeficiency virus type 1 and to foot-and-mouth disease virus genes confirms the action of positive selection on previously described regions as well as on new regions.

Base Sequence↗

Molecular mechanism of ferricsiderophore passage through the outer membrane receptor proteins of Escherichia coli.

Iron is an essential nutrient for all microorganisms with a few exceptions. Microorganisms use a variety of systems to acquire iron from the surrounding environment. One such system includes production of an organic molecule known as a siderophore by many bacteria and fungi. Siderophores have the capacity to specifically chelate ferric ions. The ferricsiderophore complex is then transported into the cell via a specific receptor protein located in the outer membrane. This is an energy dependent process and is the subject of investigation in many research laboratories. The crystal structures of three outer membrane ferricsiderophore receptor proteins FepA, FhuA and FecA from Escherichia coli and two FpvA and FptA from Pseudomonas aeruginosa have recently been solved. Four of them, FhuA, FecA, FpvA and FptA have been solved in ligand-bound forms, which gave insight into the residues involved in ligand binding. The structures are similar and show the presence of similar domains; for example, all of them consist of a 22 strand-beta-barrel formed by approximately 600 C-terminal residues while approximately 150 N-terminal residues fold inside the barrel to form a plug domain. The plug domain obstructs the passage through the barrel; therefore our research focuses on the mechanism through which the ferricsiderophore complex is transported across the receptor into the periplasm. There are two possibilities, one in which the plug domain is expelled into the periplasm making way for the ferricsiderophore complex and the second in which the plug domain undergoes structural rearrangement to form a channel through which the complex slides into the periplasm. Multiple alignment studies involving protein sequences of a large number of outer membrane receptor proteins that transport ferricsiderophores have identified several conserved residues. All of the conserved residues are located within the plug and barrel domain below the ligand binding site. We have substituted a number of these residues in FepA and FhuA with either alanine or glutamine resulting in substantial changes in the chemical properties of the residues. This was done to study the effect of the substitutions on the transport of ferricsiderophores. Another strategy used was to create a disulfide bond between the residues located on two adjacent beta-strands of the plug domain or between the residues of the plug domain and the beta-barrel in FhuA by substituting appropriate residues with cysteine. We have looked for the variants where the transport is affected without altering the binding. The data suggest a distinct role of these residues in the mechanism of transport. Our data also indicate that these transporters share a common mechanism of transport and that the plug remains within the barrel and possibly undergoes rearrangement to form a channel to transport the ferricsiderophore from the binding site to the periplasm.

Bacterial Outer Membrane Proteins↗

DNA sequence motifs are associated with aberrant homologous recombination in the mouse macrophage migration inhibitory factor (Mif) locus.

Homologous recombination is a precise genetic event that can introduce specific alteration in the genome. A planned targeted disruption by homologous recombination of the macrophage migration inhibitory factor (Mif) locus in mouse embryonic stem (ES) cells yielded the targeted clones, some of which had genomic rearrangements inconsistent with the expected homologous recombination event. A detailed characterization of the recombination breakpoints in two of these clones revealed several sequence motifs with possible roles in recombination. These motifs included short regions of sequence identity that may promote DNA alignment, multiple 5'-AAGG/TTCC-3' tetrameres, topoisomerase I consensus sites, and AT-rich sequences that can promote DNA cleavage and recombination. A retrovirus-like intracisternal-A particle (IAP) family sequence was also identified upstream of the Mif gene, and the LTR of this IAP was involved in one of the recombinations. Identification and characterization of such sequence motifs will be valuable for the gene targeting experiments.

Adenine Nucleotides↗

Bioinformatics and molecular modeling in chemical enzymology. Active sites of hydrolases.

Comparison and multiple alignments of amino acid sequences of a representative number of related enzymes demonstrate the existence of certain positions of amino acid residues which are permanently reproducible in all members of the whole family. The use of the bioinformatic approach revealed conservative residues in each of the related enzymes and ranked amino acid conservatism for the overall enzymatic catalysis. Glycine and aspartic acid residues were shown to be the most essential for structure and catalytic activity of enzymes. Amino acid residues forming catalytic subsite of the active site of enzymes are always highly conservative. Analysis revealed that aspartic acid carboxyl group is the most frequently employed nucleophilic (in deprotonated form) and electrophilic (in protonated form) agent involved in activation of molecules by the mechanism of general base and acidic catalyses in the catalytic sites of enzymes. Glycine is a unique amino acid possessing the highest possibilities for rotation along C-C and C-N bonds of the polypeptide chain. The conservative fixation of the glycine residue in polypeptide chains of related enzymes provides a possibility for directed assembly of amino acid residues into the catalytic subsite structure. It is possible that the conservative glycines provide known conformational mobility of the protein and the active site. Methods of molecular modeling were used for analysis of structural substitutions of conservative and non-conservative glycines and their effects on geometry of catalytic site of typical hydrolases. The substitution of glycine(s) for alanine significantly altered the catalytic site structures.

Binding Sites↗

Mutagenesis of the fructose-6-phosphate-binding site in the 2-kinase domain of 6-phosphofructo-2-kinase/fructose-2,6-bisphosphatase.

Multiple alignment of several isozyme sequences of the bifunctional enzyme 6-phosphofructo-2-kinase/fructose-2,6-bisphosphatase revealed conserved residues in the 2-kinase domain. Among these residues, three asparagine residues (Asn76, Asn97 and Asn133; numbering refers to the liver isozyme sequence) and three threonine residues (Thr132, Thr134 and Thr135) are located near the fructose 6-phosphate-binding site in the crystal structure of the bifunctional enzyme. The role of these residues in substrate binding and catalysis in the 6-phosphofructo-2-kinase domain has been studied by mutagenesis to alanine. Since the crystal structure of 6-phosphofructo-2-kinase does not contain fructose 6-phosphate, this substrate was docked into the putative binding site by computer modelling, and its interactions with the protein were predicted. Analysis of the mutagenesis-induced changes in kinetic properties and of the substrate-docking model revealed that all these residues are directly or indirectly involved in fructose-6-phosphate binding. All the mutants displayed an increased Km for fructose 6-phosphate (10-200-fold). We propose that Asn133 stabilises Arg138, which itself makes a direct electrostatic bond with the 6-phosphate group of fructose 6-phosphate, that Asn76 interacts with the C3 hydroxyl group of fructose 6-phosphate, that Thr132 makes a hydrogen bond with the C6 oxygen of this substrate, and that Thr134 interacts with two residues involved in fructose-6-phosphate binding, Thr132 and Tyr199. On the other hand, Asn97 and Thr135 play structural roles, by maintaining the structure of the fructose-6-phosphate-binding pocket.

Base Sequence↗