Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,117 records · Page 62Linked to original sources

A 3D model of the delta opioid receptor and ligand-receptor complexes.

A model for the 3D structure of the transmembrane domain of the delta opioid receptor was predicted from the sequence divergence analysis of 42 sequences of G-protein coupled peptide hormone receptors belonging to the opioid, somatostatin and angiotensin receptor families. No template was used in the prediction steps, which include multiple sequence alignment, calculation of a variability profile of the aligned sequences, use of the variability profile to identify the boundaries of transmembrane regions, prediction of their secondary structure, optimization of the packing shape in a helix bundle, prediction of side chain conformations and structural refinement. The general shape of the model is similar to that of the low resolution rhodopsin structure in that the TM3 and TM7 helices are most buried in the bundle and the TM1 and TM4 helices are most exposed to the lipid phase. An initial assessment of this model was made by determining to what extent a binding site identified using four structurally disparate high affinity delta opioid ligands was consistent with known mutational studies. With the assumption that the protonated amine nitrogen, a feature common to all delta opioid ligands, interacts with the highly conserved Asp127 in TM3, a pocket was found that satisfied the criteria of complementarity to the requirements for receptor recognition for these four diverse ligands, two delta selective antagonists (the fused ring naltrindole and the peptide Tyr-Tic-Phe-Phe-NH2) and the two agonists lofentanil and BW373U86 deduced from previous studies of the ligands alone. These ligands could be accommodated in a similar region of the receptor. The receptor binding site identified in the optimized complexes contained many residues in positions known to affect ligand binding in G-protein coupled receptors. These results also allowed identification of key residues as candidates for point mutations for further assessment and refinement of this model as well as preliminary indications of the requirements for recognition of this receptor.

Amino Acid Sequence↗

Homology modeling of the central catalytic domain of insertion sequence ISLC3 isolated from Lactobacillus casei ATCC 393.

The tertiary structure of the central catalytic domain of insertion sequence ISLC3 isolated from Lactobacillus casei ATCC 393 was predicted using the homology modeling approach. The novel insertion sequence was isolated by us from the template bacteriophage phiA3 of L.casei ATCC 393. The number of amino acid residues of the ISLC3 central catalytic domain was 116 and was treated as the query sequence. There were five Web-available threading methods used to find some primary structure templates for the query sequence. These primary templates were further screened using the SWISS-MODEL Protein Modeling Server and the default parameter settings therein to give six final structure templates. All of these final structure templates were the integrase (IN) protein of retroviruses. Multiple sequence alignment using these IN sequences against the query one revealed the signature DDE motif. Based on the structures of these final templates, the structure of the query sequence was constructed using the InsightII/Discover/Homology programs. A metal ion, Mg(2+), was inserted into the center of the putative catalytic pocket formed by the DDE residues of the predicted structure in the final rounds of refinement by molecular dynamics (MD) simulations. The structure with a metal ion included was designated with Mg and that without a metal ion was designated free Mg. The average exposed surface area of some hydrophobic residues of both the predicted free Mg and with Mg structures were computed and compared with those computed for the six structure templates. Whereas the predicted with Mg structure was slightly more exposed than the predicted free Mg structure, the former appeared to be more stable than the latter, as revealed by the lower conformation energy recorded for the former during the structure refinement by MD simulations. To verify further the predicted structures, the coordinates of both predicted structures were fed into the ERRAT Protein Verification Server. It was found that the quality of the predicted with Mg structure was much better than that of the free Mg structure. The validation results also indicated that regions of the predicted with Mg structure that can be rejected at the 95% confidence level were approximately 20% whereas those which can be rejected at the same level for the six structure templates were approximately 10%. The predicted with Mg structure was also docked into a short oligonucleotide representing the substrate of the ISLC3 transposase using the DOCK_4.0.2 program. It was found that both Glu140 and Asp68 residues of the DDE motif of the predicted with Mg structure were able to form hydrogen bonds with the DNA substrate, which was similar to what was observed in a docking study using the retrovirus IN 1asu and its DNA substrate.

Amino Acid Sequence↗

Conservation and covariance in PH domain sequences: physicochemical profile and information theoretical analysis of XLA-causing mutations in the Btk PH domain.

Mutations that cause X-linked agammaglobulinemia (XLA) appear throughout the Bruton tyrosine kinase (Btk) sequence, including the pleckstrin homology (PH) domain. To analyze the basis of this disease with respect to protein structure, we studied the relationships between PH domain sequences and structures by comparing sequence-based profiles of physicochemical properties and solvent accessibility profiles. The diversity of the distribution of amino acids was measured by calculating entropies for sequences containing mutations at different positions in multiple sequence alignments. Mutual information was calculated to quantify positional covariation. Eight conserved extrema were apparent in all profiles. The majority of the XLA disease-causing mutations in the Btk PH domain were found at positions having significant mutual information, indicating that there are covariant constraints for both structure and function. Together with additional structural analyses, all the XLA mutations that were analyzed could be explained at the molecular level. The method developed here is applicable to the design of mutations for protein engineering.

Agammaglobulinemia↗

Automated design of degenerate codon libraries.

Degenerate codon libraries are frequently used in protein engineering and evolution studies but are often limited to targeting a small number of positions to adequately limit the search space. To mitigate this, codon degeneracy can be limited using heuristics or previous knowledge of the targeted positions. To automate design of libraries given a set of amino acid sequences, an algorithm (LibDesign) was developed that generates a set of possible degenerate codon libraries, their resulting size, and their score relative to a user-defined scoring function. A gene library of a specified size can then be constructed that is representative of the given amino acid distribution or that includes specific sequences or combinations thereof. LibDesign provides a new tool for automated design of high-quality protein libraries that more effectively harness existing sequence-structure information derived from multiple sequence alignment or computational protein design data.

Algorithms↗

Statistical analysis and prediction of functional residues effective for GPCR-G-protein coupling selectivity.

One of the important issues in G-protein-coupled receptor (GPCR) functional analysis is the mechanism of GPCR-G-protein coupling selectivity. G-proteins are classified into Gi/o, Gq/11 and Gs families. Although several experimental and computational analyses have been attempted, the mechanism remains unknown to this day. In this study, we have analyzed the multiple sequence alignments of GPCRs of known coupling selectivities by mapping onto the tertiary structure of rhodopsin. We identified several functional residue sites in GPCRs related to coupling selectivity, which are located mainly at the intracellular loops, and found that the occurrence of positively/negatively charged amino acids of the characteristic residues varies depending on the G-protein coupling selectivity. Especially, the occurrence of positively charged amino acids in receptors coupling to Gs family is less than that in receptors coupling to Gi/o and Gq/11 families. It is interesting that some characteristic residues are located near the extracellular terminus of transmembrane helices, which is far from the GPCR/G-protein binding interface. In most of the receptors coupling to Gs family, the occurrence of proline on the position corresponding to the 170th residue on rhodopsin is rare. These findings are vital to improving our understanding of the mechanism of G-protein coupling selectivity.

Amino Acid Sequence↗

A novel class of elicitin-like genes from Phytophthora infestans.

Elicitins are a family of structurally related proteins that induce hypersensitive response in specific plant species. Two Phytophthora infestans cDNAs, inf2A and inf2B, potentially encoding novel elicitin-like proteins, were isolated from a cDNA library made from infected potato tissue. Multiple sequence alignments and phylogenetic analyses of 19 elicitins and elicitin-like proteins from nine Phytophthora spp. and from Pythium vexans suggest that there are at least five distinct classes within the elicitin family.

Algal Proteins↗

Cloning and functional characterization of a gonadal luteinizing hormone receptor complementary DNA from the African catfish (Clarias gariepinus).

A cDNA encoding a putative African catfish (Clarias gariepinus) gonadal LH receptor (cfLH-R) has been cloned. Multiple sequence alignment of the deduced amino acid sequence revealed that the cfLH-R had the highest identity with vertebrate LH receptors (>50%). Overall sequence identity between the cfLH-R and the African catfish FSH receptor (cfFSH-R) is 47%. Sequence analysis of part of the cfLH-R gene revealed the presence of an intron typically found in other vertebrate LH-R genes. Abundant cfLH-R mRNA expression was detected in ovary and testis as well as in head-kidney (the adrenal homologue in fish). Other tissues, such as muscle, brain, cerebellum, stomach, heart, and seminal vesicles, also contained detectable cfLH-R mRNA. Transient expression of the cfLH-R in HEK-T 293 cells resulted in significantly increased basal cAMP levels in the absence of gonadotropic hormone. The cAMP levels could be further elevated in response to catfish LH, salmon LH, human LH, human choriogonadotropin, and human FSH. Salmon FSH and human TSH, however, were inactive. We conclude that we have cloned a cDNA encoding the LH-R of the African catfish. This receptor displays constitutive activity but is still responsive to additional ligand-induced activation.

Amino Acid Sequence↗

Structural and evolutionary relationships among the immunophilins: two ubiquitous families of peptidyl-prolyl cis-trans isomerases.

The immunophilins, protein receptors for the immunosuppressing drugs cyclosporin A and FK506 and related proteins from plants, fungi, and bacteria, have been analyzed structurally and evolutionarily. The cyclosporin A binding proteins (cyclophilins) represent one ubiquitous family of homologous proteins, and the FK506- and rapamycin-binding proteins (FKBPs) constitute a second, unrelated family. Multiple sequence alignments of members of each of these two protein families define the highly conserved residues that are likely to play important structural and functional roles, and mutations in representative members of these two families that abolish or alter function have been evaluated. FKBPs have undergone greater evolutionary divergence than the cyclophilins. Evolutionary trees were constructed using two distinct programs, and these trees establish the structural relationships that allow division of each of these families into subgroups. The results lead to the suggestion that several genes encoding isozymic forms of the FKBPs and possibly also of the cyclophilins existed in prokaryotes before the emergence of eukaryotes on earth and that representatives of these genes were transmitted to both kingdoms to give rise to current subfamilies of these proteins. By contrast, compartmentalization of both classes of immunophilins appears to have arisen independently in prokaryotes and eukaryotes, late in evolutionary history.

Amino Acid Isomerases↗

Genomic divergence of an HIV-2 from a German AIDS patient probably infected in Mali.

The complete nucleotide sequence of an HIV-2 isolate derived from a German AIDS patient with predominantly neurological symptoms is reported. The HIV-2BEN sequence is highly divergent from those of previously described HIV-2 and SIV strains. Evolutionary tree analysis of eight HIV-2 sequences reveals the existence of three HIV-2 groups. HIV-2BEN belongs to a group with two isolates from Ghana and The Gambia. Based on a comparison of HIV-2BEN with six HIV-2 isolates, SIVsmm and SIVmac, the variability of the structural env and gag proteins is similar within the HIV-2/SIVsmm/mac and HIV-1 groups. In contrast, the regulatory HIV-1 proteins are more highly conserved than those from HIV-2 strains. Multiple sequence alignments reveal that some domains of the envelope and regulatory proteins are well conserved among HIV-1, HIV-2/SIVsmm/mac, SIVagm and SIVmnd. The identification of conserved domains within the external glycoprotein could help to develop broadly active vaccines.

Acquired Immunodeficiency Syndrome↗

Phylogenetic analysis of gag genes from 70 international HIV-1 isolates provides evidence for multiple genotypes.

OBJECTIVE: To determine the extent of genetic variation among internationally collected HIV-1 isolates, to analyse phylogenetic relationships and the geographic distribution of different variants. DESIGN: Phylogenetic comparison of 70 HIV-1 isolates collected in 15 countries on four continents. METHODS: To sequence the complete gag genome of HIV-1 isolates, build multiple sequence alignments and construct phylogenetic trees using distance matrix methods and maximum parsimony algorithms. RESULTS: Phylogenetic tree analysis identified seven distinct genotypes. The seven genotypes were evident by both distance matrix methods and maximum parsimony analysis, and were strongly supported by bootstrap resampling of the data. The intra-genotypic gag distances averaged 7%, whereas the inter-genotypic distances averaged 14%. The geographic distribution of variants was complex. Some genotypes have apparently migrated to several continents and many areas harbor a mixture of genotypes. Related variants may cluster in certain areas, particularly isolates from a single city collected over a short time. CONCLUSIONS: The genetic variation among HIV-1 isolates is more extensive than previously appreciated. At least seven distinct HIV-1 genotypes can be identified. Diversification, migration and establishment of local, temporal 'blooms' of particular variants may all occur concomitantly.

Africa↗

Site-directed mutagenesis of the lipoate acetyltransferase of Escherichia coli.

Remote but significant similarities between the primary and predicted secondary structures of the chloramphenicol acetyltransferases (CAT) and lipoate acyltransferase subunits (LAT, E2) of the 2-oxo acid dehydrogenase complexes, have suggested that both types of enzyme may use similar catalytic mechanisms. Multiple sequence alignments for CAT and LAT have highlighted two conserved motifs that contain the active-site histidine and serine residues of CAT. Site-directed replacement of Ser550 in the E2p subunit (LAT) of the pyruvate dehydrogenase complex of Escherichia coli, deemed to be equivalent to the active-site Ser148 of CAT, supported the CAT-based model of LAT catalysis. The effects of other substitutions were also consistent with the predicted similarity in catalytic mechanism although specific details of active-site geometry may not be conserved.

Acetyltransferases↗

A phylogenetic analysis of Borrelia burgdorferi sensu lato based on sequence information from the hbb gene, coding for a histone-like protein.

We describe a phylogenetic investigation of Borrelia burgdorferi sensu lato, the causative agent of Lyme disease, based on a DNA sequence analysis of the hbb gene, which encodes protein HBb, a member of the family of histone-like proteins. Because of their intimate contact with the DNA molecule, these proteins are believed to be fairly conserved through evolution. In this study we proved that the hbb gene is suitable for phylogenetic inference in the genus Borrelia. The hbb gene, which is 327 bp long and encodes 108 amino acids, was sequenced for 39 strains, including 37 strains of B. burgdorferi sensu lato, 1 strain of Borrelia turicatae, and 1 strain of Borrelia parkeri. Genetic variability was determined at the sequence level by computational analysis. Briefly, 81 substitutions were scored at the DNA level. Only 25 of these substitutions were responsible for amino acid substitutions at the translational level. The signature region for bacterial histone-like proteins was found in hbb. Although variable at the nucleotide level, it was highly conserved at the deduced amino acid level. A phylogenetic tree for the genus Borrelia that was generated from multiple sequence alignments was consistent with previously published data derived from DNA-DNA hybridization and multilocus enzyme electrophoresis analyses. The subdivision of B. burgdorferi sensu lato into five species (B. burgdorferi sensu stricto, Borrelia garinii, Borrelia afzelii, Borrelia japonica, and "Borrelia andersonii") and at least four genomic groups (groups PotiB2, VS116, CA2, and DN127) was confirmed.

Amino Acid Sequence↗

Evaluation of intraspecies genetic variation within the 60 kDa heat-shock protein gene (groEL) of Bartonella species.

A phylogenetic investigation was done on the members of the genus Bartonella, based on the DNA sequence analysis of the groEL gene, which encodes the 60 kDa heat-shock protein GroEL. Nucleotide sequence data were determined for a near full-length fragment (1368 bp) of the groEL gene of the established Bartonella species and used to infer intraspecies phylogenetic relationships. Phylogenetic trees were inferred from multiple sequence alignments by using both distance and parsimony methods, which demonstrated an architecture composed of six well-supported lineages. The results are consistent with relationships deduced from recent sequence analysis studies based upon citrate synthase (gItA) and previously observed genotypic and phenotypic characteristics; however, they showed greater statistical support at the intragenus level. This suggests that groEL may be a more robust tool for phylogenetic analysis of Bartonella lineages.

Bartonella↗

Identification of a gag protein epitope conserved among all four groups of primate immunodeficiency viruses by using monoclonal antibodies.

Five monoclonal antibodies (MAbs) were raised against the gag proteins of simian immunodeficiency virus (SIV) from African green monkey (SIVagmTYO-7). Two MAbs reacted with the matrix protein p17 and the other three with the core protein p24. Studies on the cross-reactivity of the MAbs revealed that the anti-p24 MAbs detected an epitope shared by the viruses belonging to the human immunodeficiency virus type 2 (HIV-2)/SIVmac group and SIVagmTYO-7 and SIVagmTYO-5. The anti-p17 MAbs recognized an epitope present on all these viruses and on SIVagmTYO-1, HIV-1 and SIVmnd. This finding demonstrates for the first time that the matrix protein, p17 or p18, respectively, of all nine HIV and SIV isolates tested in this study expresses at least one conserved immunogenic epitope recognized serologically. By using synthetic peptides, this epitope was identified at the N terminus of p17. Furthermore, this epitope was analysed by multiple sequence alignments of the peptide with homologous sequences of HIV and SIV p17.

Amino Acid Sequence↗

Sequence and structural analysis of murine adenovirus type 1 hexon.

The genomic region encoding the major capsid protein (hexon) of murine adenovirus type 1 (MAV-1) has been isolated and sequenced. The sequence predicts a 908 residue MAV-1 hexon protein and is flanked by a portion of the upstream pVI gene and the downstream endoproteinase gene. The order of these genes and their location in the middle of the genome are the same as those found in other adenoviruses sequenced to date. Multiple sequence alignment with the other five known hexon protein sequences reveals an overall residue identity of 51% and residue conservation of 66%. In comparison with human adenovirus type 2 (Ad2), MAV-1 hexon has major deletions between residues 141 to 170, 270 to 284 and 446 to 455. Since these regions in the Ad2 hexon are partially exposed on the outer surface of the virion, they may represent type-specific antigenic determinants. The MAV-1 hexon sequence has been modelled using the known three-dimensional structure of the Ad2 hexon. The variable regions in which the mutations, deletions and insertions occur are located in the l1 and l2 loops of the molecule that form the protruding hexon towers on the external surface of the virion.

Amino Acid Sequence↗

Genome characterization and taxonomy of Plantago asiatica mosaic potexvirus.

The complete nucleotide sequence of Plantago asiatica mosaic virus (P1AMV) genomic RNA has been determined. The 6128 nucleotide sequence contains five open reading frames (ORFs) coding for proteins of M(r) 156K (ORF1), 25K (ORF2), 12K (ORF3), 13K (ORF4) and 22K (ORF5). The sequences of these P1AMV proteins exhibit strong homology to the proteins of the other potexviruses. Phylogenetic trees based on the multiple sequence alignments of three conserved domains in ORF1 product and capsid protein reveal a close relationship of P1AMV to papaya mosaic virus and clover yellow mosaic virus. The P1AMV genomic RNA and a major subgenomic RNA (sgRNA) of 0.9 kb have been detected in infected leaves by Northern blot hybridization. The latter sgRNA is the messenger for virus capsid protein and its 5' terminus has been located 23 nucleotides upstream of the initiator codon of the coat protein gene. The P1AMV virion RNA and RNA transcript resembling the 0.9 kb sgRNA have been translated in vitro giving rise to a single major 170K product and a major 22K product, respectively.

Amino Acid Sequence↗

Sequence polymorphism in the 5'NTR and in the P1 coding region of potato virus Y genomic RNA.

Potato virus Y (PVY) the type member of the genus Potyvirus, occurs world-wide as isolates which differ in host range and the type of symptoms caused. The sequences of a 5' segment of viral RNA overlapping the 5' non-translated region (5'NTR) alone (ten isolates) or the 5'NTR and the adjacent P1 coding region (eight isolates) were established. These data were used to quantify the polymorphism in the 5'-terminal part of the PVY genome. Nucleotide sequence identity between isolates ranged from 66-100% in the 5'NTR and from 70-100% in the P1 coding region. The lowest amino acid sequence similarity between PVY P1 was 77%, illustrating the high variability of this protein in the PVY species. Phylogenetic trees based on either 5'NTR or P1 sequence analyses resulted in the same clustering of the studied isolates into three groups. Group I comprises potato isolates all inducing 'tobacco veinal necrosis' symptoms. Group II contains isolates inducing either 'tobacco veinal necrosis' or mosaic symptoms in tobacco. Group III contains mainly pepper or tomato isolates inducing mosaic symptoms in tobacco and shows a geographical clustering of the Tunisian isolates. This clustering into three groups is discussed in comparison with phylogenetic trees previously obtained from capsid gene or 3'NTR sequence analysis in the PVY species. Multiple sequence alignment indicated conserved motifs potentially involved in viral functions.

Amino Acid Sequence↗

Adelaide River virus nucleoprotein gene: analysis of phylogenetic relationships of ephemeroviruses and other rhabdoviruses.

The nucleotide sequence of the Adelaide River virus (ARV) genome was determined from the 3' terminus to the end of the nucleoprotein (N) gene. The 3' leader sequence comprises 50 nucleotides and shares a common terminal trinucleotide (3' UGC-), a conserved U-rich domain and a variable AU-rich domain with other animal rhabdoviruses. The N gene comprises 1355 nucleotides from the transcription start sequence (AACAGG) to the poly(A) sequence [CATG(A)7] and encodes a polypeptide of 429 amino acids. The N protein has a calculated molecular mass of 49429 Da and a pI of 5.4 and, like the bovine ephemeral fever virus (BEFV) N protein, features a highly acidic C-terminal domain. Analysis of amino acid sequence relationships between all available rhabdovirus N proteins indicated that ARV and BEFV are closely related viruses (48.3% similarity) which share higher sequence similarity to vesiculoviruses than to lyssaviruses. Phylogenetic trees based on a multiple sequence alignment of all available rhabdovirus N protein sequences demonstrated clustering of viruses according to genome organization, host range and established taxonomic relationships.

Amino Acid Sequence↗