Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,261 records · Page 70Linked to original sources

Inferring property selection pressure from positional residue conservation.

In this study, we attempt to understand and explain positional selection pressure in terms of underlying physical and chemical properties. We propose a set of constraining assumptions about how these pressures behave, then describe a procedure for analysing and explaining the distribution of residues at a particular position in a multiple sequence alignment. In contrast to previous approaches, our model takes into account both amino acid frequencies and a large number of physical-chemical properties. By analysing each property separately, it is possible to identify positions where distinct conservation patterns are present. In addition, the model can easily incorporate sequence weights that adjust for bias in the sample sequences. Finally, a test of statistical significance is provided for our conservation measure. The applicability of this method is demonstrated on two HIV-1 proteins: Nef and Env. The tools, data and results presented in this article are available at http://flan.blm.cs.cmu.edu.

Algorithms↗

Prediction of protein secondary structure using improved two-level neural network architecture.

In this paper we propose constructing an improved two-level neural network to predict protein secondary structure. Firstly, we code the whole protein composition information as the inputs to the first-level network besides the evolutionary information. Secondly, we calculate the reliability score for each residue position based on the output of the first-level network, and the role of the second-level network is to take full advantage of the residues with a higher reliability score to impact the neighboring residues with a lower one for improving the whole prediction accuracy. Thirdly, considering it is indeed a problem that the target protein can be lost in the multiple sequence alignment we propose to code single sequence into the second-level network. The experimental results show that our proposed method can efficiently improve the prediction accuracy.

Algorithms↗

Conformational changes preceding amyloid-fibril formation of amyloid-beta and stefin B; parallels in pH dependence.

Amyloid beta (A beta) protein is the key component of amyloid plaques in Alzheimer's disease brain whereas stefin B is an intracellular cysteine proteinase inhibitor, broadly distributed in different tissue and recently reported to form amyloid fibrils in vitro. By reducing the pH to 4.6, the native conformation of both polypeptides are changed into less ordered metastable intermediates that are stabilized by formation of the more stable fibrils. In A beta, the Glu at position 11 was found to be responsible for the conformational change at pH 4.6. Metal ions, including copper and zinc, could also induce conformational changes of A beta at neutral pH. The acid modified A beta conformer exhibited protease K resistance, preferential internalization and accumulation in the human glial cells. In stefin B, reducing the pH to pH 3.3 results in another intermediate of the molten-globule type which also leads to amyloid fibril formation. Multiple sequence alignment revealed distinct similarities of A beta (1-42) peptide, stefin B (13 to 61 residues) and prion fragment (90 to 144 residues).

Amino Acid Sequence↗

Analysis of fish IL-1beta and derived peptide sequences indicates conserved structures with species-specific IL-1 receptor binding: implications for pharmacological design.

A large number of IL-1 protein sequences have become available recently from a range of vertebrate species and especially from bony fish. However, 3D structures are still only known for mammalian IL-1. In this review, we use a multiple sequence alignment of all published non-mammalian vertebrate IL-1beta proteins to locate the structurally important residues critical for maintaining the beta-trefoil fold and we investigate the degree to which functionally important residues involved in receptor binding are conserved across vertebrate species. We find that although there is a high level of variability of positions involved in receptor binding, the mode of binding and overall shape of the ligand-receptor complex is probably maintained. This implies that each species has evolved its own unique interleukin-1 signalling system through ligand-receptor co-evolution. Nonetheless, the IL-1beta processing mechanism in non-mammalian vertebrates remains unclear because, with the exception of three bony fish, all non-mammalian IL-1beta sequences discovered so far lack an ICE (Interleukin Converting Enzyme) cut site. The IL-1 system has become an important drug target because of its significance in inflammatory diseases. Research on peptides derived from IL-1beta has identified peptides that possess agonist activity in humans and in trout, and peptides with antagonist activity. The agonist peptides map to two distinct loop regions of IL-1beta that are known to interact with the flexible domain III of the corresponding receptor. Further analysis of the IL-1 system may prove useful in engineering IL-1 with improved features and in suggesting new avenues for therapeutic intervention.

Amino Acid Sequence↗

In Ssarch of new anti-bacterial target genes: a comparative/structural genomics approach.

We outline a joint academic/industrial (CNRS/AVENTIS) functional genomics project aiming at the discovery of new anti-bacterial gene targets. Starting from all publicly available bacterial genomes, a subset of the most evolutionary conserved protein-coding genes has been identified. We retained genes with clear homolog in E. coli and at least one gram-positive bacterium among B.subtilis, M. tuberculosis, L. lactis or S. pyogenes. This subset was further reduced to genes encoding non-membrane proteins of unknown or hypothetical functions. The 221 E. coli Open Reading Frames (ORFs) identified through this comprehensive bioinformatic analysis are now submitted to a systematic 3-D structure determination protocol including cloning, protein expression and purification, crystallisation and X-ray diffraction. Our strategy was designed to focus on promising wide-spectrum targets as well as original biochemical pathways. Bioinformatics is used throughout all phases of project, including the initial large-scale comparative genomics analyses, the purification/expression and crystallisation stages for the detection of helpful sequence-specific features (e.g. cofactor binding motifs, non-structured N- or C- term extremities, etc ), and finally for the interpretation of the structures in conjunction with multiple sequence alignments for the identification of key residues, interaction areas on molecular surfaces, and overall function predictions.

Anti-Infective Agents↗

The prediction of amphiphilic alpha-helices.

A number of sequence-based analyses have been developed to identify protein segments, which are able to form membrane interactive amphiphilic alpha-helices. Earlier techniques attempted to detect the characteristic periodicity in hydrophobic amino acid residues shown by these structure and included the Molecular Hydrophobic Potential (MHP), which represents the hydrophobicity of amino acid residues as lines of isopotential around the alpha-helix and analyses based on Fourier transforms. These latter analyses compare the periodicity of hydrophobic residues in a putative alpha-helical sequence with that of a test mathematical function to provide a measure of amphiphilicity using either the Amphipathic Index or the Hydrophobic Moment. More recently, the introduction of computational procedures based on techniques such as hydropathy analysis, homology modelling, multiple sequence alignments and neural networks has led to the prediction of transmembrane alpha-helices with accuracies of the order of 95% and transmembrane protein topology with accuracies greater than 75%. Statistical approaches to transmembrane protein modeling such as hidden Markov models have increased these prediction levels to an even higher level. Here, we review a number of these predictive techniques and consider problems associated with their use in the prediction of structure / function relationships, using alpha-helices from G-coupled protein receptors, penicillin binding proteins, apolipoproteins, peptide hormones, lytic peptides and tilted peptides as examples.

Amino Acid Sequence↗

The fortuitous cloning of retroelement-like sequences from wheat and rye as by-products of a specific polymerase chain reaction.

Cloning of by-products of a specific PCR reaction, directed to the Em genes of wheat and rye, has resulted in the identification of ten sequences with homology to the known Tyl-copia-like retroelements WIS 2-1A from wheat and BARE-1 from barley. These sequences were amplified by only one of the primers due to the presence of an inverted repeat. Nine sequences are ca. 740 bp long and contain part of the left LTR, the adjacent primer-binding site and part of the leader sequence, whereas one shorter sequence (535 bp) consists of part of the leader sequence only. The dendrogram, constructed from the multiple sequence alignment, classified the isolated sequences into two narrowly related groups that belong to the WIS-2 family of cereal retroelements.

Base Sequence↗

5'-flanking regions of camel milk genes are highly similar to homologue regions of other species and can be divided into two distinct groups.

The concentrations of individual casein and whey proteins in camel milk differ markedly to respective protein concentrations in bovine milk. The ratio of beta-casein to kappa-casein is considerably higher in camel milk. beta-Lactoglobulin is absent, but whey acidic protein and peptidoglycan recognition protein have been detected. Genomic sequences upstream to milk-protein genes, which are known to regulate the expression of milk proteins to a great extent, were determined for 10 camel milk-protein genes and compared to respective sequences in other mammals. Multiple sequence alignment showed closest relationships to homologous sequences from other mammals. Comparison of milk protein regulative regions revealed two distantly related groups with pronouncedly different transcription factor site probabilities. The GC-content in sequences of the first group was considerably higher than in sequences of the second group and combined occurrence of CAAT and TATAA boxes was rare, suggesting that the first group represented mostly the housekeeping gene type, probably regulated by cellular signal transduction pathways, whereas the second group helped to regulate genes specifically expressed in terminally differentiated cells of the lactating alveolar epithelium. A core region of the composite response element, which primarily controls milk protein gene activity, was found by a search for elements conserved within all 5'-flanking sequences analyzed, and it is assumed, that the presence of this element determines gene expression in the lactating mammary gland, and binding sites for general activator and repressor factors, surrounding the milk protein gene specific element, are important for regulation of gene activity.

Animals↗

Bridging PCR and partially overlapping primers for novel allergen gene cloning and expression insert decoration.

AIM: To obtain the entire gene open reading frame (ORF) and to construct the expression vectors for recombinant allergen production. METHODS: Gene fragments corresponding to the gene specific region and the cDNA ends of pollen allergens of short ragweed (Rg, Ambrosia artemisiifolia L.) were obtained by pan-degenerate primer-based PCR and rapid amplification of the cDNA ends (RACE), and the products were mixed to serve as the bridging PCR (BPCR) template. The full-length gene was then obtained. Partially overlapping primer-based PCR (POP-PCR) method was developed to overcome the other problem, i.e., the non-specific amplification of the ORF with routine long primers for expression insert decoration. Northern blot was conducted to confirm pollen sources of the gene. The full-length coding region was evaluated for its gene function by homologue search in GenBank database and Western blotting of the recombinant protein Amb a 8(D106) expressed in Escherichia coli pET-44 system. RESULTS: The full-length cDNA sequence of Amb a 8(D106) was obtained by using the above procedure and deduced to encode a 131 amino acid polypeptide. Multiple sequence alignment exhibited the gene D106 sharing a homology as high as 54-89% and 79-89% to profilin from pollen and food sources, respectively. The expression vector of the allergen gene D106 was successfully constructed by employing the combined method of BPCR and POP-PCR. Recombinant allergen rAmb a 8(D106) was then successfully generated. The allergenicity was hallmarked by immunoblotting with the allergic serum samples and its RNA source was confirmed by Northern blot. CONCLUSION: The combined procedure of POP-PCR and BPCR is a powerful method for full-length allergen gene retrieval and expression insert decoration, which would be useful for recombinant allergen production and subsequent diagnosis and immunotherapy of pollen and food allergy.

Allergens↗

CD150 association with either the SH2-containing inositol phosphatase or the SH2-containing protein tyrosine phosphatase is regulated by the adaptor protein SH2D1A.

CD150 (SLAM/IPO-3) is a cell surface receptor that, like the B cell receptor, CD40, and CD95, can transmit positive or negative signals. CD150 can associate with the SH2-containing inositol phosphatase (SHIP), the SH2-containing protein tyrosine phosphatase (SHP-2), and the adaptor protein SH2 domain protein 1A (SH2D1A/DSHP/SAP, also called Duncan's disease SH2-protein (DSHP) or SLAM-associated protein (SAP)). Mutations in SH2D1A are found in X-linked lymphoproliferative syndrome and non-Hodgkin's lymphomas. Here we report that SH2D1A is expressed in tonsillar B cells and in some B lymphoblastoid cell lines, where CD150 coprecipitates with SH2D1A and SHIP. However, in SH2D1A-negative B cell lines, including B cell lines from X-linked lymphoproliferative syndrome patients, CD150 associates only with SHP-2. SH2D1A protein levels are up-regulated by CD40 cross-linking and down-regulated by B cell receptor ligation. Using GST-fusion proteins with single replacements of tyrosine at Y269F, Y281F, Y307F, or Y327F in the CD150 cytoplasmic tail, we found that the same phosphorylated Y281 and Y327 are essential for both SHP-2 and SHIP binding. The presence of SH2D1A facilitates binding of SHIP to CD150. Apparently, SH2D1A may function as a regulator of alternative interactions of CD150 with SHP-2 or SHIP via a novel TxYxxV/I motif (immunoreceptor tyrosine-based switch motif (ITSM)). Multiple sequence alignments revealed the presence of this TxYxxV/I motif not only in CD2 subfamily members but also in the cytoplasmic domains of the members of the SHP-2 substrate 1, sialic acid-binding Ig-like lectin, carcinoembryonic Ag, and leukocyte-inhibitory receptor families.

Amino Acid Sequence↗

Hybrid 'Sinta' papaya exhibits unique ACC synthase 1 cDNA isoforms.

Five ripening-related ACC synthase cDNA isoforms were cloned from 80% ripe papaya cv. 'Sinta' by reverse transcription-PCR using gene-specific primers. Clone 2 had the longest transcript and contained all common exons and three alternative exons. Clones 3 and 4 contained common exons and one alternative exon each, while clone 1, the most common transcript, contained only the common exons. Clone 5 could be due to cloning artifacts and might not be a unique cDNA fragment. Thus, there are only four isoforms of ACC synthase mRNA. Southern blot analysis indicates that all five clones came from only one gene existing as a single copy in the 'Sinta' papaya genome. Multiple sequence alignment indicates that the four isoforms arise from a single gene, possibly through alternative splicing mechanisms. All the putative alternative exons were present at the 5'-end of the gene comprising the N-terminal region of the protein. 'Sinta' ACC synthase cDNAs were of the capacs 1 type and are most closely related to a 1.4 kb capacs 1-type DNA (AJ277160) from Eksotika papaya. No capacs 2-type cDNAs were cloned from 'Sinta' by RT-PCR. This is the first report of possible alternative splicing mechanism in ripening-related ACC synthase genes in hybrid papaya, possibly to modulate or fine-tune gene expression relevant to fruit ripening.

Amino Acid Sequence↗

Molecular cloning, phylogenetic analysis, expressional profiling and in vitro studies of TINY2 from Arabidopsis thaliana.

A cDNA that was rapidly induced upon abscisic acid, cold, drought, mechanical wounding and to a lesser extent, by high salinity treatment, was isolated from Arabidopsis seedlings. It was classified as DREB subfamily member based on multiple sequence alignment and phylogenetic characterization. Since it encoded a protein with a typical ERF/AP2 DNA-binding domain and was closely related to the TINY gene, we named it TINY2. Gel retardation assay revealed that TINY2 was able to form a specific complex with the previously characterized DRE element while showed only residual affinity to the GCC box. When fused to the GAL4 DNA-binding domain, either full-length or its C-terminus functioned effectively as a trans-activator in the yeast one-hybrid assay while its N-terminus was completely inactive. Our data indicate that TINY2 could be a new member of the AP2/EREBP transcription factor family involved in activation of down-stream genes in response to environmental stress.

Abscisic Acid↗

Rapidly evolving aphid gall effector proteins exhibit saposin-like folds.

Many insects manipulate plants by injecting effector proteins. In one extreme example of this molecular "hijacking", Hormaphis cornu aphids inject bicycle proteins into Hamamelis virginiana (Witch Hazel), contributing to the development of novel organs called galls. Bicycle proteins share no amino acid sequence similarity with proteins of known function. Here, we report the crystal structures of two divergent bicycle proteins. Both proteins contain saposin-like folds: one with multiple disulfide bonds exhibits a helix swap; the other has no disulfide bonds and possesses two tandem domains. To explore the structural evolution of bicycle proteins, we predicted bicycle protein structures with Alphafold2 (AF2). While AF2 did not recover the two experimental structures using existing databases, it succeeded after we provided multiple sequence alignments (MSAs) containing protein sequences encoded in new genome sequences from closely related aphid species. Using this customized approach at scale, we generated 2400 high-confidence predictions for bicycle proteins from seven aphid species. This dataset revealed that bicycle proteins without cysteines are outliers in fold space and appear to have evolved from ancestral proteins with disulfide-bonded saposin-like folds. While all bicycle proteins contain predicted saposin-like folds, they display a vast diversity of structural and physicochemical properties. While this diversity thwarts prediction of conserved functions encoded in structure, it suggests that bicycle proteins have evolved to target diverse plant processes and/or to evade plant immune surveillance.

AlphaFold predictions↗

Major structural determinants of transmembrane proteins identified by principal component analysis.

We identify amino acid characteristics important in determining the secondary structures of transmembrane proteins, and compare them with characteristics important for cytoplasmic proteins. Using information derived from multiple sequence alignments, we perform a principal component analysis (PCA) to identify the directions in the 20-dimensional amino acid frequency space that comprise the most variance within each protein secondary structure. These vectors represent the important position-specific properties of the amino acids for coils, turns, beta sheets, and alpha helices. As expected, the most important axis for most of the datasets was hydrophobicity. Additional axes, distinct from hydrophobicity, are surprising, especially in the case of transmembrane alpha helices, where the effects of aromaticity and beta-branching are the next two most significant characteristics. The axis representing beta-branching also has equal importance in cytoplasmic and transmembrane helices, a finding that contrasts with some experimental results in membrane-like environments. In a further analysis, we examine trends for some of the PCA axes over averaged transmembrane alpha helices, and find interesting results for aromaticity.

Amino Acids↗

Role of evolutionary information in predicting the disulfide-bonding state of cysteine in proteins.

A neural network-based predictor is trained to distinguish the bonding states of cysteine in proteins starting from the residue chain. Training is performed by using 2,452 cysteine-containing segments extracted from 641 nonhomologous proteins of well-resolved three-dimensional structure. After a cross-validation procedure, efficiency of the prediction scores were as high as 72% when the predictor is trained by using protein single sequences. The addition of evolutionary information in the form of multiple sequence alignment and a jury of neural networks increases the prediction efficiency up to 81%. Assessment of the goodness of the prediction with a reliability index indicates that more than 60% of the predictions have an accuracy level greater than 90%. A comparison with a statistical method previously described and tested on the same database shows that the neural network-based predictor is performing with the highest efficiency. Proteins 1999;36:340-346.

Binding Sites↗

Multi-domain, cell-envelope proteinases of lactic acid bacteria.

The multi-domain, cell-envelope proteinases encoded by the genes prtB of Lactobacillus delbrueckii subsp. bulgaricus, prtH of Lactobacillus helveticus, prtP of Lactococcus lactis, scpA of Streptococcus pyogenes and csp of Streptococcus agalactiae have been compared using multiple sequence alignment, secondary structure prediction and database homology searching methods. This comparative analysis has led to the prediction of a number of different domains in these cell-envelope proteinases, and their homology, characteristics and putative function are described. These domains include, starting from the N-terminus, a pre-pro-domain for secretion and activation, a serine protease domain (with a smaller inserted domain), two large middle domains A and B of unknown but possibly regulatory function, a helical spacer domain, a hydrophilic cell-wall spacer or attachment domain, and a cell-wall anchor domain. Not all domains are present in each cell-envelope proteinase, suggesting that these multi-domain proteins are the result of gene shuffling and domain swapping during evolution.

Amino Acid Sequence↗

Molecular modeling of the oxytocin receptor/bioligand interactions.

Oxytocin is a nonapeptide hormone (CYIQNCPLG-NH2, OT), controlling labor and lactation in mammalian females, via interactions with specific cellular membrane receptors (OTRs). The native hormone is cyclized via a 1-6 disulfide and its receptor belongs to the GTP-binding (G) protein-coupled receptor (GPCR) family, also known as heptahelical transmembrane (7TM) or serpentine receptors. Using a technique combining multiple sequence alignments with available experimental constraints, a reliable OTR model was built. Subsequently, the OTR complexes with a selective agonist [Thr4,Gly7]OT, a selective cyclohexapeptide antagonist L-366,948 and oxytocin itself were modeled and relaxed using a constrained simulated annealing (CSA) protocol. All three ligands seem to prefer similar modes of binding to the receptor, manifested by repeating receptor residues which directly interact with the ligands. Those involved in the three complexes are putative helices: TM3: R113, K116, Q119, M123; TM4: Q171, and TM5: I201 and T205. Most of them are the equivalent residues/positions to those found in our earlier studies, regarding related vasopressin V2 receptor/bioligand interactions.

Amino Acid Sequence↗

Prediction of structural and functional relationships of Repeat 1 of human interphotoreceptor retinoid-binding protein (IRBP) with other proteins.

PURPOSE: We compared the structure and function of interphotoreceptor retinoid-binding protein (IRBP) related proteins and predicted domain and secondary structure within each repeat of IRBP and its relatives. We tested whether tail specific protease (Tsp), which bears sequence similarity to IRBP Domain B, binds fatty acids or retinoids, and whether IRBP possessed protease activity resembling Tsp's catalytic function. These tests helped us to learn whether the primary sequence similarities of family members extended to higher order structural and functional levels. METHODS: Predictions derived from multiple sequence alignments among IRBP and Tsp family members and secondary structure computer programs were carried out. The first repeat of human IRBP (EcR1) and Tsp were expressed, purified, and tested for binding properties. Tsp was examined for fluorescence enhancement of retinol or 16-anthroyloxy-palmitic acid (16-AP) to test for ligand binding. IRBP was tested for protease activity. RESULTS: Tsp did not exhibit fluorescence enhancement with retinol or 16-AP. IRBP did not exhibit protease activity. The positions of critical residues needed for the ligand binding properties of retinol were predicted. Primary sequence and three-dimensional similarity was found between Domain A of IRBP Repeat 3 and eglin c. CONCLUSIONS: The sequence similarity of Tsp and IRBP raised the possibility that each might share the function of the other protein: IRBP might possess protease activity or Tsp might possess retinoid or fatty acid binding activity. Our studies do not support such a shared function hypothesis, and suggest that the sequence similarity is the result of maintenance of structure. The finding of similarity to eglin c in Domain A suggests the possibility of a tight interaction between Domain A and Domain B, possibly implying the need for Domain A in retinoid-binding, and suggesting that both Domains should be present in testing mutations. The positions of predicted critical amino acids suggest models in which a large binding pocket holds the retinoid or fatty acid ligand. These predictions are tested in a companion paper.

Cluster Analysis↗