Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,567 records · Page 87Linked to original sources

[Development and preparation of recombinant gD antigen of the herpes simplex type 1 (HSV-1) virus].

The most potent antigen among HSV-1 proteins are glycoproteins gB(UL27) and gD(US6). Multiple amino acid sequence alignment of these proteins shows that gD protein is the most specific for HSV-1. Analysis of gD protein epitopes detected the main antigenic determinants not cross-reactive with antigens of other viruses. Virus was isolated and genome DNA was prepared from morphological elements of a patient with herpes simplex infection. US6 gene fragment was cloned in pUC19 vector. Cloning in bacterial expression vectors helped obtain beta-galactosidase-fused recombinant HSV-1 gD protein with 6-histidines affine target for high-performance chromatography purification. ELISA with a set of HSV-1-positive and negative donor sera and a commercial panel of HSV-1 sera (Vektor-Best) showed that recombinant gD can be used as an antigen to HSV-1-specific IgG.

Amino Acid Sequence↗

[Molecular screening of MC4R gene and association with fat traits in pig resource family].

Melanocortin-4 Receptor (MC4R) plays an important role in the regulation of human obesity. It can cooperate with leptin, neuropeptide Y(NPY) and melanocyte-stimulating hormone (MSH) to regulate body weight and feeding. Inactivation of this receptor by gene targeting in mice results in a maturity onset obesity syndrome associated with hyperphagic, hyperinsulinemia, hyperglycemia, as well as decreased linear growth and adult obesity. Multiple alignments of the sequences from individuals of several pig lines identified a single nucleotide substitution(G-->A) at position 298 of the seventh transmembrane domain. In present study, polymorphism distribution of MC4R gene fragment in resource population was studied using PCR-RFLP method based on the enzyme Taq I. The genotype was analyzed with the phenotype of the slaughtered individuals. The results showed that the frequencies of MC4R genotype varied in different breeds. The correlation analysis demonstrated the genotype of MC4R was in significant relation with back-fat thickness on thorax-waist, buttock and the average back-fat thickness, as well as with the width and area of longissmus dorsi (LD), and the percentage of skin. MC4R gene plays a role mainly in the pattern of dominant effect, and all the additive effects were not significant.

Animals↗

Molecular dissection of the beta subunit of F1-ATPase into peptide fragments.

Partial digestion of the native beta subunit of F1-ATPase from the thermophilic Bacillus strain PS3 by three different proteases produced a limited number of peptide fragments. In most cases, the peptides remained associated, and the gross structure of the beta subunit was not destroyed. Furthermore, most peptides were able to reassociate into the form of the beta subunit after denaturating urea treatment. Therefore, the cleaved sites are most likely located in water-exposed loop regions in the tertiary structure of the protein. Almost all peptides were analyzed, and 17 cleaved sites were determined. From the analysis of the distribution of cleaved sites and deletions or insertions in the multiple amino acid sequence alignment of proteins homologous to the beta subunit, locations of five loops and four candidate loops in the beta subunit are suggested. There are two large loops in the central region of the beta subunit sequence, and dicyclohexylcarbodiimide-reactive Glu190 is located in one of them. Tyr341, involved in putative catalytic ATP binding, is also found in one of the loops. Then, taking cleaved sites as a reference, two kinds of expression plasmids, each of which carried genes of two complementary peptide fragments, 1-193 and 198-473 or 1-284 and 285-473, were constructed and expressed in Escherichia coli. For each plasmid, two peptides were coexpressed, associated into a stable beta subunit form in E. coli cells, and purified without dissociation. When these beta subunits were denatured by urea and applied to polyacrylamide gel without denaturant, a protein band with the same mobility as that of the beta subunit appeared, indicating that reassociation of peptide fragments into the form of the beta subunit occurred upon removal of urea. These beta subunits retained the ability to reconstitute the alpha 3 beta 3 gamma complexes even though the efficiency of reconstitution and the recovered ATPase activities were decreased. These complexes were stable at high or low temperature, and ATPase activities were sensitive to inhibition by N3-.

Amino Acid Sequence↗

[Use of structural MNA descriptors for designing profiles of protein families].

A new approach to constructing the profiles of protein families is proposed, which uses only structural similarity of amino acid residues. We derived multiple alignments of protein sequences from 3D superpositions of the protein structures and constructed protein family profiles using structural molecular MNA descriptors. MNA (Multilevel Neighborhoods of Atoms) descriptors were developed earlier and are successfully applied for predicting the biological activity in drug-like compounds. In our approach, each aligned position was described by a set of MNA descriptors calculated for each amino acid residue in the alignment column. In this study, we constructed MNA profiles for trypsin, subtilase, and cytochrome P450 protein families and scanned SWISSPROT with some fragments of these profiles. We also calculated the Independence Accuracy of Prediction for each profile fragment. It was shown that the approach developed could be applied to predict protein function.

Amino Acid Sequence↗

[Cloning and characterization of a full-length HIV-1 genome of a prevalent subtype B-Thai strain in Henan Province].

OBJECTIVE: To clone, identify and phylogenetically characterize a clade B-Thai HIV isolate representing the most prevalent virus in Henan province. METHODS: Peripheral blood mononuclear cells (PBMCs) from an HIV-1 infected patient in Henan Province were separated, and co-cultivated with phytohemagglutinin-stimulated healthy donor PBMCs. Proviral DNA was extracted from productively infected PBMCs. The full-length HIV-1 genome was amplified by using the LA Tag long template PCR system. Primers were positioned in conserved regions within the HIV-1 long terminal repeats. Purified PCR products were T-A ligated into a pWSK29-T vector(CNHN 24 clone). Three recombinant clones containing virtually full-length HIV-1 genome were identified by PCR. The full-length genome was sequenced by using the primer-walking approach. Nucleotide sequence similarities were calculated by the local-homology algorithm. Phylogenetic trees of gag, pol and env reading frames were constructed using the Phylip software. RESULTS: HIV-1 C3V4 sequences indicate that the epidemic in this area was B-Thai subtype. V3 loop multiple amino acid sequence alignments showed amino acid alterations at nine positions. The 9,010 bp genomic sequence derived from isolate CNHN 24 contained all known structural and regulatory genes of an HIV-1 genome. No major deletions, insertions, or rearrangements were found. The highest homologies of the gag, pol, vpr, and vif reading frames to the corresponding clade B-Thai RL 42 sequences were 95.42%-97.08%. Phylogenetic trees showed the closest relationship of CNHN 24 and RL 42. CONCLUSION: The cloning and characterization of a virtually full-length HIV-1 B-Thai subtype in central China was completed in our laboratory. The data should be helpful to future studies on the genetic diversity of HIV-1.

Amino Acid Sequence↗

PCR-based detection of Mycoplasma species.

In this study, we describe our newly-developed sensitive two-stage PCR procedure for the detection of 13 common mycoplasmal contaminants (M. arthritidis, M. bovis, M. fermentans, M. genitalium, M. hominis, M. hyorhinis, M. neurolyticum, M. orale, M. pirum, M. pneumoniae, M. pulmonis, M. salivarium, U. urealyticum). For primary amplification, the DNA regions encompassing the 16S and 23S rRNA genes of 13 species were targeted using general mycoplasma primers. The primary PCR products were then subjected to secondary nested PCR, using two different primer pair sets, designed via the multiple alignment of nucleotide sequences obtained from the 13 mycoplasmal species. The nested PCR, which generated DNA fragments of 165-353 bp, was found to be able to detect 1-2 copies of the target DNA, and evidenced no cross-reactivity with the genomic DNA of related microorganisms or of human cell lines, thereby confirming the sensitivity and specificity of the primers used. The identification of contaminated species was achieved via the performance of restriction fragment length polymorphism (RFLP) coupled with Sau3AI digestion. The results obtained in this study furnish evidence suggesting that the employed assay system constitutes an effective tool for the diagnosis of mycoplasmal contamination in cell culture systems.

Animals↗

Porcine aromatases: studies on tissue-specific, functionally distinct isozymes from a single gene?

Aromatase cytochrome P450 (P450arom) is expressed in a variety of tissues. Pigs express P450arom as bilaminar blastocysts in utero, and thereafter in the gonads, adrenal glands and placenta. Our studies also demonstrate the existence of porcine isozymes of P450arom which differ substantially in their amino acid composition and function. The placental isoform, most similar to P450arom in other mammals, consists of 503 amino acids. The ovarian isoform, expressed in both theca and granulosa cells, is a 501 amino acid protein exhibiting less than 20% of the activity of the placental isozyme. Furthermore, it is inhibited not only by CGS16949A but also by etomidate which does not inhibit the placental P450arom. Partial sequences generated by the rapid amplification of the cDNA ends (RACE) procedure indicate that the expression of a third isoform in the blastocyst is switched to the placental isozyme during differentiation of the fetal membranes. In addition, these transcripts, and others from the theca, granulosa, testes, adrenal glands and placenta demonstrate differences in the 5'-untranslated region (putative exon I) suggestive of tissue-specific alternative splicing. An identical 5'-untranslated sequence was obtained from transcripts expressed in the theca and granulosa. Testes and adrenal transcripts also have identical 5' ends, which differ substantially from the ovarian sequence. Blastocyst and placenta 5'-untranslated sequences differ from each other and from those expressed in the gonads and adrenals. Several tissue-specific transcripts thus encode porcine P450arom. Interestingly, distinct 5' sequences exist for ovarian and testes P450arom mRNAs, suggesting different promoters and therefore regulation in the male and female gonads. The molecular origins of the functional isoforms and the tissue-specific transcripts are uncertain, however partial genomic sequence and other genetic analyses suggest the existence of multiple genes. However, sequence alignment of the placental and ovarian isoforms indicates complete conservation of putative exon III, so that complex splicing remains a possibility. Clearly, the regulation of P450arom expression is more complex in the pig than in other vertebrates investigated to date.

Amino Acid Sequence↗

Malate dehydrogenase: distribution, function and properties.

Malate dehydrogenase (MDH) (EC 1.1.1.37) catalyzes the conversion of oxaloacetate and malate. This reaction is important in cellular metabolism, and it is coupled with easily detectable cofactor oxidation/reduction. It is a rather ubiquitous enzyme, for which several isoforms have been identified, differing in their subcellular localization and their specificity for the cofactor NAD or NADP. The nucleotide binding characteristics can be altered by a single amino acid change. Multiple amino acid sequence alignments of MDH show that there is a low degree of primary structural similarity, apart from several positions crucial for catalysis, cofactor binding and the subunit interface. Despite the low amino acids sequence identity their 3-dimensional structures are very similar. MDH is a group of multimeric enzymes consisting of identical subunits usually organized as either dimer or tetramers with subunit molecular weights of 30-35 kDa. MDH has been isolated from different sources including archaea, eubacteria, fungi, plant and mammals.

Animals↗

Using multiple alignments and phylogenetic trees to detect RNA secondary structure.

We describe a statistical method to determine if a pair of columns in a multiple alignment of a homologous family of RNA sequences shows evidence of being base paired. The method makes explicit use of a given phylogenetic tree for the sequences in the alignment. It is tested on a multiple alignment of 16S rRNA sequences with good results.

Base Composition↗

fRagmentomics: an R package for integrating cell-free DNA fragment features with mutational status to support liquid biopsy interpretation.

SUMMARY: Liquid biopsy offers a non-invasive approach to study tumor-derived genetic material circulating in plasma. Beyond genetic alterations, the fragmentomic features of cell-free DNA-such as fragment size, genomic position, and end-motifs-provide valuable insights into the biological and clinical context of DNA release. fRagmentomics is a user-friendly R package designed to characterize cfDNA fragments overlapping one or multiple small mutations of any type, starting from an aligned sequencing file (BAM). It supports multiple mutation input formats, accommodates one-based and zero-based genomic conventions, resolves mutation representation ambiguities, and accepts any reference file in FASTA format. For each fragment overlapping a mutation of interest, fRagmentomics outputs fragment-level features including its fragment size, end-motifs, and mutational status, along with additional fragment-level or read-level information. The package implements an indel-aware and optionally soft-clip-preserving fragment size computation that improves accuracy over conventional size estimates based solely on aligned positions. AVAILABILITY AND IMPLEMENTATION: fRagmentomics is licensed under GNU General Public License v3.0 and available at https://github.com/ElsaB-Lab/fRagmentomics, https://anaconda.org/elsab-lab/r-fragmentomics and https://bioconductor.org/packages/fRagmentomics, with documentation and a tutorial. CONTACT: yoann.pradat@gustaveroussy.fr, elsa.bernard@gustaveroussy.fr. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Software↗

SSMAL: similarity searching with alignment graphs.

MOTIVATION: We want to provide biologists with a fast and sensitive scanning tool for searching local alignments of a protein query sequence against databases of protein multiple alignments, such as ProDom. Conversely, we want to provide a tool for locally aligning a protein multiple alignment query against a protein database such as SWISSPROT. RESULTS: We developed the program SSMAL (Shuffling Similarities with Multiple Alignments) which utilizes features of the Blast (Altschul et al., J. Mol. Biol., 215, 403-410, 1990) algorithm and part of the Blast code. Our software allows both scanning of multiple alignments and searching with a multiple alignment. Deletions in the multiple alignment only are handled and a SSMAL search may miss some similarities found by a profile search. However, an SSMAL scan of a database such as ProDom would be 20-30 times faster that a profile scan. In the worst case, a SSMAL search is approximately 9 times faster than a profile search. AVAILABILITY: http://www.dkfz-heidelberg.de/tbi/ people/nicodeme and follow the hyperlink SSMAL. CONTACT: p.nicodeme@DKFZ-Heidelberg.de

Amino Acid Sequence↗

Automated whole-genome multiple alignment of rat, mouse, and human.

We have built a whole-genome multiple alignment of the three currently available mammalian genomes using a fully automated pipeline that combines the local/global approach of the Berkeley Genome Pipeline and the LAGAN program. The strategy is based on progressive alignment and consists of two main steps: (1) alignment of the mouse and rat genomes, and (2) alignment of human to either the mouse-rat alignments from step 1, or the remaining unaligned mouse and rat sequences. The resulting alignments demonstrate high sensitivity, with 87% of all human gene-coding areas aligned in both mouse and rat. The specificity is also high: <7% of the rat contigs are aligned to multiple places in human, and 97% of all alignments with human sequence >100 kb agree with a three-way synteny map built independently, using predicted exons in the three genomes. At the nucleotide level <1% of the rat nucleotides are mapped to multiple places in the human sequence in the alignment, and 96.5% of human nucleotides within all alignments agree with the synteny map. The alignments are publicly available online, with visualization through the novel Multi-VISTA browser that we also present.

Animals↗

PAL2NAL: robust conversion of protein sequence alignments into the corresponding codon alignments.

PAL2NAL is a web server that constructs a multiple codon alignment from the corresponding aligned protein sequences. Such codon alignments can be used to evaluate the type and rate of nucleotide substitutions in coding DNA for a wide range of evolutionary analyses, such as the identification of levels of selective constraint acting on genes, or to perform DNA-based phylogenetic studies. The server takes a protein sequence alignment and the corresponding DNA sequences as input. In contrast to other existing applications, this server is able to construct codon alignments even if the input DNA sequence has mismatches with the input protein sequence, or contains untranslated regions and polyA tails. The server can also deal with frame shifts and inframe stop codons in the input models, and is thus suitable for the analysis of pseudogenes. Another distinct feature is that the user can specify a subregion of the input alignment in order to specifically analyze functional domains or exons of interest. The PAL2NAL server is available at http://www.bork.embl.de/pal2nal.

Codon↗

Playing with blocks: some pitfalls of forcing multiple alignments.

Block alignments of multiple amino acid sequences are useful representations of regions thought to share common ancestry and function. Often the block alignments are motivated by the expectation that a protein of interest is similar in function to members of a family of proteins. However, when alignments are forced by using ad hoc methods, it is often difficult to decide whether the proposed relationship is valid. Visual examination can be deceptive, especially when alignments are not carried out in the context of controls subjected to similar procedures. Even computer-aided methods can be misleading when biases are introduced. To illustrate some of the problems that can arise, a few examples from the literature are analyzed. It is concluded that when standard methods fail to find an interesting block alignment unaided by human intervention, then the result should be regarded with caution.

Amino Acid Sequence↗

Using evolutionary Expectation Maximization to estimate indel rates.

MOTIVATION: The Expectation Maximization (EM) algorithm, in the form of the Baum-Welch algorithm (for hidden Markov models) or the Inside-Outside algorithm (for stochastic context-free grammars), is a powerful way to estimate the parameters of stochastic grammars for biological sequence analysis. To use this algorithm for multiple-sequence evolutionary modelling, it would be useful to apply the EM algorithm to estimate not only the probability parameters of the stochastic grammar, but also the instantaneous mutation rates of the underlying evolutionary model (to facilitate the development of stochastic grammars based on phylogenetic trees, also known as Statistical Alignment). Recently, we showed how to do this for the point substitution component of the evolutionary process; here, we extend these results to the indel process. RESULTS: We present an algorithm for maximum-likelihood estimation of insertion and deletion rates from multiple sequence alignments, using EM, under the single-residue indel model owing to Thorne, Kishino and Felsenstein (the 'TKF91' model). The algorithm converges extremely rapidly, gives accurate results on simulated data that are an improvement over parsimonious estimates (which are shown to underestimate the true indel rate), and gives plausible results on experimental data (coronavirus envelope domains). Owing to the algorithm's close similarity to the Baum-Welch algorithm for training hidden Markov models, it can be used in an 'unsupervised' fashion to estimate rates for unaligned sequences, or estimate several sets of rates for sequences with heterogenous rates. AVAILABILITY: Software implementing the algorithm and the benchmark is available under GPL from http://www.biowiki.org/

Algorithms↗

Blocks-based methods for detecting protein homology.

The most highly conserved regions of proteins can be represented as blocks of aligned sequence segments, typically with multiple blocks for a given protein family. The Blocks Database World Wide Web (http://blocks.fhcrc.org) and e-mail (blocks@blocks. fhcrc.org) servers provide tools to search DNA and protein queries against the Blocks+ Database of multiple alignments. We describe features for detection of distant relationships using blocks. Blocks+ includes protein families from the PROSITE, Prints, Pfam-A, ProDom and Domo databases. Other features include searching Blocks+ with the BLIMPS and NCBI's IMPALA programs, sequence logos, phylogenetic trees, three-dimensional display of blocks on PDB structures, and a polymerase chain reaction (PCR) primer design strategy based on blocks.

Amino Acid Sequence↗

Using guide trees to construct multiple-sequence evolutionary HMMs.

MOTIVATION: Score-based progressive alignment algorithms do dynamic programming on successive branches of a guide tree. The analogous probabilistic construct is an Evolutionary HMM. This is a multiple-sequence hidden Markov model (HMM) made by combining transducers (conditionally normalised Pair HMMs) on the branches of a phylogenetic tree. METHODS: We present general algorithms for constructing an Evolutionary HMM from any Pair HMM and for doing dynamic programming to any Multiple-sequence HMM. RESULTS: Our prototype implementation, Handel, is based on the Thorne-Kishino-Felsenstein evolutionary model and is benchmarked using structural reference alignments.

Algorithms↗

PowerBLAST: a new network BLAST application for interactive or automated sequence analysis and annotation.

As the rate of DNA sequencing increases, analysis by sequence similarity search will need to become much more efficient in terms of sensitivity, specificity, automation potential, and consistency in annotation. PowerBLAST was developed, in part, to address these problems. PowerBLAST includes a number of options for masking repetitive elements and low complexity subsequences. It also has the capacity to restrict the search to any level of NCBI's taxonomy index, thus supporting "comparative genomics" applications. Postprocessing of the BLAST output using the SIM series of algorithms produces optimal, gapped alignments, and multiple alignments when a region of the query sequence matches multiple database sequences. PowerBLAST is capable of processing sequences of any length because it divides long query sequences into overlapping fragments and then merges the results after searching. The results may be viewed graphically, as a textual representation, or as an HTML page with links to GenBank and Entrez. For matching database sequences, annotated features are superimposed on the aligned query sequence in the output, thus greatly increasing the ease of interpretation. Such features may be used for automated annotation of new sequence because PowerBLAST output in ASN.1 form may be "dragged and dropped" into NCBI's Sequin program for sequence annotation and submission. PowerBLAST is capable of analyzing and annotating a 100-kb query in 60 min on NCBI's BLAST server.

Amino Acid Sequence↗