Search PubMedSearch

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Consistency of optimal sequence alignments.

Pairwise optimal alignments between three or more sequences are not necessarily consistent as a whole, but consistent and inconsistent residues are usually distributed in clusters. An efficient method has been developed for locating consistent regions when each pairwise alignment is given in the form of a "skeletal representation" (Bull. math. Biol. 52, 359-373). This method is further extended so that the combination of pairwise alignments that gives the greatest consistency is found when possibly many alignments are equally optimal for each pairwise comparison. A method for acceleration of simultaneous multiple sequence alignment is proposed in which consistent regions serve as "anchor points" limiting application of direct multi-way alignment to the rest of "inconsistent" regions.

Algorithms

Phylogenetic relationships of three porcine mycoplasmas, Mycoplasma hyopneumoniae, Mycoplasma flocculare, and Mycoplasma hyorhinis, and complete 16S rRNA sequence of M. flocculare.

The nucleotide sequence of the 16S rRNA gene of Mycoplasma flocculare was determined and was compared with the sequence of a related porcine mycoplasma, Mycoplasma hyopneumoniae. While the overall level of DNA-DNA homology was approximately 11%, sequence alignment of the two 16S rRNA genes yielded a homology value of more than 95%, emphasizing the highly conserved nature of the 16S rRNA gene. Multiple sequence alignments with other mollicutes indicated that M. flocculare, M. hyopneumoniae, and Mycoplasma hyorhinis form a subcluster within the fermentans phylogroup, and this subcluster is distinct from the Mycoplasma pneumoniae phylogroup. Thus, the three mycoplasmas isolated from porcine respiratory systems exhibit phylogenetic similarities.

Animals

Whole genome sequencing of unusual Hepatitis C virus subtypes and drug resistance analysis during direct-acting antiviral therapy in India.

INTRODUCTION AND OBJECTIVES: Pangenotypic direct-acting antivirals (DAA) are effective against highly prevalent Hepatitis C virus (HCV) subtypes, but have been clinically validated almost exclusively in high-income countries. Unusual HCV subtypes may carry natural polymorphisms, potentially impacting DAA susceptibility. We conducted full-genome characterization and resistance analysis of unusual HCV subtypes in patients receiving DAA treatment. PATIENTS AND METHODS: In this prospective hospital-based study, eligible patients were screened for anti-HCV antibodies and active infection was confirmed by diagnostic 5'NCR-based HCV RNA detection. Genotyping was performed by core region sequencing, and viral load quantified by real-time PCR. For whole genome sequencing, multiplex primers were designed using alignments of global reference sequences. Sequencing was carried out using the Oxford Nanopore Technology platform. Phylogenetic analysis used multiple sequence alignment and the HCV-GLUE resource for resistance-associated substitution (RAS) analysis. RESULTS: Predominant genotype was genotype 3 in 64.3% (n = 45); genotype 6 in 21.4% (n = 15); and genotype 1 in 14.2% (n = 10). Unusual HCV subtype 6xa was detected in two patients and showed no NS5A resistance mutations. One genotype 3b patient relapsed at 24 weeks post-DAA treatment completion and carried NS5A resistance-associated substitutions 30 K and 31 M both at baseline and at relapse, conferring high-level resistance to NS5A inhibitors. CONCLUSION: This is the first report from India of whole genome sequencing of HCV subtype 6xa. The identification of NS5A resistance mutations in the 3b relapse case underscores challenges for global HCV elimination strategies.

Humans

Suboptimal sequence alignment in molecular biology. Alignment with error analysis.

A molecular sequence alignment algorithm based on dynamic programming has been extended to allow the computation of all pairs of residues that can be part of optimal and suboptimal sequence alignments. The uncertainties inherent in sequence alignment can be displayed using a new form of dot plot. The method allows the qualitative assessment of whether or not two sequences are related, and can reveal what parts of the alignment are better determined than others. It also permits the computation of representative optimal and suboptimal alignments. The relation between alignment reliability and alignment parameters is discussed. Other applications are to cyclical permutations of sequences and the detection of self-similarity. An application to multiple sequence alignment is noted.

Algorithms

Amino acid sequence analysis of the annexin super-gene family of proteins.

The annexins are a widespread family of calcium-dependent membrane-binding proteins. No common function has been identified for the family and, until recently, no crystallographic data existed for an annexin. In this paper we draw together 22 available annexin sequences consisting of 88 similar repeat units, and apply the techniques of multiple sequence alignment, pattern matching, secondary structure prediction and conservation analysis to the characterisation of the molecules. The analysis clearly shows that the repeats cluster into four distinct families and that greatest variation occurs within the repeat 3 units. Multiple alignment of the 88 repeats shows amino acids with conserved physicochemical properties at 22 positions, with only Gly at position 23 being absolutely conserved in all repeats. Secondary structure prediction techniques identify five conserved helices in each repeat unit and patterns of conserved hydrophobic amino acids are consistent with one face of a helix packing against the protein core in predicted helices a, c, d, e. Helix b is generally hydrophobic in all repeats, but contains a striking pattern of repeat-specific residue conservation at position 31, with Arg in repeats 4 and Glu in repeats 2, but unconserved amino acids in repeats 1 and 3. This suggests repeats 2 and 4 may interact via a buried saltbridge. The loop between predicted helices a and b of repeat 3 shows features distinct from the equivalent loop in repeats 1, 2 and 4, suggesting an important structural and/or functional role for this region. No compelling evidence emerges from this study for uteroglobin and the annexins sharing similar tertiary structures, or for uteroglobin representing a derivative of a primordial one-repeat structure that underwent duplication to give the present day annexins. The analyses performed in this paper are re-evaluated in the Appendix, in the light of the recently published X-ray structure for human annexin V. The structure confirms most of the predictions and shows the power of techniques for the determination of tertiary structural information from the amino acid sequences of an aligned protein family.

Algorithms

An ATPase domain common to prokaryotic cell cycle proteins, sugar kinases, actin, and hsp70 heat shock proteins.

The functionally diverse actin, hexokinase, and hsp70 protein families have in common an ATPase domain of known three-dimensional structure. Optimal superposition of the three structures and alignment of many sequences in each of the three families has revealed a set of common conserved residues, distributed in five sequence motifs, which are involved in ATP binding and in a putative interdomain hinge. From the multiple sequence alignment in these motifs a pattern of amino acid properties required at each position is defined. The discriminatory power of the pattern is in part due to the use of several known three-dimensional structures and many sequences and in part to the "property" method of generalizing from observed amino acid frequencies to amino acid fitness at each sequence position. A sequence data base search with the pattern significantly matches sugar kinases, such as fuco-, glucono-, xylulo-, ribulo-, and glycerokinase, as well as the prokaryotic cell cycle proteins MreB, FtsA, and StbA. These are predicted to have subdomains with the same tertiary structure as the ATPase subdomains Ia and IIa of hexokinase, actin, and Hsc70, a very similar ATP binding pocket, and the capacity for interdomain hinge motion accompanying functional state changes. A common evolutionary origin for all of the proteins in this class is proposed.

Actins

Modeling Alternative Conformational States in CASP16.

The CASP16 Ensemble Prediction experiment assessed advances in methods for modeling proteins, nucleic acids, and their complexes in multiple conformational states. Targets included systems with experimental structures determined in two or three states, evaluated by direct comparison to experimental coordinates, as well as domain-linker-domain (D-L-D) targets assessed against statistical models from NMR and SAXS data. This paper focuses on the former class of multi-state targets. Ten ensembles were released as community challenges, including ligand-induced conformational changes, protein-DNA complexes, a trimeric protein, a stem-loop RNA, and multiple oligomeric states of a single RNA. For five targets, some groups produced reasonably accurate models of both reference states (best TM-score >0.75). However, with the exception of one protein-ligand complex (T1214), where an apo structure was available as a template, predictors generally failed to capture key structural details distinguishing the states. Overall, accuracy was significantly lower than for single-state targets in other CASP experiments. The most successful approaches generated multiple AlphaFold2 models using enhanced multiple sequence alignments and sampling protocols, followed by model quality based selection. While the AlphaFold3 server performed well on several targets, individual groups outperformed it in specific cases. By contrast, predictions for one protein-DNA complex, three RNA targets, and multiple oligomeric RNA states consistently fell short (TM-score <0.75). These results highlight both progress and persistent challenges in multi-state prediction. Despite recent advances, accurate modeling of conformational ensembles, particularly RNA and large multimeric assemblies, remains a critical frontier for structural biology.

AlphaFold2

Genetic diversity and recombination of&#xa0;NA-PRRSV field strains in Vietnam: Implications for vaccine efficacy.

Porcine reproductive and respiratory syndrome (PRRS) causes severe reproductive losses in pregnant sows and piglets, resulting in substantial economic impact on the swine industry worldwide. However, due to the significant genetic diversity and rapid evolutionary changes of the pathogen, continuous surveillance and detailed genetic analysis of circulating strains are essential. The current study aimed to evaluate the genetic diversity of the hypervariable (HV) region of non-structural protein 2 (nsp2) among North American PRRSV strains isolated from swine farms in Vietnam. Phylogenetic analysis and multiple sequence alignment were conducted to determine subtype classification and assess genetic variability. A total of 48 field isolates were obtained, of which 12.5% belonged to classical NA-PRRSV, 16.6% to NADC30-like and 70.9% to HP-PRRSV, primarily distributed across sublineages 1.4, 5.1, 8.7 and 8.9. Amino acid comparisons found multiple insertions, deletions and substitutions at various positions within the hypervariable region of nsp2. The study revealed substantial genetic variation in the HV region of nsp2 among NA-PRRSV field strains, largely associated with recombination and immune escape. These findings highlight epidemiological risks to vaccine efficacy and underscore the need for continuous molecular surveillance to support effective PRRSV control in Vietnam.

PRRSV

Development and epidemiological investigation of a TaqMan-based multiplex real-time quantitative PCR assay for simultaneous detection of five bovine viruses (BVDV, AKAV, BNoV, BEV, and BCoV).

INRODUCTION: Infectious diseases caused by bovine viral diarrhea virus (BVDV), Akabane virus (AKAV), bovine norovirus (BNoV), bovine enterovirus (BEV), and bovine coronavirus (BCoV) significantly threaten the cattle industry, resulting in substantial economic losses. These pathogens often present similar clinical signs, such as diarrhea, vomiting, and reproductive disorders in pregnant cattle, and frequent covert or mixed infections further complicate accurate diagnosis. Therefore, rapid, sensitive, and field&#x2011;deployable diagnostic methods are essential for effective disease surveillance and control in the cattle industry. METHODS: In this study, we report for the first time the establishment of a TaqMan&#x2011;based real&#x2011;time quantitative PCR (qPCR) assay that enables simultaneous detection of these five bovine viruses. Multiple sequence alignment of conserved genomic regions was performed, and virus&#x2011;specific primers and probes were designed and optimized using Beacon Designer 7 software. Subsequently, a TaqMan&#x2011;based multiplex real&#x2011;time qPCR assay was established for simultaneous detection of BVDV, AKAV, BNoV, BEV, and BCoV. The established detection method was applied to 200 clinical samples collected from 10 farms in multiple regions of Jilin Province. RESULTS: The results showed that the detection rates for BVDV, AKAV, BNoV, BEV, and BCoV were 33.50%, 0.50%, 4.50%, 7.50%, and 12.00%, respectively. Mixed infections were detected in 9 samples co&#x2011;infected with two of the five pathogens, with an overall mixed infection rate of 4.50%. Compared with conventional PCR, coincidence rates were 100% for BVDV, AKAV, BNoV, BEV, and BCoV. DISCUSSION: These findings indicate that the TaqMan multiplex real&#x2011;time qPCR assay developed here demonstrates favorable specificity, sensitivity, and reproducibility. This assay enables efficient detection and surveillance of bovine viruses, offering a reliable technical tool for the diagnosis and control of corresponding viral diseases in cattle.

Akabane virus (AKAV)

The evolution of rhodopsins and neurotransmitter receptors.

Rhodopsins share a limited number of amino acid identities with a variety of other integral membrane proteins. Most of these proteins have seven putative transmembrane segments and are likely to play a role in transmembrane signaling. We have undertaken a systematic series of comparisons of primary and secondary structure in order to clarify the functional and evolutionary significance of these sequence similarities. On the basis of consistently high similarity scores, we find that the most internally consistent definition of the rhodopsin gene family would include vertebrate rhodopsins, alpha- and beta-adrenergic receptors, M1 and M2 muscarinic acetylcholine receptors, substance K receptors, and insect rhodopsins, while excluding bacteriorhodopsin, the mass human oncogene, vertebrate and insect nicotinic acetylcholine receptors, and the yeast STE2 and STE3 peptide receptors. The rhodopsin gene family is highly diverged at the primary sequence level but has maintained a conserved secondary structure, including a previously unidentified hierarchy of transmembrane segment hydrophobicity. We have developed new computer algorithms for progressive multiple sequence alignment and the analysis of local conservation of protein domains, and we have used these algorithms to examine the phylogeny of the rhodopsin gene family and the changing domains of sequence conservation. The results show striking differences and similarities in the conserved domains in each of the three main branches of the rhodopsin gene family, and indicate that color vision arose independently in the lines of descent leading to modern humans and fruit flies.

Algorithms

Structural and functional relationships of human DNA polymerases.

A continuing theme of our laboratory has been the understanding of human DNA polymerases at the structural level. We have purified DNA polymerases delta, epsilon and alpha from human placenta. Monoclonal antibodies to these polymerases were isolated and used as tools to study their immunochemical relationships. These studies have shown that while DNA polymerases delta, epsilon and alpha are discrete proteins, they must share common structural features by virtue of the ability of several of our monoclonal antibodies to exhibit cross-reactivity. A second approach we have taken is the molecular cloning of human DNA polymerase delta and epsilon. We have cloned the DNA polymerase delta cDNA, and this has allowed us to compare its primary structure to those of human polymerase alpha and other members of this polymerase family. Multiple sequence alignments have revealed that human DNA polymerase delta is also closely related to the herpes virus family of DNA polymerases. In situ hybridization has shown that the human DNA polymerase delta gene is localized to chromosome 19 q13.3-q13.4. In order to further determine the functional regions of the DNA polymerase delta structure we are currently expressing human pol delta in E. coli and baculovirus systems. Other work in our laboratory is directed toward examining the expression of DNA polymerase delta during the cell cycle.

Amino Acid Sequence

Molecular Cloning, Recombinant Expression, and In Silico Structural Analysis of Cu/Zn-Superoxide Dismutase from Trachyspermum ammi.

Superoxide dismutase (SOD) is an essential antioxidant metalloenzyme that is critical for the cellular defense against oxidative damage, as it scavenges superoxide radicals and maintains the redox status. Cytosolic Cu/Zn-SOD is particularly important in the regulation of oxidative stress among different isoforms in higher plants. While Cu/Zn-SODs from several plant species have been characterized, molecular information is limited for Trachyspermum ammi, a medicinally important member of a family Apiaceae with antioxidant potential.In the present study, an integrated molecular and in silico approach has been taken to clone and analyze a Cu/Zn type SOD gene from T. ammi to get insight into its structural and evolutionary characteristics. PCR amplification yielded an open reading frame of 456&#xa0;bp encoding a protein of 152 amino acids. Sequence analysis showed that plant Cu/Zn-SODs, especially those from Daucus carota, were highly similar to one another (about 90-95%).Multiple sequence alignment confirmed the presence of conserved catalytic motifs and metal-binding histidine residues, both of which are crucial for enzymatic function. Physicochemical analysis predicted the protein to be stable, hydrophilic and compatible with cytosolic localization. The analysis of secondary structure indicated a predominance of &#x3b2;-strands, consistent with the conserved &#x3b2;-barrel architecture of plant Cu/Zn-SODs.The three-dimensional structure was built by homology modeling using a closely related plant Cu/Zn-SOD template with high sequence identity. Structural validation demonstrated an acceptable stereochemical quality with 86.3% residues in the favored region of Ramachandran plot, satisfactory ERRAT and Verify3D scores, and a low RMSD value of 0.104&#xa0;&#xc5; on structural superimposition. Phylogenetic analysis placed the enzyme in the Apiaceae lineage, suggesting evolutionary conservation among related plant species. In conclusion, this study presents the first molecular and structural characterization of Cu/Zn-SOD from T. ammi and confirms the existence of a conserved structural framework typical of plant Cu/Zn-SODs. These results provide a basis for further studies concerning recombinant expression, enzymatic validation and potential relevance in antioxidant and plant stress biology.

Cloning, Molecular

Sequence and localization of human NASP: conservation of a Xenopus histone-binding protein.

In this study the sequence and localization of human testicular NASP (nuclear autoantigenic sperm protein) are reported. NASP cDNA contains 2561 nt encoding a protein of 787 amino acids. The open reading frame contains 2446 nt followed by an ochre stop codon (TAA) and 104 nucleotides of untranslated sequence containing a poly(A) addition signal 10 bases upstream of the poly(A) tail. Northern blot analysis of human testis poly(A) mRNA indicates a message of approximately 3.2 kb. Multiple sequence alignment (MSA) analysis of the encoded human NASP amino acid sequence with the sequence for the Xenopus histone-binding protein N1/N2 and the rabbit NASP amino acid sequence demonstrates that the human sequence and the Xenopus sequence have extensive amino acid homology upstream of the rabbit initiation codon. Significantly, there is an 85% identity between the human and the rabbit NASP sequences when the alignment starts at the N-terminal of the rabbit sequence and at amino acid 101 of the human sequence. The nuclear translocation signal found in N1/N2 and rabbit NASP is completely conserved in human NASP. The first histone-binding domain of Xenopus is 70% identical and 90% similar to the human NASP domain. The second histone-binding domain of Xenopus is 48% identical and 71% similar to the human NASP domain. MSA analysis of the three sequences generated an unrooted ancestral tree with two branches, indicating that fewer amino acid changes have occurred between the Xenopus and the human sequences than between the Xenopus and the rabbit sequences. In the human testis, NASP is localized predominantly in primary spermatocytes and round spermatids. Spermatogonia, Sertoli cells, Leydig cells, peritubular cells, and other somatic cells do not stain. Human spermatozoa contain NASP in the acrosomal region. Following the acrosome reaction, some NASP remains in the equatorial and postacrosomal regions. We propose that mammalian testes and sperm contain a histone-binding protein which may play a role in regulating the early events of spermatogenesis.

Amino Acid Sequence

Conservation analysis and structure prediction of the SH2 family of phosphotyrosine binding domains.

Src homology 2 (SH2) regions are short (approximately 100 amino acids), non-catalytic domains conserved among a wide variety of proteins involved in cytoplasmic signaling induced by growth factors. It is thought that SH2 domains play an important role in the intracellular response to growth factor stimulation by binding to phosphotyrosine containing proteins. In this paper we apply the techniques of multiple sequence alignment, secondary structure prediction and conservation analysis to 67 SH2 domain amino acid sequences. This combined approach predicts seven core secondary structure regions with the pattern beta-alpha-beta-beta-beta-beta-alpha, identifies those residues most likely to be buried in the hydrophobic core of the native SH2 domain, and highlights patterns of conservation indicative of secondary structural elements. Residues likely to be involved in phosphotyrosine binding are shown and orientations of the predicted secondary structures suggested which could enable such residues to cooperate in phosphate binding. We propose a consensus pattern that encapsulates the principal conserved features of the SH2 domains. Comparison of the proposed SH2 domain of akt to this pattern shows only 12/40 matches, suggesting that this domain may not exhibit SH2-like properties.

Amino Acid Sequence

Analysis of insertions/deletions in protein structures.

An analysis of insertions and deletions (indels) occurring in a databank of multiple sequence alignments based on protein tertiary structure is reported. Indels prefer to be short (1 to 5 residues). The average intervening sequence length between them versus the percentage of residue identity in pairwise alignments shows an exponential behaviour, suggesting a stochastic process such that nearly every loop in an ancestral structure is a possible target for indels during evolution. The results also suggest a limit to the average size of indels accommodated by protein structures. The preferred indel conformations are reverse turn and coil as are the preferred conformations at the indel edges (N- and C-terminal sides). Interruptions in helices and strands were observed as very rare events.

Amino Acid Sequence

A proteolytically sensitive region common to several rat liver cytochromes P450: effect of cleavage on substrate binding.

Limited proteolysis of rat liver microsomes was used to probe the topography and structure of cytochrome P450 bound to the endoplasmic reticulum. Three cytochromes P450 from two families were examined. Monoclonal antibodies to cytochrome P450 forms 1A1, 2B1, and 2E1 were used to immunopurify these proteolyzed cytochromes P450 from microsomes from rats treated with 3-methylcholanthrene, phenobarbital, and acetone, respectively. Electrophoretic and immunoblot analysis of tryptic fragments revealed a highly sensitive cleavage site in all three cytochromes P450. N-Terminal sequencing was performed on the fragments after transfer onto poly(vinylidene difluoride) membranes and showed that this preferential cleavage site is at amino acid position 298 of P450 1A1, position 277 of P450 2B1, and position 278 of P450 2E1. Multiple sequence alignment revealed that these positions are at the amino terminal of a highly conserved region of these cytochromes P450. The important functional role implied by primary sequence conservation along with the proteolytic sensitivity at its amino terminal suggests that this region is a protein domain. Comparison with the known structure of the bacterial cytochrome P450cam predicts that this proteolytically sensitive site is within an interhelical turn region connected to the distal helix that partially encompasses the heme-containing active site. Substrate binding to the cleaved cytochromes P450 was examined in order to determine whether the newly added conformational freedom near the cleavage site functionally altered these cytochromes P450. Cleavage of P450 2B1 abolished benzphetamine binding, which indicates that the cleavage site contains an important structural determinant for binding this substrate. However, cleavage did not affect benzo[a]pyrene binding to P450 1A1.

Amino Acid Sequence

Rapidly evolving aphid gall effector proteins exhibit saposin-like folds.

Many insects manipulate plants by injecting effector proteins. In one extreme example of this molecular "hijacking," Hormaphis cornu aphids inject bicycle proteins into Hamamelis virginiana, contributing to the development of novel organs called galls. Bicycle proteins share no amino acid sequence similarity with proteins of known function. Here, we report the crystal structures of two divergent bicycle proteins. Both proteins contain saposin-like folds: one with multiple disulfide bonds exhibits a swapped domain topology; the other has no disulfide bonds and possesses two distinct, tandem domains. To explore the structural evolution of bicycle proteins, we attempted to predict bicycle protein structures with Alphafold2 (AF2) and other deep learning programs. While AF2 did not recover the two experimental structures using existing databases, it succeeded when provided with multiple sequence alignments (MSAs) of protein sequences from newly sequenced closely related species. Using this approach, we generated 2,400 high-confidence bicycle protein predictions from seven aphid species. While all aphid bicycle proteins contain predicted saposin-like folds, they display a vast diversity of structural and physicochemical properties. While this diversity thwarts prediction of conserved functions encoded in structure, it suggests that bicycle proteins have evolved to target diverse plant processes and/or to evade plant immune surveillance. Our extension of AF2 with custom MSAs of proteins from closely related species provides a generalizable, powerful approach for predicting structures of rapidly evolving protein families.

Animals

FUSE-PhyloTree: linking functions and sequence conservation modules of a protein family through phylogenomic analysis.

SUMMARY: FUSE-PhyloTree is a phylogenomic analysis software for identifying local sequence conservation associated with the different functions of a multi-functional (e.g. paralogous or multi-domain) protein family. FUSE-PhyloTree introduces an original approach that combines advanced sequence analysis with phylogenetic methods. First, local sequence conservation modules within the family are identified using partial local multiple sequence alignment. Next, the evolution of the detected modules and known protein functions is inferred within the family's phylogenetic tree using three-level phylogenetic reconciliation and ancestral state reconstruction. As a result, FUSE-PhyloTree provides a gene tree annotated with both predicted sequence modules and ancestral gene functions, enabling the association of functions with specific sequence regions based on their co-emergence. AVAILABILITY AND IMPLEMENTATION: FUSE-PhyloTree is provided as Docker and Singularity images including all the required software tools. Images, source code, test data, and documentation are available at https://github.com/OcMalde/fuse-phylotree and https://zenodo.org/records/15855068.

Phylogeny