Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,639 records · Page 91Linked to original sources

DIVERGE: phylogeny-based analysis for functional-structural divergence of a protein family.

SUMMARY: DetectIng Variability in Evolutionary Rates among GEnes (DIVERGE) is a software system to study functional divergence of a protein family by detecting site-specific change in evolutionary rate using a multiple alignment of amino acid sequences for a given phylogenetic tree. The program first conducts a statistical test for site-specific rate shifts along the tree, and predicting candidate amino acid residues responsible for functional divergence based on posterior analysis. These results can then be mapped on the 3D protein structure if available. AVAILABILITY: DIVERGE is available free of charge from http://xgu1.zool.iastate.edu/. Distribution packages for both Linux and Microsoft Windows operating systems are available, including manual and example files.

Amino Acid Sequence↗

aPhyloGeo: a Python application for correlating genetic and climatic conditions.

MOTIVATION: Environmental variation and its influence on genetic diversity is a central topic in evolutionary biology and phylogeography. Accurate correlations between genetic and climatic datasets to understand the genetic adaptations of different species to specific environments. It requires integrated and reproducible workflows. RESULTS: We developed aPhyloGeo, an open-source and multiplatform application implemented in Python, for investigating correlations between genetic variation and environmental data within a phylogenetic framework. The workflow integrates multiple analytical steps, including sequence alignment, sliding window phylogenetic inference, and statistical approaches such as the Mantel test and the Procrustean randomization test. These analyses enable the identification of mutation hotspots that exhibit strong associations with environmental variables. In addition, aPhyloGeo supports multicore data processing and provides a fully reproducible pipeline for evaluating localized relationships between genomic variation and climatic distributions. AVAILABILITY AND IMPLEMENTATION: aPhyloGeo is freely available on GitHub at: https://github.com/tahiri-lab/aPhyloGeo, as both a PyPI package and as Python scripts for Linux, macOS, and Windows.

Software↗

Distribution and molecular characterization of integron classes from Escherichia coli and Klebsiella pneumoniae isolates in Sulaymaniyah province of Iraq.

UNLABELLED: The environmental pollution from the misuse of antimicrobial drugs is fueling selection pressure in bacteria, thereby exacerbating the threat to global health. In Iraq, the situation is made worse by the poor implementation of the World Health Organization's Global Antimicrobial Resistance and Use Surveillance System (WHO-GLASS). Consequently, this study aimed to increase surveillance of the spread of antimicrobial resistance in Sulaymaniyah, Iraq. A total of 296 Enterobacteriaceae comprising 147 Klebsiella pneumoniae and 149 Escherichia coli were isolated from humans, poultry, and dairy farms. The isolates were screened using multiplex PCR to assess the prevalence of the clinically important integron integrase (intI) classes and antimicrobial resistance genes (ARGs) of commonly used antibiotics. Remarkably, 81.14% of the isolates carried at least 2 ARGs, 10.47% intI1, and 3.72% intI2. No intI3 was detected. A total of 663 ARGs were identified using multiplex PCR in the two Enterobacteriaceae: beta-lactamase genes were 43%, tetracycline resistance genes 25.20%, sulfonamide resistance gene 16.10%, quinolone resistance gene 10.2%, and aminoglycoside resistance genes 5.7%. K. pneumoniae harbored more integrons and ARGs than E. coli, thus posing a higher antimicrobial resistance threat in this province. This study underscores the importance of implementing more stringent WHO-GLASS and antibiotic stewardship to end the multidrug resistance crisis in Iraq. IMPORTANCE: These data are about the prevalence of integrons and resistance genes, helping to fill a significant gap in global surveillance efforts. Results can be used by global health authorities and the World Health Organization to develop national and international antimicrobial resistance (AMR) control strategies. The study is important because integrons are key genetic platforms that capture and disseminate antibiotic resistance genes among bacteria. In addition, Escherichia coli and Klebsiella spp. are among the top causes of hospital- and community-acquired infections, especially urinary tract infections, bloodstream infections, and pneumonia. Therefore, it will be riskier when these bacteria have a high rate of integrons and resistance genes because it impedes treatments during infection. Another importance of this study is that the study was carried out in Iraq. Iraq, like many low- and middle-income countries, faces challenges with unregulated antibiotic use, leading to high rates of AMR.

Escherichia coli↗

On the inference of parsimonious indel evolutionary scenarios.

Given a multiple alignment of orthologous DNA sequences and a phylogenetic tree for these sequences, we investigate the problem of reconstructing a most parsimonious scenario of insertions and deletions capable of explaining the gaps observed in the alignment. This problem, called the Indel Parsimony Problem, is a crucial component of the problem of ancestral genome reconstruction, and its solution provides valuable information to many genome functional annotation approaches. We first show that the problem is NP-complete. Second, we provide an algorithm, based on the fractional relaxation of an integer linear programming formulation. The algorithm is fast in practice, and the solutions it produces are, in most cases, provably optimal. We describe a divide-and-conquer approach that makes it possible to solve very large instances on a simple desktop machine, while retaining guaranteed optimality. Our algorithms are tested and shown efficient and accurate on a set of 1.8 Mb mammalian orthologous sequences in the CFTR region.

Algorithms↗

Subsets with restricted immunoglobulin gene rearrangement features indicate a role for antigen selection in the development of chronic lymphocytic leukemia.

We recently identified a chronic lymphocytic leukemia (CLL) subgroup using the immunoglobulin variable heavy-chain (V(H)) gene V(H)3-21 with almost identical heavy-chain complementarity determining region 3s (HCDR3s) and preferential variable light-chain (V(L)) gene usage, suggesting recognition of a common antigen epitope in this subset. To further explore the B-cell receptors (BCRs) in CLL, we characterized 407 V(H) rearrangements amplified from 346 CLLs regarding V(H), diversity (D), and joining (J(H)) gene usage and performed multiple alignment of the HCDR3 sequences. These analyses revealed 3 small subsets (2 V(H)1-69 groups, 7 cases; and 1 V(H)1-2 group, 5 cases) with highly restricted HCDR3 features including identical V(H)/D/J(H) usage, HCDR3 lengths, and shared N-sequences, in addition to the V(H)3-21 group (22 cases). Furthermore, another 3 groups (9 V(H)1-3(+) cases, 3 V(H)1-18(+) cases, and 5 V(H)4-39(+) cases) had essentially identical V(H)/D/J(H) use and similar HCDR3 lengths but less conserved N-regions. Analysis in all 6 of these subgroups showed restriction in V(L) gene use, whereas no association between V(H) and V(L) usage was found in cases without HCDR3 similarities. Altogether, structurally similar HCDR3s associated with preferential V(L) gene usage implies selection of BCRs, especially in subsets showing high HCDR3 similarities, thus pointing to restricted antigen recognition sites and possibly involvement of specific antigens in CLL development.

Adult↗

A web-based program for the prediction of average hydropathy, average amphipathicity and average similarity of multiply aligned homologous proteins.

We designed a web-based program, AveHAS, to determine and plot the average hydropathy, average amphipathicity and average similarity for a clustal X-derived multiple alignment of homologous protein sequences. This method is based on the TREEMOMENT and Hydro programs. It has a user-friendly interface, a convenient input format and an improved algorithm.

Algorithms↗

[Development of "Amplisens-HCV-genotype" reagent set for identification of hepatitis C virus genotypes 1a, 1b, 2a and 3a].

Multiple alignments of 119 nucleotide sequences of isolates of hepatitis C virus (HCV) were carried out to choose the type-specific primers for the 5'-ultra-core fragment of viral genome for the purpose of detecting the HCV 1a, 1b, 2a, and 3a subtypes. A PCR kit of reagents was designed for the amplification of cDNA HCV with selected type-specific primers and for making the electrophoresis in agarous gel. The kit comprises the positive control samples, i.e. HCV genome fragments, subtypes 1a, 1b, 2a and 3a, cloned in the plasmid vector. 440 cDNAHCV samples were simultaneously tested by using the worked out reagents' set and according to the method of Ohno et al. The results were found to be concordant in 336 cases, and were discordant in 4 samples. A sequencing of the PCR products and phylogenetic analysis showed that 1 sample belonged to subtype 4a, 2 samples belonged to subtypes 2k and 1 sample--to subtype 31.

DNA Primers↗

CBCAnalyzer: inferring phylogenies based on compensatory base changes in RNA secondary structures.

The CBCAnalyzer (CBC=compensatory base change) is a custom written software toolbox consisting of three parts, CTTransform, CBCDetect, and CBCTree. CTTransform reads several ct-file formats, and generates a so called "bracket-dot-bracket" format that typically is used as input for other tools such as RNAforester, RNAmovie or MARNA. The latter one creates a multiple alignment based on primary sequences and secondary structures that now can be used as input for CBCDetect. CBCDetect counts CBCs in all against all of the aligned sequences. This is important in detecting species that are discriminated by their sexual incompatibility. The count (distance) matrix obtained by CBCDetect is used as input for CBCTree that reconstructs a phylogram by using the algorithm of BIONJ. In this note we describe the features of the toolbox as well as application examples. The toolbox provides a graphical user interface. It is written in C++ and freely available at: http://cbcanalyzer.bioapps.biozentrum.uni-wuerzburg.de.

Algorithms↗

[Model of genes expression regulation in bacteria by means of formation of secondary RNA structures].

In this article a model, first, classical attenuation RNA regulation of gene expression by means of transcription termination is offered. The model bases on representation about a macrostate of secondary structure in RNA regulatory region between a ribosome and a RNA polymerase, on the formulas of a resonant type defining the value of deceleration of a RNA polymerase by a set of hairpins in the same region. The special attention is given to selection of parameters of model. To check of model the computer simulation is carried out and the dependences of transcription termination probability from the value of concentration charged tRNA are obtained, in particular, and from concentration of amino acid for many regulatory regions in genomes of bacteria (here data are presented for trpE genes in Streptomyces spp., Bradyrhizobium japonicum and Escherichia coli) and at various values of three parameters, which authors consider as the main. The obtained dependences are compounded with the accessible experimental data; including, under the form of the graphs concerning to activity of an enzyme depending on concentration of amino acid (for example, anthranilate synthase from tryptophan in S. venezuela). One possible usage: now attenuation is predicted usually by means of multiple alignment, it needs some sequences; the obtaining with the help of model on an individual sequence characteristic for attenuation or its absence of a curve at approaching parameters could be considered as argument for the benefit of presence or absence of attenuation.

Bacteria↗

Comparable gene structure of the immunoglobulin heavy chain variable region between multiple myeloma and normal bone marrow lymphocytes.

To characterize multiple myeloma (MM) from the viewpoint of the immunoglobulin (Ig) gene structure, we compared the transcripts of the Ig heavy chain variable region from 23 MM samples with 221 clones of the gamma, alpha and mu chain transcripts amplified by the reverse transcriptase-polymerase chain reaction (RT-PCR) from normal bone marrow (BM) cells. The usage of D and JH gene segments and the length of the N regions were the same between MM and the normal gamma, alpha and mu transcripts. Compared with the known germline VH genes, the frame work regions (FWRs) and complementarity determining regions (CDRs) of the VH segments mutated at rates of 8.3 +/- 4.7% and 15.9 +/- 7.7%, respectively, which were the same as the normal gamma and alpha (gamma/alpha) transcripts and higher than the normal mu transcripts. The replacement/silent (R/S) ratios of the mutations in FWRs and CDRs were 1.9 +/- 1.3 and 2.7 +/- 1.8, respectively, which were the same as the gamma/alpha and mu transcripts. On the other hand, we detected the clone-specific mu transcripts by RT-PCR using the primers corresponding to the each respective CDR-III and the constant region of the mu chain in three of the studied six MM samples, suggesting the involvement of a pre-switched B cell in some cases of MM. These findings suggested that the cellular origin of MM is heterogeneous, but that the Ig structure in MM reflects normal B cell maturation to plasma cell through mutation and selection.

Adult↗

On global sequence alignment.

We present a dynamic programming algorithm for computing a best global alignment of two sequences. The proposed algorithm is robust in identifying any of several global relationships between two sequences. The algorithm delivers a best alignment of two sequences in linear space and quadratic time. We also describe a multiple alignment algorithm based on the pairwise algorithm. Both algorithms have been implemented as portable C programs. Experimental results indicate that for a commonly used set of gap penalties, the new programs produce more satisfactory alignments on sequences of various lengths than some existing pairwise and multiple programs based on the dynamic programming algorithm of Needleman and Wunsch.

Algorithms↗

Fowlpox virus encodes a protein related to human deoxycytidine kinase: further evidence for independent acquisition of genes for enzymes of nucleotide metabolism by different viruses.

It is demonstrated that fowlpox virus (FPV) protein FP26 located in the HindIII D fragment of the genome is related to the human deoxycytidine kinase (dCK) and probably possesses the same enzymatic activity. A homologous protein is not encoded by vaccinia virus. A multiple alignment of the amino acid sequences of the human and FPV dCKs, the thymidine kinases (TK) of herpesviruses, and cellular and vaccinia virus thymidylate kinases (ThyK) was generated and the conserved motifs, at least two of which are implicated in ATP binding, were characterized. An apparent duplication of ATP-binding motif B in the dCKs was revealed, leading to the reassignment of one of the catalytic residues. Phylogenetic analysis based on the multiple alignment suggested that the putative dCK of FPV probably has diverged from the common ancestor with the human dCK at a later stage of evolution than the herpesvirus TKs, with the ThyKs being peripheral members of the family. These results are compatible with hypothesis that genes for enzymes of nucleotide metabolism could be acquired independently by different DNA viruses (Koonin, E.V. and Senkevich, T.G., Virus Genes 6:187-196, 1992).

Amino Acid Sequence↗

Comparative modeling of CASP4 target proteins: combining results of sequence search with three-dimensional structure assessment.

Comparative modeling aims at constructing molecular models for proteins of unknown structure, by using known structures of related proteins as templates. To test the comparative modeling approach reported here, predictions for 13 target proteins were submitted during the fourth round of "blind" protein structure prediction experiment (CASP4; http://PredictionCenter.llnl.gov/casp4). Sequence identity between these target proteins and the closest known structures ranged from 13 to 58%, indicating a broad spectrum of prediction difficulty. Although this broad difficulty range required addressing a variety of issues, the most important proved to be sequence-structure alignment for distant homology targets. The alignment step was based on structure-based evaluation of alignment variants produced mainly with PSI-BLAST intermediate sequence search procedure (PSI-BLAST-ISS). Although a fraction of correctly aligned residues in resulting models was markedly better than the average in all cases, for distant homology targets it was still considerably below the estimated achievable level. Results with CASP4 targets show that, along with the correctness of sequence-structure alignments, effective use of multiple template structures may significantly increase accuracy of the model structure. Improvement in this area should also result in more accurate loop modeling and side-chain prediction.

Bacterial Proteins↗

A cSNP map and database for human chromosome 21.

Single nucleotide polymorphisms (SNPs) are likely to contribute to the study of complex genetic diseases. The genomic sequence of human chromosome 21q was recently completed with 225 annotated genes, thus permitting efficient identification and precise mapping of potential cSNPs by bioinformatics approaches. Here we present a human chromosome 21 (HC21) cSNP database and the first chromosome-specific cSNP map. Potential cSNPs were generated using three approaches: (1) Alignment of the complete HC21 genomic sequence to cognate ESTs and mRNAs. Candidate cSNPs were automatically extracted using a novel program for context-dependent SNP identification that efficiently discriminates between true variation, poor quality sequencing, and paralogous gene alignments. (2) Multiple alignment of all known HC21 genes to all other human database entries. (3) Gene-targeted cSNP discovery. To date we have identified 377 cSNPs averaging ~1 SNP per 1.5 kb of transcribed sequence, covering 65% of known genes in the chromosome. Validation of our bioinformatics approach was demonstrated by a confirmation rate of 78% for the predicted cSNPs, and in total 32% of the cSNPs in our database have been confirmed. The database is publicly available at http://csnp.unige.ch or http://csnp.isb-sib.ch. These SNPs provide a tool to study the contribution of HC21 loci to complex diseases such as bipolar affective disorder and allele-specific contributions to Down syndrome phenotypes.

Base Composition↗

The genome of model malaria parasites, and comparative genomics.

The field of comparative genomics of malaria parasites has recently come of age with the completion of the whole genome sequences of the human malaria parasite Plasmodium falciparum and a rodent malaria model, Plasmodium yoelii yoelii. With several other genome sequencing projects of different model and human malaria parasite species underway, comparing genomes from multiple species has necessitated the development of improved informatics tools and analyses. Results from initial comparative analyses reveal striking conservation of gene synteny between malaria species within conserved chromosome cores, in contrast to reduced homology within subtelomeric regions, in line with previous findings on a smaller scale. Genes that elicit a host immune response are frequently found to be species-specific, although a large variant multigene family is common to many rodent malaria species and Plasmodium vivax. Sequence alignment of syntenic regions from multiple species has revealed the similarity between species in coding regions to be high relative to non-coding regions, and phylogenetic footprinting studies promise to reveal conserved motifs in the latter. Comparison of non-synonymous substitution rates between orthologous genes is proving a powerful technique for identifying genes under selection pressure, and may be useful for vaccine design. This is a stimulating time for comparative genomics of model and human malaria parasites, which promises to produce useful results for the development of antimalarial drugs and vaccines.

Animals↗

A comprehensive approach to clustering of expressed human gene sequence: the sequence tag alignment and consensus knowledge base.

The expressed human genome is being sequenced and analyzed by disparate groups producing disparate data. The majority of the identified coding portion is in the form of expressed sequence tags (ESTs). The need to discover exonic representation and expression forms of full-length cDNAs for each human gene is frustrated by the partial and variable quality nature of this data delivery. A highly redundant human EST data set has been processed into integrated and unified expressed transcript indices that consist of hierarchically organized human transcript consensi reflecting gene expression forms and genetic polymorphism within an index class. The expression index and its intermediate outputs include cleaned transcript sequence, expression, and alignment information and a higher fidelity subset, SANIGENE. The STACK_PACK clustering system has been applied to dbEST release 121598 (GenBank version 110). Sixty-four percent of 1,313, 103 Homo sapiens ESTs are condensed into 143,885 tissue level multiple sequence clusters; linking through clone-ID annotations produces 68,701 total assemblies, such that 81% of the original input set is captured in a STACK multiple sequence or linked cluster. Indexing of alignments by substituent EST accession allows browsing of the data structure and its cross-links to UniGene. STACK metaclusters consolidate a greater number of ESTs by a factor of 1. 86 with respect to the corresponding UniGene build. Fidelity comparison with genome reference sequence AC004106 demonstrates consensus expression clusters that reflect significantly lower spurious repeat sequence content and capture alternate splicing within a whole body index cluster and three STACK v.2.3 tissue-level clusters. Statistics of a staggered release whole body index build of STACK v.2.0 are presented.

Algorithms↗