Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 973 records · Page 54Linked to original sources

FunnyBase: a systems level functional annotation of Fundulus ESTs for the analysis of gene expression.

BACKGROUND: While studies of non-model organisms are critical for many research areas, such as evolution, development, and environmental biology, they present particular challenges for both experimental and computational genomic level research. Resources such as mass-produced microarrays and the computational tools linking these data to functional annotation at the system and pathway level are rarely available for non-model species. This type of "systems-level" analysis is critical to the understanding of patterns of gene expression that underlie biological processes. RESULTS: We describe a bioinformatics pipeline known as FunnyBase that has been used to store, annotate, and analyze 40,363 expressed sequence tags (ESTs) from the heart and liver of the fish, Fundulus heteroclitus. Primary annotations based on sequence similarity are linked to networks of systematic annotation in Gene Ontology (GO) and the Kyoto Encyclopedia of Genes and Genomes (KEGG) and can be queried and computationally utilized in downstream analyses. Steps are taken to ensure that the annotation is self-consistent and that the structure of GO is used to identify higher level functions that may not be annotated directly. An integrated framework for cDNA library production, sequencing, quality control, expression data generation, and systems-level analysis is presented and utilized. In a case study, a set of genes, that had statistically significant regression between gene expression levels and environmental temperature along the Atlantic Coast, shows a statistically significant (P < 0.001) enrichment in genes associated with amine metabolism. CONCLUSION: The methods described have application for functional genomics studies, particularly among non-model organisms. The web interface for FunnyBase can be accessed at http://genomics.rsmas.miami.edu/funnybase/super_craw4/. Data and source code are available by request at jpaschall@bioinfobase.umkc.edu.

Animals↗

Identification and characterization of Crumbs homolog 2 gene at human chromosome 9q33.3.

Drosophila Crumbs (Crb)--Stardust (Sdt)--Discs lost (Dlt) complex plays a pivotal role in the establishment and the maintenance of epithelial polarity. CRB1 and CRB3 are human homologs of Drosophila Crb, MPP1-MPP7 are human homologs of Drosophila Sdt, INADL/PATJ and MPDZ/MUPP1 are human homologs of Drosophila Dlt. Here, we identified and characterized a novel Crumbs family gene, Crumbs homolog 2 (CRB2), by using bioinformatics. CRB2 isoform 1 was assembled by adding nucleotide position 1-3353 of FLJ16786 cDNA (AK123000) to the 5'-end of 5'-truncated FLJ38464 cDNA (NM_173689.1), while that of CRB2 isoform 2 was derived from FLJ16786 cDNA. CRB2 isoform 1, consisting of exon 1-13, encoded a 1285-aa transmembrane protein. CRB2 isoform 2, consisting of exon 1-10 and intron 10, encoded a 1176-aa secreted protein. CRB2 gene was found to encode transmembrane protein as well as secreted protein due to alternative splicing. CRB2 isoform 1, showing 24.4% total amino-acid identity with CRB1, was type I transmembrane protein with 14 extracellular EGF-like domains, 3 extracellular Laminin G-like domains and the Crb cytoplasmic tail (CCT) domain. CCT domain, functioning as the binding site for PDZ domain of Sdt homologs, was conserved among human CRB1, CRB2, CRB3, mouse Crb1, Crb3, Drosophila Crb, and C. elegans crb. Comparative genomics revealed that CRB2-KIAA1608-LHX2-NEK6 locus at human chromosome 9q33.3 and CRB1-MGC27044-LHX9-NEK7 locus at human chromosome 1q31.3 were paralogous regions within the human genome. This is the first report on identification and characterization of the CRB2 gene.

Alternative Splicing↗

Comparative genomics on mammalian Fgf3-Fgf4 locus.

The CCND1-ORAOV1-FGF19-FGF4-FGF3-TMEM16A-FADD-PPFIA1-CTTN (EMS1) locus at human chromosome 11q13.3 is amplified in head and neck tumors, esophageal cancer, Kaposi's sarcoma, bladder tumors, breast cancer, and liver cancer. Fgf4 mRNA is expressed in embryonic stem (ES) cells depending on Sox2 and Pou5f1 (Oct3/Oct4) transcription factors, and in myotomes and limb bud AER depending on MyoD (or Myf5) and GATA transcription factors. Here, rat Fgf3 and Fgf4 complete coding sequences were determined by using bioinformatics. Multiple errors, including one-base insertion and 22-base deletion, were identified within the coding region of rat Fgf4 RefSeq (NM_053809.1 or AB079673.1). Rat Fgf3 and Fgf4 genes, consisting of three exons, were clustered in tail-to-head manner with an interval of about 16 kb. CUTL1 (CCAAT-displacement protein, CDP) and NKX2-5 binding sites and TATA box within 5'-flanking promoter region were conserved among human, rat and mouse Fgf3 orthologs. MYOD and MYOG (Myogenin) binding sites and TATA box within 5'-flanking promoter region as well as GATA, MYOD, SOX2 and POU5F1 binding sites within exon 3 were conserved among mammalian Fgf4 orthologs. Human FGF3 and FGF4 genes were clustered in tail-to-head manner with an interval of about 35 kb. Major repetitive sequence (FGF34Rep1) and minor repetitive sequence (FGF34Rep2) were identified within human FGF3-FGF4 gene cluster. FGF34Rep1 were clustered within the FGF3-FGF4 locus as well as around the IL28RA locus (1p36.11) and the NFAM1 locus (22q13.2). FGF34Rep2 was characterized by the CCA(T/C) repeats. This is the first report on comparative genomics analyses on the Fgf3-Fgf4 locus within human, rat and mouse genomes.

Amino Acid Sequence↗

BBP: Brucella genome annotation with literature mining and curation.

BACKGROUND: Brucella species are Gram-negative, facultative intracellular bacteria that cause brucellosis in humans and animals. Sequences of four Brucella genomes have been published, and various Brucella gene and genome data and analysis resources exist. A web gateway to integrate these resources will greatly facilitate Brucella research. Brucella genome data in current databases is largely derived from computational analysis without experimental validation typically found in peer-reviewed publications. It is partially due to the lack of a literature mining and curation system able to efficiently incorporate the large amount of literature data into genome annotation. It is further hypothesized that literature-based Brucella gene annotation would increase understanding of complicated Brucella pathogenesis mechanisms. RESULTS: The Brucella Bioinformatics Portal (BBP) is developed to integrate existing Brucella genome data and analysis tools with literature mining and curation. The BBP InterBru database and Brucella Genome Browser allow users to search and analyze genes of 4 currently available Brucella genomes and link to more than 20 existing databases and analysis programs. Brucella literature publications in PubMed are extracted and can be searched by a TextPresso-powered natural language processing method, a MeSH browser, a keywords search, and an automatic literature update service. To efficiently annotate Brucella genes using the large amount of literature publications, a literature mining and curation system coined Limix is developed to integrate computational literature mining methods with a PubSearch-powered manual curation and management system. The Limix system is used to quickly find and confirm 107 Brucella gene mutations including 75 genes shown to be essential for Brucella virulence. The 75 genes are further clustered using COG. In addition, 62 Brucella genetic interactions are extracted from literature publications. These results make possible more comprehensive investigation of Brucella pathogenesis. Other BBP features include publication email alert service, Brucella researchers' contact database, and discussion forum. CONCLUSION: BBP is a gateway for Brucella researchers to search, analyze, and curate Brucella genome data originated from public databases and literature. Brucella gene mutations and genetic interactions are annotated using Limix leading to better understanding of Brucella pathogenesis.

Algorithms↗

Automatic recognition of topic-classified relations between prostate cancer and genes using MEDLINE abstracts.

BACKGROUND: Automatic recognition of relations between a specific disease term and its relevant genes or protein terms is an important practice of bioinformatics. Considering the utility of the results of this approach, we identified prostate cancer and gene terms with the ID tags of public biomedical databases. Moreover, considering that genetics experts will use our results, we classified them based on six topics that can be used to analyze the type of prostate cancers, genes, and their relations. METHODS: We developed a maximum entropy-based named entity recognizer and a relation recognizer and applied them to a corpus-based approach. We collected prostate cancer-related abstracts from MEDLINE, and constructed an annotated corpus of gene and prostate cancer relations based on six topics by biologists. We used it to train the maximum entropy-based named entity recognizer and relation recognizer. RESULTS: Topic-classified relation recognition achieved 92.1% precision for the relation (an increase of 11.0% from that obtained in a baseline experiment). For all topics, the precision was between 67.6 and 88.1%. CONCLUSION: A series of experimental results revealed two important findings: a carefully designed relation recognition system using named entity recognition can improve the performance of relation recognition, and topic-classified relation recognition can be effectively addressed through a corpus-based approach using manual annotation and machine learning techniques.

Abstracting and Indexing↗

Identification and characterization of ASXL2 gene in silico.

Drosophila Asx is a Polycomb group gene. Because Drosophila Asx mutations exhibit anterior and posterior transformations, Drosophila Asx is one of the ETP (Enhancers of trithorax and Polycomb) genes with dual functions in transcriptional activation and silencing. ASXL1 is one of human homologs of Drosophila Asx. Here, we searched for ASXL1-related gene within the human genome by using bioinformatics, and identified the ASXL2 gene. Nucleotide sequence of human ASXL2 cDNA was determined by assembling the nucleotide sequences of human EST AI797346, and partial cDNAs MGC44431 (BC042999) and KIAA1685 (AB051472). Nucleotide sequence of mouse Asxl2 was derived from uncharacterized mouse cDNA 9930017F14 (AK036839). Human ASXL2 (1435 aa) showed 79.4% total-amino-acid identity with mouse Asxl2 (1370 aa), and 29.8% total-amino-acid identity with human ASXL1. ASXN domain (codon 1-86 of ASXL2), ASXM domain (codon 269-380 of ASXL2), and PHD domain (codon 1400-1431 of ASXL2) were conserved between human ASXL2 and ASXL1. Human ASXL2 gene, consisting of at least 13 exons, was mapped to human chromosome 2p23.3, one of recombination hot spots or fragile sites associated with carcinogenesis. The DNMT3A-ASXL2-KIF3C locus on human chromosome 2p23.3 and the DNMT3B-ASXL1-KIF3B locus on human chromosome 20q11.21 were paralogous regions within the human genome. Polycomb group and trithorax group proteins are implicated in embryogenesis and carcinogenesis due to transcriptional regulation of target genes through histone modification and chromatin remodeling. Based on functional conservation and human chromosomal localization, ASXL2 and ASXL1 genes were predicted cancer-associated genes.

Amino Acid Sequence↗

Comprehensive in silico functional specification of mouse retina transcripts.

BACKGROUND: The retina is a well-defined portion of the central nervous system (CNS) that has been used as a model for CNS development and function studies. The full specification of transcripts in an individual tissue or cell type, like retina, can greatly aid the understanding of the control of cell differentiation and cell function. In this study, we have integrated computational bioinformatics and microarray experimental approaches to classify the tissue specificity and developmental distribution of mouse retina transcripts. RESULTS: We have classified a set of retina-specific genes using sequence-based screening integrated with computational and retina tissue-specific microarray approaches. 33,737 non-redundant sequences were identified as retina transcript clusters (RTCs) from more than 81,000 mouse retina ESTs. We estimate that about 19,000 to 20,000 genes might express in mouse retina from embryonic to adult stages. 39.1% of the RTCs are not covered by 60,770 RIKEN full-length cDNAs. Through comparison with 2 million mouse ESTs, spectra of neural, retinal, late-generated retinal, and photoreceptor -enriched RTCs have been generated. More than 70% of these RTCs have data from biological experiments confirming their tissue-specific expression pattern. The highest-grade retina-enriched pool covered almost all the known genes encoding proteins involved in photo-transduction. CONCLUSION: This study provides a comprehensive mouse retina transcript profile for further gene discovery in retina and suggests that tissue-specific transcripts contribute substantially to the whole transcriptome.

Animals↗

Statistical potential-based amino acid similarity matrices for aligning distantly related protein sequences.

Aligning distantly related protein sequences is a long-standing problem in bioinformatics, and a key for successful protein structure prediction. Its importance is increasing recently in the context of structural genomics projects because more and more experimentally solved structures are available as templates for protein structure modeling. Toward this end, recent structure prediction methods employ profile-profile alignments, and various ways of aligning two profiles have been developed. More fundamentally, a better amino acid similarity matrix can improve a profile itself; thereby resulting in more accurate profile-profile alignments. Here we have developed novel amino acid similarity matrices from knowledge-based amino acid contact potentials. Contact potentials are used because the contact propensity to the other amino acids would be one of the most conserved features of each position of a protein structure. The derived amino acid similarity matrices are tested on benchmark alignments at three different levels, namely, the family, the superfamily, and the fold level. Compared to BLOSUM45 and the other existing matrices, the contact potential-based matrices perform comparably in the family level alignments, but clearly outperform in the fold level alignments. The contact potential-based matrices perform even better when suboptimal alignments are considered. Comparing the matrices themselves with each other revealed that the contact potential-based matrices are very different from BLOSUM45 and the other matrices, indicating that they are located in a different basin in the amino acid similarity matrix space.

Amino Acids↗

EMBL Nucleotide Sequence Database: developments in 2005.

The EMBL Nucleotide Sequence Database (www.ebi.ac.uk/embl) at the EMBL European Bioinformatics Institute, UK, offers a comprehensive set of publicly available nucleotide sequence and annotation, freely accessible to all. Maintained in collaboration with partners DDBJ and GenBank, coverage includes whole genome sequencing project data, directly submitted sequence, sequence recorded in support of patent applications and much more. The database continues to offer submission tools, data retrieval facilities and user support. In 2005, the volume of data offered has continued to grow exponentially. In addition to the newly presented data, the database encompasses a range of new data types generated by novel technologies, offers enhanced presentation and searchability of the data and has greater integration with other data resources offered at the EBI and elsewhere. In stride with these developing data types, the database has continued to develop submission and retrieval tools to maximise the information content of submitted data and to offer the simplest possible submission routes for data producers. New developments, the submission process, data retrieval and access to support are presented in this paper, along with links to sources of further information.

Animals↗

An open-source clinical bioinformatics pipeline for real-world NGS implementation: translating genomic variants into actionable treatment strategies in oncology.

BACKGROUND: Next-Generation Sequencing (NGS) has become a cornerstone technology in clinical practice, yet its adoption presents significant challenges. Physicians and oncologists must manage vast amounts of genome-scale data and transform it into actionable insights for complex decision-making. While commercial systems exist to synthesize data from NGS experiments into clinical reports, many are hindered by limitations such as closed-source designs that restrict transparency and customization. Additionally, some fail to leverage publicly available genomic databases, missing opportunities to integrate valuable external data. Furthermore, the rigidity of many tools in accommodating diverse NGS panels limits their applicability across varied clinical scenarios. METHODS: To address these limitations, we developed OncoReport, an open-source tool that generates comprehensive reports from NGS analyses. By integrating publicly accessible databases, OncoReport provides a robust, user-friendly environment equipped with essential tools for NGS analysis. This design aims to enhance data interpretation and support informed clinical decision-making. RESULTS: Rigorous testing has demonstrated OncoReport&#x2019;s effectiveness in producing detailed, actionable reports that are clear and easy to use. By automating key aspects of the workflow, the tool significantly reduces manual effort and expedites the synthesis and interpretation of NGS results, making genomic insights more accessible to clinicians. CONCLUSION: OncoReport offers a transparent, flexible, and efficient framework for clinicians to analyze and apply genomic data in patient care. By streamlining workflows and leveraging open-source principles, it empowers healthcare professionals to make informed, data-driven decisions. OncoReport is freely available at https://oncoreport.atlas.dmi.unict.it, with source code and issue tracking on GitHub: https://github.com/knowmics-lab/oncoreport .

Humans↗

Molecular classification of liver cirrhosis in a rat model by proteomics and bioinformatics.

Liver cirrhosis is a worldwide health problem. Reliable, noninvasive methods for early detection of liver cirrhosis are not available. Using a three-step approach, we classified sera from rats with liver cirrhosis following different treatment insults. The approach consisted of: (i) protein profiling using surface-enhanced laser desorption/ionization (SELDI) technology; (ii) selection of a statistically significant serum biomarker set using machine learning algorithms; and (iii) identification of selected serum biomarkers by peptide sequencing. We generated serum protein profiles from three groups of rats: (i) normal (n=8), (ii) thioacetamide-induced liver cirrhosis (n=22), and (iii) bile duct ligation-induced liver fibrosis (n=5) using a weak cation exchanger surface. Profiling data were further analyzed by a recursive support vector machine algorithm to select a panel of statistically significant biomarkers for class prediction. Sensitivity and specificity of classification using the selected protein marker set were higher than 92%. A consistently down-regulated 3495 Da protein in cirrhosis samples was one of the selected significant biomarkers. This 3495 Da protein was purified on-chip and trypsin digested. Further structural characterization of this biomarkers candidate was done by using cross-platform matrix-assisted laser desorption/ionization mass spectrometry (MALDI-MS) peptide mass fingerprinting (PMF) and matrix-assisted laser desorption/ionization time of flight/time of flight (MALDI-TOF/TOF) tandem mass spectrometry (MS/MS). Combined data from PMF and MS/MS spectra of two tryptic peptides suggested that this 3495 Da protein shared homology to a histidine-rich glycoprotein. These results demonstrated a novel approach to discovery of new biomarkers for early detection of liver cirrhosis and classification of liver diseases.

Algorithms↗

Bioinformatic discovery of microRNA precursors from human ESTs and introns.

BACKGROUND: MicroRNAs (miRNAs) function in many physiological processes, and their discovery is beneficial for further studying their physiological functions. However, many of the miRNAs predicted from genomic sequences have not been experimentally validated to be authentic expressed RNA transcripts, thereby decreasing the reliability of miRNA discovery. To overcome this problem, we examined expressed transcripts - ESTs and intronic sequences - to identify novel miRNAs as well as their target genes. RESULTS: To facilitate our approach, we developed our scanning method using criteria based on the features of 207 known human pre-miRNAs to discriminate miRNAs from random sequences. We identified 208 candidate hairpins in human ESTs and human reference gene intronic sequences, 52 of which are known pre-miRNAs. The discovery pipeline performance was further assessed using 130 newly updated pre-miRNA and randomly selected sequences. We achieved sensitivity of 85% (110/130) and overall specificity of 49.7% using this method. Because miRNAs are evolutionarily conserved regulators of gene expression, it is expected that their host genes and target genes should have respective phylogenetic orthologs. Our results confirmed that, in certain mammals, the host genes carrying the same miRNAs are orthologs, as previously reported. Moreover, this observation is also the case for some of the miRNA target genes. CONCLUSION: We have predicted 208 human pre-miRNA candidates and over 10,000 putative human target genes. Using sequence information from ESTs and introns ensures that the predicted pre-miRNA candidates are expressed and the combined expression transcription information from ESTs and introns makes our prediction results more decisive with regard to expressed pre-miRNAs.

3' Untranslated Regions↗

A strategy for the rapid identification of phosphorylation sites in the phosphoproteome.

Edman phosphate ((32)P) release sequencing provides a high sensitivity means of identifying phosphorylation sites in proteins that complements mass spectrometry techniques. We have developed a bioinformatic assessment tool, the cleavage of radiolabeled protein (CRP) program, which enables experimental identification of phosphorylation sites via (32)P labeling and Edman degradation of cleaved proteins obtained at femtomole levels. By observing the Edman cycle(s) in which radioactivity is found, candidate phosphorylation sites are identified by determining which residues occur at the observed number of cycles downstream from a peptide cleavage site. In cases where more than one residue could be responsible for the observed radioactivity, additional experiments with cleavage reagents having alternative specificities may resolve the ambiguity. Given a protein sequence and a cleavage site, CRP performs these experiments in silico, identifying resolved sites based on user-supplied experimental data, as well as suggesting combinations of reagents for additional analyses. Analysis of the PhosphoBase protein sequence database suggests that CRP data from two cleavage experiments can be used to identify unambiguously 60% of known phosphorylation sites. Data from additional cleavage experiments may increase the overall coverage to 70% of known sites. By comparing theoretical data obtained from the CRP program with (32)P release data obtained from an Edman sequencer, a known phosphorylation site was identified unambiguously and correctly. In addition, our results show that in vivo phosphorylation sites can be determined routinely by differential proteolysis analysis and Edman cycling with less than 1 fmol of protein and 1000 cpm.

Amino Acid Sequence↗

Are grammatical representations useful for learning from biological sequence data?--a case study.

This paper investigates whether Chomsky-like grammar representations are useful for learning cost-effective, comprehensible predictors of members of biological sequence families. The Inductive Logic Programming (ILP) Bayesian approach to learning from positive examples is used to generate a grammar for recognising a class of proteins known as human neuropeptide precursors (NPPs). Collectively, five of the co-authors of this paper, have extensive expertise on NPPs and general bioinformatics methods. Their motivation for generating a NPP grammar was that none of the existing bioinformatics methods could provide sufficient cost-savings during the search for new NPPs. Prior to this project experienced specialists at SmithKline Beecham had tried for many months to hand-code such a grammar but without success. Our best predictor makes the search for novel NPPs more than 100 times more efficient than randomly selecting proteins for synthesis and testing them for biological activity. As far as these authors are aware, this is both the first biological grammar learnt using ILP and the first real-world scientific application of the ILP Bayesian approach to learning from positive examples. A group of features is derived from this grammar. Other groups of features of NPPs are derived using other learning strategies. Amalgams of these groups are formed. A recognition model is generated for each amalgam using C4.5 and C4.5rules and its performance is measured using both predictive accuracy and a new cost function, Relative Advantage (RA). The highest RA was achieved by a model which includes grammar-derived features. This RA is significantly higher than the best RA achieved without the use of the grammar-derived features. Predictive accuracy is not a good measure of performance for this domain because it does not discriminate well between NPP recognition models: despite covering varying numbers of (the rare) positives, all the models are awarded a similar (high) score by predictive accuracy because they all exclude most of the abundant negatives.

Bayes Theorem↗

Anatomics: the intersection of anatomy and bioinformatics.

Computational resources are now using the tissue names of the major model organisms so that tissue-associated data can be archived in and retrieved from databases on the basis of developing and adult anatomy. For this to be done, the set of tissues in that organism (its anatome) has to be organized in a way that is computer-comprehensible. Indeed, such formalization is a necessary part of what is becoming known as systems biology, in which explanations of high-level biological phenomena are not only sought in terms of lower-level events, but are articulated within a computational framework. Lists of tissue names alone, however, turn out to be inadequate for this formalization because tissue organization is essentially hierarchical and thus cannot easily be put into tables, the natural format of relational databases. The solution now adopted is to organize the anatomy of each organism as a hierarchy of tissue names and linking relationships (e.g. the tibia is PART OF the leg, the tibia IS-A bone) within what are known as ontologies. In these, a unique ID is assigned to each tissue and this can be used within, for example, gene-expression databases to link data to tissue organization, and also used to query other data sources (interoperability), while inferences about the anatomy can be made within the ontology on the basis of the relationships. There are now about 15 such anatomical ontologies, many of which are linked to organism databases; these ontologies are now publicly available at the Open Biological Ontologies website (http://obo.sourceforge.net) from where they can be freely downloaded and viewed using standard tools. This review considers how anatomy is formalized within ontologies, together with the problems that have had to be solved for this to be done. It is suggested that the appropriate term for the analysis, computer formulation and use of the anatome is anatomics.

Adult↗

Characteristic attributes in cancer microarrays.

Rapid advances in genome sequencing and gene expression microarray technologies are providing unprecedented opportunities to identify specific genes involved in complex biological processes, such as development, signal transduction, and disease. The vast amount of data generated by these technologies has presented new challenges in bioinformatics. To help organize and interpret microarray data, new and efficient computational methods are needed to: (1) distinguish accurately between different biological or clinical categories (e.g., malignant vs. benign), and (2) identify specific genes that play a role in determining those categories. Here we present a novel and simple method that exhaustively scans microarray data for unambiguous gene expression patterns. Such patterns of data can be used as the basis for classification into biological or clinical categories. The method, termed the Characteristic Attribute Organization System (CAOS), is derived from fundamental precepts in systematic biology. In CAOS we define two types of characteristic attributes ('pure' and 'private') that may exist in gene expression microarray data. We also consider additional attributes ('compound') that are composed of expression states of more than one gene that are not characteristic on their own. CAOS was tested on three well-known cancer DNA microarray data sets for its ability to classify new microarray samples. We found CAOS to be a highly accurate and robust class prediction technique. In addition, CAOS identified specific genes, not emphasized in other analyses, that may be crucial to the biology of certain types of cancer. The success of CAOS in this study has significant implications for basic research and the future development of reliable methods for clinical diagnostic tools.

Acute Disease↗

A statistical score for assessing the quality of multiple sequence alignments.

BACKGROUND: Multiple sequence alignment is the foundation of many important applications in bioinformatics that aim at detecting functionally important regions, predicting protein structures, building phylogenetic trees etc. Although the automatic construction of a multiple sequence alignment for a set of remotely related sequences cause a very challenging and error-prone task, many downstream analyses still rely heavily on the accuracy of the alignments. RESULTS: To address the need for an objective evaluation framework, we introduce a statistical score that assesses the quality of a given multiple sequence alignment. The quality assessment is based on counting the number of significantly conserved positions in the alignment using importance sampling method in conjunction with statistical profile analysis framework. We first evaluate a novel objective function used in the alignment quality score for measuring the positional conservation. The results for the Src homology 2 (SH2) domain, Ras-like proteins, peptidase M13, subtilase and beta-lactamase families demonstrate that the score can distinguish sequence patterns with different degrees of conservation. Secondly, we evaluate the quality of the alignments produced by several widely used multiple sequence alignment programs using a novel alignment quality score and a commonly used sum of pairs method. According to these results, the Mafft strategy L-INS-i outperforms the other methods, although the difference between the Probcons, TCoffee and Muscle is mostly insignificant. The novel alignment quality score provides similar results than the sum of pairs method. CONCLUSION: The results indicate that the proposed statistical score is useful in assessing the quality of multiple sequence alignments.

Algorithms↗

Multiple sequence alignment accuracy and evolutionary distance estimation.

BACKGROUND: Sequence alignment is a common tool in bioinformatics and comparative genomics. It is generally assumed that multiple sequence alignment yields better results than pair wise sequence alignment, but this assumption has rarely been tested, and never with the control provided by simulation analysis. This study used sequence simulation to examine the gain in accuracy of adding a third sequence to a pair wise alignment, particularly concentrating on how the phylogenetic position of the additional sequence relative to the first pair changes the accuracy of the initial pair's alignment as well as their estimated evolutionary distance. RESULTS: The maximal gain in alignment accuracy was found not when the third sequence is directly intermediate between the initial two sequences, but rather when it perfectly subdivides the branch leading from the root of the tree to one of the original sequences (making it half as close to one sequence as the other). Evolutionary distance estimation in the multiple alignment framework, however, is largely unrelated to alignment accuracy and rather is dependent on the position of the third sequence; the closer the branch leading to the third sequence is to the root of the tree, the larger the estimated distance between the first two sequences. CONCLUSION: The bias in distance estimation appears to be a direct result of the standard greedy progressive algorithm used by many multiple alignment methods. These results have implications for choosing new taxa and genomes to sequence when resources are limited.

Algorithms↗