Search PubMed⌕ Search

Biomedical subjects

Mauno Vihinen

Publications and source records attributed to Mauno Vihinen.

At least 19 recordsLinked to original sources

A statistical score for assessing the quality of multiple sequence alignments.

BACKGROUND: Multiple sequence alignment is the foundation of many important applications in bioinformatics that aim at detecting functionally important regions, predicting protein structures, building phylogenetic trees etc. Although the automatic construction of a multiple sequence alignment for a set of remotely related sequences cause a very challenging and error-prone task, many downstream analyses still rely heavily on the accuracy of the alignments. RESULTS: To address the need for an objective evaluation framework, we introduce a statistical score that assesses the quality of a given multiple sequence alignment. The quality assessment is based on counting the number of significantly conserved positions in the alignment using importance sampling method in conjunction with statistical profile analysis framework. We first evaluate a novel objective function used in the alignment quality score for measuring the positional conservation. The results for the Src homology 2 (SH2) domain, Ras-like proteins, peptidase M13, subtilase and beta-lactamase families demonstrate that the score can distinguish sequence patterns with different degrees of conservation. Secondly, we evaluate the quality of the alignments produced by several widely used multiple sequence alignment programs using a novel alignment quality score and a commonly used sum of pairs method. According to these results, the Mafft strategy L-INS-i outperforms the other methods, although the difference between the Probcons, TCoffee and Muscle is mostly insignificant. The novel alignment quality score provides similar results than the sum of pairs method. CONCLUSION: The results indicate that the proposed statistical score is useful in assessing the quality of multiple sequence alignments.

Algorithms↗

Profiling genetic variation along the androgen biosynthesis and metabolism pathways implicates several single nucleotide polymorphisms and their combinations as prostate cancer risk factors.

Several candidate genes along androgen pathway have been suggested to affect prostate cancer risk but no single gene seems to be overwhelmingly important for a large fraction of the patients. In this study, we first screened for variants in candidate genes and then chose to explore the association between 18 variants and prostate cancer risk by genotyping DNA samples from unselected (n = 847) and familial (n = 121) prostate cancer patients and population controls (n = 923). We identified a novel single nucleotide polymorphism (SNP) in the CYP19A1 gene, T201M, with a mild significant association with prostate cancer [odds ratio (OR), 2.04; 95% confidence interval (95% CI), 1.03-4.03; P = 0.04]. Stratified analysis revealed that this risk was most apparent in patients with organ-confined (T(1)-T(2)) and low-grade (WHO grade 1) tumors (OR, 5.42; 95% CI, 2.33-12.6; P < 0.0001). In contrast, CYP17A1 -34T>C alteration was associated with moderate to poorly differentiated (WHO grade 2-3) organ-confined disease (OR, 1.42; 95% CI, 1.09-1.83; P = 0.007). We also tested a multigenic model of prostate cancer risk by calculating the joint effect of CYP19A1 T201M with five other common SNPs. Individuals carrying both the CYP19A1 and KLK3 -252A>G variant alleles had a significantly increased risk for prostate cancer (OR, 2.87; 95% CI, 1.10-7.49; P = 0.03). In conclusion, our results suggest that several SNPs along the androgen pathway, especially in CYP19A1 and CYP17A1, may influence prostate cancer development and progression. These genes may have different contributions to distinct clinical subsets as well as combinatorial effects in others illustrating that profiling and joint analysis of several genes along each pathway may be needed to understand genetic contributions to prostate cancer etiology.

Aged↗

Immunodeficiency mutation databases (IDbases).

Primary immunodeficiencies (IDs) are a heterogenic group of inherited disorders of the immune system. Immunodeficiency patients have increased susceptibility to recurrent and persistent, even life-threatening infections. Mutations in a large number of genes can cause defects in different cellular functions and lead to impaired immune response. To date, approximately 150 IDs and more than 100 affected genes have been identified. ID-related genes are distributed throughout the genome, and diseases can be inherited in an X-linked, an autosomal recessive, or an autosomal dominant way. We have collected ID mutation data into locus-specific patient-related mutation databases, IDbases (http://bioinf.uta.fi/IDbases). Mutations are described at DNA, mRNA, and protein levels with links to reference sequences and reference articles. The mutation data has been collated into entries along with some clinical information. IDbases offer an easy way, e.g., to find recently identified mutations, to reveal genotype-phenotype correlations, and to discover a specific mutation or to examine the most common mutations in a single immunodeficiency related gene. At the moment we have databases for 107 ID genes with 4,140 public patient entries. An exhaustive statistical analysis of mutation data from the IDbases was made. Missense and nonsense mutations are the most common mutation types, and the most common single substitution is a nonsense mutation from tryptophan to a stop codon. Arginine is the most mutated as well as the most abundant mutant amino acid.

Amino Acid Sequence↗

Bioinformatic analysis of protein structure-function relationships: case study of leukocyte elastase (ELA2) missense mutations.

Cyclic and congenital neutropenia are caused by mutations in the human neutrophil elastase (HNE) gene (ELA2), leading to an immunodeficiency characterized by decreased or oscillating levels of neutrophils in the blood. The HNE mutations presumably cause loss of enzyme activity, consequently leading to compromised immune system function. To understand the structural basis for the disease, we implemented methods from bioinformatics to analyze all the known HNE missense mutations at both the sequence and structural level. Our results demonstrate that the 32 different mutations have diverse effects on HNE structure and function, affecting structural disorder and aggregation tendencies, stability maintaining contacts, and electrostatic properties. A large proportion of the mutations are located at conserved amino acids, which are usually essential in determining protein structure and function. The majority of the disease-causing HNE missense mutations lead to major structural changes and loss of stability in the protein. A few mutations also affect functional residues, leading into decreased catalytic activity or altered ligand binding. Our analysis reveals the putative effects of all known missense mutations in HNE, thus allowing the structural basis of cyclic and congenital neutropenia to be elucidated. We have employed and analyzed a set of some 30 different methods for predicting the effects of amino acid substitutions. We present results and experience from the analysis of the applicability of these methods in the analysis of numerous genes, proteins, and diseases to reveal protein structure-function relationships and disease genotype-phenotype correlations.

Antigens, Surface↗

BTKbase: the mutation database for X-linked agammaglobulinemia.

X-linked agammaglobulinemia (XLA) is a hereditary immunodeficiency caused by mutations in the gene encoding Bruton tyrosine kinase (BTK). XLA patients have a decreased number of mature B cells and a lack of all immunoglobulin isotypes, resulting in susceptibility to severe bacterial infections. XLA-causing mutations are collected in a mutation database (BTKbase), which is available at http://bioinf.uta.fi/BTKbase. For each patient the following information is given (when available): the identification of the entry, a plain English description of the mutation followed by a reference, formal characterization of the mutation, and the various parameters from the patient. BTKbase is implemented with the MUTbase program suite, which provides an easy, interactive, and quality controlled submission of information to mutation databases. BTKbase version 8 lists mutation entries of 1,111 patients from 973 unrelated families showing 602 unique molecular events. The localization of the mutations on the gene and protein for BTK can be analyzed by clicking sequences on the web pages. The distribution of the mutations in the five structural domains is approximately proportional to the length of the domains, except for the Tec homology (TH) domain. The most frequently affected sites are CpG dinucleotides. The majority of the missense mutations are structural-disturbing Bruton tyrosine kinase (Btk) folding or decreasing stability. Many of the mutations affect functionally significant, conserved residues. The structural consequences of the mutations in all the domains have been studied based on crystallographic and nuclear magnetic resonance (NMR) structures as well as computer-aided molecular modeling.

Agammaglobulinemia↗

Proteome analysis of B-cell maturation.

Proteins affected by anti-mIgM stimulation during B-cell maturation were identified using 2-DE-based proteomics. We investigated the proteome profiles of stimulated and nonstimulated Ramos B-cells at eight time points during 5 d and compared the obtained proteomic data to the corresponding data from DNA-microarray studies. Anti-mIgM stimulation of the cells resulted in significant differences (> or =twofold) in the protein abundance close to 100 proteins and differences in post-translational protein modifications. Forty-eight up- or down-regulated proteins were identified by mass spectrometric methods and database searches. The identities of a further nine proteins were revealed by comparing their positions to the known proteins in other lymphocyte 2-DE databases. Several of the proteins are directly related to the functional and morphological characteristics of B-cells, such as cytoskeleton rearrangement and intracellular signalling triggered by the crosslinking of B-cell receptors. In addition to proteins known to be involved in human B-cell maturation, we identified several proteins that were not previously linked to lymphocyte differentiation. The results provide deeper insights into the process of B-cell maturation and may lead to novel therapeutic strategies for immunodeficiencies. An interactive 2-DE reference map is available at http://bioinf.uta.fi/BcellProteome.

B-Lymphocytes↗

Characterization of CA XV, a new GPI-anchored form of carbonic anhydrase.

The main function of CAs (carbonic anhydrases) is to participate in the regulation of acid-base balance. Although 12 active isoenzymes of this family had already been described, analyses of genomic databases suggested that there still exists another isoenzyme, CA XV. Sequence analyses were performed to identify those species that are likely to have an active form of this enzyme. Eight species had genomic sequences encoding CA XV, in which all the amino acid residues critical for CA activity are present. However, based on the sequence data, it was apparent that CA XV has become a non-processed pseudogene in humans and chimpanzees. RT-PCR (reverse transcriptase PCR) confirmed that humans do not express CA XV. In contrast, RT-PCR and in situ hybridization performed in mice showed positive expression in the kidney, brain and testis. A prediction of the mouse CA XV structure was performed. Phylogenetic analysis showed that mouse CA XV is related to CA IV. Therefore both of these enzymes were expressed in COS-7 cells and studied in parallel experiments. The results showed that CA XV shares several properties with CA IV, i.e. it is a glycosylated glycosylphosphatidylinositol-anchored membrane protein, and it binds CA inhibitor. The catalytic activity of CA XV is low, and the correct formation of disulphide bridges is important for the activity. Both specific and non-specific chaperones increase the production of active enzyme. The results suggest that CA XV is the first member of the alpha-CA gene family that is expressed in several species, but not in humans and chimpanzees.

Amino Acid Sequence↗

Dynamic covariation between gene expression and proteome characteristics.

BACKGROUND: Cells react to changing intra- and extracellular signals by dynamically modulating complex biochemical networks. Cellular responses to extracellular signals lead to changes in gene and protein expression. Since the majority of genes encode proteins, we investigated possible correlations between protein parameters and gene expression patterns to identify proteome-wide characteristics indicative of trends common to expressed proteins. RESULTS: Numerous bioinformatics methods were used to filter and merge information regarding gene and protein annotations. A new statistical time point-oriented analysis was developed for the study of dynamic correlations in large time series data. The method was applied to investigate microarray datasets for different cell types, organisms and processes, including human B and T cell stimulation, Drosophila melanogaster life span, and Saccharomyces cerevisiae cell cycle. CONCLUSION: We show that the properties of proteins synthesized correlate dynamically with the gene expression profile, indicating that not only is the actual identity and function of expressed proteins important for cellular responses but that several physicochemical and other protein properties correlate with gene expression as well. Gene expression correlates strongly with amino acid composition, composition- and sequence-derived variables, functional, structural, localization and gene ontology parameters. Thus, our results suggest that a dynamic relationship exists between proteome properties and gene expression in many biological systems, and therefore this relationship is fundamental to understanding cellular mechanisms in health and disease.

Animals↗

Genome-wide selection of unique and valid oligonucleotides.

Functional genomics methods are used to investigate the huge amount of information contained in genomes. Numerous experimental methods rely on the use of oligo- or polynucleotides. Nucleotide strand hybridization forms the underlying principle for these methods. For all these techniques, the probes should be unique for analyzed genes. In addition to being unique for the studied genes, the probes should fulfill a large number of criteria to be usable and valid. The criteria include for example, avoidance of self-annealing, suitable melting temperature and nucleotide composition. We developed a method for searching unique and valid oligonucleotides or probes for genes so that there is not even a similar (approximate) occurrence in any other location of the whole genome. By using probe size 25, we analyzed 17 complete genomes representing a wide range of both prokaryotic and eukaryotic organisms. More than 92% of all the genes in the investigated genomes contained valid oligonucleotides. Extensive statistical tests were performed to characterize the properties of unique and valid oligonucleotides. Unique and valid oligonucleotides were relatively evenly distributed in genes except for the beginning and end, which were somewhat overrepresented. The flanking regions in eukaryotes were clearly underrepresented among suitable oligonucleotides. In addition to distributions within genes, the effects on codon and amino acid usage were also studied.

Amino Acids↗

Distribution of immunodeficiency fact files with XML--from Web to WAP.

BACKGROUND: Although biomedical information is growing rapidly, it is difficult to find and retrieve validated data especially for rare hereditary diseases. There is an increased need for services capable of integrating and validating information as well as proving it in a logically organized structure. A XML-based language enables creation of open source databases for storage, maintenance and delivery for different platforms. METHODS: Here we present a new data model called fact file and an XML-based specification Inherited Disease Markup Language (IDML), that were developed to facilitate disease information integration, storage and exchange. The data model was applied to primary immunodeficiencies, but it can be used for any hereditary disease. Fact files integrate biomedical, genetic and clinical information related to hereditary diseases. RESULTS: IDML and fact files were used to build a comprehensive Web and WAP accessible knowledge base ImmunoDeficiency Resource (IDR) available at http://bioinf.uta.fi/idr/. A fact file is a user oriented user interface, which serves as a starting point to explore information on hereditary diseases. CONCLUSION: The IDML enables the seamless integration and presentation of genetic and disease information resources in the Internet. IDML can be used to build information services for all kinds of inherited diseases. The open source specification and related programs are available at http://bioinf.uta.fi/idml/.

Database Management Systems↗

KinMutBase: a registry of disease-causing mutations in protein kinase domains.

A large number of disease-causing mutations have been identified from several protein kinases. KinMutBase is a comprehensive knowledge base for human disease-related mutations in protein kinase domains (http://bioinf.uta.fi/KinMutBase/). The latest version contains 582 different mutations for 1,790 cases in 1,322 families. KinMutBase entries are described on the DNA, mRNA, and protein level. Numbers for affected patients and families are also provided. KinMutBase has extensive amount of links and cross-references to literature, other databases, and information sources. There are numerous interactive pages about sequences, structures, mutation statistics, and diseases. Detailed statistical study was done on frequencies of different types of mutations both on the DNA and protein level in serine/threonine kinase (PSK) and tyrosine kinase (PTK). Three-dimensional structures indicate clustering of disease-related mutations mainly to conserved subdomains, and substrate and coligand binding amino acids, although mutations appear throughout the sequences. CpG containing codons, especially for arginine, constitute the majority of mutational hotspots. There are certain clear differences in mutation patterns and types between PSKs and PTKs.

Amino Acid Sequence↗

B cells.

B cells are an important component of adaptive immunity. They produce and secrete millions of different antibody molecules, each of which recognizes a different (foreign) antigen. The fact that humans express a very large repertoire of antibodies is due to the complex mechanism of V(D)J recombination of immunoglobulin (Ig) genes as well as other processes including somatic hypermutation, gene conversion and class switching. The B cell receptor (BCR) is an integral membrane protein complex that is composed of two Ig heavy chains, two Ig light chains and two heterodimers of Igalpha and Igbeta. To eliminate foreign antigens, B cells cooperate with other cells of the immune system including macrophages, dendritic cells and T cells. B cell development is a tightly controlled process in which over 75% of the developing cells become apoptotic because of inappropriate immunoglobulin gene rearrangements or recognition of self antigens by Igs. Hence, the majority of B cell-associated disorders are caused by the incorrect function of genes/proteins involved in B cell development.

Animals↗

On exact string matching of unique oligonucleotides.

Unique, gene-specific oligonucleotides are used for many genetic investigations such as polymerase chain reaction, gene cloning, microarray technology and antisense DNA studies. It is a computationally demanding task to extract these oligonucleotides from DNA databases. We studied the problem from the point of view of the string matching problem. We implemented and tested several exact string matching algorithms and modified the implementations to be as effective as possible. Ten different implementations were tested on yeast genomic sequence data. The run times for the best algorithms were significantly improved compared to conventional approaches, while in principle, i.e. in respect of theoretical time complexity, these algorithms do not actually differ essentially from each other.

Algorithms↗

Bruton's tyrosine kinase: cell biology, sequence conservation, mutation spectrum, siRNA modifications, and expression profiling.

Bruton's tyrosine kinase (Btk) is encoded by the gene that when mutated causes the primary immunodeficiency disease X-linked agammaglobulinemia (XLA) in humans and X-linked immunodeficiency (Xid) in mice. Btk is a member of the Tec family of protein tyrosine kinases (PTKs) and plays a vital, but diverse, modulatory role in many cellular processes. Mutations affecting Btk block B-lymphocyte development. Btk is conserved among species, and in this review, we present the sequence of the full-length rat Btk and find it to be analogous to the mouse Btk sequence. We have also analyzed the wealth of information compiled in the mutation database for XLA (BTKbase), representing 554 unique molecular events in 823 families and demonstrate that only selected amino acids are sensitive to replacement (P < 0.001). Although genotype-phenotype correlations have not been established in XLA, based on these findings, we hypothesize that this relationship indeed exists. Using short interfering-RNA technology, we have previously generated active constructs downregulating Btk expression. However, application of recently established guidelines to enhance or decrease the activity was not successful, demonstrating the importance of the primary sequence. We also review the outcome of expression profiling, comparing B lymphocytes from XLA-, Xid-, and Btk-knockout (KO) donors to healthy controls. Finally, in spite of a few genes differing in expression between Xid- and Btk-KO mice, in vivo competition between cells expressing either mutation shows that there is no selective survival advantage of cells carrying one genetic defect over the other. We conclusively demonstrate that for the R28C-missense mutant (Xid), there is no biologically relevant residual activity or any dominant negative effect versus other proteins.

Agammaglobulinaemia Tyrosine Kinase↗

Statistical methods for identifying conserved residues in multiple sequence alignment.

The assessment of residue conservation in a multiple sequence alignment is a central issue in bioinformatics. Conserved residues and regions are used to determine structural and functional motifs or evolutionary relationships between the sequences of a multiple sequence alignment. For this reason, residue conservation is a valuable measure for database and motif search or for estimating the quality of alignments. In this paper, we present statistical methods for identifying conserved residues in multiple sequence alignments. While most earlier studies examine the positional conservation of the alignment, we focus on the detection of individual conserved residues at a position. The major advantages of multiple comparison methods originate from their ability to select conserved residues simultaneously and to consider the variability of the residue estimates. Large-scale simulations were used for the comparative analysis of the methods. Practical performance was studied by comparing the structurally and functionally important residues of Src homology 2 (SH2) domains to the assignments of the conservation indices. The applicability of the indices was also compared in three additional protein families comprising different degrees of entropy and variability in alignment positions. The results indicate that statistical multiple comparison methods are sensitive and reliable in identifying conserved residues.

Journal Article↗

Conservation and covariance in PH domain sequences: physicochemical profile and information theoretical analysis of XLA-causing mutations in the Btk PH domain.

Mutations that cause X-linked agammaglobulinemia (XLA) appear throughout the Bruton tyrosine kinase (Btk) sequence, including the pleckstrin homology (PH) domain. To analyze the basis of this disease with respect to protein structure, we studied the relationships between PH domain sequences and structures by comparing sequence-based profiles of physicochemical properties and solvent accessibility profiles. The diversity of the distribution of amino acids was measured by calculating entropies for sequences containing mutations at different positions in multiple sequence alignments. Mutual information was calculated to quantify positional covariation. Eight conserved extrema were apparent in all profiles. The majority of the XLA disease-causing mutations in the Btk PH domain were found at positions having significant mutual information, indicating that there are covariant constraints for both structure and function. Together with additional structural analyses, all the XLA mutations that were analyzed could be explained at the molecular level. The method developed here is applicable to the design of mutations for protein engineering.

Agammaglobulinemia↗

Probing the alpha-complementing domain of E. coli beta-galactosidase with use of an insertional pentapeptide mutagenesis strategy based on Mu in vitro DNA transposition.

Protein structure-function relationships can be studied by using linker insertion mutagenesis, which efficiently identifies essential regions in target proteins. Bacteriophage Mu in vitro DNA transposition was used to generate an extensive library of pentapeptide insertion mutants within the alpha-complementing domain 1 of Escherichia coli beta-galactosidase, yielding mutants at 100% efficiency. Each mutant contained an accurate 15-bp insertion that translated to five additional amino acids within the protein, and the insertions were distributed essentially randomly along the target sequence. Individual mutants (alpha-donors) were analyzed for their ability to restore (by alpha-complementation) beta-galactosidase activity of the M15 deletion mutant (alpha-acceptor), and the data were correlated to the structure of the beta-galactosidase tetramer. Most of the insertions were well tolerated, including many of those disrupting secondary structural elements even within the protein's interior. Nevertheless, certain sites were sensitive to mutations, indicating both known and previously unknown regions of functional importance. Inhibitory insertions within the N-terminus and loop regions most likely influenced protein tetramerization via direct local effects on protein-protein interactions. Within the domain 1 core, the insertions probably caused either lateral shifting of the polypeptide chain toward the protein's exterior or produced more pronounced structural distortions. Six percent of the mutant proteins exhibited temperature sensitivity, in general suggesting the method's usefulness for generation of conditional phenotypes. The method should be applicable to any cloned protein-encoding gene.

Amino Acid Sequence↗

Structure-function analysis of PrsA reveals roles for the parvulin-like and flanking N- and C-terminal domains in protein folding and secretion in Bacillus subtilis.

The PrsA protein of Bacillus subtilis is an essential membrane-bound lipoprotein that is assumed to assist post-translocational folding of exported proteins and stabilize them in the compartment between the cytoplasmic membrane and cell wall. This folding activity is consistent with the homology of a segment of PrsA with parvulin-type peptidyl-prolyl cis/trans isomerases (PPIase). In this study, molecular modeling showed that the parvulin-like region can adopt a parvulin-type fold with structurally conserved active site residues. PrsA exhibits PPIase activity in a manner dependent on the parvulin-like domain. We constructed deletion, peptide insertion, and amino acid substitution mutations and demonstrated that the parvulin-like domain as well as flanking N- and C-terminal domains are essential for in vivo PrsA function in protein secretion and growth. Surprisingly, none of the predicted active site residues of the parvulin-like domain was essential for growth and protein secretion, although several active site mutations reduced or abolished the PPIase activity or the ability of PrsA to catalyze proline-limited protein folding in vitro. Our results indicate that PrsA is a PPIase, but the essential role in vivo seems to depend on some non-PPIase activity of both the parvulin-like and flanking domains.

Bacillus subtilis↗