Search PubMed⌕ Search

Biomedical subjects

Michael Kaufmann

Publications and source records attributed to Michael Kaufmann.

7 recordsLinked to original sources

DIALIGN-T: an improved algorithm for segment-based multiple sequence alignment.

BACKGROUND: We present a complete re-implementation of the segment-based approach to multiple protein alignment that contains a number of improvements compared to the previous version 2.2 of DIALIGN. This previous version is superior to Needleman-Wunsch-based multi-alignment programs on locally related sequence sets. However, it is often outperformed by these methods on data sets with global but weak similarity at the primary-sequence level. RESULTS: In the present paper, we discuss strengths and weaknesses of DIALIGN in view of the underlying objective function. Based on these results, we propose several heuristics to improve the segment-based alignment approach. For pairwise alignment, we implemented a fragment-chaining algorithm that favours chains of low-scoring local alignments over isolated high-scoring fragments. For multiple alignment, we use an improved greedy procedure that is less sensitive to spurious local sequence similarities. To evaluate our method on globally related protein families, we used the well-known database BAliBASE. For benchmarking tests on locally related sequences, we created a new reference database called IRMBASE which consists of simulated conserved motifs implanted into non-related random sequences. CONCLUSION: On BAliBASE, our new program performs significantly better than the previous version of DIALIGN and is comparable to the standard global aligner CLUSTAL W, though it is outperformed by some newly developed programs that focus on global alignment. On the locally related test sets in IRMBASE, our method outperforms all other programs that we evaluated.

Algorithms↗

Crystal structure of THEP1 from the hyperthermophile Aquifex aeolicus: a variation of the RecA fold.

BACKGROUND: aaTHEP1, the gene product of aq_1292 from Aquifex aeolicus, shows sequence homology to proteins from most thermophiles, hyperthermophiles, and higher organisms such as man, mouse, and fly. In contrast, there are almost no homologous proteins in mesophilic unicellular microorganisms. aaTHEP1 is a thermophilic enzyme exhibiting both ATPase and GTPase activity in vitro. Although annotated as a nucleotide kinase, such an activity could not be confirmed for aaTHEP1 experimentally and the in vivo function of aaTHEP1 is still unknown. RESULTS: Here we report the crystal structure of selenomethionine substituted nucleotide-free aaTHEP1 at 1.4 A resolution using a multiple anomalous dispersion phasing protocol. The protein is composed of a single domain that belongs to the family of 3-layer (alpha/beta/alpha)-structures consisting of nine central strands flanked by six helices. The closest structural homologue as determined by DALI is the RecA family. In contrast to the latter proteins, aaTHEP1 possesses an extension of the beta-sheet consisting of four additional beta-strands. CONCLUSION: We conclude that the structure of aaTHEP1 represents a variation of the RecA fold. Although the catalytic function of aaTHEP1 remains unclear, structural details indicate that it does not belong to the group of GTPases, kinases or adenosyltransferases. A mainly positive electrostatic surface indicates that aaTHEP1 might be a DNA/RNA modifying enzyme. The resolved structure of aaTHEP1 can serve as paradigm for the complete THEP1 family.

Adenosine Triphosphatases↗

PCOGR: phylogenetic COG ranking as an online tool to judge the specificity of COGs with respect to freely definable groups of organisms.

BACKGROUND: The rapidly increasing number of completely sequenced genomes led to the establishment of the COG-database which, based on sequence homologies, assigns similar proteins from different organisms to clusters of orthologous groups (COGs). There are several bioinformatic studies that made use of this database to determine (hyper)thermophile-specific proteins by searching for COGs containing (almost) exclusively proteins from (hyper)thermophilic genomes. However, public software to perform individually definable group-specific searches is not available. RESULTS: The tool described here exactly fills this gap. The software is accessible at http://www.uni-wh.de/pcogr and is linked to the COG-database. The user can freely define two groups of organisms by selecting for each of the (current) 66 organisms to belong either to groupA, to the reference groupB or to be ignored by the algorithm. Then, for all COGs a specificity index is calculated with respect to the specificity to groupA, i. e. high scoring COGs contain proteins from the most of groupA organisms while proteins from the most organisms assigned to groupB are absent. In addition to ranking all COGs according to the user defined specificity criteria, a graphical visualization shows the distribution of all COGs by displaying their abundance as a function of their specificity indexes. CONCLUSIONS: This software allows detecting COGs specific to a predefined group of organisms. All COGs are ranked in the order of their specificity and a graphical visualization allows recognizing (i) the presence and abundance of such COGs and (ii) the phylogenetic relationship between groupA- and groupB-organisms. The software also allows detecting putative protein-protein interactions, novel enzymes involved in only partially known biochemical pathways, and alternate enzymes originated by convergent evolution.

Escherichia coli↗

DIALIGN P: fast pair-wise and multiple sequence alignment using parallel processors.

BACKGROUND: Parallel computing is frequently used to speed up computationally expensive tasks in Bioinformatics. RESULTS: Herein, a parallel version of the multi-alignment program DIALIGN is introduced. We propose two ways of dividing the program into independent sub-routines that can be run on different processors: (a) pair-wise sequence alignments that are used as a first step to multiple alignment account for most of the CPU time in DIALIGN. Since alignments of different sequence pairs are completely independent of each other, they can be distributed to multiple processors without any effect on the resulting output alignments. (b) For alignments of large genomic sequences, we use a heuristics by splitting up sequences into sub-sequences based on a previously introduced anchored alignment procedure. For our test sequences, this combined approach reduces the program running time of DIALIGN by up to 97%. CONCLUSIONS: By distributing sub-routines to multiple processors, the running time of DIALIGN can be crucially improved. With these improvements, it is possible to apply the program in large-scale genomics and proteomics projects that were previously beyond its scope.

Computational Biology↗

Thermophile-specific proteins: the gene product of aq_1292 from Aquifex aeolicus is an NTPase.

BACKGROUND: To identify thermophile-specific proteins, we performed phylogenetic patterns searches of 66 completely sequenced microbial genomes. This analysis revealed a cluster of orthologous groups (COG1618) which contains a protein from every thermophile and no sequence from 52 out of 53 mesophilic genomes. Thus, COG1618 proteins belong to the group of thermophile-specific proteins (THEPs) and therefore we here designate COG1618 proteins as THEP1s. Since no THEP1 had been analyzed biochemically thus far, we characterized the gene product of aq_1292 which is THEP1 from the hyperthermophilic bacterium Aquifex aeolicus (aaTHEP1). RESULTS: aaTHEP1 was cloned in E. coli, expressed and purified to homogeneity. At a temperature optimum between 70 and 80 degrees C, aaTHEP1 shows enzymatic activity in hydrolyzing ATP to ADP + Pi with kcat = 5 x 10(-3) s(-1) and Km = 5.5 x 10(-6) M. In addition, the enzyme exhibits GTPase activity (kcat = 9 x 10(-3) s(-1) and Km= 45 x 10(-6) M). aaTHEP1 is inhibited competitively by CTP, UTP, dATP, dGTP, dCTP, and dTTP. As shown by gel filtration, aaTHEP1 in its purified state appears as a monomer. The enzyme is resistant to limited proteolysis suggesting that it consists of a single domain. Although THEP1s are annotated as "predicted nucleotide kinases" we could not confirm such an activity experimentally. CONCLUSION: Since aaTHEP1 is the first member of COG1618 that is characterized biochemically and functional information about one member of a COG may be transferred to the entire COG, we conclude that COG1618 proteins are a family of thermophilic NTPases.

Adenosine Triphosphate↗

EPPS: mining the COG database by an extended phylogenetic patterns search.

SUMMARY: EPPS runs under Microsoft Windows. It is an extended version of the phylogenetic patterns search (PPS). The output condition of PPS is the exact match of a user defined phylogenetic pattern with the pattern represented by the respective cluster of orthologous groups (COG). In contrast, the software described here is less restrictive. The user may define the accuracy of the search by the number of genomes that are allowed not to match the predefined phylogenetic pattern. Thus, EPPS has the advantage to detect COGs even if organisms defined to be included are not or organisms defined to be excluded are present in the output COGs.

Amino Acid Sequence↗

Helping the auto repair industry manage hazardous wastes: an education project in King County, Washington.

From January 1, 2000, to August 31, 2001, a team of environmental health specialists from Public Health-Seattle & King County, a partner in King County's Local Hazardous Waste Management Program, made educational visits to 981 automotive repair shops. The purpose was to give the auto repair industry technical assistance on hazardous waste management without using enforcement action. Through site inspections and interviews, the environmental health staff gathered information on the types and amounts of conditionally exempt small-quantity generator (CESQG) hazardous wastes and how they were handled. Proper methods of hazardous waste management, storage, and disposal were discussed with shop personnel. The environmental health staff measured the impact of these educational visits by noting changes made between the initial and follow-up visits. This report focuses on nine major waste streams identified in the auto repair industry. Of the 981 shops visited, 497 were already practicing proper hazardous waste management and disposal. The remaining 484 shops exhibited 741 discrepancies from proper practice. Environmental health staff visited these shops again within six months of the initial visit to assess changes in their practices. The educational visits and technical assistance produced a 76 percent correction of all the discrepancies noted.

Automobiles↗