Search PubMed⌕ Search

Biomedical subjects

Lukasz A Kurgan

Publications and source records attributed to Lukasz A Kurgan.

4 recordsLinked to original sources

Prediction of three dimensional structure of calmodulin.

Calmodulin (CaM) is an important human protein, which has multiple structures. Numerous researchers studied the CaM structures in the past, and about 50 different structures in complex with fragments derived from CaM-regulated proteins have been discovered. Discovery and analysis of existing and new CaM structures is difficult due to the inherent complexity, i.e. flexibility of 6 loops and a central linker that constitute part of the CaM structure. The extensive interest in CaM structure analysis and discovery calls for a comprehensive study, which based on the accumulated expertise would design a method for prediction and analysis of future and existing CaM structures. It is also important to find the mechanisms by which the protein adjusts its structure with respect to various factors. To this end, this paper analyzes the known CaM structures and finds four factors that influence CaM structure, which include existence of Ca2+ binding, different binding segments, measuring surroundings, and sequence mutation. The degree of influence of specific factors on different structural regions is also investigated. Based on the analysis of the relation between the four factors and the corresponding CaM structure a novel method for prediction of the CaM structure in complex with novel segments, given that the surroundings of the complex, is developed. The developed prediction method is tested on a set aside, newest CaM structure. The prediction results provide useful and accurate information about the structure verifying high quality of the proposed prediction method and performed structural analysis.

Amino Acid Sequence↗

Quantitative analysis of the conservation of the tertiary structure of protein segments.

The publication of the crystallographic structure of calmodulin protein has offered an example leading us to believe that it is possible for many protein sequence segments to exhibit multiple 3D structures referred to as multi-structural segments. To this end, this paper presents statistical analysis of uniqueness of the 3D-structure of all possible protein sequence segments stored in the Protein Data Bank (PDB, Jan. of 2003, release 103) that occur at least twice and whose lengths are greater than 10 amino acids (AAs). We refined the set of segments by choosing only those that are not parts of longer segments, which resulted in 9297 segments called a sponge set. By adding 8197 signature segments, which occur uniquely in the PDB, into the sponge set we have generated a benchmark set. Statistical analysis of the sponge set demonstrates that rotating, missing and disarranging operations described in the text, result in the segments becoming multi-structural. It turns out that missing segments do not exhibit a change of shape in the 3D-structure of a multi-structural segment. We use the root mean square distance for unit vector sequence (URMSD) as an improved measure to describe the characteristics of hinge rotations, missing, and disarranging segments. We estimated the rate of occurrence for rotating and disarranging segments in the sponge set and divided it by the number of sequences in the benchmark set which is found to be less than 0.85%. Since two of the structure changing operations concern negligible number of segment and the third one is found not to have impact on the structure, we conclude that the 3D-structure of proteins is conserved statistically for more than 98% of the segments. At the same time, the remaining 2% of the sequences may pose problems for the sequence alignment based structure prediction methods.

Amino Acid Sequence↗

Highly scalable and robust rule learner: performance evaluation and comparison.

Business intelligence and bioinformatics applications increasingly require the mining of datasets consisting of millions of data points, or crafting real-time enterprise-level decision support systems for large corporations and drug companies. In all cases, there needs to be an underlying data mining system, and this mining system must be highly scalable. To this end, we describe a new rule learner called DataSqueezer. The learner belongs to the family of inductive supervised rule extraction algorithms. DataSqueezer is a simple, greedy, rule builder that generates a set of production rules from labeled input data. In spite of its relative simplicity, DataSqueezer is a very effective learner. The rules generated by the algorithm are compact, comprehensible, and have accuracy comparable to rules generated by other state-of-the-art rule extraction algorithms. The main advantages of DataSqueezer are very high efficiency, and missing data resistance. DataSqueezer exhibits log-linear asymptotic complexity with the number of training examples, and it is faster than other state-of-the-art rule learners. The learner is also robust to large quantities of missing data, as verified by extensive experimental comparison with the other learners. DataSqueezer is thus well suited to modern data mining and business intelligence tasks, which commonly involve huge datasets with a large fraction of missing data.

Algorithms↗

Highly accurate and consistent method for prediction of helix and strand content from primary protein sequences.

OBJECTIVE: One of interesting computational topics in bioinformatics is prediction of secondary structure of proteins. Over 30 years of research has been devoted to the topic but we are still far away from having reliable prediction methods. A critical piece of information for accurate prediction of secondary structure is the helix and strand content of a given protein sequence. Ability to accurately predict content of those two secondary structures has a good potential to improve accuracy of prediction of the secondary structure. Most of the existing methods use composition vector to predict the content. Their underlying assumption is that the vector can be used to provide functional mapping between primary sequence and helix/strand content. While this is true for small sets of proteins we show that for larger protein sets such mapping are inconsistent, i.e. the same composition vectors correspond to different contents. To this end, we propose a method for prediction of helix/strand content from primary protein sequences that is fundamentally different from currently available methods. METHODS AND MATERIAL: Our method is accurate and uses a novel approach to obtain information from primary sequence based on a composition moment vector, which is a measure that includes information about both composition of a given primary sequence and the position of amino acids in the sequence. In contrast to the composition vector, we show that it provides functional mapping between primary sequence and the helix/strand content. RESULTS: A set of benchmarks involving a large protein dataset consisting of over 11,000 protein sequences from Protein Data Bank was performed to validate the method. Prediction done by a neural network had average accuracy of 91.5% for the helix and 94.5% for the strand contents. We also show that using the new measure results in about 40% reduction of error rates when compared with the composition vector results. CONCLUSIONS: The developed method has much better accuracy when compared with other existing methods, as shown on a large body of proteins, in contrast to other reported results that often target small sets of specific protein types, such as globular proteins.

Amino Acid Sequence↗