Search PubMed⌕ Search

Biomedical subjects

Thy-Hou Lin

Publications and source records attributed to Thy-Hou Lin.

15 recordsLinked to original sources

Modeling the ligand-receptor interaction for a series of inhibitors of the capsid protein of enterovirus 71 using several three-dimensional quantitative structure-activity relationship techniques.

The structure of enterovirus 71 (EV71) capsid protein VP1 has been constructed by using homology modeling and molecular dynamics simulation techniques. The ligand structures were a series of EV71 VP1 inhibitors synthesized by Shia et al. in 2002 and Chern et al. in 2004. The training set was selected by the VOLSURF4.1/PCA program and the IC50 values varied from 0.06 to 10.83 microm. Then, the training set was analyzed by the following three-dimensional quantitative structure-activity relationship techniques: CoMFA, CoMSIA, CATALYST4.9, and VOLSURF4.1/PCA. The model generated by a two-stage flexible docking procedure and without any structural alignment has far more significant statistics. Highly accurate activities for the test sets were then predicted by the top hypothesis of the CATALYST program and were compared with those predicted by CoMFA, CoMSIA, and VOLSURF. These studies identified some important clues for searching or making more potent inhibitors against the EV71 infection.

Binding Sites↗

Discovery of a novel family of SARS-CoV protease inhibitors by virtual screening and 3D-QSAR studies.

The severe acute respiratory syndrome-associated coronavirus (SARS-CoV) 3C-like protease (3CL(pro) or M(pro)) is an attractive target for the development of anti-SARS drugs because of its crucial role in the viral life cycle. In this study, a compound database was screened by the structure-based virtual screening approach to identify initial hits as inhibitors of SARS-CoV 3CL(pro). Out of the 59,363 compounds docked, 93 were selected for the inhibition assay, and 21 showed inhibition against SARS-CoV 3CL(pro) (IC(50) <or= 30 microM), with three of them having common substructures. Furthermore, a search for analogues with common substructure in the Maybridge, ChemBridge, and SPECS_SC databases led to the identification of another 25 compounds that exhibited inhibition against SARS-CoV 3CL(pro) (IC(50) = 3-1,000 microM). These compounds, 28 in total, were subjected to 3D-QSAR studies to elucidate the pharmacophore of SARS-CoV 3CL(pro).

Binding Sites↗

Association algorithm to mine the rules that govern enzyme definition and to classify protein sequences.

BACKGROUND: The number of sequences compiled in many genome projects is growing exponentially, but most of them have not been characterized experimentally. An automatic annotation scheme must be in an urgent need to reduce the gap between the amount of new sequences produced and reliable functional annotation. This work proposes rules for automatically classifying the fungus genes. The approach involves elucidating the enzyme classifying rule that is hidden in UniProt protein knowledgebase and then applying it for classification. The association algorithm, Apriori, is utilized to mine the relationship between the enzyme class and significant InterPro entries. The candidate rules are evaluated for their classificatory capacity. RESULTS: There were five datasets collected from the Swiss-Prot for establishing the annotation rules. These were treated as the training sets. The TrEMBL entries were treated as the testing set. A correct enzyme classification rate of 70% was obtained for the prokaryote datasets and a similar rate of about 80% was obtained for the eukaryote datasets. The fungus training dataset which lacks an enzyme class description was also used to evaluate the fungus candidate rules. A total of 88 out of 5085 test entries were matched with the fungus rule set. These were otherwise poorly annotated using their functional descriptions. CONCLUSION: The feasibility of using the method presented here to classify enzyme classes based on the enzyme domain rules is evident. The rules may be also employed by the protein annotators in manual annotation or implemented in an automatic annotation flowchart.

Algorithms↗

Structure of the virus capsid protein VP1 of enterovirus 71 predicted by some homology modeling and molecular docking studies.

The homology modeling technique has been used to construct the structure of enterovirus 71 (EV 71) capsid protein VP1. The protein is consisted of 297 amino acid residues and treated as the target. The amino acid sequence identity between the target protein and sequences of template proteins 1EAH, 1PIV, and 1D4M searched from NCBI protein BLAST and WorkBench protein tools were 38, 37, and 36%, respectively. Based on these template structures, the protein model was constructed by using the InsightII/Homology program. The protein model was briefly refined by energy minimization and molecular dynamics (MD) simulation steps. The protein model was validated using some web available servers such as ERRAT, PROCHECK, PROVE, and PROSA2003. However, an inconsistency between the docking scores and the measured activity was observed for a series of EV 71 VP1 inhibitors synthesized by Shia et al. (J Med Chem 2002, 45, 1644) and docked into the binding pocket of the protein model using the DOCK 4.0.2 program. The protein model with an EV 71 VP1 inhibitor docked and engulfed was then refined further by some MD simulation steps in the presence of water molecules. The docking scores obtained for these inhibitors after such a MD refinement were well correlated with the activities. The structure-activity relationships for the ligand-protein model system was also analyzed using the GRID-VOLSURF programs and the corresponding noncrossvalidated and crossvalidated (by leave-one-out) r2 and q2 were 0.99 and 0.61, respectively. The hydrophobic nature of the binding pocket of the protein model was also examined using the GRID21 program. The possibility of improving the potency of the current series of EV 71 VP1 inhibitors was discussed based on all the studies presented.

Amino Acid Sequence↗

Complete genomic sequence of the temperate bacteriophage PhiAT3 isolated from Lactobacillus casei ATCC 393.

The complete genomic sequence of a temperate bacteriophage PhiAT3 isolated from Lactobacillus (Lb.) casei ATCC 393 is reported. The phage consists of a linear DNA genome of 39,166 bp, an isometric head of 53 nm in diameter, and a flexible, noncontractile tail of approximately 200 nm in length. The number of potential open reading frames on the phage genome is 53. There are 15 unpaired nucleotides at both 5' ends of the PhiAT3 genome, indicating that the phage uses a cos-site for DNA packaging. The PhiAT3 genome was grouped into five distinct functional clusters: DNA packaging, morphogenesis, lysis, lysogenic/lytic switch, and replication. The amino acid sequences at the NH2-termini of some major proteins were determined. An in vivo integration assay for the PhiAT3 integrase (Int) protein in several lactobacilli was conducted by constructing an integration vector including PhiAT3 int and the attP (int-attP) region. It was found that PhiAT3 integrated at the tRNAArg gene locus of Lactobacillus rhamnosus HN 001, similar to that observed in its native host, Lb. casei ATCC 393.

Bacteriophages↗

Molecular dynamics study on the interaction of a mithramycin dimer with a decanucleotide duplex.

The complex of a minor groove binding drug mithramycin (MTR) and the self-complementary d(TAGCTAGCTA) 10-mer duplex was investigated by molecular dynamics (MD) simulations using the AMBER 7.0 suite of programs. There is one disaccharide and trisaccharide segment projecting from opposite ends of an aglycone chromophore of MTR. A MTR dimer complex (MTR)2Mg2+ is formed in the presence of a coordinated ion Mg2+. A NMR solution structure of two (MTR)2Mg2+ complexes bound with one DNA duplex, namely, the 2:1 duplex complex, was taken as the starting structure for the MD simulation. The partial charge on each atom was calculated using the multiple-RESP fitting procedure, and all of the missing parameters in the Parm99 force field used were adapted comparably from the literature. The length of the MD simulation was 5 ns, and the binding free energy for the formation of a 1:1 or 2:1 duplex complex was determined from the last 4 ns of the simulation. The binding free energies were decomposed to components of the contributions from different energy types, and the changes in the helical parameters of the bound DNA duplex plus the glycosidic linkages between sugar residues of the bound MTR dimer were determined. It was found that binding of the first (MTR)2Mg2+ complex with the DNA duplex to form a 1:1 duplex complex does not cause stiffening of the duplex especially in the unoccupied site of the duplex. However, the overall flexibility of the DNA duplex is reduced substantially once the second (MTR)2Mg2+ complex is bound with the unoccupied site to form the 2:1 duplex complex. The van der Waals interactions were found to be dominant in the central part of the DNA duplex where sugar residues from each bound (MTR)2Mg2+ complex were inwardly pointing and the corresponding minor groove was widened.

Base Pairing↗

A theoretical study on the activation of Ser70 in the acylation mechanism of cephalosporin antibiotics.

A computational study using some molecular modeling and quantum mechanical methods has been performed for determining the most favor activation process for Ser70 in the acylation mechanism for the cephalosporin antibiotics among the three proposed ones given in the literature. The computation is based on an X-ray structure of the B chain of the Bacillus licheniformis BS3 beta-lactamase-cefoxitin complex. The position of a catalytic water involved in one of the reaction mechanism is defined using the Grid20 and InsightII programs, while that of the truncated ligand is defined using the InsightII and FirstDiscovery programs. The geometry of structures of each reaction scheme is optimized at the HF/3-21 G level of theory, and then the single point energy for each reactive species in each reaction scheme is computed at the levels of HF/6-31 + G (3df, 2p) and B3LYP/6-31 + G (3df, 2p). The effect of solvent on each reaction scheme is also studied by comparing the calculation results for each reaction scheme either in gas phase or in solution using the HF/6-31 + G (3df, 2p) level of theory. A computation using the B3LYP/6-31 + G (3df, 2p) level of theory and the Polarized Continuum Model (PCM) and by treating water as a solvent is also conducted for each activation process. It is found that, energetically, the most favor activation process for Ser70 in the acylation mechanism is the one where a proton transfer is mediated by the catalytic water and the catalytic residues Glu166 and Ser70. This agrees with those observed in an ultrahigh resolution X-ray structure and a QM/MM theoretical study published recently on the same acylation process.

Acylation↗

Homology modeling of the central catalytic domain of insertion sequence ISLC3 isolated from Lactobacillus casei ATCC 393.

The tertiary structure of the central catalytic domain of insertion sequence ISLC3 isolated from Lactobacillus casei ATCC 393 was predicted using the homology modeling approach. The novel insertion sequence was isolated by us from the template bacteriophage phiA3 of L.casei ATCC 393. The number of amino acid residues of the ISLC3 central catalytic domain was 116 and was treated as the query sequence. There were five Web-available threading methods used to find some primary structure templates for the query sequence. These primary templates were further screened using the SWISS-MODEL Protein Modeling Server and the default parameter settings therein to give six final structure templates. All of these final structure templates were the integrase (IN) protein of retroviruses. Multiple sequence alignment using these IN sequences against the query one revealed the signature DDE motif. Based on the structures of these final templates, the structure of the query sequence was constructed using the InsightII/Discover/Homology programs. A metal ion, Mg(2+), was inserted into the center of the putative catalytic pocket formed by the DDE residues of the predicted structure in the final rounds of refinement by molecular dynamics (MD) simulations. The structure with a metal ion included was designated with Mg and that without a metal ion was designated free Mg. The average exposed surface area of some hydrophobic residues of both the predicted free Mg and with Mg structures were computed and compared with those computed for the six structure templates. Whereas the predicted with Mg structure was slightly more exposed than the predicted free Mg structure, the former appeared to be more stable than the latter, as revealed by the lower conformation energy recorded for the former during the structure refinement by MD simulations. To verify further the predicted structures, the coordinates of both predicted structures were fed into the ERRAT Protein Verification Server. It was found that the quality of the predicted with Mg structure was much better than that of the free Mg structure. The validation results also indicated that regions of the predicted with Mg structure that can be rejected at the 95% confidence level were approximately 20% whereas those which can be rejected at the same level for the six structure templates were approximately 10%. The predicted with Mg structure was also docked into a short oligonucleotide representing the substrate of the ISLC3 transposase using the DOCK_4.0.2 program. It was found that both Glu140 and Asp68 residues of the DDE motif of the predicted with Mg structure were able to form hydrogen bonds with the DNA substrate, which was similar to what was observed in a docking study using the retrovirus IN 1asu and its DNA substrate.

Amino Acid Sequence↗

Target integration by a chimeric Sp1 zinc finger domain-Moloney murine leukemia virus integrase in vivo.

A specificity protein 1 (Sp1) zinc finger domain containing two tandem zinc fingers was fused to the C terminus of the integrase (IN) protein of the Moloney murine leukemia virus (MuLV). The integrity of the MuLV IN was completely preserved, since the fusion was conducted at the last amino acid residue of the protein. The vector pMIN-Sp1, which carried the fused MuLV IN-Sp1 zinc finger domain gene, was cotransfected with a wild-type MuLV vector pMLV-K to NIH/3T3 cells. A nonradioactive reverse transcriptase assay was performed on culture supernatants collected from the cotransfected cells to confirm the production of recombinant viruses. The expression of the fusion protein and the integration of the MuLV genome by the fusion protein were confirmed by a Northern and then a Southern hybridization analysis on the total RNA or genomic DNA extracted from cells infected by viruses collected from the supernatants of the cotransfected cells. Regions of the host chromosome that were selected by the fusion protein as the integration targets were sequenced using the TOPO(TM) cloning method on a series of PCR products generated with a nested set of primers. The percentage of positive clones screened that contained the DNA-binding sequence of the fused Sp1 zinc finger domain was around 13% (5 out of 39 clones). It was found that the Sp1 DNA-binding sequence was only present in regions that were proximal to one of the long terminal repeats of the integrated viral genome, suggesting that the fusion protein could select a target sequence for integration. The host flanking sequences determined for all the positive clones were also used as queries to perform a BLAST search on the GenBank mouse EST entries. Although matching scores for sequences of some of the clones computed were more significant than others, it was difficult to judge whether or not the integration in these clones had been targeted to some gene sequences. Most of the integration sites might exist in the introns, since we found that the probability of the gene sequences containing an Sp1 DNA-binding site was low.

3' Flanking Region↗

Prediction of beta-turns in proteins using the first-order Markov models.

We present a method based on the first-order Markov models for predicting simple beta-turns and loops containing multiple turns in proteins. Sequences of 338 proteins in a database are divided using the published turn criteria into the following three regions, namely, the turn, the boundary, and the nonturn ones. A transition probability matrix is constructed for either the turn or the nonturn region using the weighted transition probabilities computed for dipeptides identified from each region. There are two such matrices constructed for the boundary region since the transition probabilities for dipeptides immediately preceding or following a turn are different. The window used for scanning a protein sequence from amino (N-) to carboxyl (C-) terminal is a hexapeptide since the transition probability computed for a turn tetrapeptide is capped at both the N- and C- termini with a boundary transition probability indexed respectively from the two boundary transition matrices. A sum of the averaged product of the transition probabilities of all the hexapeptides involving each residue is computed. This is then weighted with a probability computed from assuming that all the hexapeptides are from the nonturn region to give the final prediction quantity. Both simple beta-turns and loops containing multiple turns in a protein are then identified by the rising of the prediction quantity computed. The performance of the prediction scheme or the percentage (%) of correct prediction is evaluated through computation of Matthews correlation coefficients for each protein predicted. It is found that the prediction method is capable of giving prediction results with better correlation between the percent of correct prediction and the Matthews correlation coefficients for a group of test proteins as compared with those predicted using some secondary structural prediction methods. The prediction accuracy for about 40% of proteins in the database or 50% of proteins in the test set is better than 70%. Such a percentage for the test set is reduced to 30 if the structures of all the proteins in the set are treated as unknown.

Amino Acid Sequence↗

Classification of some active HIV-1 protease inhibitors and their inactive analogues using some uncorrelated three-dimensional molecular descriptors and a fuzzy c-means algorithm.

A fuzzy c-means algorithm was used to classify some 3D convex hull descriptors computed for 345 active HIV-1 protease inhibitors collected from literature and 437 inactive analogues searched from the MDL/ISIS database. The number of descriptors used to represent each compound was from 4 to 8, and they were uncorrelated using the principal component analysis. These uncorrelated descriptors were then divided into two groups and classified by the fuzzy c-means algorithm. The classification produced a clear-cut switch in membership functions computed for each uncorrelated descriptor at the group boundary. Compounds with nonswitching membership functions computed were treated as outliers, and they were counted for estimating the accuracy of the classification. The averaged accuracy of classification for the active inhibitor set was about 80% which was better than that directly classified by a linear discriminant function on the original 3D convex hull descriptors. The whole classification scheme was also applied to several sets of some conventional descriptors computed for each compound, but the averaged accuracy was around 58%. Further classification using some 3D convex hull descriptors searched from comparing the distribution of these descriptors was performed on a new data set composed of 289 outliers-deducted active inhibitors and 63 outliers identified from the inactive analogues through previous classification. This final classification identified 19 inactive analogues which were similar in structural and topological features to those of some highly active inhibitors classified together with them.

Algorithms↗

Implementing the Fisher's discriminant ratio in a k-means clustering algorithm for feature selection and data set trimming.

The Fisher's discriminant ratio has been used as a class separability criterion and implemented in a k-means clustering algorithm for performing simultaneous feature selection and data set trimming on a set of 221 HIV-1 protease inhibitors. The total number of molecular descriptors computed for each inhibitor is 43, and they are scaled to lie between 1 and 0 before being subjected to the feature selection process. Since the purpose is to select some of the most class sensitive descriptors, several feature evaluation indices such as the Shannon entropy, the linear regression of selected descriptors on the pKi of selected inhibitors, and a stepwise variable selection program are used to filter them. While the Shannon entropy provides the information content for each descriptor computed, more class sensitive descriptors are searched by both the linear regression and stepwise variable selection procedures. The inhibitors are divided into several different numbers of classes. They are subsequently divided into five classes due to the fact that the best feature selection result is obtained by the division. Most of the good features selected are the topological descriptors, and they are correlated well with the pKi values. The outliers or the inhibitors with less class-sensitive descriptor values computed for each selected descriptor are identified and gathered by the k-means clustering algorithm. These are the trimmed inhibitors, while the remaining ones are retained or selected. We find that 44% or 98 inhibitors can be retained when the number of good descriptors selected for clustering is three. The descriptor values of these selected inhibitors are far more class sensitive than the original ones as evidenced by substantial increasing in statistical significance when they are subjected to both the SYBYL CoMFA PLS and Cerius2 PLS regression analyses.

Journal Article↗

A ligand-based molecular modeling study on some matrix metalloproteinase-1 inhibitors using several 3D QSAR techniques.

Some three-dimensional quantitative structure-activity relationship (3D-QSAR) models have been constructed using the comparative molecular field analysis (CoMFA) and comparative molecular similarity indices (CoMSIA) for a series of 84 proline-based plus 12 structurally more diversified nonproline matrix metalloproteinase inhibitors. The structures of these inhibitors were built from a structure template extracted from the crystal structure of stromelysin. The structures built were divided into the training and test sets for both the CoMFA and CoMSIA analyses for each being composed of 60 and 24 inhibitors, respectively. The structures in the training set were aligned using some alignment rules derived from the analysis of the Ligplot program on a recent crystal structure of ligand-collagenase-1 complex. Some stepwise CoMSIA's were performed on the aligned training set on which the best CoMFA result was obtained. The best CoMSIA model was identified from the stepwise results, and the corresponding pharmacophore features were used for the construction of a pharmacophore hypothesis by the Catalyst 4.9 program. The training set was extended to include 11 structurally more diversified and nonproline inhibitors. To construct a pharmacophore hypothesis, the conformation of 60 structurally aligned proline-based inhibitors was fixed, while that of the 11 structurally more diversified nonproline inhibitors was allowed to vary during the hypothesis construction process. It was found that the predicted activities by the top hypothesis constructed for both the training and test sets were as good in statistics as those predicted by the best CoMSIA model from which the hypothesis was derived. The top hypothesis was mapped onto the structures of several highly active inhibitors selected from both the training and test sets. The goodness of mapping on each inhibitor was found to be correlated well with the activity of each inhibitor.

Ligands↗

Modeling ligand-receptor interaction for some MHC class II HLA-DR4 peptide mimetic inhibitors using several molecular docking and 3D QSAR techniques.

The ligand-receptor interaction between some peptidomimetic inhibitors and a class II MHC peptide presenting molecule, the HLA-DR4 receptor, was modeled using some three-dimensional (3D) quantitative structure-activity relationship (QSAR) methods such as the Comparative Molecular Field Analysis (CoMFA), Comparative Molecular Similarity Indices Analysis (CoMSIA), and a pharmacophore building method, the Catalyst program. The structures of these peptidomimetic inhibitors were generated theoretically, and the conformations used in the 3D QSAR studies were defined by docking them into the known structure of HLA-DR4 receptor through the GOLD, GLIDE Rigidly, GLIDE Flexible, and Xscore programs. Some of the parameters used in these docking programs were selected by docking an X-ray ligand into the receptor and comparing the root-means-square difference (RMSD) computed between the coordinates of the X-ray and docked structure. However, the goodness of a docking result for docking a series of peptidomimetic inhibitors into the HLA-DR4 receptor was judged by comparing the Spearman's rank correlation coefficient computed between each docking result and the activity data taken from the literature. The best CoMFA and CoMSIA models were constructed using the aligned structures of the best docking result. The CoMSIA was conducted in a stepwise manner to identify some important molecular features that were further employed in a pharmacophore building process by the Catalyst program. It was found that most inhibitors of the training set were accurately predicted by the best pharmacophore model, the Hypo1 hypothesis constructed. The deviation or conflict found between the actual and predicted activities of some inhibitors of both the training and the test sets were also investigated by mapping the Hypo1 hypothesis onto the corresponding structures of the inhibitors.

Binding Sites↗

Supervised feature ranking using a genetic algorithm optimized artificial neural network.

A genetic algorithm optimized artificial neural network GNW has been designed to rank features for two diversified multivariate data sets. The dimensions of these data sets are 85x24 and 62x25 for 24 or 25 molecular descriptors being computed for 85 matrix metalloproteinase-1 inhibitors or 62 hepatitis C virus NS3 protease inhibitors, respectively. Each molecular descriptor computed is treated as a feature and input into an input layer node of the artificial neural network. To optimize the artificial neural network by the genetic algorithm, each interconnected weight between input and hidden or between hidden and output layer nodes is binary encoded as a 16 bits string in a chromosome, and the chromosome is evolved by crossover and mutation operations. Each input layer node and its associated weights of the trained GNW are systematically omitted once (the self-depleted weights), and the corresponding weight adjustments due to the omission are computed to keep the overall network behavior unchanged. The primary feature ranking index defined as the sum of self-depleted weights and the corresponding weight adjustments computed is found capable of separating good from bad features for some artificial data sets of known feature rankings tested. The final feature indexes used to rank the data sets are computed as a sum of the weighted frequency of each feature being ranked in a particular rank for each data set being partitioned into numerous clusters. The two data sets are also clustered by a standard K-means method and trained by a support vector machine (SVM) for feature ranking using the computed F-scores as feature ranking index. It is found that GNW outperforms the SVM method on three artificial as well as the matrix metalloproteinase-1 inhibitor data sets studied. A clear-cut separation of good from bad features is offered by the GNW but not by the SVM method for a feature pool of known feature ranking.

Algorithms↗