Search PubMed⌕ Search

Biomedical subjects

Guoli Wang

Publications and source records attributed to Guoli Wang.

10 recordsLinked to original sources

LS-NMF: a modified non-negative matrix factorization algorithm utilizing uncertainty estimates.

BACKGROUND: Non-negative matrix factorisation (NMF), a machine learning algorithm, has been applied to the analysis of microarray data. A key feature of NMF is the ability to identify patterns that together explain the data as a linear combination of expression signatures. Microarray data generally includes individual estimates of uncertainty for each gene in each condition, however NMF does not exploit this information. Previous work has shown that such uncertainties can be extremely valuable for pattern recognition. RESULTS: We have created a new algorithm, least squares non-negative matrix factorization, LS-NMF, which integrates uncertainty measurements of gene expression data into NMF updating rules. While the LS-NMF algorithm maintains the advantages of original NMF algorithm, such as easy implementation and a guaranteed locally optimal solution, the performance in terms of linking functionally related genes has been improved. LS-NMF exceeds NMF significantly in terms of identifying functionally related genes as determined from annotations in the MIPS database. CONCLUSION: Uncertainty measurements on gene expression data provide valuable information for data analysis, and use of this information in the LS-NMF algorithm significantly improves the power of the NMF technique.

Algorithms↗

PISCES: recent improvements to a PDB sequence culling server.

PISCES is a database server for producing lists of sequences from the Protein Data Bank (PDB) using a number of entry- and chain-specific criteria and mutual sequence identity. Our goal in culling the PDB is to provide the longest list possible of the highest resolution structures that fulfill the sequence identity and structural quality cut-offs. The new PISCES server uses a combination of PSI-BLAST and structure-based alignments to determine sequence identities. Structure alignment produces more complete alignments and therefore more accurate sequence identities than PSI-BLAST. PISCES now allows a user to cull the PDB by-entry in addition to the standard culling by individual chains. In this scenario, a list will contain only entries that do not have a chain that has a sequence identity to any chain in any other entry in the list over the sequence identity cut-off. PISCES also provides fully annotated sequences including gene name and species. The server allows a user to cull an input list of entries or chains, so that other criteria, such as function, can be used. Results from a search on the re-engineered RCSB's site for the PDB can be entered into the PISCES server by a single click, combining the powerful searching abilities of the PDB with PISCES's utilities for sequence culling. The server's data are updated weekly. The server is available at http://dunbrack.fccc.edu/pisces.

Databases, Protein↗

Critical hydrophobic interactions between phosphorylation and actuator domains of Ca2+-ATPase for hydrolysis of phosphorylated intermediate.

Functional roles of seven hydrophobic residues on the interface between the actuator (A) and phosphorylation (P) domains of sarcoplasmic reticulum Ca2+-ATPase were explored by alanine and serine substitutions. The residues examined were Ile179/Leu180/Ile232 on the A domain, Val705/Val726 on the P domain, and Leu119/Tyr122 on the loop linking the A domain and M2 (the second transmembrane helix). These residues gather to form a hydrophobic cluster around Tyr122 in the crystal structures of Ca2+-ATPase in Ca2+-unbound E2 (unphosphorylated) and E2P (phosphorylated) states but are far apart in those of Ca2+-bound E1 (unphosphorylated) and E1P (phosphorylated) states. The substitution-effects were also compared with those of Ile235 on the A domain/M3 linker and those of T181GE of the A domain, since they are in the immediate vicinity of the Tyr122-cluster. All these substitutions almost completely inhibited ATPase activity without inhibiting Ca2+-activated E1P formation from ATP. Substitutions of Ile235 and T181GE blocked the E1P to E2P transition, whereas those in the Tyr122-cluster blocked the subsequent E2P hydrolysis. Substitutions of Ile235 and Glu183 also blocked EP hydrolysis. Results indicate that the Tyr122-cluster is formed during the E1P to E2P transition to configure the catalytic site and position Glu183 properly for hydrolyzing the acylphosphate. Ile235 on the A domain/M3 linker likely forms hydrophobic interactions with the A domain and thereby allowing the strain of this linker to be utilized for large motions of the A domain during these processes. The Tyr122-cluster, Ile235, and T181GE thus seem to have different roles and are critical in the successive events in processing phosphorylated intermediates to transport Ca2+.

Adenosine Triphosphatases↗

Quasi-consensus-based comparison of profile hidden Markov models for protein sequences.

A simple approach for the sensitive detection of distant relationships among protein families and for sequence-structure alignment via comparison of hidden Markov models based on their quasi-consensus sequences is presented. Using a previously published benchmark dataset, the approach is demonstrated to give better homology detection and yield alignments with improved accuracy in comparison to an existing state-of-the-art dynamic programming profile-profile comparison method. This method also runs significantly faster and is therefore suitable for a server covering the rapidly increasing structure database. A server based on this method is available at http://liao.cis.udel.edu/website/servers/modmod

Algorithms↗

Domain definition and target classification for CASP6.

Assessment of structure predictions in CASP6 was based on single domains isolated from experimentally determined structures, which were categorized into comparative modeling, fold recognition, and new fold targets. Domain definitions were defined upon visual examination of the structures with the aid of automated domain-parsing programs. Domain categorization was determined by comparison of the target structures with those in the Protein Data Bank at the time each target expired and a variety of sequence and structure-based methods to determine potential homologous relationships.

Amino Acid Sequence↗

Assessment of fold recognition predictions in CASP6.

The Sixth Community Wide Experiment on the Critical Assessment of Techniques for Protein Structure Prediction (CASP6) held in December 2004 focused on the prediction of the structures of 90 protein domains from 64 targets. Thirty-eight of these were classified as "fold recognition," defined as being similar in fold to proteins of known structure at the time of submission of the predictions. Only the "first" predictions and those longer than 20 amino acids for each domain were assessed, resulting in 4527 predictions from 165 groups. The assessment was accomplished by the use of six structure alignment programs and three scoring measures based on these alignments. The use of a variety of measures resulted in scoring insensitive to the peculiarities of any one alignment method. The top-ranked methods in the prediction of structures that were clearly homologous to proteins in the Protein Data Bank primarily used servers and other programs based on achieving a consensus of many remote homology detection and fold recognition methods. The top-ranked methods in prediction of structures less clearly related or unrelated to proteins of known structures used fragment building methods in addition to the fold recognition meta methods.

Algorithms↗

Scoring profile-to-profile sequence alignments.

Sequence alignment profiles have been shown to be very powerful in creating accurate sequence alignments. Profiles are often used to search a sequence database with a local alignment algorithm. More accurate and longer alignments have been obtained with profile-to-profile comparison. There are several steps that must be performed in creating profile-profile alignments, and each involves choices in parameters and algorithms. These steps include (1) what sequences to include in a multiple alignment used to build each profile, (2) how to weight similar sequences in the multiple alignment and how to determine amino acid frequencies from the weighted alignment, (3) how to score a column from one profile aligned to a column of the other profile, (4) how to score gaps in the profile-profile alignment, and (5) how to include structural information. Large-scale benchmarks consisting of pairs of homologous proteins with structurally determined sequence alignments are necessary for evaluating the efficacy of each scoring scheme. With such a benchmark, we have investigated the properties of profile-profile alignments and found that (1) with optimized gap penalties, most column-column scoring functions behave similarly to one another in alignment accuracy; (2) some functions, however, have much higher search sensitivity and specificity; (3) position-specific weighting schemes in determining amino acid counts in columns of multiple sequence alignments are better than sequence-specific schemes; (4) removing positions in the profile with gaps in the query sequence results in better alignments; and (5) adding predicted and known secondary structure information improves alignments.

Algorithms↗

PISCES: a protein sequence culling server.

PISCES is a public server for culling sets of protein sequences from the Protein Data Bank (PDB) by sequence identity and structural quality criteria. PISCES can provide lists culled from the entire PDB or from lists of PDB entries or chains provided by the user. The sequence identities are obtained from PSI-BLAST alignments with position-specific substitution matrices derived from the non-redundant protein sequence database. PISCES therefore provides better lists than servers that use BLAST, which is unable to identify many relationships below 40% sequence identity and often overestimates sequence identity by aligning only well-conserved fragments. PDB sequences are updated weekly. PISCES can also cull non-PDB sequences provided by the user as a list of GenBank identifiers, a FASTA format file, or BLAST/PSI-BLAST output.

Algorithms↗

Deletions of any single residues in Glu40-Ser48 loop connecting a domain and the first transmembrane helix of sarcoplasmic reticulum Ca(2+)-ATPase result in almost complete inhibition of conformational transition and hydrolysis of phosphoenzyme intermediate.

Possible roles of the Glu40-Ser48 loop connecting A domain and the first transmembrane helix (M1) in sarcoplasmic reticulum Ca(2+)-ATPase (SERCA1a) were explored by mutagenesis. Deletions of any single residues in this loop caused almost complete loss of Ca(2+)-ATPase activity, while their substitutions had no or only slight effects. Single deletions or substitutions in the adjacent N- and C-terminal regions of the loop (His32-Asn39 and Leu49-Ile54) had no or only slight effects except two specific substitutions of Asn39 found in SERCA2b in Darier's disease pedigrees. All the single deletion mutants for the Glu40-Ser48 loop and the specific Asn39 mutants formed phosphoenzyme intermediate (EP) from ATP, but their isomeric transition from ADP-sensitive EP (E1P) to ADP-insensitive EP (E2P) was almost completely or strongly inhibited. Hydrolysis of E2P formed from Pi was also dramatically slowed in these deletion mutants. On the other hand, the rates of the Ca(2+)-induced enzyme activation and subsequent E1P formation from ATP were not altered by the deletions and substitutions. The results indicate that the Glu40-Ser48 loop, with its appropriate length (but not with specific residues) and with its appropriate junction to A domain, is a critical element for the E1P to E2P transition and formation of the proper structure of E2P, therefore, most likely for the large rotational movement of A domain and resulting in its association with P and N domains. Results further suggest that the loop functions to coordinate this movement of A domain and the unique motion of M1 during the E1P to E2P transition.

Adenosine Triphosphate↗

CASA: a server for the critical assessment of protein sequence alignment accuracy.

SUMMARY: A public server for evaluating the accuracy of protein sequence alignment methods is presented. CASA is an implementation of the alignment accuracy benchmark presented by Sauder et al. (Proteins, 40, 6-22, 2000). The benchmark currently contains 39321 pairwise protein structure alignments produced with the CE program from SCOP domain definitions. The server produces graphical and tabular comparisons of the accuracy of a user's input sequence alignments with other commonly used programs, such as BLAST, PSI-BLAST, Clustal W, and SAM-T99. AVAILABILITY: The server is located at http://capb.dbi.udel.edu/casa.

Algorithms↗