Search PubMed⌕ Search

Biomedical subjects

Stephen F Altschul

Publications and source records attributed to Stephen F Altschul.

6 recordsLinked to original sources

Protein database searches using compositionally adjusted substitution matrices.

Almost all protein database search methods use amino acid substitution matrices for scoring, optimizing, and assessing the statistical significance of sequence alignments. Much care and effort has therefore gone into constructing substitution matrices, and the quality of search results can depend strongly upon the choice of the proper matrix. A long-standing problem has been the comparison of sequences with biased amino acid compositions, for which standard substitution matrices are not optimal. To address this problem, we have recently developed a general procedure for transforming a standard matrix into one appropriate for the comparison of two sequences with arbitrary, and possibly differing compositions. Such adjusted matrices yield, on average, improved alignments and alignment scores when applied to the comparison of proteins with markedly biased compositions. Here we review the application of compositionally adjusted matrices and consider whether they may also be applied fruitfully to general purpose protein sequence database searches, in which related sequence pairs do not necessarily have strong compositional biases. Although it is not advisable to apply compositional adjustment indiscriminately, we describe several simple criteria under which invoking such adjustment is on average beneficial. In a typical database search, at least one of these criteria is satisfied by over half the related sequence pairs. Compositional substitution matrix adjustment is now available in NCBI's protein-protein version of blast.

Algorithms↗

A structure-based method for protein sequence alignment.

MOTIVATION: With the continuing rapid growth of protein sequence data, protein sequence comparison methods have become the most widely used tools of bioinformatics. Among these methods are those that use position-specific scoring matrices (PSSMs) to describe protein families. PSSMs can capture information about conserved patterns within families, which can be used to increase the sensitivity of searches for related sequences. Certain types of structural information, however, are not generally captured by PSSM search methods. Here we introduce a program, Structure-based ALignment TOol (SALTO), that aligns protein query sequences to PSSMs using rules for placing and scoring gaps that are consistent with the conserved regions of domain alignments from NCBI's Conserved Domain Database. RESULTS: In most cases, the alignment scores obtained using the local alignment version follow an extreme value distribution. SALTO's performance in finding related sequences and producing accurate alignments is similar to or better than that of IMPALA; one advantage of SALTO is that it imposes an explicit gapping model on each protein family. AVAILABILITY: A stand-alone version of the program that can generate global or local alignments is available by ftp distribution (ftp://ftp.ncbi.nih.gov/pub/SALTO/), and has been incorporated to Cn3D structure/alignment viewer. CONTACT: bryant@ncbi.nlm.nih.gov.

Algorithms↗

The construction of amino acid substitution matrices for the comparison of proteins with non-standard compositions.

MOTIVATION: Amino acid substitution matrices play a central role in protein alignment methods. Standard log-odds matrices, such as those of the PAM and BLOSUM series, are constructed from large sets of protein alignments having implicit background amino acid frequencies. However, these matrices frequently are used to compare proteins with markedly different amino acid compositions, such as transmembrane proteins or proteins from organisms with strongly biased nucleotide compositions. It has been argued elsewhere that standard matrices are not ideal for such comparisons and, furthermore, a rationale has been presented for transforming a standard matrix for use in a non-standard compositional context. RESULTS: This paper presents the mathematical details underlying the compositional adjustment of amino acid or DNA substitution matrices.

Algorithms↗

The compositional adjustment of amino acid substitution matrices.

Amino acid substitution matrices are central to protein-comparison methods. In most commonly used matrices, the substitution scores take a log-odds form, involving the ratio of "target" to "background" frequencies derived from large, carefully curated sets of protein alignments. However, such matrices often are used to compare protein sequences with amino acid compositions that differ markedly from the background frequencies used for the construction of the matrices. Of course, the target frequencies should be adjusted in such cases, but the lack of an appropriate way to do this has been a long-standing problem. This article shows that if one demands consistency between target and background frequencies, then a log-odds substitution matrix implies a unique set of target and background frequencies as well as a unique scale. Standard substitution matrices therefore are truly appropriate only for the comparison of proteins with standard amino acid composition. Accordingly, we present and evaluate a rationale for transforming the target frequencies implicit in a standard matrix to frequencies appropriate for a nonstandard context. This rationale yields asymmetric matrices for the comparison of proteins with divergent compositions. Earlier approaches are unable to deal with this case in a fully consistent manner. Composition-specific substitution matrix adjustment is shown to be of utility for comparing compositionally biased proteins, including those of organisms with nucleotide-biased, and therefore codon-biased, genomes or isochores.

Amino Acid Sequence↗

Expression of a recombinant IRP-like Plasmodium falciparum protein that specifically binds putative plasmodial IREs.

Plasmodium falciparum iron regulatory-like protein (PfIRPa, accession AJ012289) has homology to a family of iron-responsive element (IRE)-binding proteins (IRPs) found in different species. We have previously demonstrated that erythrocyte P. falciparum PfIRPa binds a mammalian consensus IRE and that the binding activity is regulated by iron status. In the work we now report, we have cloned a C-terminus histidine-tagged PfIRPa and overexpressed it in a bacterial expression system in soluble form capable of binding IREs. To overexpress PfIRPa, we used the T7 promoter-driven vector, pET28a(+), in conjunction with the Rosetta(DE3)pLysS strain of E. coli, which carries extra copies of tRNA genes usually found in organisms such as P. falciparum whose genome is (A+T)-rich. The histidine-tagged recombinant protein (rPfIRPa) in soluble form was partially purified using His-bind resin. We searched the plasmodial database, plasmoDB, to identify sequences capable of forming IRE loops using a specially developed algorithm, and found three plasmodial sequences matching the search criteria. In gel retardation assays, rPfIRPa bound three 32P-labeled putative plasmodial IREs with affinity exceeding the affinity for the mammalian consensus IRE. The binding was concentration-dependent and was not inhibited by heparin, an inhibitor of non-specific binding. Immunodepletion of rPfIRPa resulted in substantial inhibition of the signal intensity in the gel retardation assays and in Western blot-determinations of rPfIRPa protein levels. Endogenous PfIRPa retained all three putative 32P-IREs at the same position on the gel as the recombinant PfIRPa.

Animals↗

Generation and initial analysis of more than 15,000 full-length human and mouse cDNA sequences.

The National Institutes of Health Mammalian Gene Collection (MGC) Program is a multiinstitutional effort to identify and sequence a cDNA clone containing a complete ORF for each human and mouse gene. ESTs were generated from libraries enriched for full-length cDNAs and analyzed to identify candidate full-ORF clones, which then were sequenced to high accuracy. The MGC has currently sequenced and verified the full ORF for a nonredundant set of >9,000 human and >6,000 mouse genes. Candidate full-ORF clones for an additional 7,800 human and 3,500 mouse genes also have been identified. All MGC sequences and clones are available without restriction through public databases and clone distribution networks (see http:mgc.nci.nih.gov).

Algorithms↗