Search PubMedSearch

Biomedical subjects

T J Gibson

Publications and source records attributed to T J Gibson.

At least 19 recordsLinked to original sources

PairWise and SearchWise: finding the optimal alignment in a simultaneous comparison of a protein profile against all DNA translation frames.

DNA translation frames can be disrupted for several reasons, including: (i) errors in sequence determination; (ii) RNA processing, such as intron removal and guide RNA editing; (iii) less commonly, polymerase frameshifting during transcription or ribosomal frameshifting during translation. Frameshifts frequently confound computational activities involving homologous sequences, such as database searches and inferences on structure, function or phylogeny made from multiple alignments. A dynamic alignment algorithm is reported here which compares a protein profile (a residue scoring matrix for one or more aligned sequences) against the three translation frames of a DNA strand, allowing frameshifting. The algorithm has been incorporated into a new package, WiseTools, for comparison of biological sequences. A protein profile can be compared against either a DNA sequence or a protein sequence. The program PairWise may be used interactively for alignment of any two sequence inputs. SearchWise can perform combinations of searches through DNA or protein databases by a protein profile or DNA sequence. Routine application of the programs has revealed a set of database entries with frameshifts caused by errors in sequence determination.

Algorithms

Three-dimensional structure and stability of the KH domain: molecular insights into the fragile X syndrome.

The KH module is a sequence motif found in a number of proteins that are known to be in close association with RNA. Experimental evidence suggests a direct involvement of KH in RNA binding. The human FMR1 protein, which has two KH domains, is associated with fragile X syndrome, the most common inherited cause of mental retardation. Here we present the three-dimensional solution structure of the KH module. The domain consists of a stable beta alpha alpha beta beta alpha fold. On the basis of our results, we suggest a potential surface for RNA binding centered on the loop between the first two helices. Substitution of a well-conserved hydrophobic residue located on the second helix destroys the KH fold; a mutation of this position in FMR1 leads to an aggravated fragile X phenotype.

Asparagine

Using CLUSTAL for multiple sequence alignments.

We have tested CLUSTAL W in a wide variety of situations, and it is capable of handling some very difficult protein alignment problems. If the data set consists of enough closely related sequences so that the first alignments are accurate, then CLUSTAL W will usually find an alignment that is very close to ideal. Problems can still occur if the data set includes sequences of greatly different lengths or if some sequences include long regions that are impossible to align with the rest of the data set. Trying to balance the need for long insertions and deletions in some alignments with the need to avoid them in others is still a problem. The default values for our parameters were tested empirically using test cases of sets of globular proteins where some information as to the correct alignment was available. The parameter values may not be very appropriate with nonglobular proteins. We have argued that using one weight matrix and two gap penalties is too simplistic to be of general use in the most difficult cases. We have replaced these parameters with a large number of new parameters designed primarily to help encourage gaps in loop regions. Although these new parameters are largely heuristic in nature, they perform surprisingly well and are simple to implement. The underlying speed of the progressive alignment approach is not adversely affected. The disadvantage is that the parameter space is now huge; the number of possible combinations of parameters is more than can easily be examined by hand. We justify this by asking the user to treat CLUSTAL W as a data exploration tool rather than as a definitive analysis method. It is not sensible to automatically derive multiple alignments and to trust particular algorithms as being capable of always getting the correct answer. One must examine the alignments closely, especially in conjunction with the underlying phylogenetic tree (or estimate of it) and try varying some of the parameters. Outliers (sequences that have no close relatives) should be aligned carefully, as should fragments of sequences. The program will automatically delay the alignment of any sequences that are less than 40% identical to any others until all other sequences are aligned, but this can be set from a menu by the user. It may be useful to build up an alignment of closely related sequences first and to then add in the more distant relatives one at a time or in batches, using the profile alignments and weighting scheme described earlier and perhaps using a variety of parameter settings. We give one example using SH2 domains. SH2 domains are widespread in eukaryotic signalling proteins where they function in the recognition of phosphotyrosine-containing peptides. In the chapter by Bork and Gibson ([11], this volume), Blast and pattern/profile searches were used to extract the set of known SH2 domains and to search for new members. (Profiles used in database searches are conceptually very similar to the profiles used in CLUSTAL W: see the chapters [11] and [13] for profile search methods.) The profile searches detected SH2 domains in the JAK family of protein tyrosine kinases, which were thought not to contain SH2 domains. Although the JAK family SH2 domains are rather divergent, they have the necessary core structural residues as well as the critical positively charged residue that binds phosphotyrosine, leaving no doubt that they are bona fide SH2 domains. The five new JAK family SH2 domains were added sequentially to the existing alignment of 65 SH2 domains using the CLUSTAL W profile alignment option. Figure 6 shows part of the resulting alignment. Despite their divergent sequences, the new SH2 domains have been aligned nearly perfectly with the old set. No insertions were placed in the original SH2 domains. In this example, the profile alignment procedure has produced better results than a one-step full alignment of all 70 SH2 domains, and in considerably less time. (ABSTRACT TRUNCATED)

Amino Acid Sequence

Dystrophin and utrophin: the missing links!

There is considerable sequence homology between dystrophin and utrophin, both at the protein and DNA level, and consequently it was assumed that their domain structures and functions would be similar. As more of the detailed biochemical and cell biological properties of these two proteins become known, so it becomes clear that there are subtle if not significant differences between them. We review recent findings and present new hypotheses into the structural and functional properties of the actin-binding domain, central coiled-coil region and regulatory/membrane protein-binding regions of dystrophin and utrophin.

Actins

Structure of the dsRNA binding domain of E. coli RNase III.

The double-stranded RNA binding domain (dsRBD) is a approximately 70 residue motif found in a variety of modular proteins exhibiting diverse functions, yet always in association with dsRNA. We report here the structure of the dsRBD from RNase III, an enzyme present in most, perhaps all, living cells. It is involved in processing transcripts, such as rRNA precursors, by cleavage at short hairpin sequences. The RNase III protein consists of two modules, a approximately 150 residue N-terminal catalytic domain and a approximately 70 residue C-terminal recognition module, homologous with other dsRBDs. The structure of the dsRBD expressed in Escherichia coli has been investigated by homonuclear NMR techniques and solved with the aid of a novel calculation strategy. It was found to have an alpha-beta-beta-beta-alpha topology in which a three-stranded anti-parallel beta-sheet packs on one side against the two helices. Examination of 44 aligned dsRBD sequences reveals several conserved, positively charged residues. These residues map to the N-terminus of the second helix and a nearby loop, leading to a model for the possible contacts between the domain and dsRNA.

Amino Acid Sequence

CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice.

The sensitivity of the commonly used progressive multiple sequence alignment method has been greatly improved for the alignment of divergent protein sequences. Firstly, individual weights are assigned to each sequence in a partial alignment in order to down-weight near-duplicate sequences and up-weight the most divergent ones. Secondly, amino acid substitution matrices are varied at different alignment stages according to the divergence of the sequences to be aligned. Thirdly, residue-specific gap penalties and locally reduced gap penalties in hydrophilic regions encourage new gaps in potential loop regions rather than regular secondary structure. Fourthly, positions in early alignments where gaps have been opened receive locally reduced gap penalties to encourage the opening up of new gaps at these positions. These modifications are incorporated into a new program, CLUSTAL W which is freely available.

Algorithms

Detection of dsRNA-binding domains in RNA helicase A and Drosophila maleless: implications for monomeric RNA helicases.

Searches with dsRNA-binding domain profiles detected two copies of the domain in each of RNA helicase A, Drosophila maleless and C. elegans ORF T20G5-11 (of unknown function). RNA helicase A is unusual in being one of the few characterised DEAD/DExH helicases that are active as monomers. Other monomeric DEAD/DExH RNA helicases (p68, NPH-II) have domains that match another RNA-binding motif, the RGG repeat. The DEAD/DExH domain appears to be insufficient on its own to promote helicase activity and additional RNA-binding capacity must be supplied either as domains adjacent to the DEAD/DExH-box or by bound partners as in the eIF-4AB dimer. The presence or absence of extra RNA-binding domains should allow classification of DEAD/DExH proteins as monomeric or multimeric helicases.

Amino Acid Sequence

Evidence for a protein domain superfamily shared by the cyclins, TFIIB and RB/p107.

Cyclins, TFIIB and RB play major roles in cell cycle and/or gene regulation. Earlier work has suggested common ancestry for the TFIIB repeats and RB pocket B which share 20% sequence identity. We now report that database searches with profiles based on a multiple alignment of cyclin core regions (the 'cyclin box') detect the TFIIB repeats with equivalent scores to divergent cyclins. Several features of the sequences support the notion of common ancestry: e.g. cyclins A/B, C and D share approximately 20-30% identity but each have approximately 15-20% identity with vertebrate TFIIB, showing that conserved cyclin features underlie the match. These results suggest the presence of a domain superfamily, which we term the TR domain, in nuclear regulatory proteins belonging to the TFIIB, cyclin and RB families, that has been duplicated many times during eukaryotic evolution. The TR domain appears to function in protein-protein interactions.

Amino Acid Sequence

The evolution of titin and related giant muscle proteins.

Titin and twitchin are giant proteins expressed in muscle. They are mainly composed of domains belonging to the fibronectin class III and immunoglobulin c2 families, repeated many times. In addition, both proteins have a protein kinase domain near the C-terminus. This paper explores the evolution of these and related muscle proteins in an attempt to determine the order of events that gave rise to the different repeat patterns and the order of appearance of the proteins. Despite their great similarity at the level of sequence organization, titin and twitchin diverged from each other at least as early as the divergence between vertebrates and nematodes. Most of the repeating units in titin and twitchin were estimated to derive from three original domains. Chicken smooth-muscle myosin light-chain kinase (smMLCK) also has a kinase domain, several immunoglobulin domains, and a fibronectin domain. From a comparison of the kinase domains, titin is predicted to have appeared first during the evolution of the family, followed by twitchin and with the vertebrate MLCKs last to appear. The so-called C-protein from chicken is also a member of this family but has no kinase domain. Its origin remains unclear but it most probably pre-dates the titin/twitchin duplication.

Animals

Improved sensitivity of profile searches through the use of sequence weights and gap excision.

Position-specific substitution matrices, known as profiles, derived from multiple sequence alignments are currently used to search sequence databases for distantly related members of protein families. The performance of the database searches is enhanced by using (i) a sequence weighting scheme which assigns higher weights to more distantly related sequences based on branch lengths derived from phylogenetic trees, (ii) exclusion of positions with mainly padding characters at sites of insertions or deletions and (iii) the BLOSUM62 residue comparison matrix. A natural consequence of these modifications is an improvement in the alignment of new sequences to the profiles. However, the accuracy of the alignments can be further increased by employing a similarity residue comparison matrix. These developments are implemented in a program called PROFILEWEIGHT which runs on Unix and Vax computers. The only input required by the program is the multiple sequence alignment. The output from PROFILEWEIGHT is a profile designed to be used by existing searching and alignment programs. Test results from database searches with four different families of proteins show the improved sensitivity of the weighted profiles.

Algorithms

The KH domain occurs in a diverse set of RNA-binding proteins that include the antiterminator NusA and is probably involved in binding to nucleic acid.

New findings are presented for the approximately 50 residue KH motif, a domain recently discovered in RNA-binding proteins. The conserved sequence is approximately 10 residues larger than previously reported. Profile searches have revealed new members of this family, including two, E. coli NusA and human GAP-associated p62 phosphoprotein, for which RNA-binding data exists. A nusA homolog was detected in the RNA polymerase gene complex of six archaebacterial species and may encode an antiterminator. All KH-containing proteins are linked with RNA and the KH motif most probably functions as a nucleic acid binding domain.

Bacterial Proteins

Proposed structure for the DNA-binding domain of the helix-loop-helix family of eukaryotic gene regulatory proteins.

A modelled tertiary structure for the dimeric HLH domain of the E47 protein is presented. Structural information was obtained from the aligned sequences of > 40 members of the HLH family. The information was used to model each monomer as an alpha-helical hairpin, with knobs-into-holes packing of side-chains as found in antiparallel coiled-coil. The dimer forms a four-helix bundle with additional knobs-into-holes packing at the dimer interface. The size and electrostatic properties of core-forming residues are all accounted for in the model. The model does not violate any known properties of protein structure. The monomers are related by two-fold rotational symmetry, in agreement with the observed DNA-binding sites which are imperfect inverted repeats. The N-terminal basic region, in which DNA binding and base specificity reside, forms the first part of helix 1. A prediction based on the model structure is that the HLH domains do not bind to DNA in its B form but require a partially unwound conformation in order to enter the major groove.

Amino Acid Sequence