Search PubMed⌕ Search

Biomedical subjects

S Henikoff

Publications and source records attributed to S Henikoff.

At least 73 records · Page 4Linked to original sources

Embedding strategies for effective use of information from multiple sequence alignments.

We describe a new strategy for utilizing multiple sequence alignment information to detect distant relationships in searches of sequence databases. A single sequence representing a protein family is enriched by replacing conserved regions with position-specific scoring matrices (PSSMs) or consensus residues derived from multiple alignments of family members. In comprehensive tests of these and other family representations, PSSM-embedded queries produced the best results overall when used with a special version of the Smith-Waterman searching algorithm. Moreover, embedding consensus residues instead of PSSMs improved performance with readily available single sequence query searching programs, such as BLAST and FASTA. Embedding PSSMs or consensus residues into a representative sequence improves searching performance by extracting multiple alignment information from motif regions while retaining single sequence information where alignment is uncertain.

Algorithms↗

A helix-turn-helix DNA-binding motif predicted for transposases of DNA transposons.

A helix-turn-helix (HTH) DNA-binding motif is identified in transposase sequences in Tc1, mariner and pogo DNA transposum. The findings are supported by results of various sequence analysis methods. Tc1 transposases are also predicted to contain another DNA-binding region. These findings are in accord with experimental evidence obtained from Tc1A, Tc3A and pogo transposases. The pogo family transposases, but not the pogo-type transcription factors, contain the HTH motif, suggesting that HTH structures are essential for Tc1/mariner/pogo transposition. Analysis of multiple sequence alignments enabled the identification of the HTH motif in distantly related protein sequences.

Amino Acid Sequence↗

Nuclear organization and gene expression: homologous pairing and long-range interactions.

Genetic studies have demonstrated that pairing interactions between homologous chromosomes and long-range associations between nonhomologous sites can influence gene expression. Recent work has revealed that such influences are widespread in eukaryotes and that chromosome architecture is likely to be of fundamental importance for nuclear structure and function.

Alleles↗

Heterochromatic trans-inactivation of Drosophila white transgenes.

Position effect variegation of most Drosophila melanogaster genes, including the white eye pigment gene is recessive. We find that this is not always the case for white transgenes. Three examples are described in which a lesion causing variegation is capable of silencing the white transgene on the paired homologue (trans-inactivation). These examples include two different transgene constructs inserted at three distinct genomic locations. The lesions that cause variegation of white minimally disrupt the linear order of genes on the chromosomes, permitting close homologous pairing. At one of these sites, trans-inactivation has also been extended to include a vital gene in the vicinity of the white transgene insertion. These findings suggest that many Drosophila genes, in many positions in the genome, can sense the heterochromatic state of a paired homologue.

ATP-Binding Cassette Transporters↗

Transgene repeat arrays interact with distant heterochromatin and cause silencing in cis and trans.

Tandem repeats of Drosophila transgenes can cause heterochromatic variegation for transgene expression in a copy-number and orientation-dependent manner. Here, we demonstrate different ways in which these transgene repeat arrays interact with other sequences at a distance, displaying properties identical to those of a naturally occurring block of interstitial heterochromatin. Arrays consisting of tandemly repeated white transgenes are strongly affected by proximity to constitutive heterochromatin. Moving an array closer to heterochromatin enhanced variegation, and enhancement was reverted by recombination of the array onto a normal sequence chromosome. Rearrangements that lack the array enhanced variegation of white on a homologue bearing the array. Therefore, silencing of white genes within a repeat array depends on its distance from heterochromatin of the same chromosome or of its paired homologue. In addition, white transgene arrays cause variegation of a nearby gene in cis, a hallmark of classical position-effect variegation. Such spreading of heterochromatic silencing correlates with array size. Finally, white transgene arrays cause pairing-dependent silencing of a non-variegating white insertion at the homologous position.

ATP-Binding Cassette Transporters↗

Genetic modification of heterochromatic association and nuclear organization in Drosophila.

Heterochromatin is the highly compact, usually pericentromeric, region of eukaryotic chromosomes. Unlike the more gene-rich euchromatin, heterochromatin remains condensed during interphase, when it is sequestered to the periphery of the nucleus. Here we show, by using fluorescent in situ hybridization to interphase diploid nuclei of Drosophila, that the insertion of heterochromatin into a euchromatic gene, which results in position-effect variegation (PEV), also causes the aberrant association of the gene and its homologous copy with heterochromatin. In correlation with the gene's mutant variegating phenotype, the cytological association of the heterochromatic region is affected by chromosomal distance from heterochromatin and by genic modifiers of PEV. Proteins that are thought to be involved in the formation of heterochromatin can therefore influence the interphase nuclear position of a chromosomal region. This suggests that heterochromatin and proteins involved in its formation provide a structural framework for the interphase nucleus.

Animals↗

The Blocks database--a system for protein classification.

The Blocks Database contains multiple alignments of conserved regions in protein families. The database can be searched by e-mail and World Wide Web(WWW) servers (http://blocks.fhcrc.org/help) to classify protein and nucleotide sequences.

Amino Acid Sequence↗

Dosage-dependent modification of position-effect variegation in Drosophila.

Many loci in Drosophila exhibit dosage effects on single phenotypes. In the case of modifiers of position-effect variegation, increases and decreases in dosage can have opposite effects on variegating phenotypes. This is seemingly paradoxical: if each locus encodes a limiting gene product sensitive to dosage decreases, then increasing the dosage of any one should have no effect, because the others should remain limiting. An earlier model put forward to resolve this paradox suggested that dosage-dependent modifiers encode protein subunits of a macromolecular complex that is sensitive to mass action equilibrium conditions. Because chemical equilibria are dynamic, however, such hypothetical complexes will be unstable to an extent that is inconsistent with the known properties of molecules that make up chromatin. An alternative model accounts for the dosage effects in terms of interactions between structural proteins that bind at multiple linked sites. These might include indirect interactions occurring between regulatory proteins and genes for structural proteins or their protein products. The large number of direct and inverse regulatory genes which are known to exist in Drosophila could account for the apparent genetic complexity that is seen for modifiers of position-effect variegation and for other systems of phenotypic modification.

Animals↗

Introduction of a DNA methyltransferase into Drosophila to probe chromatin structure in vivo.

The dam DNA methyltransferase gene from Escherichia coli was introduced into Drosophila in order to probe chromatin structure in vivo. Expression of the gene caused no visible defects or developmental delay even at high levels of active methylase. About half of each target site was found to be methylated in vivo, apparently reflecting a general property of chromatin packaged in nucleosomes. Although site-specific differences were detected, most euchromatic and heterochromatic sites showed comparable degrees of methylation, at least at high methylase levels. Methylase accessibility of a lacZ reporter gene subject to position-effect variegation throughout development was only slightly reduced, consistent with studies of chromatin accessibility in vitro. Silencing of lacZ during development differed from silencing of an adjacent white eye pigment reporter gene in the adult, as though chromatin structure can undergo dynamic alterations during development.

Animals↗

Blocks database and its applications.

Protein blocks consist of multiply aligned sequence segments without gaps that represent the most highly conserved regions of protein families. A database of blocks has been constructed by successive application of the fully automated PROTOMAT system to lists of protein family members obtained from Prosite documentation. Currently, Blocks 8.0 based on protein families documented in Prosite 12 consists of 2884 blocks representing 770 families. Searches of the Blocks Database are carried out using protein or DNA sequence queries, and results are returned with measures of significance for both single and multiple block hits. The databse has also proved useful for derivation of amino acid substitution matrices (the Blosum series) and other sets of parameters. WWW and E-mail servers provide access to the database and associated functions, including a block maker for sequences provided by the user.

Amino Acid Sequence↗

Scores for sequence searches and alignments.

Every sequence comparison method requires a set of scores. For aligning protein sequences, substitution scores are based on models of amino acid conservation and properties, and matrices of these scores have substantially improved in recent years. Position-specific scoring matrices provide representations of sequence families that are capable of detecting subtle similarities. Comprehensive evaluations can effectively guide the choice of scores for sequence alignment and searching applications, including those that aid in the prediction of protein structures.

Amino Acid Sequence↗

Using substitution probabilities to improve position-specific scoring matrices.

Each column of amino acids in a multiple alignment of protein sequences can be represented as a vector of 20 amino acid counts. For alignment and searching applications, the count vector is an imperfect representation of a position, because the observed sequences are an incomplete sample of the full set of related sequences. One general solution to this problem is to model unobserved sequences by adding artificial 'pseudo-counts' to the observed counts. We introduce a simple method for computing pseudo-counts that combines the diversity observed in each alignment position with amino acid substitution probabilities. In extensive empirical tests, this position-based method out-performed other pseudo-count methods and was a substantial improvement over the traditional average score method used for constructing profiles.

Amino Acid Sequence↗

Copy number and orientation determine the susceptibility of a gene to silencing by nearby heterochromatin in Drosophila.

The classical phenomenon of position-effect variegation (PEV) is the mosaic expression that occurs when a chromosomal rearrangement moves a euchromatic gene near heterochromatin. A striking feature of this phenomenon is that genes far away from the junction with heterochromatin can be affected, as if the heterochromatic state "spreads." We have investigated classical PEV of a Drosophila brown transgene affected by a heterochromatic junction approximately 60 kb away. PEV was enhanced when the transgene was locally duplicated using P transposase. Successive rounds of P transposase mutagenesis and phenotypic selection produced a series of PEV alleles with differences in phenotype that depended on transgene copy number and orientation. As for other examples of classical PEV, nearby heterochromatin was required for gene silencing. Modifications of classical PEV by alterations at a single site are unexpected, and these observations contradict models for spreading that invoke propagation of heterochromatin along the chromosome. Rather, our results support a model in which local alterations affect the affinity of a gene region for nearby heterochromatin via homology-based pairing, suggesting an alternative explanation for this 65-year-old phenomenon.

Animals↗

Automated construction and graphical presentation of protein blocks from unaligned sequences.

Protein blocks consist of multiply aligned sequence segments that correspond to the most highly conserved regions of protein families. Typically, a set of related proteins has more than one region in common and their relationship can be represented as a series of ungapped blocks separated by unaligned regions. Blockmaker is an automated system available by electronic mail (blockmaker@howard.fhcrc.org) and the World Wide Web (http://www.blocks.fhcrc.org4) that finds blocks in a group of related protein sequences submitted by the user. It adapts and extends existing algorithms to make them useful to biologists looking for conserved regions in a group of related proteins sequences. Two sets of blocks are returned, one in which candidate blocks are detected using the MOTIF algorithm and the other using a Gibbs sampler algorithm that has been adapted for full automation. This use of two block-finding methods based on completely different principles provides a 'reality check,' whereby a block detected by both methods is considered to be correct. Resulting blocks can be displayed using the information-based 'sequence logo' method, adapted to incorporate sequence weights, which provides an intuitive visual description of both the residue and the conservation information at each position. Blocks generated by this system are useful in diverse applications, such as searching databases and designing degenerate PCR primers. As an example, blocks made from amino acid sequences related to Caenorhabditis elegans Tc1 transposase were used to search GenBank, revealing that several fish and amphibian genomic sequences harbor previously unreported Tc1 homologs.

Algorithms↗