Search PubMed⌕ Search

Biomedical subjects

S Karlin

Publications and source records attributed to S Karlin.

At least 91 records · Page 5Linked to original sources

An efficient algorithm for identifying matches with errors in multiple long molecular sequences.

An efficient algorithm is described for finding matches, repeats and other word relations, allowing for errors, in large data sets of long molecular sequences. The algorithm entails hashing on fixed-size words in conjunction with the use of a linked list connecting all occurrences of the same word. The average memory and run time requirement both increase almost linearly with the total sequence length. Some results of the program's performance on a database of Escherichia coli DNA sequences are presented.

Algorithms↗

Assessment of inhomogeneities in an E. coli physical map.

A statistical method based on r-fragments, sums of distances between (r + 1) consecutive restriction enzyme sites, is introduced for detecting nonrandomness in the distribution or too markers in sequence data. The technique is applicable whenever large numbers of markers are available and will detect clumping, excessive dispersion or too much evenness of spacing of the markers. It is particularly adapted to varying the scale on which inhomogeneities can be detected, from nearest neighbor interactions to more distant interactions. The r-fragment procedure is applied primarily to the Kohara et al. (1) physical map of E. coli. Other applications to DAM methylation sites in E. coli and NotI sites in human chromosome 21 are presented. Restriction sites for the eight enzymes used in (1) appear to be randomly distributed, although at widely differing densities. These conclusions are substantially in agreement with the analysis of Churchill et al. (3). Extreme variability in the density of the eight restriction enzyme sites cannot be explained by variability in mono-, di- or trinucleotide frequencies.

Base Sequence↗

Very long charge runs in systemic lupus erythematosus-associated autoantigens.

Systemic lupus erythematosus and other chronic systemic autoimmune diseases are associated with circulating autoantibodies reactive with a limited set of mostly nuclear proteins. Using rigorous statistical methods we have identified segments of highly significant charge concentration in the majority of the characteristic nuclear and cytoplasmic autoantigens. Extremely long runs of charged residues, including some sequences of greater than 20 consecutive charged residues (purely acidic or mixed basic and acidic), occur in about a third of these proteins, whereas equivalent runs are found in less than 3% of other mammalian proteins. The other sequences have less extreme charge clusters, the type and location of which are often conserved between several otherwise nonsimilar antigens. We propose that supercharged surfaces render the targeted host proteins strongly immunogenic and that antinuclear antibody profiles might result from chronic exposure to intracellular contents, possibly in conjunction with crossreactive viral products. The limited number of potential systemic autoantigens may partly be due to the rarity of requisite charge properties.

Amino Acid Sequence↗

Evidence for selective evolution in codon usage in conserved amino acid segments of human alphaherpesvirus proteins.

The genomes of human viruses herpes simplex 1 (HSV1) and varicella zoster (VZV), although similar in biology, largely concordant in gene order, and identical in many amino acid segments, differ widely in their genomic G + C (abbreviated S) content, which is high in HSV1 (68%) and low in VZV (46%). This paper analyzes several striking codon usage contrasts. The S difference in coding regions is dramatically large in codon site 3, S3, about 42%. The large difference in S3 is maintained at the same level in a subset of closely similar genes and even in corresponding identical amino acid blocks. A similar difference in S levels in silent site 1 (S1) is found in leucine and arginine. The difference in S3 levels occurs in every gene and in every multicodon amino acid form. The S difference also exists in amino acid usage, with HSV1 using significantly more codon types SSN, while VZV uses more codon types WWN (where W stands for A or T). The nonoverlapping and narrow histograms of S3 gene frequencies in both viruses suggest that the difference has arisen and been maintained by a process of selective rather than nonselective effects. This is in sharp contrast to the relatively large variance seen for highly similar genes in the human versus yeast analysis. Interpretations and hypotheses to explain the HSV1 vs VZV codon usage disparity relate to virus-host interactions, to the role of viral genes in DNA metabolism, to availability of molecular resources (molecular Gause exclusion principle), and to differences in genomic structure.

Amino Acids↗

Global convergence properties in multilocus viability selection models: the additive model and the Hardy-Weinberg law.

A natural coordinate system is introduced for the analysis of the global stability of the Hardy-Weinberg (HW) polymorphism under the general multilocus additive viability model. A global convergence criterion is developed and used to prove that the HW polymorphism is globally stable when each of the loci is diallelic, provided the loci are overdominant and the multilocus recombination is positive. As a corollary the multilocus Hardy-Weinberg law for neutral selection is derived.

Genetics, Population↗

Evolution of sexual preferences in quantitative characters.

An analysis of equilibria and dynamics of the means, variances, and covariances of female mating preference for a quantitative male secondary sexual character following a Gaussian model is presented. For many combinations of viability and sexual selection parameters the evolving Gaussian distribution of phenotypes can diverge. The results on the cases of convergence and their limiting forms suggest some reinterpretations of Fisher's "runaway" process of sexual selection. One possibility is to interpret Fisher's postulated "initial advantage not due to female preference" as a shift in viability selection where runaway evolution occurs if the mean preferred trait evolves beyond its new viability optimum (due to sexual selection). This definition is contrasted with situations in which the new viability optimum is undershot. The quantitative and qualitative conclusions differ from models that approximate genetic covariance evolution involving a constant covariance.

Biological Evolution↗

Levels of multiallelic overdominance fitness, heterozygote excess and heterozygote deficiency.

Concepts and results on selection balance in multiallelic systems are described. These include a multidimensional concept of heterozygote excess and heterozygote deficiency, a hierarchy of means of assessment of heterozygote advantage, comparisons and contrasts of allelic versus gametic polymorphic states, and conditions defining stable equilibria of complementary gametic sets. The concepts are illustrated in the context of viability selection and behavioral models of kin selection and for two major categories of multilocus selection regimes.

Alleles↗

Methods for assessing the statistical significance of molecular sequence features by using general scoring schemes.

An unusual pattern in a nucleic acid or protein sequence or a region of strong similarity shared by two or more sequences may have biological significance. It is therefore desirable to know whether such a pattern can have arisen simply by chance. To identify interesting sequence patterns, appropriate scoring values can be assigned to the individual residues of a single sequence or to sets of residues when several sequences are compared. For single sequences, such scores can reflect biophysical properties such as charge, volume, hydrophobicity, or secondary structure potential; for multiple sequences, they can reflect nucleotide or amino acid similarity measured in a wide variety of ways. Using an appropriate random model, we present a theory that provides precise numerical formulas for assessing the statistical significance of any region with high aggregate score. A second class of results describes the composition of high-scoring segments. In certain contexts, these permit the choice of scoring systems which are "optimal" for distinguishing biologically relevant patterns. Examples are given of applications of the theory to a variety of protein sequences, highlighting segments with unusual biological features. These include distinctive charge regions in transcription factors and protooncogene products, pronounced hydrophobic segments in various receptor and transport proteins, and statistically significant subalignments involving the recently characterized cystic fibrosis gene.

Amino Acid Sequence↗

Contrasts in codon usage of latent versus productive genes of Epstein-Barr virus: data and hypotheses.

Epstein-Barr virus (EBV) has two different modes of existence: latent and productive. There are eight known genes expressed during latency (and hardly at all during the productive phase) and about 70 other ("productive") genes. It is shown that the EBV genes known to be expressed during latency display codon usage strikingly different from that of genes that are expressed during lytic growth. In particular, the percentage of S3 (G or C in codon site 3) is persistently lower (about 20%) in all latent genes than in nonlatent genes. Moreover, S3 is lower in each multicodon amino acid form. Also, the percentage of S in silent codon sites 1 of leucine and arginine is lower in latent than in nonlatent genes. The largest absolute differences in amino acid usage between latent and nonlatent genes emphasize codon types SSN and WWN (W means nucleotide A or T and N is any nucleotide). Two principal explanations to account for the EBV latent versus productive gene codon disparity are proposed. Latent genes have codon usage substantially different from that of host cell genes to minimize the deleterious consequences to the host of viral gene expression during latency. (Productive genes are not so constrained.) It is also proposed that the latency genes of EBV were acquired recently by the viral genome. Evidence and arguments for these proposals are presented.

Amino Acid Sequence↗

Charge configurations in oncogene products and transforming proteins.

Statistically significant charge clusters are of infrequent occurrence in all kinds of proteins. In the six standard classes of proto-oncogene products, all of the nuclear class contain a significant charge cluster and several, but not all, of the transmembrane class do, whereas significant charge clusters or patterns are not found in protooncogenes of primarily cytoplasmic location, nor in membrane-bound (src-like) proto-oncogenes, nor in those of the ras family. Among nuclear oncogene families, such as myc, jun, fos, myb, or ets-related, and among homologous proteins across species, the significant charge clusters are part of the most conserved region. These gene families generally have similar charge distributions embodying a significant charge cluster, not of an invariant sign, preceded by a substantial uncharged stretch of predominantly polar residues. The nuclear transforming proteins p53 and p68 also contain significant charge clusters together with long uncharged segments, suggestive of a modular structure of these proteins. The transmembrane oncogene c-mas contains a mixed charge cluster and c-fms displays an unusual (0, +)7 pattern, in both cases positioned within their intracellular activating domain. Distinctive charge configurations for excreted proto-oncogenes are of a mixed character. Possible functions, mechanisms, and associated experimental procedures for studying proteins with anomalous charge distributions are discussed.

Amino Acid Sequence↗

A method to identify distinctive charge configurations in protein sequences, with application to human herpesvirus polypeptides.

Charge interactions are of great importance for protein function and structure, and for a variety of cellular and biochemical processes. We present a systematic approach to the detection of distinctive clusters, runs and periodic patterns of charged residues in a protein sequence. Criteria and formulae are set forth to assess statistical significance of these charge configurations. For the 80-odd proteins potentially encoded by the Epstein-Barr virus, only the major nuclear antigens of the latent state and the transactivator of the lytic cycle contain separated charge clusters of opposite sign as well as periodic charge patterns. From our studies of the polypeptides of the human herpesviruses and of a broad collection of human and other viral protein sequences, distinctive charge configurations appear to be associated with viral capsid and core proteins (positive clusters or runs, mostly at the carboxyl terminus), with many viral glycoproteins and membrane-associated proteins (negative charge clusters), and with transactivators and transforming proteins (multiple charge structures). The statistics developed in this paper apply more generally to other than charge properties of a protein and should aid in the evaluation of a large variety of sequence features.

Cytomegalovirus↗

Association of charge clusters with functional domains of cellular transcription factors.

Using rigorous statistical methods, we have identified and evaluated unusual properties of the distribution of charged residues within the sequences of eukaryotic cellular transcription factors. Virtually all transcription factors, including GAL4, c-Jun, C/EBP, CREB, Oct-1, Oct-2, Sp1, Egr-1, CTF-1, steroid and thyroid hormone receptors, and others, carry one or more highly significant charge clusters. For the most part these clusters (conserved within families of homologous proteins) are of positive net charge but contain also substantial numbers of acidic residues. Predominantly basic charge clusters are often, but not exclusively, associated with DNA-binding domains, and vice versa. Negative charge clusters of note occur only in the yeast protein PHO4 and in the proteins encoded at the Drosophila loci zeste (zeta) and knrl. This dearth of statistically significant negative charge clusters raises questions with respect to the generality of acidic activation domains. A number of sequences (Oct-1, Oct-2, zeste, Dhr23, E75, and knrl) contain multiple charge clusters together with one or more significantly long uncharged regions. The occurrence of multiple charge clusters is a rare phenomenon (found in less than 3% of all proteins, mainly in Drosophila developmental control proteins and in transactivators of eukaryotic DNA viruses). Most of the proteins with zinc-binding "fingers" carry a mixed charge cluster centered at the zinc-finger motif preceded by a long uncharged stretch, suggestive of a modular structure for these proteins.

Animals↗

Distinctive charge configurations in proteins of the Epstein-Barr virus and possible functions.

The protein products of several open reading frames (ORFs) of the Epstein-Barr virus (EBV) are remarkable in their distribution of charged residues. The nuclear antigen proteins EBNA1-EBNA4 of the EBV latent state contain separate significant clusters of charge of each sign. They (excepting EBNA4) also feature distinctive periodic charge patterns [e.g., (+, O)8, (O, -, -)7] and significant tandem repeats. None of the other ORFs (about 80) of the genome possess the conjunction of these properties. Only the protein encoded from BMLF1, the first immediate early transactivator protein, contains significant multiple charge clusters and periodic charge patterns. All proteins that contain significant repeats also contain at least one significant charge cluster of a single sign. These include EBNA5 and LYDMA produced during latency and BZLF1, whose expression terminates latency and initiates productive growth. It is reasonable to conclude that these aggregate significant charge configurations and repeats are important functionally for the latent existence and for the initiation of the lytic cycle and may be characteristic of these conditions. We discuss how large multimeric protein structures bound together by clusters of unlike charge may provide a mechanism for regulation of the expression of these proteins.

Amino Acid Sequence↗

Charge configurations in viral proteins.

The spatial distribution of the charged residues of a protein is of interest with respect to potential electrostatic interactions. We have examined the proteins of a large number of representative eukaryotic and prokaryotic viruses for the occurrence of significant clusters, runs, and periodic patterns of charge. Clusters and runs of positive charge are prominent in many capsid and core proteins, whereas surface (glyco)proteins frequently contain a negative charge cluster. Significant charge configurations are abundant in regulatory proteins implicated in transcriptional transactivation and cellular transformation. Proteins with charge structures are much more predominant in animal DNA viruses as compared to animal RNA viruses and prokaryotic viruses. This contrast might reflect the role of protein charge structures in facilitating competitive virus-host interactions involving the cellular transcription, translation, protein sorting, and transport apparatus.

Adenoviridae↗

Efficient algorithms for molecular sequence analysis.

Efficient (linear time) algorithms are described for identifying global molecular sequence features allowing for errors including repeats, matches between sequences, dyad symmetry pairings, and other sequence patterns. A multiple sequence alignment algorithm is also described. Specific applications are given to hepatitis B viruses and the J5-C (J, joining; C, constant) region of the immunoglobulin kappa gene.

Algorithms↗