Search PubMed⌕ Search

Biomedical subjects

S Karlin

Publications and source records attributed to S Karlin.

At least 109 records · Page 6Linked to original sources

Algorithms for identifying local molecular sequence features.

Efficient algorithms are described for identifying local molecular sequence features including repeats, dyad symmetry pairings and aligned matches between sequences, while allowing for errors. Specific applications are given to the genomic sequences of the Epstein-Barr virus, Varicella-Zoster virus and the bacteriophages lambda and T7.

Algorithms↗

A model for the development of the tandem repeat units in the EBV ori-P region and a discussion of their possible function.

This paper presents an analysis of the repeat units of the ori-P region of the Epstein-Barr virus (EBV) genome. These repeat units are well-conserved palindromes. The pattern of these repeats, their lengths, phases, and the distribution of the relatively few substitutions are explained by a scenario that gives a reasonable course for the evolutionary development of the pattern. The scenario suggests a model for the production of an initiating 3/2 palindrome from a moderately lengthy sequence. The palindromic units are then multiplied in judicious combinations by mechanisms of unequal crossing-over events associated with some point substitutions and a few instances of slippage replication. The potential secondary structures of the two separated tandem palindromic repeat regions in ori-P are contrasted. Possible modes of binding of Epstein-Barr nuclear antigen (EBNA) 1 protein to these hairpins are discussed. A number of possibilities for the origin and development of the ori-P region in relation to viral and cellular function are considered.

Base Sequence↗

Permutation analyses of familial association arrays for lipoprotein concentrations in families of the Stanford Five City Project.

Permutation models are introduced as a formal method for assigning significance to association matrices that assess the correlation of spouse, parent-offspring, and sibling similarity over an entire class of data transformations (usually, the class of all increasing functions). Analysis of 218 nuclear families who participated in the Stanford Five City Project revealed that parent and offspring triglyceride concentrations correlated more strongly when data transformations emphasized contrasts among low to moderate levels, and that high density lipoprotein (HDL) cholesterol correlated more strongly between family members with relatively higher HDL cholesterol concentrations. Application of family weights to the association matrices revealed a tendency for greater correlation among sibling triglyceride concentrations in larger families. Parent-child, mother-child, father-child, parent-daughter, and sibling total cholesterol concentrations correlated significantly for all monotonically increasing transformations (designated strong association), and father-daughter and parent-son cholesterol concentrations correlated significantly for most increasing transformations of the data (moderate association). There were fewer significant associations for plasma triglyceride concentrations: parent-child and sibling (both strong), parent-daughter and mother-daughter (both moderate), and mother-child (weak). HDL cholesterol showed no strong or moderate familial associations and was weakly associated only among siblings. Thus, concordance in familial lipoprotein levels appears to be restricted to a narrower range of values for triglycerides and HDL cholesterol than total cholesterol levels, possibly reflecting in part the influences of diet or other environmental factors on specific regions of the HDL cholesterol or triglyceride distributions in casual blood samples.

Adolescent↗

Significant potential secondary structures in the Epstein-Barr virus genome.

This paper identifies all statistically significant dyad symmetry combinations in the Epstein-Barr virus genome. The distribution of long dyad symmetry pairings emphasizes two regions, the 5' third of the 3.1-kilobase-pair (kbp) repeat and the oriP region, the latter essential for Epstein-Barr virus replication during latency. A 600-base-pair (bp) stretch in the 3.1-kbp repeat can establish an extended hairpin loop of stem length in excess of 208 bp of predominantly G + C stacking. Moreover, the 3.1-kbp repeat has the potential to form a wide variety of secondary structures based on juxtapositions of sizable palindromes, close dyad symmetry pairings, and direct repeats. The 3.1-kbp repeat presents several features that portend it as an important control region. The oriP region contains an abundance of statistically significant dyad symmetry combinations that strongly correlate with the "21 X 30 bp" tandem repeat units and four truncated copies of this repeat unit 1 kbp downstream. Each of the units centers on the same approximately 30-bp palindrome. Contrasts in the content and the secondary structure formations associated with the 3.1-kbp repeat units versus those of the oriP region are discussed in relation to viral or cellular function.

Antigens, Viral↗

The use of multiple alphabets in kappa-gene immunoglobulin DNA sequence comparisons.

Comparisons within and between the human, mouse and rabbit immunoglobulin-kappa gene (J-C region) DNA sequences are carried out in terms of three two-letter nucleotide alphabets: (i) S-W alphabet (W = A or T; S = G or C); (ii) P-Q alphabet which distinguishes purines (P = A or G) from pyrimidines (Q = C or T); and (iii) a 'control' E-F alphabet (E = A or C; F = G or T). All statistically significant direct repeats within each of the three sequences and all significant block identities (a set of consecutive matching letters) shared by two or more sequences are determined for each alphabet. By contrast to the S-W and E-F alphabets, the P-Q alphabet comparisons reveal an abundance of statistically significant block identities not seen at the nucleotide level. Various interpretations of these P-Q structures with respect to control and functional roles are considered.

Animals↗

DNA sequence patterns in human, mouse, and rabbit immunoglobulin kappa-genes.

DNA sequences of the human, mouse, and rabbit immunoglobulin kappa-gene (J-C regions) are compared with respect to various DNA patterns, including dyad symmetry pairings, runs of nucleotides, repeat clusters, and repeats that occur with unusually high frequency. The significant dyad symmetry pairings within each of the sequences emphasize the two "control-enhancer" elements of the J5-C intron. Dyad symmetry pairs between the J-C region and a number of kappa variable (V)-gene domains suggest differences in the affinities between the V and J segments. It is the "consensus heptamer" rather than the "consensus nonamer" that embodies the longest V-J dyad symmetry combinations. In the rabbit there are long runs and repeat clusters of the sequences that identify regions of high duplication; these regions are absent in the human and mouse sequences. High-frequency oligonucleotides feature the consensus nonamer 5' to the J segments, especially in the mouse sequence.

Animals↗

Comparative statistics for DNA and protein sequences: single sequence analysis.

Four categories of data representations are used to help interpret structures and similarities of nucleic acid and protein sequences. Statistical significance of the observed relationships revealed by these representations are assessed by a hierarchy of permutation procedures and by comparisons with theoretical random models. Applications are presented for various DNA sequences including papovaviruses, Epstein-Barr virus, mitochondrial genomes, and several globin and immunoglobulin genes.

Amino Acid Sequence↗

Comparative statistics for DNA and protein sequences: multiple sequence analysis.

Concepts and methods [Karlin, S. & Ghandour, G. (1985) Proc. Natl. Acad. Sci. USA 82, 5800-5804] for the analysis of patterns and relationships are extended to multiple DNA and protein sequences. Functionals include multiple sequence common word occurrence distributions, characterizations of high frequency shared words, and ascertainment of long block identities. Various comparisons of sequences using natural alphabets obtained from grouping nucleotides or amino acids by their chemical and functional characteristics are described. Specific applications are given to globin genes, mitochondrial genomes, and a variety of mammalian viruses.

Amino Acid Sequence↗

Multiple-alphabet amino acid sequence comparisons of the immunoglobulin kappa-chain constant domain.

We compare the amino acid sequences of the constant domains of the immunoglobulin kappa chain of human, mouse, and rabbit by using four classification schemes ("alphabets") of the 20 amino acids based on their chemical, functional, charge, and structural properties. The comparison reveals three regions of pronounced similarity across the three species, independent of allotype. Two of these regions (residues 65-73 and 99-103) entail a high degree of identity at the DNA level and are distinguished from the rest of the constant domain in codon usage and in the dinucleotide sequence at abutting sites of adjacent codons. Residues 22-29 are highly conserved among the three species in the chemical and functional alphabets but do not show any three-sequence significant amino acid block identities. These results are discussed in terms of transcript processing, effector functions, and structural interactions within the constant domain and with the heavy chain.

Amino Acid Sequence↗

Permutation methods for the structured exploratory data analysis (SEDA) of total cholesterol measured in five Israeli populations.

Three structured exploratory data analysis-functionals are applied to plasma total cholesterol concentrations measured for 2,480 young men and women aged 17-18 years and living in Jerusalem, and for their parents. These triad families are divided into five groups according to whether both parents were born in Asia, North Africa, Europe-America, or Israel or whether they were of mixed "origins." The significances of the functionals were determined by a spectrum of permutation techniques that selectively shuffled the trait values across families in order to systematically alter certain family structure relationships while keeping other familial relationships intact. These analyses suggest that generational differences and various distributional effects influence patterns of spouse and parent-offspring interactions within these families and that the nature and forms of these effects and interactions may differ according to the origin of the parents. Results are discussed in relationship to historical and cultural differences among groups.

Adolescent↗

DNA sequence comparisons of the human, mouse, and rabbit immunoglobulin kappa gene.

A comparative analysis between human, mouse, and rabbit immunoglobulin (Ig) kappa-gene DNA sequences is presented. New formulas for determining the expected length and variance of the longest block identity (a succession of matching nucleotides) between multiple random sequences are given and are used to establish statistical criteria for ascertaining the significance of block identities shared in r out of s sequences. The statistically significant block identities within and between the Ig-kappa-gene sequences are ascertained, and alignment maps based on these similarities are constructed. The human and rabbit sequences (especially in the noncoding regions) and the human and mouse sequences (on the coding regions) show a similarity much stronger than that between the mouse and rabbit sequences. The existence of several highly significant shared oligonucleotides occurring in alignment with each other or with respect to the J- and C-gene segments suggests a configuration of multiple control sites. Discussion and interpretations of the form and distribution of the block identities are given.

Animals↗

Alignment maps and homology analysis of the J-C intron in human, mouse, and rabbit immunoglobulin kappa gene.

The statistically significant shared oligonucleotides (block identities) of the intervening region (J5-C) in the human, mouse, and rabbit immunoglobulin (Ig)-kappa gene were determined. These identities maintain their order (do not cross) and never connect with any Ig-kappa segment external to the intron region. The two regions of pronounced similarity are (1) the vicinity of the established enhancer element (Queen and Baltimore 1983) and (2) a 200-bp region approximately 1 kb upstream that we have labeled the second enhancer element. Similarity is strong between the human and mouse sequences in the neighborhood of the first enhancer element and more pronounced between the human and rabbit sequences in the vicinity of the second enhancer region. The overall extent of similarity between the mouse and rabbit sequences is less than that between the human and mouse sequences and that between the human and rabbit sequences. All close and large dyad-symmetry pairings were ascertained and their possible relations to the enhancer elements discussed.

Animals↗

On the optimal sex-ratio: a stability analysis based on a characterization for one-locus multiallele viability models.

Theoretical one-locus multiallele sex-determination models are found to admit even sex ratio equilibrium surfaces besides the equilibria for corresponding one-locus multiallele viability models. Both types of equilibria can be defined in terms of a single spectral radius function, the former corresponding to level surfaces and the latter to critical points. The stable equilibria in the corresponding viability models are associated with the local maxima, and the equilibrium structures for the sex-determination models can be fully described. Several optimality properties of the even-sex-ratio equilibrium surfaces can be deduced.

Alleles↗

On the evolution of altruism by kin selection.

A general model for the evolution of altruism is formulated. Central to the model is a pair of local fitness functions, which prescribe the fitness of the altruist and selfish phenotypes as functions of the composition of local groups into which prereproductives are subdivided. When the local groups are sibships or other kin groups, the model is one for kin selection. Functions for cost and benefit of altruism are derived from the fitness functions. Conditions for evolution of altruism are then determined in terms of cost and benefit. It is shown that the Hamilton rule has quantitative validity only in the special case of linear fitness functions. Sufficient conditions are found for qualitative validity of the Hamilton rule. Qualitative violation of the rule is also possible.

Journal Article↗

Comparative analysis of human and bovine papillomaviruses.

A method is presented for the analysis and comparison of nucleic acid and protein sequences utilizing all identity blocks (the term "identity block" refers to a set of consecutive matches between two sequences) above a prescribed length. Moreover, such identity blocks are determined for various groupings of amino acids according to chemical, functional, charge, and hydrophobic classifications. Alignment maps based on these classifications and containing all statistically significant identity blocks between two or more sequences are constructed. New theoretical results for determining the expected length of the longest identity block between sequences are also presented and are used, along with permutation procedures, to ascertain the significance of sequence identity blocks. As an example of the type of information that can be obtained, comparison has been made of the complete DNA sequences and the E1, E2, L1, and L2 genes of human and bovine papillomaviruses based on the classification schemes described above.

Amino Acids↗