Search PubMed⌕ Search

Biomedical subjects

S Karlin

Publications and source records attributed to S Karlin.

At least 73 records · Page 4Linked to original sources

Why is CpG suppressed in the genomes of virtually all small eukaryotic viruses but not in those of large eukaryotic viruses?

Dinucleotide over- and underrepresentation is evaluated in all available completely sequenced DNA or RNA viral genomes, ranging in size from 3 to 250 kb (available RNA viruses fall into the small-virus category). The dinucleotide CpG is statistically underrepresented (suppressed) in all but four of the small viruses (more than 75 with lengths of < 30 kb) but has normal relative abundances in most large viruses (> or = 30 kb). Most retrotransposons in eukaryotic species also show low CpG relative abundances. Interpretations, especially in some cases of DNA viruses or viruses with a DNA intermediate, might relate to methylation effects and modes of viral integration and excision. Other possible contributing factors relate to dinucleotide stacking energies, special mutation mechanisms, and evolutionary events.

Arginine↗

Computational DNA sequence analysis.

This paper reviews several new developments in computer and statistical analysis of DNA and protein sequences. We present criteria and describe means for assessing and interpreting genomic inhomogeneities within and between sequences. These include: (a) characterizations of short oligonucleotide biases and general compositional tendencies; (b) molecular evolutionary reconstructions based on dinucleotide relative abundance distance measures and partial orderings; and (c) the application of r-scan statistics, quantile distributions, and score-based analyses to identify clustering, overdispersion, and excessive evenness in the distribution of a marker array along a sequence. These apply, for example, to restriction sites, microsatellite runs, regulatory motifs, and nucleosome placements. Furthermore, (d) the definition and determination of rare and frequent oligonucleotides and peptides provides another perspective on sequence heterogeneity, and (e) score methods are also applied in exon and gene locations. Most of the ideas and methods are illustrated with respect to bacteriophage genomes, to megabase amounts of several eukaryotic sequences, to a diverse collection of bacterial sets, to mitochondrial chromosomes, and to a broad assembly of viral genomes.

Amino Acid Sequence↗

Comparative DNA sequence features in two long Escherichia coli contigs.

The recent sequencing of two relatively long (approximately 100 kb) contigs of E.coli presents unique opportunities for investigating heterogeneity and genomic organization of the E.coli chromosome. We have evaluated a number of common and contrasting sequence features in the two new contigs with comparisons to all available E.coli sequences (> 1.6 Mb). Our analyses include assessments of: (i) counts and distributions of restriction sites, special oligonucleotides (e.g., Chi sites, Dam and Dcm methylase targets), and other marker arrays; (ii) significant distant and close direct and inverted repeat sequences; (iii) sequence similarities between the long contigs and other E.coli sequences; (iv) characterization and identification of rare and frequent oligonucleotides; (v) compositional biases in short oligonucleotides; and (vi) position-dependent fluctuations in sequence composition. The two contigs reveal a number of distinctive features, including: a cluster of five repeat/dyad elements with very regular spacings resembling a transcription attenuator in one of the contigs; REP elements, ERICs, and other long repeats; distinction of the Chi sequence as the most frequent oligonucleotide; regions of clustering, overdispersion, and regularity of certain restriction sites and short palindromes; and comparative domains of inhomogeneities in the two long contigs. These and other features are discussed in relation to the organization of the E.coli chromosome.

Base Sequence↗

Unusual charge configurations in transcription factors of the basic RNA polymerase II initiation complex.

A systematic analysis of the primary sequences of the polymerase II initiation complex has revealed unusual charge features in the TFII family proteins. In particular, the proteins TFIIA alpha, TFIIE alpha, and TFIIF carry multiple charge clusters and hyper charge runs, sequence features occurring in < 4% of all (available) eukaryotic proteins. Possible implications for these charge structures are discussed in relation to the assembly and function of the polymerase II transcriptional complex.

Amino Acid Sequence↗

Applications and statistics for multiple high-scoring segments in molecular sequences.

Score-based measures of molecular-sequence features provide versatile aids for the study of proteins and DNA. They are used by many sequence data base search programs, as well as for identifying distinctive properties of single sequences. For any such measure, it is important to know what can be expected to occur purely by chance. The statistical distribution of high-scoring segments has been described elsewhere. However, molecular sequences will frequently yield several high-scoring segments for which some combined assessment is in order. This paper describes the statistical distribution for the sum of the scores of multiple high-scoring segments and illustrates its application to the identification of possible transmembrane segments and the evaluation of sequence similarity.

Amino Acid Sequence↗

Significant dispersed recurrent DNA sequences in the Escherichia coli genome. Several new groups.

New computer and statistical methods were used to determine significant direct and inverted repeats in the Escherichia coli contig sequence collection of aggregate 1.6 x 10(6) base-pairs. Eight groups of mostly new structural repeat identities were uncovered. Apart from the high statistical significance of these repeat sequences, there are suggestive relationships of the group matches in terms of neighboring genes, of genomic distributions, of their texts, and of their potentials for secondary structure. Four of these groups are relatively numerous, 11 to 26 members, one is in coding sequences and three are in non-coding. The coding group consists of the ATP-activated transmembrane component of a typical high-affinity protein-binding transport system. One of the non-coding groups consists of a special rho-independent transcription termination signal closely following an operon. The gene neighbors of this group often appear to be involved in some way in processing RNA or DNA. A second non-coding group has, for one or both neighboring genes, a component of a system responding to stress or starvation for some nutrient.

Algorithms↗

Assessments of DNA inhomogeneities in yeast chromosome III.

With the sequencing of the first complete eukaryotic chromosome, III of yeast (YCIII) of length 315 kb, several types of questions concerning chromosomal organization and the heterogeneity of eukaryotic DNA sequences can be approached. We have undertaken extensive analysis of YCIII with the goals of: (1) discerning patterns and anomalies in the occurrences of short oligonucleotides; (2) characterizing the nature and locations of significant direct and inverted repeats; (3) delimiting regions unusually rich in particular base types (e.g., G+C, purines); and (4) analyzing the distributions of markers of interest, e.g., delta (delta) elements, ARS (autonomous replicating sequences), special oligonucleotides, close repeats and close dyad pairings, and gene sequences. YCIII reveals several distinctive sequence features, including: (i) a relative abundance of significant local and global repeats highlighting five genes containing substantial close or tandem DNA repeats; (ii) an anomalous distribution of delta elements involving two clusters and a long gap; (iii) a significantly even distribution of ARS; (iv) a relative increase in the frequency of T runs and AT iterations downstream of genes and A runs upstream of genes; and (v) two regions of complex repetitive sequences and anomalous DNA composition, 29000-31000 and 291000-295000, the latter centered at the HMRa locus. Interpretations of these findings for chromosomal organization and implications for regulation of gene expression are discussed.

Chromosome Mapping↗

Patchiness and correlations in DNA sequences.

The highly nonrandom character of genomic DNA can confound attempts at modeling DNA sequence variation by standard stochastic processes (including random walk or fractal models). In particular, the mosaic character of DNA consisting of patches of different composition can fully account for apparent long-range correlations in DNA.

Analysis of Variance↗

A comparative analysis of distinctive features of yeast protein sequences.

The recently published sequence of yeast chromosome III (YCIII) provides the longest continuous stretch of a eukaryotic DNA molecule sequenced to date (315 kb). The sequence contains 116 distinct AUG-initiated open reading frames of at least 200 codons in length, more than 50 of which had not been described previously nor bear significant similarity to known proteins. We have analysed the YCIII known and putative protein sequences with respect to significant statistical features which might reflect on structural and functional characteristics. The YCIII proteins have striking similarities and differences in their sequence attribute distributions compared to the corresponding distributions for all available yeast sequences and other protein collections. Nine examples of YCIII proteins with distinctive sequence features are discussed in detail.

Amino Acid Sequence↗

Correlation analysis of amino acid usage in protein classes.

We present a comparative study of residue usage correlations of various organism protein sets of diverse phylogenetic species and of open reading frames of several large human viral genomes. Our correlation analysis reveals three major tendencies: (i) charge compensation reflected by the high correlation of basic with acidic residues; (ii) the positive correlations of functionally and structurally similar amino acids including many pairs of hydrophobic amino acids, all pairs of aromatic amino acids, the anionic pair (glutamate and aspartate), but not the cationic pair (lysine and arginine), moderately the hydroxyl pair (serine and threonine), the small amino acids (glycine and alanine), and many (but not all) of those having high values in the Dayhoff substitutability matrix (characteristics such as amino acid polarity or codon usage agreement, except for the wobble position, do not necessarily imply significant positive correlations); (iii) a widespread negative correlation of the aggregate strong codon group amino acids (Ala, Gly, Pro) versus the weak codon group amino acids (Lys, Ile, Tyr, Asn, Phe). Discussion and speculations relate amino acid usage correlations to protein function/structure, cellular localization, proximity in amino acid biosynthetic pathways, amino acid relative abundances, tRNA and aminoacyl synthetase availabilities, and evolutionary processes.

Amino Acids↗

Chance and statistical significance in protein and DNA sequence analysis.

Statistical approaches help in the determination of significant configurations in protein and nucleic acid sequence data. Three recent statistical methods are discussed: (i) score-based sequence analysis that provides a means for characterizing anomalies in local sequence text and for evaluating sequence comparisons; (ii) quantile distributions of amino acid usage that reveal general compositional biases in proteins and evolutionary relations; and (iii) r-scan statistics that can be applied to the analysis of spacings of sequence markers.

Amino Acid Sequence↗

Human cytomegalovirus origin of DNA replication (oriLyt) resides within a highly complex repetitive region.

A global analysis of the 230-kilobase-pair (kbp) human cytomegalovirus genome revealed three regions that were very rich in repeated sequences. The region with the highest content of inverted and direct repeats lies between 92,100 and 93,500 bp, upstream of the gene encoding the single-stranded DNA binding protein. Cloned restriction fragments containing this region were able to replicate when trans-acting factors were provided by virus infection in a transient replication assay. With this assay, the region between 92,210 and 93,715 bp on the viral genome was defined as the minimal replication origin, oriLyt. The sequence composition and repeats within oriLyt were used to divide the region into two domains that may be important in origin function. Sequences flanking either the left or right side of the minimal oriLyt contributed to efficient replication; however, these sequences were not essential for origin function. Thus, the region of the viral genome with the most striking concentration of direct and inverted repeats corresponds to the oriLyt of human cytomegalovirus.

Base Sequence↗

Statistical analyses of counts and distributions of restriction sites in DNA sequences.

Counts and spacings of all 4- and 6-bp palindromes in DNA sequences from a broad range of organisms were investigated. Both 4- and 6-bp average palindrome counts were significantly low in all bacteriophages except one, probably as a means of avoiding restriction enzyme cleavage. The exception, T4 of normal 4- and 6-palindrome counts, putatively derives protection from modification of cytosine to hydroxymethylcytosine plus glycosylation. The counts and distributions of 4-bp and of 6-bp restriction sites in bacterial species are variable. Bacterial cells with multiple restriction systems for 4-bp or 6-bp target specificities are low in aggregate 4- or 6-bp palindrome counts/kb, respectively, but bacterial cells lacking exact 4-cutter enzymes generally show normal or high counts of 4-bp palindromes when compared with random control sequences of comparable nucleotide frequencies. For example, E. coli, apparently without an exact 4-bp target restriction endonuclease (see text), contains normal aggregate 4-palindrome counts/kb, while B. subtilis, which abounds with 4-bp restriction systems, shows a significant under-representation of 4-palindrome counts. Both E. coli and B. subtilis have many 6-bp restriction enzymes and concomitantly diminished aggregate 6-palindrome counts/kb. Eukaryote, viral, and organelle sequences generally have aggregate 4- and 6-palindromic counts/kb in the normal range. Interpretations of these results are given in terms of restriction/methylation regimes, recombination and transcription processes, and possible structural and regulatory roles of 4- and 6-bp palindromes.

Animals↗

Methods and algorithms for statistical analysis of protein sequences.

We describe several protein sequence statistics designed to evaluate distinctive attributes of residue content and arrangement in primary structure. Considered are global compositional biases, local clustering of different residue types (e.g., charged residues, hydrophobic residues, Ser/Thr), long runs of charged or uncharged residues, periodic patterns, counts and distribution of homooligopeptides, and unusual spacings between particular residue types. The computer program SAPS (statistical analysis of protein sequences) calculates all the statistics for any individual protein sequence input and is available for the UNIX environment through electronic mail on request to V.B. (volker/genomic@stanford.edu).

Algorithms↗

Over- and under-representation of short oligonucleotides in DNA sequences.

Strand-symmetric relative abundance functionals for di-, tri-, and tetranucleotides are introduced and applied to sequences encompassing a broad phylogenetic range to discern tendencies and anomalies in the occurrences of these short oligonucleotides within and between genomic sequences. For dinucleotides, TA is almost universally under-represented, with the exception of vertebrate mitochondrial genomes, and CG is strongly under-represented in vertebrates and in mitochondrial genomes. The traditional methylation/deamination/mutation hypothesis for the rarity of CG does not adequately account for the observed deficiencies in certain sequences, notably the mitochondrial genomes, yeast, and Neurospora crassa, which lack the standard CpG methylase. Homodinucleotides (AA.TT, CC.GG) and larger homooligonucleotides are over-represented in many organisms, perhaps due to polymerase slippage events. For trinucleotides, GCA.TGC tends to be under-represented in phage, human viral, and eukaryotic sequences, and CTA.TAG is strongly under-represented in many prokaryotic, eukaryotic, and viral sequences. The CCA.TGG triplet is ubiquitously over-represented in human viral and eukaryotic sequences. Among the tetranucleotides, several four-base-pair palindromes tend to be under-represented in phage sequences, probably as a means of restriction avoidance. The tetranucleotide CTAG is observed to be rare in virtually all bacterial genomes and some phage genomes. Explanations for these over- and under-representations in terms of DNA/RNA structures and regulatory mechanisms are considered.

Animals↗

Significant similarity and dissimilarity in homologous proteins.

Common practice emphasizes significant sequence similarities between different members of protein families. These similarities presumably reflect on evolutionary conservation of structurally and functionally essential residues. The nonconserved regions, on the other hand, may be either selectively neutral or differentiated. We propose several distributional sequence statistics (e.g., clustering of charged residues, compositional biases, and repetitive patterns) as indicators of differentiation events. These ideas are illustrated with various examples, including comparisons among G protein-coupled receptors, herpesvirus proteins, and GTPase-activating proteins.

GTP-Binding Proteins↗

Quantile distributions of amino acid usage in protein classes.

A comparative study of the compositional properties of various protein sets from both cellular and viral organisms is presented. Invariants and contrasts of amino acid usages have been discerned for different protein function classes and for different species using robust statistical methods based on quantile distributions and stochastic ordering relationships. In addition, a quantitative criterion to assess amino acid compositional extremes relative to a reference protein set is proposed and applied. Invariants of amino acid usage relate mainly to the central range of quantile distributions, whereas contrasts occur mainly in the tails of the distributions, especially contrasts between eukaryote and prokaryote species. Influences from genomic constraint are evident, for example, in the arginine:lysine ratios and the usage frequencies of residues encoded by G + C-rich versus A + T-rich codon types. The structurally similar amino acids, glutamate versus aspartate and phenylalanine versus tyrosine, show stochastic dominance relationships for most species protein sets favoring glutamate and phenylalanine respectively. The quantile distribution of hydrophobic amino acid usages in prokaryote data dominates the corresponding quantile distribution in human data. In contrast, glutamate, cysteine, proline and serine usages in human proteins dominate the corresponding quantile distributions in Escherichia coli. E. coli dominates human in the use of basic residues, but no dominance ordering applies to acidic residues. The discussion centers on commonalities and anomalies of the amino acid compositional spectrum in relation to species, function, cellular localization, biochemical and steric attributes, complexity of the amino acid biosynthetic pathway, amino acid relative abundances and founder effects.

Amino Acids↗