Search PubMed⌕ Search

Biomedical subjects

S Karlin

Publications and source records attributed to S Karlin.

At least 55 records · Page 3Linked to original sources

Trinucleotide repeats and long homopeptides in genes and proteins associated with nervous system disease and development.

Several human neurological disorders are associated with proteins containing abnormally long runs of glutamine residues. Strikingly, most of these proteins contain two or more additional long runs of amino acids other than glutamine. We screened the current human, mouse, Drosophila, yeast, and Escherichia coli protein sequence data bases and identified all proteins containing multiple long homopeptides. This search found multiple long homopeptides in about 12% of Drosophila proteins but in only about 1.7% of human, mouse, and yeast proteins and none among E. coli proteins. Most of these sequences show other unusual sequence features, including multiple charge clusters and excessive counts of homopeptides of length > or = two amino acid residues. Intriguingly, a large majority of the identified Drosophila proteins are essential developmental proteins and, in particular, most play a role in central nervous system development. Almost half of the human and mouse proteins identified are homeotic homologs. The role of long homopeptides in fine-tuning protein conformation for multiple functional activities is discussed. The relative contributions of strand slippage and of dynamic mutation are also addressed. Several new experiments are proposed.

Amino Acid Sequence↗

Evolutionary conservation of RecA genes in relation to protein structure and function.

Functional and structural regions inferred from the Escherichia coli R ecA protein crystal structure and mutation studies are evaluated in terms of evolutionary conservation across 63 RecA eubacterial sequences. Two paramount segments invariant in specific amino acids correspond to the ATP-binding A site and the functionally unassigned segment from residues 145 to 149 immediately carboxyl to the ATP hydrolysis B site. Not only are residues 145 to 149 conserved individually, but also all three-dimensional structural neighbors of these residues are invariant, strongly attesting to the functional or structural importance of this segment. The conservation of charged residues at the monomer-monomer interface, emphasizing basic residues on one surface and acidic residues on the other, suggests that RecA monomer polymerization is substantially mediated by electrostatic interactions. Different patterns of conservation also allow determination of regions proposed to interact with DNA, of LexA binding sites, and of filament-filament contact regions. Amino acid conservation is also compared with activities and properties of certain RecA protein mutants. Arginine 243 and its strongly cationic structural environment are proposed as the major site of competition for DNA and LexA binding to RecA. The conserved acidic and glycine residues of the disordered loop L1 and its proximity to the RecA acidic monomer interface suggest its involvement in monomer-monomer interactions rather than DNA binding. The conservation of various RecA positions and regions suggests a model for RecA-double-stranded DNA interaction and other functional and structural assignments.

Amino Acid Sequence↗

How are close residues of protein structures distributed in primary sequence?

Structurally neighboring residues are categorized according to their separation in the primary sequence as proximal (1-4 positions apart) and otherwise distal, which in turn is divided into near (5-20 positions), far (21-50 positions), very far ( > 50 positions), and interchain (from different chains of the same structure). These categories describe the linear distance histogram (LDH) for three-dimensional neighboring residue types. Among the main results are the following: (i) nearest-neighbor hydrophobic residues tend to be increasingly distally separated in the linear sequence, thus most often connecting distinct secondary structure units. (ii) The LDHs of oppositely charged nearest-neighbors emphasize proximal positions with a subsidiary maximum for very far positions. (iii) Cysteine-cysteine structural interactions rarely involve proximal positions. (iv) The greatest numbers of interchain specific nearest-neighbors in protein structures are composed of oppositely charged residues. (v) The largest fraction of side-chain neighboring residues from beta-strands involves near positions, emphasizing associations between consecutive strands. (vi) Exposed residue pairs are predominantly located in proximal linear positions, while buried residue pairs principally correspond to far or very far distal positions. The results are principally invariant to protein sizes, amino acid usages, linear distance normalizations, and over- and underrepresentations among nearest-neighbor types. Interpretations and hypotheses concerning the LDHs, particularly those of hydrophobic and charged pairings, are discussed with respect to protein stability and functionality. The pronounced occurrence of oppositely charged interchain contacts is consistent with many observations on protein complexes where multichain stabilization is facilitated by electrostatic interactions.

Amino Acid Sequence↗

Statistical significance of sequence patterns in proteins.

I discuss three recent developments in sequence analysis by the statistical method of scores. First is the identification of segments of high aggregate score in a single protein sequence. Charge clusters and hyper-charge runs are prime examples. Proteins containing hyper-charge runs are principally associated with DNA and RNA processing, chromatin structure, ion storage and exchange, and protein complex assembly. Second is the protein sequence comparisons identifying common segments having high total similarity scores. These are illustrated by comparisons within the family of prokaryotic heat shock 70 kDa proteins. Third is the scoring protocols applied to the inverse folding problem.

Animals↗

Dinucleotide relative abundance extremes: a genomic signature.

Early biochemical experiments established that the set of dinucleotide odds ratios or 'general design' is a remarkably stable property of the DNA of an organism, which is essentially the same in protein-coding DNA, bulk genomic DNA, and in different renaturation rate and density gradient fractions of genomic DNA in many organisms. Analysis of currently available genomic sequence data has extended these earlier results, showing that the general designs of disjoint samples of a genome are substantially more similar to each other than to those of sequences from other organisms and that closely related organisms have similar general designs. From this perspective, the set of dinucleotide odds ratio (relative abundance) values constitute a signature of each DNA genome, which can discriminate between sequences from different organisms. Dinucleotide-odds ratio values appear to reflect not only the chemistry of dinucleotide stacking energies and base-step conformational preferences, but also the species-specific properties of DNA modification, replication and repair mechanisms.

Animals↗

Bacterial classifications derived from recA protein sequence comparisons.

RecA protein sequences from 62 eubacterial sources were compared with one another and relative to one archaebacterial RecA-like and a number of eukaryotic RecA-like sequences. Pairwise similarity scores were determined by a novel method based on significant segment pair alignment. The sequences of different species were grouped on the basis of mutually high similarity scores within groups and consistency of score ranges in comparison to other groups. Following this protocol, the gamma-proteobacteria can be subclassified into two major groups, those of mostly vertebrate hosts and those of mostly soil habitat. The alpha-proteobacterial sequences also divide into two distinct groups, whereas classification of the beta-proteobacteria is more complex. The gram-positive bacterial sequences split into three groups of low and three groups of high G+C genome content. However, neither the combined low-G+C-content nor the combined high-G+C-content group nor the aggregate of all gram-positive bacteria form homogeneous groups. The mycoplasma sequences score best with the Bacillus subtilis sequence, consistent with their presumed origin from a gram-positive ancestor. The eukaryotic RAD proteins generally show a single high-scoring segment pair with the proteobacterial RecA sequences around the ATP-binding domain. The bacteriophage T4 UvsX protein aligns best with RecA sequences on two segments disjoint from the ATP-binding domain. The distribution of the most highly conserved regions shared between RecA and noneubacterial RecA-like sequences suggests a mosaic character and evolution of RecA. The discussion considers some questions on the validity and consistency of bacterial classifications derived from RecA sequence comparisons.

Amino Acid Sequence↗

Comparisons of eukaryotic genomic sequences.

A method for assessing genomic similarity based on relative abundances of short oligonucleotides in large DNA samples is introduced. The method requires neither homologous sequences nor prior sequence alignments. The analysis centers on (i) dinucleotide (and tri- and tetra-) relative abundance extremes in genomic sequences, (ii) distances between sequences based on all dinucleotide relative abundance values, and (iii) a multidimensional partial ordering protocol. The emphasis in this paper is on assessments of general relatedness of genomes as distinguished from phylogenetic reconstructions. Our methods demonstrate that the relative abundance distances almost always differ more for genomic interspecific sequence comparisons than for genomic intraspecific sequence comparisons, indicating congruence over different genome sequence samples. The genomic comparisons are generally concordant with accepted phylogenies among vertebrate and among fungal species sequences. Several unexpected relationships between the major groups of metazoa, fungal, and protist DNA emerge, including the following. (i) Schizosaccharomyces pombe and Saccharomyces cerevisiae in dinucleotide relative abundance distances are as similar to each other as human is to bovine. (ii) S. cerevisiae, although substantially far from, is significantly closer to the vertebrates than are the invertebrates (Drosophila melanogaster, Bombyx mori, and Caenorhabditis elegans). This phenomenon may suggest variable evolutionary rates during the metazoan radiations and slower changes in the fungal divergences, and/or a polyphyletic origin of metazoa. (iii) The genomic sequences of D. melanogaster and Trypanosoma brucei are strikingly similar. This DNA similarity might be explained by some molecular adaptation of the parasite to its dipteran (tsetse fly) host, a host-parasite gene transfer hypothesis. Robustness of the methods may be due to a genomic signature of dinucleotide relative abundance values reflecting DNA structures related to dinucleotide stacking energies, constraints of DNA curvature, and mechanisms attendant to replication, repair, and recombination.

Animals↗

Heterogeneity of genomes: measures and values.

Genomic homogeneity is investigated for a broad base of DNA sequences in terms of dinucleotide relative abundance distances (abbreviated delta-distances) and of oligonucleotide compositional extremes. It is shown that delta-distances between different genomic sequences in the same species are low, only about 2 or 3 times the distance found in random DNA, and are generally smaller than the between-species delta-distances. Extremes in short oligonucleotides include underrepresentation of TpA and overrepresentation of GpC in most temperate bacteriophage sequences; underrepresentation of CTAG in most eubacterial genomes; underrepresentation of GATC in most bacteriophage; CpG suppression in vertebrates, in all animal mitochondrial genomes, and in many thermophilic bacterial sequences; and overrepresentation of GpG/CpC in all animal mitochondrial sets and chloroplast genomes. Interpretations center on DNA structures (dinucleotide stacking energies, DNA curvature and superhelicity, nucleosome organization), context-dependent mutational events, methylation effects, and processes of replication and repair.

Animals↗

Which bacterium is the ancestor of the animal mitochondrial genome?

We present considerable data supporting the hypothesis that a Sulfolobus- or Mycoplasma-like endosymbiont, rather than an alpha-proteobacterium, is the ancestor of animal mitochondrial genomes. This hypothesis is based on pronounced similarities in oligonucleotide relative abundance extremes common to animal mtDNA, Sulfolobus, and Mycoplasma capricolum and pronounced discrepancies of these relative abundance values with respect to alpha-proteobacteria. In addition, genomic dinucleotide relative abundance measures place Sulfolobus and M. capricolum among the closest to animal mitochondrial genomes, whereas the classical eubacteria, especially the alpha-proteobacteria, are at excessive distances. There are also considerable molecular and cellular phenotypic analogies among mtDNA, Sulfolobus, and M. capricolum.

Base Composition↗

Geometry of interplanar residue contacts in protein structures.

The relative spatial disposition of interacting side-chain planar groups (aromatic, guanidinium, amide, carboxyl, imidazole) is analyzed for 186 non-homologous well-resolved protein structures. The dihedral angle of amide or carboxyl planar groups with other planar groups accords with a random distribution of planes. By contrast, the dihedral angle of the planes between close aromatic rings or of the histidine ring interacting with aromatic residues is significantly nonrandom, showing an approximately uniform distribution. Our results indicate that edge-to-edge and edge-to-center spatial dispositions of residue planar sections are prevalent, while complete stacking configurations are uncommon. The hypothesis that electrostatic forces are a major determinant of the geometry of interactions between side-chain planar groups is discussed.

Mathematics↗

Statistical studies of biomolecular sequences: score-based methods.

The massive accumulation of DNA and protein sequence data poses challenges and opportunities in terms of interpretation and analysis. This presentation reviews the method of score-based sequence analysis with the objectives of discerning distinctive segments in single sequences and identifying significant common segments in sequence comparisons. A number of new results are described here for both the theory and its applications. These include distributional theory involving several high scoring segments in single sequences, distribution formulas for general scoring regimes in multiple sequence comparisons, bounds for periodic scoring assignments, sensitivity analysis of genome composition and refinements on predicting exons and genes in DNA sequences.

Amino Acid Sequence↗

Measuring residue associations in protein structures. Possible implications for protein folding.

We propose a number of distance measures between residues in protein structures based on average, minimum and maximum distances of all atom (backbone and side-chain) coordinates or with respect to side-chain atom coordinates only. The d1-distance (D1-distance) refers to the average distance between side-chain (backbone and side-chain) atoms of a residue pair in a given structure. The dm-distance (Dm-distance) refers to the minimum distance between side-chain atoms (non-trivial minimum distance between all atoms of a residue pair). For each distance measure, averaging and normalizing over representative protein structures, association values and closeness orderings for all amino acid types are determined. The expected associations of side-chain interactions between oppositely charged residues, among hydrophobic residues and of cysteine with cysteine are confirmed. Several surprising associations are observed relative to (1) the aromatic residues tyrosine and tryptophan, but not phenylalanine; (2) multiple histidine residues; (3) asymmetries of arginine versus lysine, aspartate versus glutamate, alanine versus glycine, and asparagine versus glutamine; (4) absence of correlations of alpha-carbon distances with side-chain distances. The all atoms D1-distance attractions are dominated by steric relationships, with glycine and alanine significantly close to all amino acids, whereas large residues are under-associated with all residue types. In contrast, for the closeness ordering corresponding to the minimum side-chain dm-distance, glycine and alanine are among the least associated. However, in the d1-distance alanine is significantly close to all hydrophobic residues with the exception of tryptophan. The dm-distance preferences display a pervasive attraction for tyrosine by almost all residue types, the prominence of tyrosine and tryptophan in cation-aromatic interactions, and the versatility of histidine in functionality. The principal findings suggest a new perspective on the early and intermediate stages of protein folding.

Amino Acid Sequence↗

Pervasive CpG suppression in animal mitochondrial genomes.

All available complete mitochondrial genomes (21 species) are evaluated for dinucleotide over- and under-representation. The CpG dinucleotide is pervasively under-represented in all animal mitochondria, but it is of variable relative abundance in fungal, protist, and plant mitochondrial genomes. Interpretations and hypotheses are considered relative to mitochondrial genome organization, methylation, structural specificities, directed mutation, and evolutionary events. In particular, our results support Mycoplasma capricolum or a close relative as the most likely bacterial ancestor of the mitochondria.

Base Composition↗

Theoretical recombination processes incorporating interference effects.

With the acquisition of genetic, physical, and sequence maps, linkage relationships among genes (markers) may be more accurately approached in terms of global models for the distribution of recombination events that take into account interference. There are two principal analytical methods used for ascertaining linkage relationships. The first method, the Haldane-Kosambi differential equation approach, has the limitation that all of its calculations rest on consideration of only three gene markers, where recombination depends only on the physical distance between markers. In this formulation the resulting map function is in general not feasible for use with multiple markers. The second method starts with a model of the crossover process from which recombination values are determined. The best studied global recombination processes are based on sequential (renewal) crossover formation processes, the count-location crossover structure, and crossovers evolving by a cascade mechanism. This paper, containing both review and new results, concentrates on two aspects of recombination structures: (i) classifications and characterizations of multimarker crossover distributions; and (ii) analysis of regular and higher order crossover interference forms. In eucaryotic species, the general impression is that positive interference prevails, while in procaryotic and viral organisms, there may be circumstances of negative interference. We would propose in estimating the crossover formation process a binomial count distribution or any other count distribution satisfying property (a) of Theorem 9.1 and a location distribution fitted by the data. It is also reasonable to try one or more obligate crossover points superimposed on independent Poisson processes determining other crossover points. This latter model also generates a situation of positive interference (Theorem 3.1).

Animals↗

Applications of statistical criteria in protein sequence analysis: case study of yeast RNA polymerase II subunits.

We have recently proposed statistical techniques to identify unusual protein sequence features. Extensive mapping of these features to particular groups of proteins may afford new ways of protein classification. Here we present a case study of such analysis by discussing special features of the amino acid sequences of yeast RNA polymerase II, the first eukaryotic RNA polymerase for which all subunits have been sequenced. Specific new suggestions derived from this analysis include: (i) based on unusual charge configurations in some of the sequences, electrostatic forces may play a significant role in subunit interactions; (ii) RPB4, on account of similar charge distribution, may well be grouped together with RNA polymerase II transcription initiation factors.

Amino Acid Sequence↗

Molecular evolution of herpesviruses: genomic and protein sequence comparisons.

Phylogenetic reconstruction of herpesvirus evolution is generally founded on amino acid sequence comparisons of specific proteins. These are relevant to the evolution of the specific gene (or set of genes), but the resulting phylogeny may vary depending on the particular sequence chosen for analysis (or comparison). In the first part of this report, we compare 13 herpesvirus genomes by using a new multidimensional methodology based on distance measures and partial orderings of dinucleotide relative abundances. The sequences were analyzed with respect to (i) genomic compositional extremes; (ii) total distances within and between genomes; (iii) partial orderings among genomes relative to a set of sequence standards; (iv) concordance correlations of genome distances; and (v) consistency with the alpha-, beta-, gammaherpesvirus classification. Distance assessments within individual herpesvirus genomes show each to be quite homogeneous relative to the comparisons between genomes. The gammaherpesviruses, Epstein-Barr virus (EBV), herpesvirus saimiri, and bovine herpesvirus 4 are both diverse and separate from other herpesvirus classes, whereas alpha- and betaherpesviruses overlap. The analysis revealed that the most central genome (closest to a consensus herpesvirus genome and most individual herpesvirus sequences of different classes) is that of human herpesvirus 6, suggesting that this genome is closest to a progenitor herpesvirus. The shorter DNA distances among alphaherpesviruses supports the hypothesis that the alpha class is of relatively recent ancestry. In our collection, equine herpesvirus 1 (EHV1) stands out as the most central alphaherpesvirus, suggesting it may approximate an ancestral alphaherpesvirus. Among all herpesviruses, the EBV genome is closest to human sequences. In the DNA partial orderings, the chicken sequence collection is invariably as close as or closer to all herpesvirus sequences than the human sequence collection is, which may imply that the chicken (or other avian species) is a more natural or more ancient host of herpesviruses. In the second part of this report, evolutionary relationships among the 13 herpesvirus genomes are evaluated on the basis of recent methods of amino acid alignment applied to four essential protein sequences. In this analysis, the alignment of the two betaherpesviruses (human cytomegalovirus versus human herpesvirus 6) showed lower scores compared with alignments within alphaherpesviruses (i.e., among EHV1, herpes simplex virus type 1, varicella-zoster virus, pseudorabies virus type 1 and Marek's disease virus) and within gammaherpesviruses (EBV versus herpesvirus saimiri).(ABSTRACT TRUNCATED AT 400 WORDS)

Adenoviridae↗