Search PubMed⌕ Search

Biomedical subjects

Tamar Schlick

Publications and source records attributed to Tamar Schlick.

At least 19 recordsLinked to original sources

The influence of 10n and 10n+5 linker lengths on chromatin fiber topologies explored by mesoscale modeling.

The structural organization of chromatin is intricately influenced by the length of linker DNA connecting nucleosomes. Some studies have suggested preferred linker lengths of 10n and 10n+5 base pairs (bp) (n = integer). Because these lengths dictate the rotational orientation of successive nucleosomes in the fiber axis, they can markedly affect chromatin fiber compaction and topology. Using a refined mesoscale chromatin model with 5-bp resolution, we investigate the influence of linker DNA periodicity, linker histone density, salt concentration, and starting fiber topology on chromatin architecture for regular fibers versus "life-like" fibers, the latter with irregular spacing between nucleosomes. Our results reveal that regular fibers with 10n linkers exhibit compact zigzag configurations, whereas 10n+5 linkers generate more open and flexible structures. However, these effects are pronounced only for short linker lengths, as longer linkers are more heterogeneous. Moreover, increased linker histone density further enhances compaction for long linker lengths, and lower salt concentration modifies chromatin topologies, diminishing periodicity-driven effects. In addition, any periodicity effect in tightly packed solenoid configurations is much less pronounced. All these trends for regular fibers are reduced in life-like fibers with irregularly spaced nucleosomes, despite having the same average spacing. Moreover, the trend details depend highly on specific features of the fiber architecture as designed in experiments and simulations. Overall, our study highlights how reported differences depend on modeling details and emphasizes the role of linker DNA length in regulating chromatin fiber architecture and its potential implications for genome accessibility and expression.

Chromatin↗

Molecular dynamics simulations reveal subtle consequences of H3K9 and H3K27 tri-methylation on chromatin constituents.

Epigenetic modifications of histone tails are key mechanisms of genome regulation. In particular, tri-methylation of lysines (K) 9 and K27 of the histone H3 tail is important for genome silencing. In this work, we explore, using all-atom molecular dynamics simulations, the effect of these two epigenetic marks on the structure and interactions of the H3 tail in several contexts: isolated tails, nucleosomes, chromatosomes, and stacked nucleosomes. Overall, we find that although the isolated tails do not show significant conformational changes upon methylation, a more flexible and extended H3 tail compared to the native tail results in the nucleosome systems, with K9 methylation effects more pronounced. This change could facilitate the interaction of the tail with protein readers like heterochromatin protein 1 or Polycomb group. We also observe that both methylations increase the interactions of the H3 tail with the linker DNA in the context of the chromatosome, producing a chromatosome with tighter linker DNA, which could favor chromatin compaction. For stacked nucleosomes mimicking i±2 zigzag interactions, methylation of either K9 or K27 reduces the interactions of one of the H3 tails with its parental nucleosome and increases its interactions with the nonparental nucleosome, which could also help compact the chromatin fiber. In the three nucleosome-containing systems, we observe an asymmetry between the two tails, especially in the chromatosome, where one tail extends to interact with the linker DNA. This asymmetry modulates the effect that methylation has on each tail. Thus, overall, methylations of K9 and K27 have a subtle but notable impact on the H3 tail structure and its interactions within the chromatin fiber. These results help explain how this epigenetic modification compacts chromatin fibers and promotes longer-range interactions; these changes also guide how to approximate these effects in coarse-grained chromatin models.

Histones↗

Regulation of DNA repair fidelity by molecular checkpoints: "gates" in DNA polymerase beta's substrate selection.

With an increasing number of structural, kinetic, and modeling studies of diverse DNA polymerases in various contexts, a complex dynamical view of how atomic motions might define molecular "gates" or checkpoints that contribute to polymerase specificity and efficiency is emerging. Such atomic-level information can offer insights into rate-limiting conformational and chemical steps to help piece together mechanistic views of polymerases in action. With recent advances, modeling and dynamics simulations, subject to the well-appreciated limitations, can access transition states and transient intermediates along a reaction pathway, both conformational and chemical, and such information can help bridge the gap between experimentally determined equilibrium structures and mechanistic enzymology data. Focusing on DNA polymerase beta (pol beta), we present an emerging view of the geometric, energetic, and dynamic selection criteria governing insertion rate and fidelity mechanisms of DNA polymerases, as gleaned from various computational studies and based on the large body of existing kinetic and structural data. The landscape of nucleotide insertion for pol beta includes conformational changes, prechemistry, and chemistry "avenues", each with a unique deterministic or stochastic pathway that includes checkpoints for selective control of nucleotide insertion efficiency. For both correct and incorrect incoming nucleotides, pol beta's conformational rearrangements before chemistry include a cascade of slow and subtle side chain rearrangements, followed by active site adjustments to overcome higher chemical barriers, which include critical ion-polymerase geometries; this latter notion of a prechemistry avenue fits well with recent structural and NMR data. The chemical step involves an associative mechanism with several possibilities for the initial proton transfer and for the interaction among the active site residues and bridging water molecules. The conformational and chemical events and associated barriers define checkpoints that control enzymatic efficiency and fidelity. Understanding the nature of such active site rearrangements can facilitate interpretation of existing data and stimulate new experiments that aim to probe enzyme features that contribute to fidelity discrimination across various polymerases via such geometric, dynamic, and energetic selection criteria.

DNA Polymerase beta↗

Role of histone tails in chromatin folding revealed by a mesoscopic oligonucleosome model.

The role of each histone tail in regulating chromatin structure is elucidated by using a coarse-grained model of an oligonucleosome incorporating flexible histone tails that reproduces the conformational and dynamical properties of chromatin. Specifically, a tailored configurational-bias Monte Carlo method that efficiently samples the possible conformational states of oligonucleosomes yields positional distributions of histone tails around nucleosomes and illuminates the nature of tail/core/DNA interactions at various salt milieus. Analyses indicate that the H4 histone tails are most important in terms of mediating internucleosomal interactions, especially in highly compact chromatin with linker histones, followed by H3, H2A, and H2B tails in decreasing order of importance. In addition to mediating internucleosomal interactions, the H3 histone tails crucially screen the electrostatic repulsion between the entering/exiting DNA linkers. The H2A and H2B tails distribute themselves along the periphery of chromatin fibers and are important for mediating fiber/fiber interactions. A delicate balance between tail-mediated internucleosomal attraction and repulsion among linker DNAs allows the entering/exiting linker DNAs to align perpendicular to each other in linker-histone deficient chromatin, leading to the formation of an irregular zigzag-folded fiber with dominant pair-wise interactions between nucleosomes i and i +/- 4.

Computational Biology↗

Correct and incorrect nucleotide incorporation pathways in DNA polymerase beta.

Tracking the structural and energetic changes in the pathways of DNA replication and repair is central to the understanding of these important processes. Here we report favorable mechanisms of the polymerase-catalyzed phosphoryl transfer reactions corresponding to correct and incorrect nucleotide incorporations in the DNA by using a novel protocol involving energy minimizations, dynamics simulations, quasi-harmonic free energy calculations, and mixed quantum mechanics/molecular mechanics dynamics simulations. Though the pathway proposed may not be unique and invites variations, geometric and energetic arguments support the series of transient intermediates in the phosphoryl transfer pathways uncovered here for both the G:C and G:A systems involving a Grotthuss hopping mechanism of proton transfer between water molecules and the three conserved aspartate residues in pol beta's active-site. In the G:C system, the rate-limiting step is the initial proton hop with a free energy of activation of at least 17 kcal/mol, which corresponds closely to measured k(pol) values. Fidelity discrimination in pol beta can be explained by a significant loss of stability of the closed ternary complex of the enzyme in the G:A system and much higher activation energy of the initial step of nucleophilic attack, namely deprotonation of terminal DNA primer O3'H group. Thus, subtle differences in the enzyme active-site between matched and mismatched base pairs generate significant differences in catalytic performance.

Binding Sites↗

Sequential side-chain residue motions transform the binary into the ternary state of DNA polymerase lambda.

The nature of conformational transitions in DNA polymerase lambda (pol lambda), a low-fidelity DNA repair enzyme in the X-family that fills short nucleotide gaps, is investigated. Specifically, to determine whether pol lambda has an induced-fit mechanism and open-to-closed transition before chemistry, we analyze a series of molecular dynamics simulations from both the binary and ternary states before chemistry, with and without the incoming nucleotide, with and without the catalytic Mg(2+) ion in the active site, and with alterations in active site residues Ile(492) and Arg(517). Though flips occurred for several side-chain residues (Ile(492), Tyr(505), Phe(506)) in the active site toward the binary (inactive) conformation and partial DNA motion toward the binary position occurred without the incoming nucleotide, large-scale subdomain motions were not observed in any trajectory from the ternary complex regardless of the presence of the catalytic ion. Simulations from the binary state with incoming nucleotide exhibit more thumb subdomain motion, particularly in the loop containing beta-strand 8 in the thumb, but closing occurred only in the Ile(492)Ala mutant trajectory started from the binary state with incoming nucleotide and both ions. Further connections between active site residues and the DNA position are also revealed through our Ile(492)Ala and Arg(517)Ala mutant studies. Our combined studies suggest that while pol lambda does not demonstrate large-scale subdomain movements as DNA polymerase beta (pol beta), significant DNA motion exists, and there are sequential subtle side chain and other motions-associated with Arg(514), Arg(517), Ile(492), Phe(506), Tyr(505), the DNA, and again Arg(514) and Arg(517)-all coupled to active site divalent ions and the DNA motion. Collectively, these motions transform pol lambda to the chemistry-competent state. Significantly, analogs of these residues in pol beta (Lys(280), Arg(283), Arg(258), Phe(272), and Tyr(271), respectively) have demonstrated roles in determining enzyme efficiency and fidelity. As proposed for pol beta, motions of these residues may serve as gate-keepers by controlling the evolution of the reaction pathway before the chemical reaction.

Computer Simulation↗

Flexible histone tails in a new mesoscopic oligonucleosome model.

We describe a new mesoscopic model of oligonucleosomes that incorporates flexible histone tails. The nucleosome cores are modeled using the discrete surface-charge optimization model, which treats the nucleosome as an electrostatic surface represented by hundreds of point charges; the linker DNAs are treated using a discrete elastic chain model; and the histone tails are modeled using a bead/chain hydrodynamic approach as chains of connected beads where each bead represents five protein residues. Appropriate charges and force fields are assigned to each histone chain so as to reproduce the electrostatic potential, structure, and dynamics of the corresponding atomistic histone tails at different salt conditions. The dynamics of resulting oligonucleosomes at different sizes and varying salt concentrations are simulated by Brownian dynamics with complete hydrodynamic interactions. The analyses demonstrate that the new mesoscopic model reproduces experimental results better than its predecessors, which modeled histone tails as rigid entities. In particular, our model with flexible histone tails: correctly accounts for salt-dependent conformational changes in the histone tails; yields the experimentally obtained values of histone-tail mediated core/core attraction energies; and considers the partial shielding of electrostatic repulsion between DNA linkers as a result of the spatial distribution of histone tails. These effects are crucial for regulating chromatin structure but are absent or improperly treated in models with rigid histone tails. The development of this model of oligonucleosomes thus opens new avenues for studying the role of histone tails and their variants in mediating gene expression through modulation of chromatin structure.

Binding Sites↗

Stereochemistry and position-dependent effects of carcinogens on TATA/TBP binding.

The TATA-box binding protein (TBP) is required by eukaryotic RNA polymerases to bind to the TATA box, an eight-basepair DNA promoter element, to initiate transcription. Carcinogen adducts that bind to the TATA box can hamper this important process. Benzo[a]pyrene (BP) is a representative chemical carcinogen that can be metabolically converted to highly reactive benzo[a]pyrene diol epoxides (BPDE), which in turn can form chemically stereoisomeric BP-DNA adducts. Depending on the TATA-bound adduct's location and stereochemistry, TATA/TBP binding can be decreased or increased. Our previous study interpreted the location-dependent effect in terms of conformational freedom and major-groove space available to BP. Here we further explore specific structural changes of the TATA/TBP complex to help interpret the stereochemical effect in terms of the flexibility of the TATA bases that frame the intercalated adduct. Thermodynamic analyses using molecular mechanics Poisson-Boltzmann surface area (MM-PBSA) yield large standard deviations, which make the computed binding free energies the same within the error bars and point to current limitations of free energy calculations of large and highly charged systems like DNA/protein complexes.

Base Sequence↗

Subtle but variable conformational rearrangements in the replication cycle of Sulfolobus solfataricus P2 DNA polymerase IV (Dpo4) may accommodate lesion bypass.

The possible conformational changes of DNA polymerase IV (Dpo4) before and after the nucleotidyl-transfer reaction are investigated at the atomic level by dynamics simulations to gain insight into the mechanism of low-fidelity polymerases and identify slow and possibly critical steps. The absence of significant conformational changes in Dpo4 before chemistry when the incoming nucleotide is removed supports the notion that the "induced-fit" mechanism employed to interpret fidelity in some replicative and repair DNA polymerases does not exist in Dpo4. However, significant correlated movements in the little finger and finger domains, as well as DNA sliding and subtle catalytic-residue rearrangements, occur after the chemical reaction when both active-site metal ions are released. Subsequently, Dpo4's little finger grips the DNA through two arginine residues and pushes it forward. These metal ion correlated movements may define subtle, and possibly characteristic, conformational adjustments that operate in some Y-family polymerase members in lieu of the prominent subdomain motions required for catalytic cycling in other DNA polymerases like polymerase beta. Such subtle changes do not easily provide a tight fit for correct incoming substrates as in higher-fidelity polymerases, but introduce in low-fidelity polymerases different fidelity checks as well as the variable conformational-mobility potential required to bypass different lesions.

Arginine↗

Predicting candidate genomic sequences that correspond to synthetic functional RNA motifs.

Riboswitches and RNA interference are important emerging mechanisms found in many organisms to control gene expression. To enhance our understanding of such RNA roles, finding small regulatory motifs in genomes presents a challenge on a wide scale. Many simple functional RNA motifs have been found by in vitro selection experiments, which produce synthetic target-binding aptamers as well as catalytic RNAs, including the hammerhead ribozyme. Motivated by the prediction of Piganeau and Schroeder [(2003) Chem. Biol., 10, 103-104] that synthetic RNAs may have natural counterparts, we develop and apply an efficient computational protocol for identifying aptamer-like motifs in genomes. We define motifs from the sequence and structural information of synthetic aptamers, search for sequences in genomes that will produce motif matches, and then evaluate the structural stability and statistical significance of the potential hits. Our application to aptamers for streptomycin, chloramphenicol, neomycin B and ATP identifies 37 candidate sequences (in coding and non-coding regions) that fold to the target aptamer structures in bacterial and archaeal genomes. Further energetic screening reveals that several candidates exhibit energetic properties and sequence conservation patterns that are characteristic of functional motifs. Besides providing candidates for experimental testing, our computational protocol offers an avenue for expanding natural RNA's functional repertoire.

Algorithms↗

Mismatch-induced conformational distortions in polymerase beta support an induced-fit mechanism for fidelity.

Molecular dynamics simulations of DNA polymerase (pol) beta complexed with different incorrect incoming nucleotides (G x G, G x T, and T x T template base x incoming nucleotide combinations) at the template-primer terminus are analyzed to delineate structure-function relationships for aberrant base pairs in a polymerase active site. Comparisons, made to pol beta structure and motions in the presence of a correct base pair, are designed to gain atomically detailed insights into the process of nucleotide selection and discrimination. In the presence of an incorrect incoming nucleotide, alpha-helix N of the thumb subdomain believed to be required for pol beta's catalytic cycling moves toward the open conformation rather than the closed conformation as observed for the correct base pair (G x C) before the chemical reaction. Correspondingly, active-site residues in the microenvironment of the incoming base are in intermediate conformations for non-Watson-Crick pairs. The incorrect incoming nucleotide and the corresponding template residue assume distorted conformations and do not form Watson-Crick bonds. Furthermore, the coordination number and the arrangement of ligands observed around the catalytic and nucleotide binding magnesium ions are mismatch specific. Significantly, the crucial nucleotidyl transferase reaction distance (P(alpha)-O3') for the mismatches between the incoming nucleotide and the primer terminus is not ideally compatible with the chemical reaction of primer extension that follows these conformational changes. Moreover, the extent of active-site distortion can be related to experimentally determined rates of nucleotide misincorporation and to the overall energy barrier associated with polymerase activity. Together, our studies provide structure-function insights into the DNA polymerase-induced constraints (i.e., alpha-helix N conformation, DNA base pair bonding, conformation of protein residues in the vicinity of dNTP, and magnesium ions coordination) during nucleotide discrimination and pol beta-nucleotide interactions specific to each mispair and how they may regulate fidelity. They also lend further support to our recent hypothesis that additional conformational energy barriers are involved following nucleotide binding but prior to the chemical reaction.

Base Pair Mismatch↗

In silico studies of the African swine fever virus DNA polymerase X support an induced-fit mechanism.

The African swine fever virus DNA polymerase X (pol X), a member of the X family of DNA polymerases, is thought to be involved in base excision repair. Kinetics data indicate that pol X catalyzes DNA polymerization with low fidelity, suggesting a role in viral mutagenesis. Though pol X lacks the fingers domain that binds the DNA in other members of the X family, it binds DNA tightly. To help interpret details of this interaction, molecular dynamics simulations of free pol X at different salt concentrations and of pol X bound to gapped DNA, in the presence and in the absence of the incoming nucleotide, are performed. Anchors for the simulations are two NMR structures of pol X without DNA and a model of one NMR structure plus DNA and incoming nucleotide. Our results show that, in its free form, pol X can exist in two stable conformations that interconvert to one another depending on the salt concentration. When gapped double stranded DNA is introduced near the active site, pol X prefers an open conformation, regardless of the salt concentration. Finally, under physiological conditions, in the presence of both gapped DNA and correct incoming nucleotide, and two divalent ions, the thumb subdomain of pol X undergoes a large conformational change, closing upon the DNA. These results predict for pol X a substrate-induced conformational change triggered by the presence of DNA and the correct incoming nucleotide in the active site, as in DNA polymerase beta. The simulations also suggest specific experiments (e.g., for mutants Phe-102Ala, Val-120Gly, and Lys-85Val that may reveal crucial DNA binding and active-site organization roles) to further elucidate the fidelity mechanism of pol X.

African Swine Fever Virus↗

Fidelity discrimination in DNA polymerase beta: differing closing profiles for a mismatched (G:A) versus matched (G:C) base pair.

Understanding fidelity-the faithful replication or repair of DNA by polymerases-requires tracking of the structural and energetic changes involved, including the elusive transient intermediates, for nucleotide incorporation at the template/primer DNA junction. We report, using path sampling simulations and a reaction network model, strikingly different transition states in DNA polymerase beta's conformational closing for correct dCTP versus incorrect dATP incoming nucleotide opposite a template G. The cascade of transition states leads to differing active-site assembly processes toward the "two-metal-ion catalysis" geometry. We demonstrate that these context-specific pathways imply different selection processes: while active-site assembly occurs more rapidly with the correct nucleotide and leads to primer extension, the enzyme remains open longer, has a more transient closed state, and forms product more slowly when an incorrect nucleotide is present. Our results also suggest that the rate-limiting step in pol beta's conformational closing is not identical to that for overall nucleotide insertion and that the rate-limiting step in the overall nucleotide incorporation process for matched as well as mismatched systems occurs after the closing conformational change.

Base Pair Mismatch↗

Electrostatic mechanism of nucleosomal array folding revealed by computer simulation.

Although numerous experiments indicate that the chromatin fiber displays salt-dependent conformations, the associated molecular mechanism remains unclear. Here, we apply an irregular Discrete Surface Charge Optimization (DiSCO) model of the nucleosome with all histone tails incorporated to describe by Monte Carlo simulations salt-dependent rearrangements of a nucleosomal array with 12 nucleosomes. The ensemble of nucleosomal array conformations display salt-dependent condensation in good agreement with hydrodynamic measurements and suggest that the array adopts highly irregular 3D zig-zag conformations at high (physiological) salt concentrations and transitions into the extended "beads-on-a-string" conformation at low salt. Energy analyses indicate that the repulsion among linker DNA leads to this extended form, whereas internucleosome attraction drives the folding at high salt. The balance between these two contributions determines the salt-dependent condensation. Importantly, the internucleosome and linker DNA-nucleosome attractions require histone tails; we find that the H3 tails, in particular, are crucial for stabilizing the moderately folded fiber at physiological monovalent salt.

Computer Simulation↗

Conformational transition pathway of polymerase beta/DNA upon binding correct incoming substrate.

The closing conformational transition of wild-type polymerase beta bound to DNA template/primer before the chemical step (nucleotidyl transfer reaction) is simulated using the stochastic difference equation (in length version, "SDEL") algorithm that approximates long-time dynamics. The order of the events and the intermediate states during pol beta's closing pathway are identified and compared to a separate study of pol beta using transition path sampling (TPS) (Radhakrishnan, R.; Schlick, T. Proc. Natl. Acad. Sci. USA 2004, 101, 5970-5975). Results highlight the cooperative and subtle conformational changes in the pol beta active site upon binding the correct substrate that may help explain DNA replication and repair fidelity. These changes involve key residues that differentiate the open from the closed conformation (Asp192, Arg258, Phe272), as well as residues contacting the DNA template/primer strand near the active site (Tyr271, Arg283, Thr292, Tyr296) and residues contacting the beta and gamma phosphates of the incoming nucleotide (Ser180, Arg183, Gly189). This study compliments experimental observations by providing detailed atomistic views of the intermediates along the polymerase closing pathway and by suggesting additional key residues that regulate events prior to or during the chemical reaction. We also show general agreement between two sampling methods (the stochastic difference equation and transition path sampling) and identify methodological challenges involved in the former method relevant to large-scale biomolecular applications. Specifically, SDEL is very quick relative to TPS for obtaining an approximate path of medium resolution and providing qualitative information on the sequence of events; however, associated free energies are likely very costly to obtain because this will require both successful further refinement of the path segments close to the bottlenecks and large computational time.

DNA↗

Modular RNA architecture revealed by computational analysis of existing pseudoknots and ribosomal RNAs.

Modular architecture is a hallmark of RNA structures, implying structural, and possibly functional, similarity among existing RNAs. To systematically delineate the existence of smaller topologies within larger structures, we develop and apply an efficient RNA secondary structure comparison algorithm using a newly developed two-dimensional RNA graphical representation. Our survey of similarity among 14 pseudoknots and subtopologies within ribosomal RNAs (rRNAs) uncovers eight pairs of structurally related pseudoknots with non-random sequence matches and reveals modular units in rRNAs. Significantly, three structurally related pseudoknot pairs have functional similarities not previously known: one pair involves the 3' end of brome mosaic virus genomic RNA (PKB134) and the alternative hammerhead ribozyme pseudoknot (PKB173), both of which are replicase templates for viral RNA replication; the second pair involves structural elements for translation initiation and ribosome recruitment found in the viral internal ribosome entry site (PKB223) and the V4 domain of 18S rRNA (PKB205); the third pair involves 18S rRNA (PKB205) and viral tRNA-like pseudoknot (PKB134), which probably recruits ribosomes via structural mimicry and base complementarity. Additionally, we quantify the modularity of 16S and 23S rRNAs by showing that RNA motifs can be constructed from at least 210 building blocks. Interestingly, we find that the 5S rRNA and two tree modules within 16S and 23S rRNAs have similar topologies and tertiary shapes. These modules can be applied to design novel RNA motifs via build-up-like procedures for constructing sequences and folds.

Algorithms↗

In vitro RNA random pools are not structurally diverse: a computational analysis.

In vitro selection of functional RNAs from large random sequence pools has led to the identification of many ligand-binding and catalytic RNAs. However, the structural diversity in random pools is not well understood. Such an understanding is a prerequisite for designing sequence pools to increase the probability of finding complex functional RNA by in vitro selection techniques. Toward this goal, we have generated by computer five random pools of RNA sequences of length up to 100 nt to mimic experiments and characterized the distribution of associated secondary structural motifs using sets of possible RNA tree structures derived from graph theory techniques. Our results show that such random pools heavily favor simple topological structures: For example, linear stem-loop and low-branching motifs are favored rather than complex structures with high-order junctions, as confirmed by known aptamers. Moreover, we quantify the rise of structural complexity with sequence length and report the dominant class of tree motifs (characterized by vertex number) for each pool. These analyses show not only that random pools do not lead to a uniform distribution of possible RNA secondary topologies; they point to avenues for designing pools with specific simple and complex structures in equal abundance in the goal of broadening the range of functional RNAs discovered by in vitro selection. Specifically, the optimal RNA sequence pool length to identify a structure with x stems is 20x.

Computational Biology↗

Candidates for novel RNA topologies.

Because the functional repertiore of RNA molecules, like proteins, is closely linked to the diversity of their shapes, uncovering RNA's structural repertoire is vital for identifying novel RNAs, especially in genomic sequences. To help expand the limited number of known RNA families, we use graphical representation and clustering analysis of RNA secondary structures to predict novel RNA topologies and their abundance as a function of size. Representing the essential topological properties of RNA secondary structures as graphs enables enumeration, generation, and prediction of novel RNA motifs. We apply a probabilistic graph-growing method to construct the RNA structure space encompassing the topologies of existing and hypothetical RNAs and cluster all RNA topologies into two groups using topological descriptors and a standard clustering algorithm. Significantly, we find that nearly all existing RNAs fall into one group, which we refer to as "RNA-like"; we consider the other group "non-RNA-like". Our method predicts many candidates for novel RNA secondary topologies, some of which are remarkably similar to existing structures; interestingly, the centroid of the RNA-like group is the tmRNA fold, a pseudoknot having both tRNA-like and mRNA-like functions. Additionally, our approach allows estimation of the relative abundance of pseudoknot and other (e.g. tree) motifs using the "edge-cut" property of RNA graphs. This analysis suggests that pseudoknots dominate the RNA structure universe, representing more than 90% when the sequence length exceeds 120 nt; the predicted trend for <100 nt agrees with data for existing RNAs. Together with our predictions for novel "RNA-like" topologies, our analysis can help direct the design of functional RNAs and identification of novel RNA folds in genomes through an efficient topology-directed search, which grows much more slowly in complexity with RNA size compared to the traditional sequence-based search.

Algorithms↗